The Oracle Without an Explanation: Machine Learning, Predictive Supremacy, and the Philosophical Abandonment of Scientific Causation
A New Kind of Knowing
For most of the history of Western science, prediction and explanation were understood as complementary—ideally unified—epistemic achievements. To explain a phenomenon was, in the classical Hempelian account, to subsume it under a law; and a law, once established, enabled prediction. The covering-law model of scientific explanation assumed that understanding why something happened and being able to anticipate that it would happen again were two aspects of the same cognitive accomplishment. A theory that predicted without explaining was, within this framework, philosophically incomplete—a useful instrument, perhaps, but not genuine knowledge.
That assumption is now under sustained pressure from within scientific practice itself. The rise of machine learning as a dominant methodology across multiple scientific disciplines has introduced a new epistemic paradigm in which prediction and explanation are not merely separable but routinely traded against each other—and in which prediction increasingly wins. The implications of this trade are profound, and they have not yet received the philosophical scrutiny they deserve.
Genomics and the Unreadable Map
Consider the case of polygenic risk scores in human genomics. These statistical constructs, derived from genome-wide association studies and refined through machine learning algorithms, can predict with meaningful accuracy an individual's relative risk of developing conditions ranging from type 2 diabetes to schizophrenia. They are genuine achievements of computational science, and their clinical utility is real. They are also, in a philosophically important sense, explanatorily empty.
A polygenic risk score aggregates the statistical contributions of thousands or millions of genetic variants, most of which individually account for a vanishingly small fraction of the phenotypic variance being predicted. The algorithm that produces the score does not model the biological mechanisms by which those variants influence the trait in question. It identifies correlations in large datasets and assigns weights to variants based on their predictive utility, without any commitment to—or even interest in—the causal pathways that connect genotype to phenotype.
This is not a temporary limitation awaiting resolution by future research. It is a structural feature of the approach. The dimensionality of the genomic data involved, combined with the complexity of gene-environment interactions, makes mechanistic modeling of the relevant pathways computationally intractable by current methods. Machine learning fills the gap not by illuminating the mechanism but by bypassing the requirement for one. The result is a predictive instrument that functions without a corresponding explanation—an oracle that tells you what will happen without telling you why.
The philosophical question this raises is not merely technical. If science's purpose is understanding—if explanation, not merely prediction, is the goal—then polygenic risk scores represent a significant departure from that purpose, however useful they may be as clinical instruments. The departure is rarely acknowledged as such, because the vocabulary of scientific achievement has quietly shifted to accommodate it: success is now measured in predictive accuracy metrics, and the absence of mechanistic explanation is treated as a secondary concern rather than a fundamental limitation.
Climate Modeling and the Retreat from Mechanism
The tension between prediction and explanation is equally visible in contemporary climate science, though its contours differ. Earth system models have historically been constructed around physical principles—fluid dynamics, thermodynamics, radiative transfer—that provide genuine explanatory purchase on the phenomena they describe. A modeler who can tell you that global mean temperatures will rise by 2.5 degrees Celsius under a given emissions scenario, and who can trace that prediction through the physical mechanisms of the model, is doing something philosophically different from a system that arrives at the same number through pattern-matching on historical climate data.
The introduction of machine learning into climate modeling has begun to blur this distinction in ways that are scientifically productive and philosophically troubling in equal measure. Neural networks trained on observational data can emulate the outputs of computationally expensive physical models with remarkable fidelity, enabling faster projections across a wider range of scenarios. They can also identify patterns in climate data that elude conventional analysis. But the patterns they identify are not explained—they are detected. The network does not know why the pattern exists; it knows only that it does, and that it has predictive utility.
When machine learning components are embedded within larger physical models, the explanatory status of the resulting hybrid becomes difficult to assess. Which outputs reflect genuine physical understanding, and which reflect statistical regularities in training data whose causal basis remains opaque? The question is not merely academic. Policy decisions about emissions targets, infrastructure investment, and adaptation strategies are made on the basis of climate model outputs. The philosophical status of those outputs—whether they represent causal understanding or sophisticated pattern recognition—matters for how much confidence we should place in their projections under genuinely novel conditions, precisely the conditions that climate change is producing.
Drug Discovery and the Molecule That Cannot Be Understood
In pharmaceutical research, machine learning has been deployed with considerable enthusiasm for the task of identifying candidate drug molecules—compounds predicted to bind effectively to target proteins, exhibit favorable pharmacokinetic properties, and avoid known toxicity profiles. Systems such as AlphaFold, which predicts protein structures with remarkable accuracy, and various generative chemistry platforms represent genuine advances in the efficiency of early-stage drug discovery.
But the molecules these systems identify are frequently ones that no human chemist would have proposed on mechanistic grounds. They are optimized for predicted outcomes across a high-dimensional parameter space in ways that resist intuitive interpretation. When such a molecule enters clinical development and fails—as the majority of drug candidates do—the failure is often as opaque as the original prediction. The model that predicted success does not provide a causal account of why the molecule failed, because it never provided a causal account of why it was expected to succeed.
This opacity has consequences for the iterative, learning-based character of pharmaceutical science. Drug discovery has historically been a domain in which failures are as scientifically valuable as successes, because understanding why a molecule fails illuminates the biology of the target and guides subsequent development. When failures are generated by systems that cannot explain their predictions, that iterative learning is disrupted. The failure becomes data for the next training cycle, but not understanding for the next generation of researchers.
What We Lose When We Stop Asking Why
The philosopher of science Peter Lipton argued that the best scientific explanations are not merely predictively accurate but genuinely illuminating—they produce what he called the "aha" of understanding, a felt sense that one now grasps something about the world that was previously opaque. This is not merely a psychological criterion; it tracks something real about the difference between knowing that and knowing why.
Machine learning's predictive achievements are genuine and in many cases extraordinary. But they do not, in general, produce Lipton's "aha." They produce confidence in predictions without the accompanying sense of having grasped an underlying structure. And when a discipline's primary epistemic achievements are of this character—when its most celebrated results are outputs of systems whose internal operations are opaque even to their designers—there is reason to ask whether the discipline is still doing science in the fullest philosophical sense of that term, or whether it has become something else: a sophisticated form of empirically grounded prophecy.
This is not an argument against machine learning in science. It is an argument for philosophical clarity about what machine learning provides and what it does not—and for institutional structures that preserve space for the causal, mechanistic inquiry that prediction-focused methods tend to crowd out. The question of why things happen is not a luxury that science can afford to defer indefinitely. It is the question that gives scientific knowledge its depth, its transferability, and its capacity to generate genuine understanding rather than merely useful forecasts.