IHPST Review All articles
Philosophy of Science

Code in the Dark: The Epistemological Erasure of Scientific Software

IHPST Review
Code in the Dark: The Epistemological Erasure of Scientific Software

The Instrument That Signs Nothing

When a research team publishes a genomics study, the paper will list its authors, acknowledge its funders, and cite its data sources. What it will almost certainly not do is formally credit the alignment algorithm that processed the raw sequencing reads, the statistical pipeline that filtered for significance, or the visualization library that rendered every figure the peer reviewers actually examined. These contributions—without which the findings would not exist in any recognizable form—are absorbed into the background of scientific practice, treated as infrastructure rather than intellect.

This is not a minor administrative oversight. It is a structural feature of how contemporary science represents itself, and it carries consequences that reach into the foundations of epistemology. The question of what counts as a knowledge-producing agent, and who or what deserves recognition for having produced knowledge, is among the oldest in the philosophy of science. Software has forced that question open again, and the discipline has been slow to answer it.

From Calculator to Co-Investigator

The transformation of computational tools in scientific practice did not happen suddenly. For much of the twentieth century, software occupied a role analogous to laboratory glassware: necessary, standardized, and epistemically subordinate. A statistical package performed arithmetic at scale; the scientist interpreted the result. The boundary between calculation and cognition seemed clear enough.

That boundary has since dissolved. Contemporary scientific software does not merely accelerate computation—it makes decisions. Machine learning classifiers in medical imaging determine which tissue patterns warrant clinical attention. Natural language processing tools in social science research identify conceptual categories that human coders would never have isolated. Simulation environments in climate science generate the very phenomena that researchers then describe and publish. In each of these cases, the software is not executing a researcher's pre-formed judgment; it is constituting the epistemic object under investigation.

Philosophers of science have long distinguished between instruments that extend perception and instruments that transform it. A microscope reveals what the eye cannot reach but does not alter the underlying logic of observation. A convolutional neural network trained on labeled images does something categorically different: it encodes a theory of relevance, learned from data, that shapes what the researcher is permitted to see. The philosophical category of "instrument" struggles to contain this.

Attribution's Blind Spot

American scientific publishing operates under attribution norms developed largely in the mid-twentieth century, when the primary knowledge-producing agents were understood to be human beings working with physical materials. The International Committee of Medical Journal Editors' criteria for authorship, which remain enormously influential across disciplines, require intellectual contribution, drafting or critical revision, and accountability—none of which map naturally onto software.

The result is a systematic exclusion. Software that shapes findings is typically mentioned, if at all, in a methods section that specifies a version number and perhaps a citation to the original publication describing the tool. This is the epistemic equivalent of crediting a book's printing press while ignoring its author. The version number is meaningful—reproducibility depends on it—but it is not a substitute for the kind of transparent accounting that would allow a reader to understand how the tool's design choices constrained or enabled the conclusions drawn.

The reproducibility crisis that has roiled fields from psychology to cancer biology has multiple causes, but the opacity of computational pipelines is among the least examined. When a study cannot be replicated, investigators typically scrutinize sample sizes, statistical choices, and researcher degrees of freedom. Rarely do they subject the underlying software to equivalent scrutiny—partly because the norms for doing so do not yet exist, and partly because the software is not understood as a site of epistemic decision-making in the first place.

Responsibility Without a Face

The philosophical stakes extend beyond attribution into the domain of epistemic responsibility. When a finding later proves erroneous, the machinery of scientific accountability looks for human agents who can answer for what went wrong. Authorship, in this sense, is not merely honorific; it is a mechanism for distributing responsibility across a community of knowers.

Software complicates this mechanism in ways that are only beginning to be theorized. Consider a widely used preprocessing package that contains a bug—undiscovered for years—that systematically biases the results of hundreds of studies that employed it. The researchers who used the package are, in the conventional sense, responsible for their findings. But they may have had no means of detecting the error, no access to the source code, and no institutional incentive to audit tools that the field had collectively accepted as reliable. The responsibility is real but diffusely distributed across developers, maintainers, journal editors, and the broader community that normalized the tool's use without demanding transparency.

This is not a hypothetical scenario. Cases of consequential software errors in published science have been documented across fields including genomics, neuroimaging, and econometrics. In each instance, the error was eventually traced not to a researcher's analytical decision but to a computational layer that had been treated as epistemically inert. The invisibility of software in attribution systems is, in this light, also an invisibility of risk.

Toward a Philosophy of Computational Authorship

Addressing this problem requires more than updating citation guidelines, though that would be a start. It requires a more fundamental philosophical reckoning with what it means to produce scientific knowledge in a computational era.

Some scholars have proposed treating software as a form of methodology that should be subject to the same standards of transparency and peer evaluation as any other methodological choice. Others have argued for extending something like authorship to software by requiring that the humans responsible for its design and maintenance be formally acknowledged in any publication that relies on it. Still others have suggested that the very concept of authorship is too individualist and too bound to print-era assumptions to serve the needs of networked, computational science.

Each of these proposals carries its own difficulties. Requiring acknowledgment of software developers in every paper that uses their tools would generate lists of hundreds of names in fields like genomics, where pipelines stack dozens of packages. Treating software as pure methodology sidesteps the question of whether the tool's design embodies theoretical commitments that should themselves be subject to scrutiny.

What seems clear is that the current arrangement—in which software is simultaneously central to knowledge production and absent from knowledge attribution—is philosophically untenable. Science's self-representation depends on the premise that findings can be traced to accountable agents and reproducible procedures. When the most consequential steps in that procedure are performed by code that is unnamed, unreviewed, and sometimes proprietary, that premise is quietly undermined with every paper published.

The Archive That Forgot Its Own Methods

Future historians of science will face a peculiar challenge when they turn to the computational literature of the early twenty-first century. They will find papers that describe findings in meticulous detail, attribute credit to named researchers, and gesture toward methods sections that specify software versions now long obsolete and unavailable. The actual epistemic machinery that produced those findings will be, in many cases, irrecoverable.

This is not merely an archival inconvenience. It is a rupture in the chain of reasoning that allows science to be understood as a cumulative, self-correcting enterprise. Recognizing software as a genuine collaborator in knowledge production—rather than a neutral conduit through which human intelligence flows—is a prerequisite for closing that rupture. The philosophy of science has the conceptual tools to begin this work. What remains is the institutional will to apply them.

All Articles

Related Articles

When the Gavel Defined the Data: Litigation, Legal Discovery, and the Remaking of Scientific Evidence

When the Gavel Defined the Data: Litigation, Legal Discovery, and the Remaking of Scientific Evidence

Measured Into Existence: How the Standardized Test Became Science's Epistemological Blueprint

Measured Into Existence: How the Standardized Test Became Science's Epistemological Blueprint

Enclosing the Scientific Commons: How Patent Expansion Turned Fundamental Research into a Credentialed Marketplace

Enclosing the Scientific Commons: How Patent Expansion Turned Fundamental Research into a Credentialed Marketplace