Skip to content

Reproducibility Is a Verification Problem, and AI Helps Only If You Point It at Checking

Science's replication crisis is a failure of checking, not producing. AI helps a checking problem only when aimed at verification — point it at production and it industrializes the noise.

By Mehdi9 min read
Share
On this page

Science has a reproducibility crisis, AI is about to collide with it, and whether that collision heals the crisis or industrializes it comes down to a single distinction: reproducibility is a verification problem, and AI helps a verification problem only when you aim it at checking rather than at producing. Point the same models at generating papers, hypotheses, and results and you do not accelerate science — you accelerate the manufacture of confident, plausible, unfalsified claims, which is precisely the disease. The tool is genuinely double-edged. The edge you get depends entirely on which side you point at reality.

I want to argue this from inside the field where I feel it most, computational biology, and then generalize the mechanism, because the mechanism is the whole point.

The crisis is real, and it is a failure of checking

Start with the numbers, stated carefully. The Reproducibility Project: Psychology (Open Science Collaboration, 2015) attempted 100 replications and got significant effects in the same direction as the original in roughly 36% of cases, with replication effect sizes about half the originals on average. In preclinical biology the picture is worse: Begley and Ellis (2012) reported that Amgen scientists could reproduce only 6 of 53 "landmark" oncology studies; a Bayer group (Prinz et al., 2011) reported reproducing well under half of its target set. The Reproducibility Project: Cancer Biology, completing in 2021, could not even attempt many experiments because the original papers did not describe their methods in enough detail to repeat them — and where it could, effect sizes ran dramatically smaller than published.

Note what these failures are and are not. They are not, mostly, fraud. They are the predictable output of a publication system that rewards the production of surprising positive results and under-rewards the verification of them. The specific mechanisms are well characterized. Small samples produce high-variance estimates, so the significant findings that survive are disproportionately the ones that got lucky — the "winner's curse" inflating effect sizes. Batch effects in high-dimensional assays confound biology with the day the sample was run, the reagent lot, the technician. P-hacking and the garden of forking paths — Gelman and Loken's term for the fact that a researcher can make dozens of defensible analysis choices after seeing the data and still honestly believe they made only one — turn a nominal p < 0.05 into something with far less than a 5% false-positive rate. Publication bias then filters the literature so that the file drawer full of null results never appears, and the meta-analyst sees only the survivors.

Every one of those is a checking failure. The finding got produced. What did not happen was adequate verification before it entered the record and got cited as if it were load-bearing. This is the frame that makes the AI question tractable: we are not short of scientific claims. We are drowning in them. We are short of trustworthy checks.

There is a related and deeper point about what a "result" even is, which I've argued elsewhere: in biology the ground truth is the bottleneck, not the analysis. A model or an assay can be internally immaculate and still be measuring an artifact. Reproducibility is the crude, essential social technology we use to distinguish signal from artifact after the fact. It is verification standing in for the ground truth we couldn't nail down the first time.

The optimistic case: AI as verifier

Aim these systems at the checking problem and the upside is large and, importantly, mostly not speculative — some of it already works.

Statistical auditing is the cleanest example because it is nearly deterministic. The GRIM test (Brown and Heathers, 2017) checks whether a reported mean is even arithmetically possible given the sample size and the fact that the underlying data are integers — a startling number of published means are not. Tools in the statcheck lineage (Nuijten et al., 2016) recompute p-values from reported test statistics and degrees of freedom and flag internal inconsistencies; applied at scale they found reporting errors in roughly half of psychology papers and decision-flipping errors in around one in eight. None of that needs a large language model. But LLMs extend it enormously: they can read the messy prose of a methods section, extract the design, and reason about whether the reported analysis matches it — the part that used to require a human referee with time nobody has.

Image forensics is the second front. Duplicated, spliced, and reused Western blots and microscopy images are a large share of retractions; Elisabeth Bik's manual work demonstrated the scale, and it is exactly the kind of pattern-matching that vision models do at volume no human can match. Detecting the garden of forking paths is subtler but approachable: given the data and the analysis code, a model can enumerate the reasonable alternative specifications and report how fragile the headline result is to defensible choices — a multiverse analysis run automatically instead of never.

Then there is synthesis across the literature. A system that can read ten thousand papers and actually track which effects replicate, which methods differ, and where the file drawer is leaking can do meta-analysis at a scale and consistency no research group sustains by hand. And on the front end, AI can enforce the boring hygiene that prevents crises: standardizing method reporting so the next Reproducibility Project can actually rerun the experiment, containerizing analysis code so "it ran on my machine" stops being a defense, and attempting computational reproduction — rerunning the authors' pipeline on the authors' data before a paper is accepted — as a routine gate.

Notice the common structure. In every case the AI is checking an existing claim against something — against arithmetic, against the raw image, against the code, against the rest of the literature. It is doing verification work. And verification has a property that makes AI unusually well suited to it: the answer is checkable. You can tell whether the recomputed p-value matches. That is the whole game.

The dangerous case: AI as producer

Now point the identical capability at production and watch the sign flip.

The literature can be flooded with AI-generated papers and AI-generated peer reviews faster than any check can absorb them. Paper mills already existed as a manual industry; generative models drop their marginal cost toward zero. Peer review, the one human verification bottleneck we have, gets attacked from both ends — reviewers quietly delegating to a model that hallucinates approval, and submissions engineered to pass that model. When both the claim and its check are generated by the same kind of system, the check verifies nothing; it launders fluency into apparent legitimacy.

The sharper danger is fabricated data that is statistically correct. The old forgeries failed because humans are bad at faking the low-level texture of real data — Benford's Law violations, too-clean digit distributions, variances that are too tidy. A generative model can produce synthetic datasets that pass exactly those forensic tests, because reproducing the statistical texture of real data is the thing generative models are built to do. GRIM and statcheck catch the careless; they do not catch a model that has learned what real data looks like. So the very forensic advances of the last decade risk being neutralized by a system trained to satisfy them. Detection and generation are locked in the same escalation, and generation has the structural advantage because it only has to fool the current detector.

The most seductive failure is automated hypothesis-and-result generation: a pipeline that proposes a hypothesis, "analyzes," and emits a confident finding — with no experiment in the loop. This produces claims with perfect surface plausibility and zero contact with reality, and it is worth being precise about why that is fatal rather than merely incomplete. Science is not the production of plausible claims. On the Popperian account — and I think this is the part of the tradition that survives Kuhn's and Lakatos's amendments intact — a claim earns scientific standing by surviving a test it could have failed. Falsification requires an experiment that can lose. A language model does not run experiments; it models the distribution of text about experiments. Ask it for a result and it returns the most probable-sounding result, which is optimized for plausibility — the exact quantity that is uncorrelated with truth once you're in the regime where wrong answers look right.

This is why the dream of the fully automated scientist is a category error rather than merely an engineering challenge not yet solved. Knowledge, in the justified-true-belief tradition and everything after Gettier that tried to repair it, requires the belief to be connected to the fact in the right way — by evidence that tracks reality. A generated result is a belief with the form of justification (it cites, it hedges, it reports a p-value) and none of the substance, because nothing in its causal history touched the world it purports to describe. It is a Gettier case manufactured at industrial scale: statements that happen to look justified and true but are true, when they are, by accident. That is not knowledge. It is testimony from a source that has never observed anything — and the epistemology of testimony (Hume, Coady, and the modern debate) turns entirely on the reliability of the source. An automated producer is a maximally fluent, maximally unreliable witness.

The asymmetry decides everything

Line the two cases up and the governing principle is stark.

AI as verifier AI as producer
What it does checks a claim against arithmetic, images, code, data, the literature emits new claims, data, reviews, findings
Is the output checkable? yes — recompute and compare not without independent contact with reality
Effect on the record filters noise out adds noise, hard-to-detect noise
Error direction false alarms, which humans adjudicate confident falsehoods that look like signal

This is the same asymmetry that shows up everywhere in how these systems create or destroy value: generation is cheap and unreliable; verification is where trust is actually manufactured. A proof is hard to find and easy to check. A replication is expensive to run and, once run, dispositive. AI is a lever, and a lever amplifies whatever you attach it to. Attach it to checking and you get more trustworthy science per unit of human attention. Attach it to producing and you get more claims per unit of reality, which in a system already failing at the ratio of claims to reality is not progress — it is the crisis with a bigger engine.

The practical instruction follows without much interpretation, and I'll state my recommendations as recommendations. For scientists and for journals: deploy AI aggressively on the verification side and treat AI on the production side as an adversarial input to defend against. Run automated statistical audits — GRIM, statcheck-style consistency, distributional sanity checks, multiverse fragility — on every submission as a gate, not a courtesy. Require raw data and executable analysis code, and attempt computational reproduction before acceptance rather than trusting the reported result. Make provenance the load-bearing defense rather than post-hoc detection, because provenance wins where forensics loses: cryptographically signed, timestamped raw data and analysis pipelines establish that a dataset existed and was collected, which is exactly what a generative model cannot forge and a detector cannot reliably confirm. And screen generated text and images as what they are — potential attacks on the record — while never trusting a generated result as evidence of anything about the world.

I'll label the forecast as mine: I think the near-term equilibrium gets worse before it gets better, because production is being deployed faster than verification, and the incentives that built the crisis — reward the surprising positive, ignore the null — point the cheap new engine straight at the production side. The optimistic outcome is available but it is not the default. It requires a deliberate institutional choice to spend the AI dividend on checking rather than on output, against a gradient that pushes the other way.

The reassuring thing, and the reason I am not simply pessimistic, is that the asymmetry is not a matter of opinion. Verification is checkable and generation is not, and no amount of scaling changes which side of that line a given deployment sits on. Science's trust problem was never a shortage of ideas. It was, and remains, a shortage of contact with reality — and the only version of AI that helps is the version pointed at the reality, not at the page.

Frequently asked questions

Is the reproducibility crisis real, or overstated?
It is real and well-documented, though its size varies by field. The Reproducibility Project: Psychology (Open Science Collaboration, 2015) successfully replicated only about 36% of 100 studies. In preclinical cancer biology, Amgen scientists reported reproducing 6 of 53 landmark studies (Begley and Ellis, 2012), and a Bayer group reported similarly low rates (Prinz et al., 2011). The Reproducibility Project: Cancer Biology (2021) found effect sizes on average far smaller than originally reported. The exact fraction is contested and field-dependent, but that a large share of published findings fail to replicate is not seriously in dispute.
Can AI actually detect fabricated data or image manipulation?
Partially, and the balance is shifting. Tools already flag duplicated or spliced Western blot images, impossible statistical distributions, and GRIM-test inconsistencies (whether reported means are arithmetically possible given sample size). But the same generative capacity that detects fabrication can also produce fabrications with correct low-level statistics, which is why detection cannot be the whole strategy. Provenance — signed, timestamped raw data and analysis code — is more durable than post-hoc forensics.
Why won't an 'automated scientist' just solve reproducibility by doing better science?
Because the binding constraint in a replication crisis is contact with reality, not idea generation. A system that generates hypotheses and confident results without running them against the world produces fluent, plausible claims that have never been tested — exactly the failure mode reproducibility is meant to catch. Falsification requires an experiment that can lose. Language models do not run experiments; they model text. Automating the production of claims without automating their exposure to disconfirmation makes the crisis worse, not better.
What should a journal or lab actually do differently?
Point AI at the checking side and treat AI on the production side as a threat model. Concretely: run automated statistical audits (GRIM, statcheck-style consistency, distributional sanity) on every submission; require raw data and executable analysis code with cryptographic provenance; use AI to attempt computational reproduction of results before acceptance; and screen for AI-generated text and images as adversarial inputs. The asymmetry is the design principle — verification, not generation.

Filed under Cross-Disciplinary Deep Essays. Where biology, computation, markets, and philosophy collide.

Essays like this, in your inbox.

Thoughtful essays. No spam. Unsubscribe anytime.

Cross-Disciplinary Deep Essays

"AGI by Year X" Is Unanswerable Until You Name the Definition

"Are we close to AGI?" is incoherent because AGI names at least four incompatible criteria that come apart in practice. Separate them and the timeline debate dissolves into concrete, checkable questions.

9 min read