The scientific literature passed the limit of human comprehension a long time ago, and everyone in research quietly works around it. No living immunologist has read immunology. No one holds oncology, or even a subfield of it, in a single head. My bet — and it is a forecast, so I will argue for it rather than assert it — is that the most valuable thing research agents will do in the next few years is not design experiments or propose theories, but read the whole corpus carefully and synthesize it: find what is already known but sits unconnected across two subfields, surface contradictions no single reviewer ever saw together, and map the actual state of the evidence. That is a job no human can do at scale, and it is the one closest to being real.
The catch is a discipline, not a capability, and it decides whether the same tool serves truth or industrializes noise. I will get to it. First, the size of the gap.
The corpus outgrew the reader decades ago
Do the arithmetic, because it is the whole premise. PubMed indexes north of 35 million citations and takes in well over a million new ones a year — on the order of a few thousand biomedical papers a day. A specialized subfield — say, epigenetic aging clocks, which is part of what I work on — might hold 50,000 papers. A diligent researcher who read ten papers a day, every day, with no weekends and no experiments to run, would clear that subfield in about fourteen years, by which point it would have roughly tripled. Nobody does this. What people actually do is read a few hundred papers deeply, track a few dozen groups, and infer the rest from reviews, talks, and the ambient sense of the field.
That inference layer is where knowledge goes missing. A review article is one author's compression of a slice they happened to see, written on a two-year lag. The "ambient sense of the field" is a social artifact — it tracks who is loud, who is at the right institutions, and which results were fun to repeat, none of which is the same as what the evidence says. The literature is not a body of knowledge that humans hold. It is a body of knowledge that no one holds, sampled thinly and unrepresentatively by each of us.
So a great deal of scientific value is not locked inside any paper. It is locked in the gaps between papers that nobody read together.
The value is in the gaps, and there is a real precedent for mining them
The clearest historical proof that these gaps are real and exploitable comes from Don R. Swanson, an information scientist at the University of Chicago who in 1986 coined the term undiscovered public knowledge. His argument: if one literature establishes that A affects B, and a separate literature — one that does not cite the first and is read by different people — establishes that B affects C, then "A may affect C" is a testable hypothesis already implied by the published record but stated by no one, because no one read both literatures. He worked the example by hand. Dietary fish oil was known to reduce blood viscosity and platelet aggregation; Raynaud's syndrome was known to involve high blood viscosity and vascular reactivity; the two literatures did not talk to each other. Swanson proposed, from reading alone, that fish oil might help Raynaud's. It was later supported clinically. He did the same for magnesium and migraine.
Swanson found these by manually cross-reading disjoint literatures — an artisanal process he could run a handful of times in a career. The field of literature-based discovery grew out of it, and it has stayed niche for one boring reason: the reading does not scale on human hands. This is precisely the shape of task that changes character when an agent can hold the whole corpus. The forecast is not that agents will invent a new kind of insight. It is that they will make Swanson's move cheap and routine instead of heroic.
The gaps come in at least three flavors, and each is a distinct near-term product:
- Cross-subfield connection. A result in one community answers an open question in another that never cites it. This is Swanson's case, and it is a retrieval-and-inference problem at heart — exactly what large-scale reading is good at.
- Contradiction detection. Study X reports an effect; studies Y and Z fail to find it under slightly different conditions; no reviewer ever placed all three side by side because they sit in different journals under different keywords. Surfacing the contradiction — and the boundary condition that might reconcile it — is pure collation at scale.
- Meta-patterns. A regularity visible only across thousands of methods sections: a reagent that correlates with a class of results, a covariate that flips a sign whenever it is included, a p-value distribution that betrays selective reporting across a whole subfield. No individual reader has the sample size to see these. An agent that has read all of it does.
Systematic review is the existing, respectable version of this work, and it shows both the demand and the ceiling. A Cochrane-style review screens thousands of abstracts to include a few dozen studies, takes a year or more, costs a great deal, and is partly stale by publication. What agents threaten to change is not the rigor of that process but its throughput — turning a year into a week, and letting the review be re-run the day a new trial posts. That alone would reshape what "knowing the literature" means.
The discipline: read the primary source, keep the caveats
Here is where the whole thing turns, and it is the point I keep returning to. A summary is lossy compression, and the loss is not random — it deletes exactly the caveats, effect sizes, and boundary conditions that decide whether a claim is true. I have argued elsewhere that reading the primary source is becoming a superpower precisely because the codec of summarization is tuned to keep the memorable headline and drop the machinery of doubt. That argument does not soften when you point it at an agent. It sharpens into a design constraint.
A synthesis agent that reads abstracts is worse than useless. Not weaker — actively harmful. The abstract is already the authors' own lossy compression, written to maximize citations, and it keeps "intervention reverses epigenetic age" while dropping that n was small, the work was in vitro on one cell type, the clock was never validated on that tissue, and the measured reversal sits inside the clock's own error band. Every one of those caveats is fatal to the headline, and none survives to the abstract. Now run synthesis across ten thousand such abstracts. The agent aggregates ten thousand headlines and discards ten thousand sets of caveats, then hands you a confident map of "consensus" that is really a map of overclaims with the qualifications stripped out. It has laundered lossiness into authority. A human skimming abstracts at least knows they are skimming; an agent that produces a polished cross-paper synthesis erases that signal, and the synthesis gets cited as though it read the papers.
So the requirement is concrete and non-negotiable: the agent must read the methods and the results table, not the abstract, and it must carry the caveats forward with every claim it propagates. When it tells you "fourteen studies support X," it has to attach, per study, the effect size, the sample, the population, the conditions, and the one discussion sentence conceding what could break the result. A synthesis that says "significantly associated with mortality" without carrying "hazard ratio 1.05" has told you nothing and made you feel informed, which is the worse of the two states. The value of reading the whole literature exists only if the reading goes to the layer where nobody was trying to make you feel anything.
The peril: the same tool industrializes the literature's biases
Comprehensiveness is not neutrality. The published record is a biased sample of the experiments that were actually run — positive results get published, null results sit in file drawers, and p-hacking tilts the surviving effects. An agent that faithfully synthesizes the literature will faithfully reproduce every one of those distortions, and it will do so with a fluency and completeness that makes the bias more convincing, not less. Point the same capability at production rather than verification and you have built a machine for turning publication bias into apparent consensus at scale. This is the same fork I have argued runs through AI and the reproducibility crisis: the technology helps a checking problem only when it is aimed at checking, and deepens the problem the moment it is aimed at fluent production.
Which direction you get is decided entirely by whether the agent verifies or merely aggregates. Verifying looks like weighting by study quality and sample size instead of counting papers; hunting the null results in registries and preprints that never made it to a journal; reporting effect-size heterogeneity rather than averaging a real contradiction into a comfortable mean; and — again — carrying caveats so a reader can see when the confidence is thin. Aggregating looks like summing what the corpus says and presenting the sum as the state of nature. The first fights the literature's biases. The second industrializes them. Same model, same corpus, opposite epistemics, and the difference is a posture you have to demand explicitly, because nothing in the training objective supplies it for free.
The honest limit, and what to demand
The strongest counterargument is that judging a methods section is real expertise the machines mostly do not have yet, and I think that is true. Deciding whether an adjustment set is adequate, whether a clock transfers across tissues, or whether an effect is an artifact of one lab's protocol is domain judgment, and current systems do it unreliably when they do it at all. Citation hallucination is a live failure mode; a synthesis is worthless if some of its supporting papers do not say what it claims or do not exist. None of this is a reason to dismiss the forecast, but all of it bounds it.
The bound points at the right division of labor. The near-term win is not the agent adjudicating the evidence. It is the agent doing the one thing humans genuinely cannot — comprehensive, careful reading and cross-paper collation at corpus scale — and staging the primary-source passages so a human expert makes the call. Keep the machine on the reading and the retrieval, keep the human on the verification, and insist the whole chain stay auditable back to the methods.
So the thing to demand of any literature-synthesis agent is narrow and testable: cite the primary source, quote the methods-level passage, and preserve the caveat — never the abstract-level claim, never a naked "studies show." If it cannot show you the results table it is standing on, it did not read the literature. It read the headlines, and it is selling you the field's overclaims with the doubt filed off.