Scientific authorship is not primarily a credit line. It is a chain of responsibility. An author is the person who vouches for the work and answers for it when it fails — who fields the questions, investigates the anomaly a reader flags, signs the correction and, if it comes to that, the retraction. As agents begin to do real scientific labor — designing analyses, writing and running the pipeline, drafting the results section — the institutions built on human authorship will have to adapt. The adaptation is not the one most people are bracing for. The hard problem is not how to credit the machine. It is that authorship encodes accountability, and an agent has no accountability to give. You can list a tool; you cannot make it answerable. So the humans on the paper must remain fully responsible for everything the agent touched — which means they must actually understand and verify it, the same verification burden that surfaces everywhere agents do consequential work.
"AI as co-author" is therefore a structural question wearing the costume of a formatting question. The debate gets framed as taxonomy — which byline, which acknowledgement, which disclosure field — as if the only issue were where to put the label. The real issue is what a byline does. It assigns liability. And liability cannot be assigned to a thing that cannot be held to it.
Authorship bundles credit with liability on purpose
Start with what an author formally is, because the field already wrote it down. The most widely adopted standard in biomedical science, the ICMJE criteria, requires an author to meet all of a short set of conditions: substantial contribution to the work; drafting or critically revising it; approval of the final version; and — the one that matters here — agreement to be accountable for all aspects of the work, including the integrity and accuracy of every part, with a standing obligation to investigate and resolve questions raised about it. That last criterion is not decoration. It is the load-bearing wall. It says authorship is a promise to answer.
Notice that credit and accountability are deliberately bundled. You do not get to claim you conceived the central idea while disclaiming the botched analysis three floors down. The institution is engineered so that reward and exposure travel together, which is exactly what makes the reward mean something. A byline is worth having because it is also a bond posted against the work being real.
An agent can satisfy the first two criteria and fails the last two absolutely. It can contribute substantially — more than substantially — and it can draft. It cannot approve a final version, because approval is a commitment and it has nothing with which to commit. It cannot be accountable, because accountability is enforced through consequence, and there is no consequence a model checkpoint can suffer. You cannot depose it, sanction it, strike it off, or make it fund the correction. This is the same structural fact I have argued sets the ceiling on delegation generally: an AI agent has no skin in the game, no bond to forfeit and no license to lose, so the liability never lands on it and always relocates onto the nearest human with standing. In authorship, that human is whoever signs.
The journals reached this conclusion fast, and for the right reason. When large language models started appearing in manuscripts in early 2023, the major publishers — Nature's family, Science, and others — converged within months on a policy that looks superficially like a rule about credit but is actually a rule about accountability: an LLM cannot be listed as an author because it cannot take responsibility for the content or consent to the submission, and its use must instead be documented in the methods or acknowledgements. They did not ban the tool. They banned the pretense that the tool could stand where an author stands. That was correct, and it generalizes from chatbots-that-wordsmith to agents-that-do-experiments without softening. If anything the agent case makes it sharper, because the agent contributes more of the thing you are being asked to vouch for.
The four tensions, and which one is fake
Four pressures are converging on the authorship system. Three are real problems the institution can absorb with new norms. One is a genuine structural break. Sorting them is the whole point.
Credit. How much did the agent do, and does its contribution dilute the human's? This feels urgent and is mostly a distraction. Science already handles wildly uneven human contribution — the PI who conceived the study, the postdoc at the bench, the statistician who ran the models, the core facility that generated the sequencing. We do not resolve this by giving the mass spectrometer a byline. We resolve it by disclosure and by the convention that credit follows intellectual ownership. An agent that generates forty hypotheses is doing what a good brainstorming session does; the credit accrues to the humans who chose which one was worth testing. Credit is a coordination problem with known solutions.
Accountability. Who answers when the agent introduces an error, fabricates a citation, or hallucinates a result that survives into the published figure? This is the real break, and it does not need a new solution because the old rule already covers it, and the old rule is unforgiving. The human authors answer. All of them, for all of it. No fraction of the responsibility flows to the agent, because responsibility does not divide across parties who cannot bear it. If the agent's contribution cannot be pointed to as its own liability, then it is, in full, the liability of the people who deployed it and put their names on the output. An agent-introduced fabrication is not a lesser sin than a human-introduced one. It is the same retraction, with the same names on it.
Peer review. Reviewers, increasingly agents themselves, evaluating agent-assisted work. This is where the failure mode is most seductive, because it closes into an infinite regress of unaccountable verification: an agent drafts the paper, an agent drafts the review, an agent drafts the response to the review, and at no point does a person who has actually checked anything stand behind the judgment. Peer review is itself an accountability step — an editor and reviewers vouch that the work clears a bar — so the same constraint applies one level up. The verdict needs a named human who owns it. An agent can be a superb instrument for a reviewer, catching a statistical error or a missed prior result faster than the reviewer would. It cannot be the reviewer, for precisely the reason it cannot be the author.
Provenance. As claims flow from agent to draft to paper, the field needs to trace which claim came from where — which number came from the run the agent designed, which sentence from the model's draft, which citation from its retrieval. This is a tractable engineering problem, and it is the one worth investing in, because it is the substrate that makes accountability possible. You cannot own what you cannot locate. Provenance is not a nicety; it is the precondition for a human to verify the agent's contribution and therefore to legitimately sign for it.
Three of these — credit, peer review, provenance — resolve into norms and tooling. Accountability does not resolve at all. It relocates onto the humans and stays there, heavier than before.
Why "own it" means "verify it," with no discount
Here is the part researchers will want to wish away. If you are fully accountable for what the agent contributed, and the agent contributed a substantial chunk of the analysis, then you are on the hook for a substantial chunk of work you did not personally perform. There is exactly one way to make that a position a rational person will occupy: verify the contribution until you can honestly vouch for it. Owning is not a signature. Owning is understanding.
This is the same verification burden that appears wherever agents do consequential work, and science does not get a discount on it — if anything it pays a premium, because the currency of science is the reliability of the claim. Consider the concrete texture. An agent designs an analysis and reports a significant effect. To sign for that, you need to know what it controlled for, whether the significance survives a correction for multiple comparisons the agent may not have applied, whether the positive control ruled out what it needed to, whether the effect is an artifact of a preprocessing choice buried in code you did not read. The agent will produce a fluent, coherent account of why the result is sound — that is what these systems are best at — and coherence is not correctness. The only thing that discharges your accountability is checking, at the resolution where the error would actually live.
Which is why provenance is the actionable core, not a compliance afterthought. If the pipeline records what the agent did at each step, you can audit the steps that carry the liability. If it does not, you are signing for a black box, and signing for a black box is how a lab ends up with a retraction whose cause no one can reconstruct. Demand provenance from your tools the way you would demand a lab notebook from a student, and for the same reason: when the question comes — and in science the question always eventually comes — "the agent did it" is not an answer a court, a board, or a journal will accept, and it should not be one you accept from yourself.
What the machine cannot be handed, and why
Step back to what science actually is, because it explains why the accountability cannot migrate to the agent no matter how capable the agent gets. Science is not the generation of plausible claims. It is the disciplined killing of claims against reality, plus the judgment of which are worth the cost of the test — treating it as a generation problem is a category error. Authorship is the human institution that sits on top of that epistemic act. It exists to guarantee that behind every published claim there is a person who exercised the judgment, ran the falsification, and will answer for the result. The byline is the field's promise that the loop closed through a mind that had something at stake.
An agent can accelerate the generation and mechanize large stretches of the execution. It cannot supply the stake, and the stake is the thing authorship certifies. A byline attached to a system with nothing to lose certifies nothing. That is not a temporary limitation waiting on a better model. A perfectly calibrated, never-hallucinating agent would still be unable to be an author, because the barrier is not accuracy. It is answerability, and answerability is defined by consequence, which the agent structurally cannot have.
The forecast, labeled as one
My bet — and it is a bet about institutions, so hold it as a forecast, not a fact — is that norms will converge on a stable equilibrium: agents treated as powerful instruments whose output humans must own and verify, never as authors. Disclosure of agent contribution will become mandatory and increasingly granular, pushed there by provenance tooling that makes granularity cheap. Journals will require a named human to attest that the agent's contributions were verified, much as they now require statements of data availability and author contribution. Attempts to grant agents authorship, or to let agent reviewers close the loop unsupervised, will be tried, will produce a visible failure — a retraction no one can explain, a fabricated result no one caught — and will be rolled back. The system will settle where it always settles: on the principle that accountability cannot be delegated to a thing with no stake.
The strongest counterargument deserves a fair hearing. It runs: authorship conventions are contingent social technology, not laws of nature, and if agents come to do the overwhelming majority of scientific cognition, the institution could simply be redesigned — perhaps around organizational accountability, where a lab or a company vouches for agent-produced work the way a firm stands behind an audit, no individual byline required. This is not silly. But it does not escape the argument; it relocates it, exactly as insurance and corporate indemnity relocate liability without dissolving it. An organization is accountable only through the humans who run it, who sign, who can be sued and sanctioned. Push the responsibility to the lab and you have chosen which humans are on the hook — the PI, the officers — not removed the requirement that some human be. The agent still cannot hold the responsibility. It has only moved to a different desk.
For the researcher, the operating rule is concrete. Treat every agent contribution as your own claim, not the machine's — a claim you are about to publish under your name and defend for the rest of your career. Verify it to the depth at which its errors would hide. Demand provenance from your tools so that verification is possible and so that, when someone asks where a number came from, you can show them. And do not let anyone, including yourself, use the agent's fluency as a substitute for your judgment, because the byline was never a reward for producing the words. It was a promise to answer for them, and there is still no one to make that promise but you.