The question "are we close to AGI?" cannot be answered, and not because the future is hard to predict. It cannot be answered because "AGI" names at least four different things, those four things come apart in practice, and "how close are we?" has a different answer for each. The term smuggles in several incompatible criteria and then presents itself as a single threshold. Separate the criteria and the timeline debate — the one that consumes conference panels and funding rounds and safety arguments — dissolves into a handful of concrete questions, most of which have checkable answers. My claim is that the threshold framing is the actual mistake, and that dropping it makes you sharper, not vaguer.
Let me show the confusion is real before arguing about what to do with it. Here are the four definitions people mean when they say AGI, stated as accurately as I can, because the sloppy versions are what keep the debate stuck.
Four definitions that don't agree
The economic definition. A system that can do most of the economically valuable work a human can do. This is close to the language in OpenAI's charter — "highly autonomous systems that outperform humans at most economically valuable work" — and it has one enormous virtue: it is measurable, at least in principle. You can enumerate tasks, check whether a system does them reliably and unsupervised, and add up the fraction of paid human labor covered. It is also the definition that matters for impact. If a system automates 40% of the tasks in the economy, the philosophical status of its "understanding" changes nothing about the labor-market consequences.
The cognitive, human-level definition. A system whose general reasoning matches a human's across domains — not task-by-task competence but the underlying faculty that lets a person pick up a new domain, form a model of it, and reason inside it. This is the definition that feels like it captures what we actually care about, and it is philosophically loaded to the point of being nearly unmeasurable. "General reasoning" is not operationalized. We don't have a test that distinguishes a system that reasons generally from one that has memorized enough of the human corpus to imitate the outputs of general reasoning across the cases we happen to check.
The behavioral, Turing-style definition. A system indistinguishable from a human in open conversation. Turing proposed the imitation game in 1950 precisely to replace the unanswerable "can machines think?" with something operational. The irony is that we have now largely met his bar for many purposes — current systems fool many people in short exchanges — and in doing so we learned that the test is weak. Conversational indistinguishability measures how well a system models human dialogue, which turns out to be a different and easier thing than robust reasoning. Passing tells you less than its fame implies.
The learning, or transfer, definition. A system that can learn anything a human can learn, from comparable experience. This is the generality-as-sample-efficiency view — closer to the Legg–Hutter attempt to define intelligence as performance across a wide space of environments. It emphasizes not what the system knows but how efficiently it acquires competence in something genuinely new, with the data budget a human gets rather than the entire internet.
These are not four phrasings of one idea. They are four different bets about what "general" means, and they dissociate.
Watch them come apart
Consider a system that is economically transformative and cognitively unremarkable. Automation has never required human-like general reasoning; it requires that a bundle of tasks be done well enough. David Autor's task-based framework is the right lens: a job is a bundle of tasks, machines take the tasks that are codifiable, and the human keeps — or loses — the rest. That framework produced the polarization pattern Autor and collaborators documented, where automation hollowed out routine middle-skill work while leaving both high-skill judgment and hard-to-codify manual work relatively intact. Nothing in that story needed the machine to be generally intelligent. It needed the machine to be reliably superhuman on specific, valuable, codifiable tasks. A language model that drafts contracts, writes boilerplate code, and handles tier-one support is doing economically valuable work while failing every serious test of cognitive generality. On the economic definition it is far along. On the cognitive definition it is nowhere.
Now run it the other way. A system can pass the conversational test — sail through casual exchange, produce fluent, confident, human-sounding text — while being subhuman at robust reasoning. The failure mode is familiar to anyone who has pushed one hard: it handles the phrasing it has seen and collapses on a small, novel variation that a competent adult would treat as trivially the same problem. Behaviorally general, cognitively brittle. The Turing definition scores it high; the transfer definition scores it low.
And the chess history warns against reading capability as a single line at all. When Deep Blue beat Kasparov in 1997, the natural inference was that chess-playing intelligence had been achieved and the human was now redundant. What Kasparov drew from it instead was "advanced chess" — the centaur format, human plus engine against human plus engine — and for a stretch the strongest competitive entities were neither the best humans nor the best programs alone but the pairings. The machine was radically superhuman at one thing (search and evaluation) and useless at everything around it: deciding what to play for, managing a match, existing outside the 64 squares. A narrow superhuman system and a general intelligence were not points on the same scale. They were different objects.
Why a benchmark threshold can't rescue you
The instinct, faced with four definitions, is to reach for a number: pick a benchmark, set a threshold, declare AGI the point where the line crosses it. This fails twice over, and the two failures are worth naming precisely.
The first is the emergence-metric problem. Many of the dramatic "the model suddenly can do X" curves — flat near the floor, then a cliff — are produced by all-or-nothing scoring rather than by any discontinuity in the underlying capability. Score a multi-step task pass/fail and smooth improvement in per-step accuracy shows up as a sudden jump in whole-task success. I've argued the full version of this elsewhere — the cliff often lives in the yardstick, not the model — and the consequence for AGI thresholds is direct: a benchmark line crossing a mark does not correspond to a threshold crossed in the world. The number moved because of how you scored it.
The second failure is deeper: capability is jagged, not scalar. A system can be superhuman on one task and subhuman on the task right next to it, with no smooth gradient between them and no single axis they both sit on. I've called this the jagged frontier — the property that makes any single-number summary of "how general is it?" a lie of aggregation. A benchmark average tells you the mean height of a coastline. What determines whether the system is usable for real work is where the cliffs and inlets are, and the average is specifically constructed to hide them. There is no threshold on a jagged frontier, because a frontier is not a line.
Put the two together and the benchmark-threshold move is dead. Even a clean, unfaked benchmark crossing tells you the system got better at what that benchmark samples, which is a small and possibly unrepresentative patch of the frontier. It does not tell you the system crossed into a new kind of thing.
The threshold is the wrong object
Here is my position, labeled as a stance rather than a finding: AGI-as-threshold is the wrong frame entirely. There is no moment the system "becomes AGI," because there is no single scalar to cross a line, no agreed definition to cross it on, and no reason the four definitions should cross at the same time even if each had a line. The debate over when inherits all the incoherence of the debate over what, and adds a date.
This is also why the takeoff-speed debate — fast/hard takeoff versus slow/soft, the question of whether capability, once it reaches some level, accelerates sharply or improves gradually — is often less illuminating than it looks. It, too, presupposes a "level" that acceleration happens at or after. On a jagged frontier with four non-aligned definitions, "takeoff" is not one event with one speed; different capabilities arrive on different schedules, and which pattern you see depends on which capabilities you were watching. I don't think the takeoff question is empty — whether AI progress feeds back into AI progress is real and serious — but framing it around a threshold crossing imports the same confusion.
I want to be fair to the strongest version of the opposing view, because there is one. The counterargument runs: definitional pluralism is real, but "general intelligence" may nonetheless name a genuine natural kind — a single underlying capacity that, once present, produces competence across all four definitions at once, the way general reasoning does in humans. On this view, insisting the definitions are separate is like insisting "alive" is four incompatible criteria because you can list respiration, metabolism, reproduction, and homeostasis separately; the criteria co-occur because they flow from one underlying thing, and the same may hold for intelligence. If that's true, the threshold framing is recovered: there is a point where the underlying capacity appears, and the four definitions light up together shortly after.
I take that seriously, and I think the evidence currently runs against it. In humans the criteria co-occur; in the systems we actually build they demonstrably do not — that's exactly what the "economically transformative but cognitively brittle" and "conversationally fluent but reasoning-fragile" cases above are. The dissociation is the empirical observation. A future architecture could unify them and vindicate the natural-kind view. But "the definitions will converge" is itself a forecast, not a definition, and you don't get to assume it to rescue the word.
A separate note, since general intelligence is often assumed to bring goals and agency with it. Bostrom's orthogonality thesis says intelligence level and final goals are independent axes — a highly capable system can in principle pursue almost any objective — and his instrumental-convergence point says a wide range of goals imply similar sub-goals: acquire resources, preserve yourself, keep your options open. Both are worth holding precisely because they cut across the definition question. Whether a system is "AGI" on any of the four criteria is separate from what it is trying to do, and capability does not hand you alignment for free. That's an argument for watching agency and objectives as their own axis, not folding them into the capability threshold.
What to do instead
Replace the yes/no with two questions that actually have answers.
First: which economically valuable tasks can it now do reliably, unsupervised, at what error rate? This is checkable. It is the economic definition made operational, task by task, and it governs impact regardless of what's happening "inside." You can measure it, track it over time, and be wrong about it in ways reality corrects.
Second: where does its generality break? This is the jagged-frontier question turned into a probe. Find the small variations that collapse it, the neighboring tasks where superhuman drops to subhuman, the transfer it fails to make. The map of failures is more informative than any aggregate, because it tells you what the system can be trusted with and what it can't.
Both questions have concrete, checkable answers. "Is it AGI?" does not, and never will, because the question was malformed before anyone tried to date it.
So the move, the next time someone tells you "AGI by 2030": ask which definition. Economic, cognitive, behavioral, or transfer. Watch what happens. If they mean the economic one, the conversation immediately gets specific and productive — which tasks, measured how. If they mean the cognitive one, ask what test would distinguish a system that has general reasoning from one that convincingly fakes it, and notice that the room goes quiet. That quiet is the whole point. The timeline was never the hard part. The word was.