Skip to content

Recursion Isn't New: The Historical Base Rate for Self-Improvement Is Bounded, Not Explosive

Self-improving systems already exist: self-hosting compilers, science refining its own methods, cumulative culture. Each produced fast, compounding, bounded growth — never an explosion. That is the base rate AI must beat.

By Mehdi11 min read
Share
On this page

Self-improving systems are not new, and this is the single most useful fact for reasoning about AI that almost no one puts on the table. Compilers have compiled themselves for six decades. Science is a method whose central achievement is refining its own methods. Culture is a recursive accumulator in which every generation builds on the tools and knowledge of the last. Each of these is a working, historical instance of recursive self-improvement — and each one produced fast, compounding, transformative growth that was nonetheless bounded, gated hard by external inputs and diminishing returns, never a vertical discontinuity that left the physical world behind. That gives us something the singularity debate almost entirely lacks: a base rate. And the base rate says the burden of proof lies on the explosion.

Let me be precise about what I am arguing against, because the strong version deserves respect. The intelligence-explosion idea has a real intellectual lineage, and it is not naive.

The explosion thesis, stated fairly

The canonical statement is I.J. Good's, in his 1965 paper "Speculations Concerning the First Ultraintelligent Machine." Good's move was a single tight loop: "Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an 'intelligence explosion,' and the intelligence of man would be left far behind." The first ultraintelligent machine, he added, "is the last invention that man need ever make." Vernor Vinge sharpened this into the "technological singularity" in his 1993 essay, framing it as an event horizon past which prediction fails. Ray Kurzweil folded it into a Law of Accelerating Returns, an exponential extrapolated to a knee. Nick Bostrom, in Superintelligence, did the most careful work — distinguishing takeoff speeds (slow, moderate, fast) and flagging the hard sub-problems, including whether a system rewriting its own code preserves its goals across the rewrite (his "goal-content integrity," a version of the goal-stability-under-self-modification problem).

The mechanism underneath all of them is recursive self-improvement: a system that improves its own capacity to improve. If each increment of capability makes the next increment easier, and the effect is strong enough, you get super-exponential growth — a runaway.

The logic is valid. The question is entirely about the constant. Does the loop have gain greater than one, sustained, unbottlenecked? That is not a philosophical question. It is an empirical one, and — this is the reframe — we have run the experiment several times already, in domains nobody thinks to check because we do not usually file them under "recursive self-improvement." We should.

Case one: the compiler that builds itself

A compiler is a program that translates source code into machine code. A self-hosting compiler is one written in the very language it compiles — so it can compile its own source. Corrado Böhm described such a scheme in his 1951 dissertation; Hart and Levin built a self-compiling LISP compiler in the early 1960s; today it is routine, which is why GCC ships with a three-stage bootstrap that literally uses the compiler to rebuild the compiler and checks that the second and third stages produce byte-identical output. Ken Thompson's 1984 Turing Award lecture, "Reflections on Trusting Trust," is built entirely on the fact that this loop is real: a compiler can propagate behavior into its own successors.

Look carefully at what this is. Version N of an optimizing compiler contains optimization passes. You use it to compile the source of version N+1. If N+1's source contains a better optimization, the resulting binary generates faster code — and it can be used to compile version N+2 faster and better still. Improvements to the tool are inherited by the process that builds the next tool. That is a recursive self-improvement loop in production software, running since before most of the singularity literature was written.

So where is the explosion? There isn't one, and the reason is instructive. The loop is bottlenecked at two points. The first is human design insight: the compiler does not invent the new optimization pass — a person does. The second is the physical target: a compiler can only exploit the instruction set and memory hierarchy the hardware actually provides. The measured result is stark. Todd Proebsting's rule of thumb — Proebsting's Law — is that improvements from compiler optimization have historically doubled program performance on a timescale of roughly eighteen years. Compare that to hardware, which for decades doubled on a timescale closer to two. Run the arithmetic: over a span in which hardware improved by several orders of magnitude, the self-improving compiler loop contributed something like a single-digit multiple. The recursion is real, it compounds, and it is swamped by the external input it depends on. The tool that most literally builds better versions of itself is one of the clearest cases of bounded recursion we have.

Case two: the method that improves its own methods

Science is a recursive self-improvement system operating at the scale of an institution rather than a program, and it has been running for four centuries. This is looser than the compiler case — science is not a single agent editing its own source — but it meets the structural definition that matters: the outputs of the method feed back to improve the method itself. Better instruments produce discoveries that produce better instruments. Better statistics let you extract more from the same experiments, which motivates still better statistics. Better institutions — the journal, peer review, the controlled trial, preregistration — improve the reliability of the knowledge that is used to design the next institution.

The instrument loop alone is remarkable. The telescope and microscope extended observation past the senses; what they revealed drove the optics, the metallurgy, and eventually the electronics that produced the electron microscope and the space telescope, each of which revealed more, which funded and motivated the next. This is exactly the sense in which I have argued elsewhere that AI's real test is whether it can be a genuine new instrument of method rather than a faster way to run the old one — see AI as a New Organon. A new organon would enter this loop: not merely produce answers, but improve the method by which warranted answers are generated and checked.

And yet the loop is paced, hard, by the world it studies. You cannot compress the time an experiment takes below what the physical system allows. A clinical readout waits on human biology; a materials result waits on synthesis and characterization; an astronomical confirmation waits on the sky. Autonomous "self-driving labs" are compressing the design-build-test cycle in chemistry and materials right now, which is genuine and important — but they run into reagent costs, instrument throughput, and the sim-to-real gap the moment a simulated candidate has to survive contact with an actual reaction. Funding is finite. Replication takes real time.

The quantitative signature of this is one of the most important empirical results in the economics of ideas. Bloom, Jones, Van Reenen, and Webb, in their 2020 paper "Are Ideas Getting Harder to Find?", documented that across many fields, sustaining a constant rate of progress requires an exponentially growing input of researchers — research productivity per person is falling sharply. Sustaining Moore's Law, on their accounting, took roughly an order of magnitude more researchers by the 2010s than it did in the early 1970s. That is the mathematical fingerprint of a recursive system running against diminishing returns: it keeps compounding, but only by throwing ever more input at a frontier that keeps receding. Science is the most powerful self-improvement engine our species has ever built, and it is bounded — by experiment, by funding, and by the stubborn pace of the physical world.

Case three: culture, the original recursive accumulator

Before compilers and before science, there was cumulative culture, and it is the deepest case of all. What makes humans unusual is not raw individual intelligence but the ratchet: our capacity to transmit knowledge and technique across generations with high enough fidelity that improvements accumulate instead of being lost. Michael Tomasello named the ratchet effect; Joseph Henrich, in The Secret of Our Success, made the case that this cumulative process, not individual genius, is the actual source of human capability. Each generation inherits the tools, words, and know-how of the last and adds to them. That is recursion: the accumulated output of culture is the input to the next round of cultural production.

Its growth curve looks exponential over the long run — stone tools to agriculture to writing to spaceflight. But it is bounded by real resources, and the boundedness is visible in the failures. Henrich's discussion of Tasmania is the sharp example: a population that became small and isolated after rising seas cut it off from mainland Australia appears to have lost technologies — the ratchet ran backward, because cumulative culture is gated by population size, connectivity, and the number of minds available to carry and recombine knowledge. Cultural accumulation needs energy, materials, people, and transmission bandwidth. Remove the inputs and the compounding stalls or reverses. Fast, transformative, and firmly bounded by the physical substrate it runs on.

The pattern is the argument

Put the three cases in one frame.

System The recursive loop What bounds it
Self-hosting compiler Each version compiles and improves the next Human design insight; hardware limits; steep diminishing returns (Proebsting)
Science Method improves its own instruments, statistics, institutions Experiment time; funding; falling research productivity (Bloom et al.)
Cumulative culture Each generation builds on the last Population, energy, materials, transmission fidelity

The regularity is not subtle. Every recursive self-improvement system we can actually inspect produces fast, compounding growth that is nonetheless bounded by external inputs and diminishing returns at each level of the loop. This is not a separate objection to the explosion thesis. It is the bottleneck argument — the claim that self-improvement is rate-limited by whatever input it cannot manufacture — observed as history rather than asserted as theory. The compiler cannot manufacture faster hardware. Science cannot manufacture experimental results faster than the world yields them. Culture cannot manufacture the population and energy its ratchet consumes.

So here is the reframe, stated as my actual position. The historical base rate for recursive self-improvement is fast-but-bounded compounding. Not a singularity. Given three independent instances all showing the same shape across timescales from decades to millennia, the rational prior for a fourth recursive system — AI improving AI — is that it, too, compounds hard and then grinds against external limits: training data, experimental validation, energy, capital, and the irreducible latency of learning things about a world you have to actually poke. The prior is not neutral. It is that AI self-improvement looks like the base rate until shown otherwise, and the burden is on the explosion case to demonstrate why AI escapes a pattern that nothing else — not the one system literally engineered to rebuild itself — has ever escaped. This is the same move I made against scaling: an observed trend is not a mechanism, and a curve is not a law (Scaling Is Not a Theory of Intelligence). "It will keep going up" is a description, not an argument for where it stops.

The strongest counterargument, and why it doesn't settle it

The best rebuttal is genuinely good, and I want to state it at full strength rather than knock down a weak version. It goes: every one of your three cases was bottlenecked, ultimately, by the speed and quantity of human cognitive labor. The compiler waited on human designers. Science waits on human researchers — the entire Bloom et al. result is denominated in researchers. Culture waits on human brains and their glacial generational turnover. AI is the first recursive system in history that attacks that specific input directly. If it supplies cognition itself, at machine speed and arbitrary scale, it dissolves the very constraint that paced everything else, and my base rate — built entirely from cognition-limited systems — could be exactly the wrong reference class.

I think this is the real frontier of the disagreement, and I hold my view with correspondingly plain uncertainty. But I do not think it wins, for one reason the historical cases make concrete: cognition was never the only binding constraint, and relieving it reveals the next one rather than removing them all. An experiment still takes physical time. Data about a system you have not yet observed still has to be gathered from that system — you cannot introspect your way to a measurement you have not made, and this is precisely the wall that JEPA-style world models and model-based RL keep hitting as the sim-to-real gap. Chips are fabricated in finite quantity; energy is metered. I watch a small version of this in building Velya: the bottleneck on an AI product improving is almost never the model's cleverness in the abstract — it is the rate at which you can gather real, labeled, ground-truth signal about whether its outputs were actually good, and that rate is set by the world and the users, not by the model. Remove the cognitive bottleneck and you do not get a runaway; you get a system that now compounds against the next limit in line. Which is to say: you get the base rate again, relocated one level up.

The honest position is that the counterargument identifies the one way the base rate could break, and neither side can currently prove its case — anyone claiming certainty here is selling the singularity or dismissing it, and both are overclaiming. What history gives us is not a proof. It is a prior, and a demanding one.

So here is the concrete thing to watch for, the thing that would actually move the prior. Do not watch benchmark scores climb; the base rate already predicts fast compounding on any single axis. Watch for a self-improvement loop that stops being bottlenecked by a non-cognitive input — a system that measurably reduces its own dependence on fresh real-world data, or physical experiment, or human-supplied design insight, and keeps compounding anyway. That, and only that, is what an escape from the base rate would look like. Until you see it, the recursion is not new, and neither is the ceiling.

Frequently asked questions

Doesn't the compiler analogy break because compilers don't design their own successors — humans do?
Partly, and that is exactly the point I want to isolate. A self-hosting compiler already closes part of the loop: version N compiles the source of version N+1, and any speedup or better code generation it produces is inherited by the tool that builds the next one. What the compiler does not do is invent the new optimization pass; a human supplies the design insight. So the loop is real but the intelligence-generating step sits outside it. The interesting question about AI is whether it can internalize that step too. My argument is that even if it does, it still inherits the other bottlenecks the compiler case reveals — hardware, and the sharply diminishing returns captured by Proebsting's observation that compiler optimizations have historically doubled performance on a scale of roughly two decades, not two years. Removing the human from one link does not remove the diminishing returns from the others.
Isn't 'science improves its own methods' too loose to count as recursive self-improvement in the technical sense?
It is looser than the compiler case, and I flag it as such. Science is not a single agent modifying its own source; it is a distributed institution. But it meets the structural definition that matters: the outputs of the method (better statistics, better instruments, better institutional norms like peer review and preregistration) feed back to improve the method that produces future outputs. Better microscopes yield discoveries that yield better microscopes. That is a genuine recursive loop operating for four centuries, which is precisely why it is useful as a base rate. And its measured behavior — compounding output alongside falling research productivity per researcher, as documented by Bloom, Jones, Van Reenen and Webb — is the empirical shape I am asking you to take seriously as a prior.
If the base rate is 'bounded,' why worry about AI at all?
Bounded is not the same as small. Cumulative culture is bounded and it took humans from stone tools to spaceflight; science is bounded and it rebuilt the material conditions of the species. The claim is not that AI self-improvement will be unimpressive. It is that its shape will be fast, compounding growth that grinds against external limits — data, experiment, energy, physical validation — rather than a vertical discontinuity that leaves the physical world behind in weeks. That distinction matters enormously for what you should build, watch, and worry about. A bounded fast-compounder rewards attention to which input binds first. A hard-takeoff singularity makes that attention pointless. The base rate says bet on the former.
What is the strongest version of the case that AI escapes this base rate?
That all three historical cases were bottlenecked by the speed and quantity of human cognitive labor, and AI is the first recursive system that attacks that specific bottleneck directly. Compilers waited on human designers; science waits on human researchers; culture waits on human brains and their generational turnover. If AI supplies cognition at machine speed and scale, it removes the very constraint that paced the others, and the analogy could fail. I take this seriously and think it is the argument that deserves a real answer. My answer is that cognition was never the only binding constraint — experiments still take physical time, data about the world still has to be gathered from the world, chips and energy are finite — so relieving the cognitive bottleneck reveals the next one rather than removing all of them. But this is the frontier of the disagreement, and anyone who claims certainty on either side is overclaiming.

Filed under Cross-Disciplinary Deep Essays. Where biology, computation, markets, and philosophy collide.

Essays like this, in your inbox.

Thoughtful essays. No spam. Unsubscribe anytime.

Cross-Disciplinary Deep Essays

"AGI by Year X" Is Unanswerable Until You Name the Definition

"Are we close to AGI?" is incoherent because AGI names at least four incompatible criteria that come apart in practice. Separate them and the timeline debate dissolves into concrete, checkable questions.

9 min read