The most valuable thing AI might do for human learning is not new — it is an old idea that has finally become affordable. Give every child a personal tutor. One-to-one tutoring is the single most effective educational intervention we have ever measured, and it has always been too expensive to give to everyone. A patient, personalized, always-available AI tutor is the first plausible way to afford it at scale. Whether that promise is real or hollow comes down to one design decision, and only one: whether the tutor makes the student do the thinking, or does the thinking for them.
That single choice separates the largest positive application this technology may have from a beautifully engineered way to make people worse at learning while feeling helped. I want to make the optimistic case honestly — which means stating the condition that decides it, not burying it.
Start with the evidence, stated accurately
In 1984 the educational psychologist Benjamin Bloom published a paper with a deliberately provocative title: "The 2 Sigma Problem." His students had run controlled comparisons of three conditions — conventional classroom instruction, classroom instruction augmented with mastery-learning feedback, and one-to-one tutoring that also used mastery learning. The tutored students performed about two standard deviations better than the conventionally-taught ones. Two sigma is an enormous effect. It means the average tutored student scored above roughly 98% of the students in the conventional class. Move a median kid to the top of the room.
The word Bloom chose was problem, not triumph, and that framing is the reason the paper still matters. Tutoring works spectacularly, and tutoring at population scale is economically impossible. You cannot assign a skilled human tutor to every one of tens of millions of students for hours a week. So Bloom posed the challenge as a search: find methods of group instruction that get anywhere close to the tutoring effect at a cost society can actually bear. For forty years the honest answer was that we had partial fixes — mastery learning, better feedback, worked examples — that closed some of the gap but never the whole two sigma, and never cheaply.
A caveat on the number, because precision is the point of this whole essay. The exact 2-sigma figure came from particular studies pairing tutoring with mastery learning, and subsequent research has found a spread of effect sizes depending on subject, method, and who does the tutoring. The strong, defensible version of the claim is not "exactly two standard deviations, guaranteed." It is: individualized tutoring is reliably among the very largest effects in all of education, routinely dwarfing what typical classroom interventions produce. The direction and the rough magnitude survive even where the precise number wobbles. That is enough.
Why an AI tutor is a real answer to Bloom's challenge
Here is the genuinely under-hyped part. The reason tutoring stayed a "problem" for four decades was economic, not pedagogical. We knew what worked. We could not afford to deliver it. That specific constraint — the cost of one skilled adult's undivided attention per child — is exactly the constraint a competent AI tutor relaxes for the first time.
Consider what a good tutor actually provides, mechanism by mechanism, and how much of it an AI can supply. Infinite patience: it will re-explain the same idea the ninth time without a sigh, which matters enormously for a struggling student who has learned to expect frustration from adults. Adaptation to the individual: it can meet a specific learner at their specific confusion rather than at the class average. An inexhaustible supply of examples and problems, generated on demand at the right difficulty. And availability — the piece that dwarfs the rest — to any child with a cheap device and a connection, in any language, at eleven at night when the homework is due and no human tutor exists.
That last point is where the moral weight sits. The children who most need tutoring are precisely the ones whose families cannot buy it. Bloom's result, read cynically, is a description of how advantage compounds: wealthy kids get the two-sigma intervention, everyone else gets the classroom. An AI tutor that is genuinely good is the first technology that could break that coupling — deliver the historically-unaffordable intervention to the students who have never had access to it. I do not say this lightly, and I am allergic to education-technology hype. But the shape of the opportunity is real, and it is large.
The condition that decides everything
Now the honest part, which most of the excitement skips. Return to what a good human tutor actually does in the moment a student is stuck, because it is the opposite of what the word "helpful" suggests.
A skilled tutor, watching a student flounder on a problem, does not supply the answer. They ask a question. They let the silence sit. They let the student struggle — visibly, uncomfortably — because the struggle is where the learning happens. Then they diagnose the specific misconception, not the general topic, and offer the minimal intervention that lets the student take the next step on their own. A well-placed "what do you think happens to this term when x is zero?" instead of "the answer is four." The tutor's craft is calibrating help so that the student keeps doing the cognitive work, always operating just past the edge of what they can already do.
This is not folk wisdom. It is one of the most robust findings in the learning sciences, and it has a name. Robert Bjork and Elizabeth Bjork call them desirable difficulties: conditions that make learning feel harder and slower in the moment but produce far more durable and transferable knowledge. Effortful retrieval beats re-reading. Spacing beats cramming. Struggling to generate an answer before being told it — even struggling and getting it wrong — produces better retention than being handed the answer smoothly. The difficulty is not a cost you tolerate to reach the learning. The difficulty is the learning. Remove it and you remove the effect.
Which sets up the trap precisely. An AI tutor optimized to be maximally, immediately "helpful" — to give the answer, complete the proof, write the paragraph, do the work — delivers the exact inverse of tutoring. It removes the productive struggle that makes tutoring the two-sigma intervention in the first place. It feels wonderful. The student gets unstuck instantly, the friction vanishes, the homework gets done. And almost nothing is learned, because the cognitive work that would have built the skill was outsourced to the machine. I have argued the general version of this elsewhere — that the struggle is precisely the thing these systems are built to remove — and education is where the stakes of that removal are highest, because the entire point of the exercise was the struggle.
So the same underlying model can be built two ways that look superficially similar and produce opposite outcomes. One is engineered as an answer engine: fast, frictionless, convenience-maximizing, and quietly corrosive to the skill it claims to teach. The other is engineered as a Socratic partner: it withholds, it questions, it diagnoses, it gives the minimal hint, it makes you generate the answer. These are not different levels of quality. They are different products built toward opposite objectives, and the difference is a design choice, not a capability the model lacks.
The reliability caveat that raises the stakes
There is a second condition, and in this application it is unusually severe. A tutor that confidently teaches wrong things is worse than no tutor at all.
Ordinary software that fails announces itself — it crashes, it errors, it stops. A language model that fails does the opposite: it produces a fluent, confident, well-formed statement that happens to be false, indistinguishable in form from a true one. I have written separately about why this is structural rather than a bug to be patched away — that these systems rank outputs by plausibility, not by truth, because plausibility is the only variable a text-trained model has. In most uses a plausible-but-wrong answer is an annoyance you catch. In tutoring it is a landmine, because the student is by definition not in a position to detect the error — that is why they are being tutored. A confidently-taught falsehood does not just fail to teach; it installs a misconception that the student will carry forward and build on, and that a later teacher will have to detect and uproot. Negative learning is possible, and this is how it happens.
This makes accuracy and calibration matter more in tutoring than in almost any other consumer application. The practical implication is to scope deployment by verifiability. In domains with checkable answers — arithmetic, algebra, a program that either compiles or does not, a grammatical rule — the tutor can be held to a high bar today, and tool use grounds the claims that count: route the computation through a calculator, run the code, check against a symbolic solver rather than asserting from textual plausibility. In open-ended domains where fluent and false are the same kind of object, the risk is far higher and human oversight matters far more. Match the tool to the terrain.
The forecast, labeled as one
Here is my bet, stated as a bet and not a prediction of the inevitable. If AI tutors are built to create desirable difficulty and to hold accuracy, they could become one of the largest positive applications of this entire technology — the first real answer to a problem Bloom posed forty years ago, delivered first and most to the learners who never had access to a tutor at all. If they are built to maximize convenience — to feel maximally helpful, to remove friction, to hand over answers — they will be a wasted opportunity at best and, at worst, an engine for producing a generation that is fluent at prompting and worse at thinking.
The strongest counterargument deserves a fair hearing, because it is not obviously wrong. It runs: struggle is romanticized. Plenty of educational friction is undesirable — pointless drudgery, badly-designed problem sets, difficulty that demoralizes rather than builds. Perhaps the right frame is not the lone learner grinding against a problem but something like Kasparov's "advanced chess," where a human paired with a machine outperforms either alone. Maybe the future student is a centaur — offloading the mechanical parts to the AI and reserving human cognition for the higher-order moves — and my worry about lost struggle is nostalgia for skills that will matter as little as long division by hand.
I take this seriously, and I would note two things. First, the Bjorks' distinction is desirable versus undesirable difficulty precisely because not all struggle is good; the design task is to preserve the productive kind and cut the rest, which is exactly what a well-built tutor should do. Second, the centaur analogy cuts both ways: the human half of a strong centaur is strong because they built genuine skill first, and — worth saying plainly — in chess itself the centaur's edge has since largely evaporated as engines pulled ahead of any human-machine team. A partnership that hollows out the human's underlying competence does not produce a centaur. It produces a person who is helpless without the machine. For foundational learning — the years when a mind is being built rather than augmented — I will bet on struggle. For a working expert augmenting existing skill, the centaur frame may well win. Those are different regimes, and conflating them is the error.
What to do with this
Evaluate any AI tutor — as a parent, a teacher, a school, a builder — on two questions, and refuse to be moved by a third. Does it make the learner struggle productively: does it withhold, ask, and hint rather than hand over answers? And is it accurate, ideally grounded by tools in domains where answers are checkable? Those are the questions that determine whether it delivers Bloom's two sigma or quietly steals it.
The question that will try to seduce you, and the one to ignore, is how helpful it feels. In tutoring, the feeling of being helped and the fact of being taught point in opposite directions more often than not. The best tutor you ever had made you do the work. Build the machine that does the same, or you have built something else and mislabeled it.