Skip to content

A Calculator Automates a Step; AI Can Automate the Understanding

"Calculators didn't ruin math, so AI won't ruin learning" is a bad analogy — the calculator offloaded a mechanical sub-skill, while AI can offload the thinking itself. The test is which one your AI is doing.

By Mehdi9 min read
Share
On this page

"Calculators didn't ruin math, so AI won't ruin learning" is the most common way sensible people dismiss worry about AI in education, and it is a bad analogy. Not because the original calculator panic was correct — it wasn't — but because it succeeds for a reason that does not transfer. The calculator automated a mechanical sub-skill that was never where mathematical understanding lived. AI can automate the thinking itself. That difference is not a footnote to the analogy; it is the whole question the analogy is being used to wave away.

So let me take the analogy seriously, because it deserves it, and then show you the exact seam where it splits. By the end you should have a single test you can run on any AI use in learning — your kid's, your team's, your own — that tells you whether it is safe.

Why the calculator actually worked out

The pocket calculator arrived in classrooms in the 1970s and provoked a genuine moral panic: students would forget how to compute, mathematical ability would rot, a generation would be helpless without a machine. Decades later, the consensus is that this mostly didn't happen, and it's worth being precise about why, because the reason is doing all the load-bearing work.

Long division and multi-digit multiplication are execution skills. They are algorithms — deterministic procedures you run to turn one number into another. Crucially, understanding what division means, when a problem calls for it, how to set up a ratio, how to sanity-check a magnitude — none of that lives inside the execution. You can grasp the concept of dividing a bill among friends perfectly and still make an arithmetic slip on paper. The slip is not a failure of understanding; it is a failure of clerical labor.

The calculator offloaded the clerical labor and left the understanding untouched. Better than untouched, arguably: freed from the attentional cost of hand-computation, a student can hold more of the actual problem in working memory — the modeling, the estimation, the "does this answer even make sense." This is a real and underrated category: good offloading, where automating a low-level step frees cognition for the high-level work that was the point. The panic was wrong because it misidentified where the learning was. Arithmetic execution was a step on the way to mathematical thinking, not the thinking itself.

Hold onto that phrase. It is the entire test.

Where the analogy holds for AI

Run the same logic on AI and it holds cleanly in a large set of cases. Whenever a model automates a mechanical sub-skill that is not the learning target, it is a calculator, and the calculator precedent applies with full force.

A programmer learning system design who has the model generate boilerplate — the config file, the import statements, the getter/setter scaffolding — is offloading clerical labor. The learning target is the architecture, and the boilerplate is arithmetic. A biology student who has AI format a bibliography, convert units, or retrieve the melting point of a compound is offloading look-up and formatting. A founder learning financial modeling who lets the model write the spreadsheet formula syntax while they decide what to model is calculator-using in the purest sense. In each case the automated step is a deterministic procedure that understanding does not inhabit, and removing it frees attention for the part that does.

This is not a grudging concession. It is a genuinely good use, and the calculator history is the right guide to it. If that were all AI did in learning, the dismissers would be right and there would be no essay to write.

Where it breaks, precisely

The analogy breaks the instant the "sub-skill" being automated is the learning target — when the cognitive work the AI performs is the very work that was supposed to build the understanding.

Writing is the cleanest example, because writing is not the transcription of thought that already exists. Writing is the process by which the thought gets built. You discover that your argument has a hole only when you try to write the sentence that crosses it. You find out your two ideas don't actually connect only when the transition refuses to be written. The friction of composition is not overhead on top of thinking — it is the thinking, externalized where you can see it fail. A student who has a model write the essay has not offloaded a clerical step on the way to understanding. They have offloaded the understanding. The essay was never the deliverable; it was the exercise apparatus, and they skipped the exercise.

Same for working a problem. When a student derives why gradient descent stalls at a saddle point — writes the loss surface, computes the gradient, feels the jolt of "zero gradient but not a minimum?" resolve into structure — they build an internal model they can later recognize in a landscape they've never seen. When a student asks a model and receives a crisp paragraph about saddle points, they can say the words. Both now possess the answer. Only one possesses the understanding, because understanding is the residue that the struggle leaves behind, and you cannot get the residue without running the reaction. I've argued this at length in Productive Struggle Is the Thing AI Removes: the difficulty is not a bug in the learning process to be optimized away, it is the mechanism.

This isn't folk psychology. It's one of the most robust results in the science of learning. Robert and Elizabeth Bjork's work on desirable difficulties shows that conditions which make learning feel harder and slower in the moment — retrieval practice, spacing, generating an answer before being shown it — produce more durable, more transferable knowledge than conditions that feel easy and fluent. The feeling of ease during learning is actively anti-correlated with actually learning. An AI that hands you a finished, fluent answer optimizes for precisely the sensation that predicts you did not learn. I've traced how that fluency does its damage even when the content is correct in The Epistemic Cost of Fluency; here the point is narrower and sharper — the frictionless answer removes the friction that was the curriculum.

The one test

Strip it to a single question you can ask of any AI use in a learning context:

Does the AI automate a step on the way to understanding, or does it automate the understanding itself?

If the former, it's a calculator — it clears mechanical labor out of the way so cognition can go to the real target. Safe, often excellent. If the latter, it's a ghostwriter — it performs the exact cognitive act the exercise existed to build, and hands you the output while quietly starving the process that was supposed to grow a mind. Not safe, no matter how good the output looks.

The tell is not the tool and not the output — it's the location of the learning target relative to the automated step.

Calculator (safe) Ghostwriter (not safe)
What's automated A mechanical sub-skill The cognitive work itself
Where the learning lived Elsewhere — the step is a means In the step being automated
Effect on attention Frees it for the real target Removes the target
Writing course AI fixes your grammar AI writes your argument
Learning to code AI generates boilerplate AI solves the design problem
Learning statistics AI runs the arithmetic AI decides which test applies and why
Same output, different verdict The answer was never the point The answer was the residue of a process you skipped

Notice the last three rows swap sides depending on what the course is for. A model drafting prose is a calculator in a chemistry class where writing isn't the target and a ghostwriter in a composition class where it is. The verdict is a property not of the AI but of the relationship between the automated step and the thing you were trying to become.

The strongest counterargument, taken seriously

Here is the objection I find hardest, and it's the one worth sitting with. The skills AI automates simply stop being valuable — the way longhand square roots stopped mattering. If a machine can write the essay, maybe essay-writing is the next longhand: a skill whose economic value evaporates, so mourning its atrophy is nostalgia.

This is real, and there's serious economic backing for the shape of it. David Autor's work on task-based automation and labor-market polarization shows that technology doesn't erase jobs wholesale; it hollows out the specific tasks it can do, reshaping what human work is worth. Skills genuinely do get retired, and pretending otherwise is its own kind of denial.

But the objection conflates two categories that the calculator test pulls apart. Some skills are terminal: their entire value is the output they produce. Longhand arithmetic was terminal — nobody's inner life depended on doing it, so a machine that produced the number retired it cleanly. Other skills are developmental: their value is the cognitive structure you build by exercising them, which you then carry into everything else. Working the problem builds the model you use to recognize the next problem. Writing the essay builds the reasoning you bring to the next argument, the next negotiation, the next decision that has nothing to do with essays.

Automating a terminal skill retires an obsolete task. Automating a developmental skill removes the mechanism that grows a mind — and the output it produces is not evidence the mind was grown. This is the deeper reason human-plus-machine can beat either alone: Garry Kasparov's "advanced chess" and the centaur teams that followed worked because the human retained the strategic understanding and offloaded the calculation, the developmental skill kept and the terminal skill delegated. The centaur that offloads the understanding isn't a centaur. It's a person holding a leash attached to nothing.

The honest uncertainty: I can't tell you the exact boundary between terminal and developmental for every skill, and some will migrate as the economy reshapes around AI — that migration is a genuine forecast, not a fact, and I flag it as one. What I can tell you is that the boundary exists, that it's the same boundary the calculator respected by accident, and that "the skill will stop mattering" is an argument you have to win per skill, not assert across all of them.

What to actually do

For any AI use in learning, run the test before you run the tool. Ask what the learner would have had to build in their own head to do this step unaided. If the answer is a mechanical operation understanding doesn't live inside, let the machine have it and spend the freed attention on the real target. If the answer is the mental model the exercise exists to construct, keep it in human hands — and consider pointing the AI the other way, as a sparring partner that increases the struggle: demanding you defend the claim, poking holes, refusing to hand over the conclusion.

The calculator didn't ruin math because it never touched the part of math that was thinking. Whether AI ruins learning depends entirely on whether we let it touch the part of learning that is. That's not a question the calculator can answer for us. It's the one it was quietly keeping us from asking.

Frequently asked questions

Isn't the calculator analogy still useful for the cases where it holds?
Yes, and that is the point. Wherever AI offloads a mechanical sub-skill that is not the learning target — formatting a citation, generating boilerplate, converting units, looking up a syntax — it is a calculator, and the analogy holds cleanly. The analogy is not wrong; it is over-applied. People reach for it to wave away every worry, including the cases where AI offloads the reasoning that was supposed to be the acquisition. The distinction between those two uses is the entire question, and 'calculators were fine' answers only half of it.
How do I tell whether a specific AI use is a calculator or a ghostwriter?
Ask what the learner would have had to build in their own head to do the step without the tool. If the answer is a mechanical operation that understanding does not live inside — arithmetic, formatting, recall of a fact you could look up — it is a calculator, and offloading it frees attention for the concepts. If the answer is the mental model, argument, or problem representation the exercise exists to construct, it is a ghostwriter, and offloading it hands you the residue of understanding without the reaction that produces it. Same tool, opposite verdict, depending on where the learning target sits.
Doesn't this just mean students should never use AI while learning?
No. It means matching the tool to the target. Use AI as a calculator for the mechanical scaffolding around a task, and as a Socratic sparring partner that increases the struggle rather than removing it — asking you to defend a claim, poking holes, refusing to hand over the answer. What you should not do is let it perform the specific cognitive act the exercise was designed to build. A model that drafts your essay is dangerous in a writing course and harmless in a chemistry course where writing is not the target.
Isn't offloading fine because the skills AI automates simply stop being valuable, the way longhand arithmetic did?
That is the strongest counterargument, and it holds for terminal skills whose only value was producing an output a machine now produces. It fails for developmental skills — ones whose value is the cognitive structure you build by exercising them, not the output. Working the problem builds the model you use to recognize the next problem; the essay teaches the thinking you carry into every future essay. Automating the output there does not retire an obsolete skill, it removes the mechanism that grows a mind. The question is always whether the skill was the destination or the road.

Filed under Cross-Disciplinary Deep Essays. Where biology, computation, markets, and philosophy collide.

Essays like this, in your inbox.

Thoughtful essays. No spam. Unsubscribe anytime.

Cross-Disciplinary Deep Essays

"AGI by Year X" Is Unanswerable Until You Name the Definition

"Are we close to AGI?" is incoherent because AGI names at least four incompatible criteria that come apart in practice. Separate them and the timeline debate dissolves into concrete, checkable questions.

9 min read