The Calculator Question

The Calculator Question

We already settled how to handle a tool that automates part of the work — in math class, with the calculator. The rule we landed on is exactly the one the AI-in-school panic keeps refusing to use.

Every school is having the same fight, and almost nobody fighting it has noticed we already had it once, and won.

The fight is about AI. A student turns in an essay an AI mostly wrote, and the observers split into two camps that yell past each other: ban it, lock the laptops, back to blue books and proctors. Or embrace it, it’s the future, teaching without it is malpractice. Both camps think they’re facing a new problem. They’re facing arithmetic we’ve already solved.

We ran this experiment in math class in the 1970s and ’80s, with a machine that automated part of the work and frightened teachers in exactly the way the chatbot does now: the calculator. And we settled it so completely that the rule went invisible, lived by every teacher without a second thought. You take the calculator away when you’re testing whether the kid can add — the arithmetic is the thing you’re certifying, and a machine that does it for her ruins the test. You hand it back the instant the problem turns into distance-rate-time — now the arithmetic is just plumbing, and what you want to see is whether she knows what to do. The tool is taken away exactly when the skill it automates is the skill under examination, and goes back in when it’s just an assistant to a different skill.

That’s the whole framework, and it has been sitting in every classroom for forty years.

And forty years understates the lineage. The calculator that terrified the algebra teacher arrived with a body count of its own: the slide rule, engineering’s sidearm for a century, was dead within about four years of the first pocket scientific model; Keuffel & Esser, who had made them by the million, simply stopped. Go back further and the same fight is running in medieval ledgers. For centuries Europe did its math on the abacus, sliding beads on a counting board, and recorded the results in Roman numerals. Then Hindu-Arabic numerals, the 0 through 9 we use today, arrived carrying a new trick: the arithmetic itself could be done on paper, no beads involved. The bead-counters fought the algorists (the paper-arithmetic crowd) for generations, and in 1299 the money-changers guild of Florence banned the new numerals from account books, arguing that a written figure was too easy to doctor — a zero becomes a six with one stroke — while beads moved in the open, where everyone could watch. The new tool won anyway. It always does. Which is exactly the kind of comfort that deserves the suspicion it’s about to get.

The wrong question

“Should we allow AI in school” is the wrong question the same way “what percent of this was human” was the wrong question about authorship : a yes/no demanded of something that was never one decision. There’s no school-wide answer because there’s no school-wide skill. The proof, the lab report, the close reading, and the take-home research paper each want a different answer, and the calculator question is what separates them, asked one assignment at a time: is the thing the tool automates the thing I’m trying to measure here?

Sort the panic that way and it falls into two piles. Some assessments certify a foundational skill the student has to prove unaided — building an argument, following a proof, reading a literary page closely enough to know what it actually said. You can’t prompt your way out of not understanding; the model hands you a fluent paragraph and you can’t tell whether it’s right, which is the entire problem. For those, the laptop closes. It’s the addition test. Other assessments certify applied judgment standing on a foundation already laid — weighing sources, shaping a case, spotting which of the model’s four confident answers is the plausible-sounding wrong one. Test those with the tool banned and you’ve measured the student in the one condition she’ll never work in again: alone, from memory. I built a hiring gate that made precisely this mistake for four years, scoring people on unaided recall — the exact thing the job had stopped asking of anyone — so I’m not theorizing about how easy it is to certify the wrong skill with a straight face.

Where the calculator stops helping

Here’s where the precedent runs out, and it’s the part nobody really seems too interested in addressing.

The calculator was safe to hand back early because using one well never required you to have internalized arithmetic. You can be hopeless at long division and still nail the answer to how long it takes for delivery van A traveling at 55mph from Chicago to meet delivery van B traveling from LA at 68mph. And that’s because the calculator automates a means that sits cleanly apart from the end. The long division answer is 672 ÷ 41 = 16.39 hours; that’s the roughly 2,015 highway miles between the vans closing at a combined 123 miles an hour, knocked down by a factor of three so the long division stays polite. You’re welcome. AI doesn’t sit apart from this mechanism. Using it well demands formed judgment — taste, the ear for the wrong note, the instinct that says that citation is invented or that paragraph is confident and empty. And that judgment gets built almost entirely by the unaided struggle the tool is so good at letting you skip. The bad first draft is where you find out what you think. The proof you sweat alone is what grows the sense that later tells you a slick wrong proof is wrong. AI is nothing if not slick.

Crumpled balls of paper scattered across a school desk around one clean blank sheet with a well-worn pencil resting on it, lit by a window. The part the machine offers to skip.

So the real hazard in a classroom isn’t cheating on the test. It’s quietly skipping the formative friction that builds the person who could pass it honestly, and walking out credentialed, fluent, and unable to separate good work from garbage, because the skill that does the separating was never grown. The calculator could only ever steal the arithmetic. The model can steal the formation. Which is no doubt worse, so the alarm is understandable.

What’s still worth testing unaided

This turns teaching into a judgment call no policy can make from above, taken assignment by assignment. Is this a skill the student has to forge unaided, because the forging is the education? Protect it — tool removed, and don’t apologize for the friction. Is this applied judgment resting on a foundation already there? Hand the tool back, and notice that banning it now produces the artificial condition, not the rigorous one. A blanket ban is taking the calculator away from the entire math curriculum: you stop teaching the word problem out of fear of the addition test. A blanket allow is handing it out during the addition test: you graduate kids who can’t add and never found out how. The two failures are mirror images, and right now most schools are picking one of them to dodge the work of choosing per assignment.

If you teach and want the line drawn plainer, run the question down your own syllabus. The seventh-grader learning to build a paragraph writes it alone, because paragraph construction is the skill being certified; same for the times tables, the reading quiz, the first proof. Graded tests of foundations keep the tool out, full stop. The senior who already writes competently is a different case: let her brainstorm against the model, argue with it, have it mark up a draft, and then make her defend the result out loud and answer for every citation, because judging the machine’s output is now the skill being certified, and it’s the one she’ll use for the rest of her working life. In between, the honest default is disclosure: say what the tool did, show what you did. And no accusation should ever rest on a detector alone — a human reads the work before anyone’s name goes to the office. If those rules sound sensible, 98 high-schoolers wrote nearly the same ones into a mock congressional bill this summer, which tells you the calculator question isn’t hard to answer. It’s just hard to answer from the district office.

I got to watch someone choose right, once. In electrical engineering school I took an experimental calculus class from a professor who had decided that seeing the math was part of learning the math. Enrolling meant buying an HP-28S, one of the first calculators that could draw a graph, and we leaned on those little screens constantly, watching a function bend while its coefficients moved. That we might use the machines to cheat never seemed to cross his mind, to his credit. He had asked the calculator question and answered it out loud: the skill under examination was intuition, and the tool was hired onto the teaching side.

This is the part of the essay where I come clean and share my prompt from earlier:

How long does it take two delivery vans to meet if van A is starting
from Chicago at 55mph, and van B is starting from Los Angeles at 68mph?
Show me how you got the answer with long division plz.

That’s the uncomfortable reality of a tool that is unlike any calculator we’ve ever seen. I mostly remember how to do long division, I even have a minor in math (a reluctant requirement of that same degree), and I went straight to an LLM to solve a basic distance-rate-time problem. But keeping us on point, I already learned how to solve those once in my life, so I consider that skill at least loosely banked. I’m not being tested on my formation of learning long division in writing this, and that’s the point.

So getting back to removing the tools, none of this is a case for struggle as virtue, the back-in-my-day sermon students tune out for good reason. It’s narrower. Some struggle is mere toil and the machine should take it. Some struggle is the only process by which judgment gets built, and judgment is the whole thing a diploma is meant to vouch for. Telling those two apart was always the work; the calculator forced us to answer it once, cleanly, and we did. AI asks the same question and forces a second one underneath it. The calculator question is is the automated skill the one I’m testing. The harder one is is that skill the one that builds a student worth testing at all. When the answer to everything is one prompt away, what’s left worth protecting is the slow, unaided making of a mind that can tell whether the answer is any good. And the ability to judge the value of an essay written by a guy who will gladly use a prompt to solve a problem he once knew how to solve on his own.

Becoming Gnarly

Essays on AI, work, and choosing the harder path when it's worth it. New ones by email as they're published.

or grab the RSS feed