Show Me the Three
AI can make a student's work look better without making the student better. The real question is what the learner owns: where they started, how far they moved, and whether they can explain and defend the work.
A student walks in with almost no understanding of something. Call it a 1 out of 10. They hand the assignment to an AI and get back something that looks like a 10 — clean structure, confident language, better work than they could have produced alone in a year.
What do I actually have? A 10-looking product sitting in front of a student who is still at 1.
Now run it again. Same kid, same starting point. This time they ask the AI to explain the idea a different way, push back on an answer that seemed wrong, ask for an example, notice the part they'd skipped. They end up at a 3. What they turn in is worse than the first version. Rougher, thinner, obviously unfinished.
I want the 3, and it isn't close.
That's the entire argument. Everything else in the current panic about AI and education is downstream of confusing those two students, and we were confusing them long before anybody could generate an essay in nine seconds.
I've had to convince my own students of this, which tells you how deep the training goes. They arrive certain that the beautiful finished product is the thing I want, because for twelve years it was the thing everybody wanted. I tell them I don't expect a perfect project, I don't believe they can make one, and I'm not especially interested in whether the final artifact looks impressive. If you started at 1 and you're at 3 now, show me the 3. At first they don't believe me. Then they show me something honest and rough, I treat it as the win it is, and eventually the message becomes credible.
What the conversation actually is
The question I keep hearing is: how are you supposed to know whether the student understands any of it?
Ask them. That's it. That's the technology.
But I want to be precise about what that looks like, because people hear "ask them" and picture an oral exam — the student standing there while the professor probes for weakness, fluency deciding the grade. That's not what happens in my room, and if it were, most of the objections people raise would be fair.
Here is the actual move. Tell me about your project. The kid starts talking. Something they say is interesting, so I follow it. Something else is vague, so I ask a smaller question. If they can't find the words, that isn't a failure condition — that's my cue to come at it from a different direction until the understanding surfaces or I find the edge of it. I'm not waiting to see whether they can perform on demand. I'm pulling it out of them.
I don't need them to be good at talking. I want them to get better at talking, and at writing, and at explaining themselves under pressure — those are real capabilities and we work on them. But they are not the toll booth you have to pass through before I'm willing to believe you learned something.
And nobody fails a conversation. That's the part that dissolves most of the worry. If I find the place where it falls apart, the next sentence is: go back, fix that, make it yours, come find me. If it takes two rounds, fine. Three, fine. The conversation isn't a verdict. It's how we figure out what to do next.
Which means no single exchange is load-bearing. I can misread a kid on a Tuesday — they were mid-thought, or tired, or having a day — and it costs nothing, because I come back around. The thirty-second read is a first pass, not a file closed.
Then we blamed the tool
Set that next to what a lot of institutions are doing instead.
Generative AI has made it easy enough to produce work you don't understand that some professors and colleges are moving back toward handwritten, in-person exams. Meanwhile other universities are warning faculty that AI detectors aren't reliable enough to treat as evidence of cheating. So we have a technology that can explain a hard idea six different ways at midnight, and one of our answers is the little paper booklet.
Put away the computer. Put away the phone. Here's a pen and a clock. Now show me what you know.
I understand the appeal. Sometimes I genuinely want to see what someone can do unassisted. But why did we decide that this artificial condition was the purest form of knowing? The blue book removes the tools so we can trust the artifact. I'd rather leave the tools in place and test the person.
In actual life I want people using everything available to them. Look it up when you're unsure. Ask for another explanation when the first one didn't land. Check the calculation. Ask for criticism. Then own the result — explain it, defend it, change it when conditions change, notice when the tool is wrong, know when you don't know. That's closer to competence than whatever you can reproduce in seventy-five minutes after every useful thing has been taken away from you.
We built a space-age tool, discovered that a system supposedly devoted to student understanding had never really mastered the oldest method humans have for finding out whether someone understands something, and then blamed the space-age tool.
Fake learning is not a new product
I have very little nostalgia for the educational world AI supposedly ruined, because I spent an unreasonable amount of my life inside it. I started college in 1983 and finished in 1988. I started MBA programs and didn't finish them. I went back full time from 1998 through 2003 for a master's, then went back again for more graduate work after that.
That's roughly ten institutions across five states, community college through full research universities, in person and online, as a teenager and as an adult with a career. None of them elite, and I'll say plainly that I don't believe a seat in a more prestigious one would change what I'm about to say.
There were excellent professors and books that mattered. But let's not rewrite history. People didn't do the reading, and we held discussions about the reading anyway. Somebody said enough to establish they knew roughly what the chapter was about, somebody else built on it, and the machinery moved forward. We learned what participation looked like and how to sound like graduate students discussing graduate-school things.
Years later I became a high school teacher and recognized the pattern immediately. Better vocabulary in grad school. Older people. More practice at sounding educated. Same basic transaction: produce enough evidence of engagement to keep the machinery moving.
ChatGPT didn't invent that. What it did was make the performance dramatically better. The student who used to produce mediocre work they didn't understand can now produce excellent work they don't understand. The distance between the artifact and the person got too wide to ignore, and we decided that was an emergency rather than a diagnosis.
The number was never what we said it was
The other half of this has bothered me for years.
I've sat through teams of teachers arguing about whether a problem should be worth four points or six, whether the test needs another question on systems of equations, whether this section should count for more. Then we add it up, produce a number, sometimes carry it to a decimal place, enter it in a gradebook, and act as though we performed a measurement.
We didn't norm that test against any population. We never established that our handful of questions was a representative sample of what the student was supposed to have learned. We never showed that our weighting matched the relative importance of the capabilities. Most of us couldn't tell you what real difference exists between a 78 and an 82 beyond four points on the particular questions we happened to write that week.
Rubrics don't fix this. They make expectations clearer, which is genuinely useful, but putting judgment into boxes doesn't remove the judgment. We keep building structure around uncertain measurements and then treating the structure as proof of precision.
That doesn't make teacher-written tests worthless. I've made hundreds of them and they tell me useful things. But a useful snapshot is not a calibrated instrument, and the precision of the number has been quietly disguising the weakness of the claim.
An 82 looks exact. What did it establish? How representative were the questions, how arbitrary were the point values, how much did reading speed matter, how much did one careless mistake cost, how far does the student's understanding extend past the narrow sample we collected that day?
So education spent decades generating authoritative-looking numbers and calling them measures of learning. Now students have tools for generating authoritative-looking work of their own, and we're appalled. There's an uncomfortable symmetry in that.
"But then how do we rank them?"
This is the objection people reach for, and it's weaker than it sounds.
If grades stop meaning what we pretended they meant, how does anybody sort people? Employers will sort people the way employers have always sorted people. That was never the gradebook's achievement and it isn't the gradebook's to lose.
But I want to make a stronger claim, and I want to be honest that I can't prove it.
I've spent decades watching people get hired — in corporate work, consulting across a range of industries, and now education. Nearly every serious miss I've watched happen would have surfaced in a real conversation. I'm convinced of that. I can't demonstrate it, and I'm aware I only remember the misses. Nobody keeps a record of the candidate I'd have misjudged who turned out fine.
Here's why I believe it anyway. The thing we call an interview usually isn't a conversation. It's a list of prepared questions producing prepared answers, scored against a checklist, and nobody actually wants the candidate's real answer to "what would you do in this situation." We ask the generic question, we accept the generic answer, we fill in the boxes. Then the problems arrive six months later, and they're the same problems that were sitting visibly in the room during the interview.
The exceptions were real. A couple of the higher-end firms I worked with knew exactly what they were looking for and skipped most of the ritual. That's the pattern running through this whole essay: the places that know what they want tend to need less procedure, and the places that don't build more of it to cover the gap. In education, hiring runs mostly on impressions with a checklist stapled on top, and then we act surprised.
Somebody will say conversation is impressions too. Sure. It's also a method we've had for centuries and set down because it won't produce a number. "We can't quantify it, so we can't rely on it" is the same confusion as the 82, relocated from the gradebook to the hiring committee.
And yes, some people are good at concealment. They're rare, and they're not generally who's applying to teach ninth grade. Designing the whole process around the rare one degrades it for everybody who was never the problem — which is the blue book argument wearing a different suit.
Most people cannot hide who they are for very long, if you are actually talking to them.
Growth is the measure
Here's what I actually believe, plainly.
You cannot ask a person for more than as much growth as they can make.
If a kid arrives in my classroom in twelfth grade sitting at a 1, that is not their fault and it isn't mine either. Something happened across twelve years, or across a whole life. Some kids have been handed a much harder set of circumstances than anybody in the room is accounting for. And the response to that is not to hold up a chart and announce that a senior is supposed to be at a 7.
Where does the 7 come from? Usually a grade-level expectation, applied as though it were a requirement. Some benchmarks are built carefully, and I'm not claiming they're all invented out of nothing. But even a good one can only tell you where a student stands relative to other students or to a standard. None of them can tell you what was possible for this kid from where this kid started. That's the only question I can actually act on.
If a kid comes to me at 1 and leaves at 4, that is an enormous piece of work. If a kid comes in at 5 and leaves at 8, also enormous. Neither one is diminished by the other, and neither one is measured by where a chart says they should have been.
Does that mean the 1-to-4 student gets an A? No. They probably end up with a C, which is roughly where they'd have landed anyway, and I'm fine with that. The grade is a rough marker of where you ended up. That's all it ever was. What I'm not willing to do is let that C be the whole verdict on the kid, because the C can't see the distance traveled and the distance traveled is the part I care about.
The number records where you landed. The celebration is for how far you moved. Those are two different things, and almost everybody collapses them into one.
And if that kid walks into the next class and slides back to a 1 because the incentives there reward the artifact again — that's a real problem, and it isn't mine to solve from inside one room. It's also not a reason to have asked less of them, or more.
What we owe them
Growth-only invites a fair objection. Doesn't that mean anything counts? Senior leaves at a 4, we throw a party, everybody goes home?
No. But I'm not going to answer it by inventing a floor and announcing it either.
Here's what I actually do. I ask them. You're nine months from being out there. Do you feel ready? Are you ready to weigh a major decision — not have one handled for you, weigh it? Do you know how you'd even go about deciding? Do you know what you're walking into, or what to do about it, without a parent or somebody else standing behind you?
They say no. Nearly all of them. They know exactly what position they're in, they know what's coming, and most of them want help. So we talk about it directly. How to think through something. How to decide. How to tell whether you're being told the truth.
I didn't hand them a floor. They can already see it. That's worth sitting with, because it means the minimum isn't a chart somebody normed. It's the thing a seventeen-year-old can name out loud the moment an adult asks honestly. I've tried to describe that minimum more fully in The 18-Year-Old Floor, but the need for one isn't theoretical to these kids.
What I can't do is settle what the system's minimums ought to be, or fix twelve years in nine months. I have a hundred and ninety students, and a large share of them were failed by this thing long before it was my turn. That isn't a figure of speech and it isn't a complaint about the kids.
What bothers me most is how quiet it stays. Ask teachers one at a time and you'll hear the same thing again and again — a lot of these kids were never taught to think through a problem, and everyone can see it. Ask the same question of a room, a meeting, a public forum, and it isn't on the agenda. It's an open secret that never becomes a priority.
So no, I can't fix the floor from inside one classroom. What I can do is know roughly where it sits, work toward it, and move a kid as far as nine months allows.
I can't control what happened to a seventeen-year-old before he got to me. I can do something about where he ends up.
Growth is how anybody gets to a floor. It isn't a consolation prize handed out instead of one.
It's easy because I built it to be easy
The obvious objection to everything above: that's fine for you, you have time. A high school teacher has well over a hundred students. A professor may have hundreds. You can't run an oral defense after every assignment.
True. So I built the time in.
The broader course design is laid out on the AQR site's Why AQR page, and the AI-specific approach is on Why AI?.
For most of the period I'm not standing at the front talking. Students are working. I'm moving around the room, looking at what's in front of them, asking what they're doing and why, checking what's actually theirs. Most of those exchanges are short. Sometimes thirty seconds tells me plenty. Sometimes the first answer gives me no reason to stay, and sometimes the second one opens something up and I sit down. The assessment isn't bolted onto the end of the learning — it's happening while the learning happens.
That was a design decision, and it's the reason any of this is easy. I'm not doing something clever. I removed the conditions that make it hard.
Which means if this doesn't scale to a 200-person lecture hall, the finding isn't that conversation is inferior assessment. The finding is that the format makes good assessment impossible. Schools need methods that survive contact with large numbers, so we built scalable proxies — tests, essays, rubrics, points, grades. Then we forgot that scalability was the compromise and started treating the compromise as the superior method.
AI has broken some of those compromises. That's useful information about the compromises.
I'm not offering a fix
I should be clear about what this is.
It isn't a reform proposal. I don't have a plan to scale any of it, and I don't believe this system gets fixed at scale — it's too big, too old, and too invested in its own paperwork. I'm not trying to right the ship. I'm saying I can see what's wrong with it from where I'm standing, and I can decide what happens in the one room I control.
That's also all the AI panic is really about. We can ban the tools, lock the browsers, bring back the blue books, add the surveillance, and get the paperwork looking trustworthy again. Then we can go back to believing that the student with the good grade understands, the student with the finished assignment learned, and the institution knows what its numbers mean.
That's the delusion worth worrying about. Not the chatbot.
AI didn't invent fake learning or fake precision. It widened the gap between how something looks and what's underneath it until we couldn't keep not seeing it. And what it exposed is almost embarrassingly old-fashioned: after all the assessment theory and all the rubrics, we still have to be able to sit across from another person and tell whether they know what they're talking about.
So stop confusing a finished artifact with a capable person. Decide what you're actually going to reward in the room you control. And when you want to know whether it worked, ask the kid — then listen closely enough that the next question matters.
If you're working through this in your own classroom, school or house, I want to hear one thing in particular: how are you figuring out what the learner actually owns once AI has touched the work? Send me a note.