What Counts As Learning?

There is a chart in the OECD's Digital Education Outlook 2026 that continues to attract my attention.

It comes from a randomized controlled trial of 1,000 high school students in Türkiye, working through math exercises across six 90-minute sessions. One group studied the old-fashioned way — notes and textbook. A second group practiced with a general-purpose chatbot. A third practiced with a chatbot configured for tutoring, one built to withhold direct answers and support the process of figuring things out. During practice, the results were not close. Students using the tutoring chatbot solved 127% more problems correctly than the group working alone. Even the general-purpose chatbot group beat the unassisted group by 48%. If one were to stop reading there, one would conclude that AI had transformed math practice for the better…perhaps dramatically.

Then came the closed-book exam. Results? The tutoring-chatbot group performed about the same as the students who had studied alone. The general-purpose chatbot group performed worse — 17% worse than students who never touched an AI tool at all. The performance gains from practice had simply vanished. Brief lesson: successfully performing a task with GenAI does not automatically lead to learning. The tools made the students' work look better without the desired result of the students learning.

What happened here? A chatbot that supplies answers, or supplies enough scaffolding to get to an answer, is very good at helping a student produce a “correct” response. What it is not automatically good at is building the internal structure, such as retrieval pathways, pattern recognition, or productive struggle—items that enable a student to produce that response again, alone, under different conditions, weeks later. Researchers have a name for the mechanism most likely at work: cognitive offloading. When a tool is available to carry part of a task, people tend to let it carry that part, and the part they no longer do is often the part where the learning was happening.

Key question: what were we measuring before AI arrived? Were we ever measuring learning, or were we measuring something that merely correlated with learning until the correlation broke?

A worksheet score has never been learning. It has always been a proxy for learning, a kind of stand-in that we accept because, in a world without generative AI sitting next to every student, the proxy and the underlying thing moved together closely enough that we stopped noticing the gap. Practice performance tracked mastery well enough, for long enough, that we built entire grading systems on the assumption that they were nearly the same thing.

If a well-designed tutoring bot can produce practice scores that look exactly like mastery while leaving actual mastery untouched, then the tool has not corrupted our assessment; rather, it has exposed what our assessment was already vulnerable to: scoring the artifact of thinking rather than the thinking itself.

In other words, the difference between performance and understanding requires different kinds of evidence. If we want evidence of durable understanding (what we might well term “learning”), we need assessment conditions that actually isolate it and allow for things like transfer tasks in unfamiliar contexts (ability to reproduce in difference circumstances, over time), and the opportunity to interact with other learners (or teachers) who can pose follow-up questions that force further application of understanding. These are not new inventions; they are simply ways of assessing that are more difficult to outsource to AI, yet we’ve allowed ourselves (over time) to deprioritize them because faster, cheaper proxies were acceptable.

Independent schools, with their structural flexibility and scale, are better positioned than most to decouple practice from proof, to let students use AI within appropriate constraints where the goal is production, and to build assessment moments that actually require the offloaded work to have happened and to have mattered.

The Türkiye study's real finding is not that AI tutors are good or bad. It is that we have been coasting for a long time on the assumption that looking like learning and being learning were close enough not to bother separating. The fact is, they were never the same thing.

Next
Next

Lessons for Our Attention