Does AI Hurt Learning? Inside the MIT Cognitive Debt Study
Does AI hurt learning? An MIT study found the weakest brain engagement in AI-assisted writing, and that drafting unaided first then adding AI worked best.
Not inherently, but how and when you use it matters a great deal. A 2025 MIT Media Lab study found that people who wrote essays with an AI assistant showed the weakest brain engagement of the three groups tested. It also found that participants who drafted unaided first and brought AI in afterwards did better than those who started with the tool.
That second result is the one worth building a habit around, and it is the one most summaries of the study skip. It is easy to read the headline finding as proof that AI damages your ability to think. The evidence does not reach that far, and the thing it does reach is more actionable anyway. Below is what the study actually did, what it found, what it cannot tell you, and the single change to your process that its own data points to. If your question is about permission rather than cognition, the line between studying with AI and handing in work the AI did for you is a separate matter covered elsewhere.
What did the MIT cognitive debt study actually do?
The paper is Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task, by Nataliya Kosmyna, Eugene Hauptmann, Ye Tong Yuan, Jessica Situ, Xian-Hao Liao, Ashly Vivian Beresnitzky, Iris Braunstein, and Pattie Maes, researchers at the MIT Media Lab. It was first posted to arXiv in June 2025, with a revised version following in December 2025.
Fifty-four participants each completed three essay-writing sessions, split into three groups. One group used ChatGPT to help write. One group used a search engine and no LLM. One group, labelled Brain-only, wrote with no external tools at all. The researchers recorded EEG during the writing to measure brain connectivity, ran natural-language analysis on the essays themselves, and had the finished essays scored by both human teachers and an AI judge.
Then came the part that makes the study genuinely interesting. In a fourth session, some participants swapped conditions. Some of the LLM group were moved to writing unaided, a condition the paper calls LLM-to-Brain. Some of the Brain-only group were handed the LLM for the first time, called Brain-to-LLM. Eighteen participants completed that fourth session.
Three measurement channels are running at once here: what the brain was doing during the task, what the finished text looked like, and what the writer could still recall afterwards. A study that only graded the essays would have missed almost everything that follows.
What happened as more of the writing was handed over
Across the first three sessions the pattern was ordered, and though the direction is not surprising, it is rarely measured this directly. The Brain-only group showed the strongest and most distributed brain connectivity. The search engine group sat in the middle. The LLM group showed the weakest. Cognitive engagement scaled down roughly in proportion to how much of the task had been handed to an external tool.
Two softer results are arguably more vivid than the EEG traces. Self-reported ownership over the finished essay was lowest in the LLM group and highest in the Brain-only group: the more the tool wrote, the less the writer felt the piece belonged to them. And LLM-group participants struggled to quote accurately from essays they had just produced.
That last detail is worth sitting with. This is not a chapter someone skimmed a week ago. It is an essay submitted minutes earlier under their own name. When material passes through your hands without ever being encoded well enough to quote back, it was never really processed, and that is about the plainest description of cognitive debt available. The name echoes technical debt in software: every individual shortcut is cheap and sensible at the time, and the cost arrives later, all at once. Across the four months of the study, the authors report that LLM users consistently underperformed at neural, linguistic, and behavioural levels compared with the other groups.
Does the order you use AI in matter?
This is where the study earns its attention, and where most retellings of it stop short. The fourth session was not a repeat of the first three. It was a switch, and the two directions behaved nothing alike.
Participants who had spent three sessions with the AI and were then asked to write unaided showed reduced brain connectivity, even with no tool in front of them. The under-engagement did not lift when the assistant was taken away. Whatever pattern had settled in persisted past the thing that produced it.
Participants who had spent three sessions writing unaided and were then handed the AI showed close to the opposite: higher memory recall, and more activation across occipito-parietal and prefrontal regions, putting them nearer the search engine group than the LLM group. Same tool, same task, opposite result, separated only by what came first.
| Aspect | LLM-to-Brain: AI first, then unaided | Brain-to-LLM: unaided first, then AI |
|---|---|---|
| Sessions 1 to 3 | Wrote each essay with an AI assistant | Wrote each essay unaided, with no external tools |
| The session 4 switch | Wrote unaided for the first time | Used an AI assistant for the first time |
| Brain connectivity in session 4 | Reduced: under-engagement that showed up even with no tool present | Higher activation, including occipito-parietal and prefrontal regions |
| Memory recall in session 4 | Not reported as improving once the tool was removed | Higher, and closer to the search engine group |
| What it suggests | Three sessions of leaning on AI left something that removing the tool did not undo | Doing the thinking first appears to make AI additive rather than substitutive |
The practical translation is almost embarrassingly simple: do the thinking first, then let AI act on thinking you have already done. Drafting from your own head and then asking AI to challenge the argument, find the gap, or tighten the prose is a different cognitive event from asking it for a starting point and editing whatever comes back. In the first case, AI arrives after your model of the material has been built. In the second, it arrives instead of building one.
What one preprint can and cannot tell you
Being honest about the limits is not hedging here, because the limits change what you should do with the result.
- It is a preprint. Posted to arXiv by a credible lab, but not something to describe as settled or peer-reviewed evidence.
- The sample is small. 54 participants across the main sessions, and only 18 in the fourth session that produced the most interesting finding. Eighteen people is a hint, not a proof.
- The task was essay writing, specifically. Essay writing is a generative task where the output is the thing being assessed. That is not the same activity as using AI to quiz yourself on a chapter or unpick a concept you are stuck on. Moving from “AI-assisted essay writing showed reduced engagement in this study” to “AI hurts learning” is a bigger jump than the data supports.
- The reported results are directional. Strongest, weakest, reduced, higher. Anyone quoting you a precise percentage drop in brain activity from this study has invented it.
What survives all of that is still worth acting on. In a controlled setting with brain measurement running, the sequence in which AI entered the work changed the outcome, and it changed it in a direction that costs you nothing to respect. The researchers’ own framing is careful in the same way: they present this as grounds for concern and further research, not as a verdict.
How to keep AI on the useful side of cognitive debt
- Attempt first, always. Write the rough draft, try the problem, sketch the argument from memory. Then open the AI. This is the study finding turned into a rule, and honestly it is most of the actionable content in the whole paper.
- Make the AI test you, not tell you. Reading a good explanation produces the feeling of understanding. Answering a question you cannot look up produces the evidence. That gap is the mechanism behind active recall and spaced repetition, and it is why an AI-generated quiz behaves so differently from an AI-generated explanation.
- Run the quoting test on yourself. After a session, close everything and write down what you can reproduce unaided. If the shape of it will not come, you have not learned it yet, whatever the tool told you along the way.
- Keep AI pointed at material you already hold. A model recalling a topic from training data has room to invent a source, which is why AI confidently fabricates citations. A model working from a document you supplied has much less.
- Notice when the work stops feeling effortful. Difficulty is not a sign the method is failing. The prompting discipline in the guide to using ChatGPT to study is built around keeping the effort where it belongs, which is with you.
It is worth being precise about which task this study is a warning about, because that gets lost fast. The finding concerns AI writing on your behalf, where the output is the assessed work and the thinking is the part being outsourced. Generating practice material is a different job. When you upload your notes and get back a quiz, a set of flashcards, or recall prompts, the AI is not doing the cognitive work of the study session. It is building the prompts that force you to do it, and the retrieval still has to come out of your own head. That sits closer to what the Brain-only and search engine groups were doing well than to what the LLM group was doing badly.
That is the design GeniusPal is built around: point it at material you already have, and it produces quiz, flashcard, and recall sets you then have to answer, with a small number of study-set generations free for the life of the account. This is not a claim that any tool cancels cognitive debt, and the study is not about study-set generation at all. The shared principle is narrower and more durable than any product: active engagement beats passive outsourcing, and a tool that keeps you retrieving rather than reading is on the right side of that line.
So does AI hurt learning? On this evidence, what hurts is letting AI do the part of the work that was supposed to build the understanding. The same tool, arriving after your own attempt, looked additive rather than substitutive in the one session that tested it. That is a single small study and it deserves to be held loosely. But it costs nothing to try first and reach for AI second, and the downside of getting that order wrong may not become visible until the exam.
Frequently asked questions
- What is cognitive debt?
- Cognitive debt is the term MIT Media Lab researchers use for the accumulated cost of repeatedly letting an external tool do thinking you would otherwise have done yourself. The name echoes technical debt in software: each individual shortcut is cheap and sensible in the moment, and the bill arrives later in a lump. In their study, participants who used an AI assistant across three essay-writing sessions showed the weakest and least distributed brain connectivity of the three groups tested, reported the lowest sense of ownership over the finished essay, and struggled to quote accurately from work they had produced minutes earlier. The debt shows up as material that never got encoded deeply enough to retrieve later. Nothing feels wrong while you are accruing it, which is exactly what makes the idea worth naming.
- Does the order you use AI in matter?
- In this study, yes, and it may be the most useful thing the researchers found. In a fourth session they reassigned some participants. Those who had used the AI assistant for three sessions and then wrote unaided showed reduced brain connectivity even with no tool in front of them, a pattern of under-engagement that outlasted the tool itself. Those who had written unaided for three sessions and were then given the AI showed higher memory recall and more activation across occipito-parietal and prefrontal regions, putting them closer to the search engine group. The practical reading is that drafting from your own head first, then bringing AI in to challenge, check, or tighten that draft, produced better outcomes than starting with AI. Only 18 participants completed that session, so treat it as a strong hint rather than a settled fact.
- Is the MIT study conclusive proof that AI hurts learning?
- No, and the authors do not present it that way. It is a preprint from the MIT Media Lab, posted to arXiv in June 2025 and revised in December 2025, and it frames its results as raising concerns and calling for more research rather than delivering a verdict. Three limits matter. The sample is small: 54 participants across the main sessions, and only 18 completed the fourth session that produced the most interesting finding. The task studied was essay writing, not learning in general, so nothing here measures what happens when you use AI to quiz yourself on a chapter. And one study, however carefully instrumented, is a data point rather than a body of evidence. What it does support is narrower: in this task, AI-assisted writing showed reduced engagement, and the sequence in which the tool entered the work appeared to matter.
Keep reading
- AI & Studying
Can Turnitin Detect ChatGPT? What to Know
Turnitin does try to detect ChatGPT, but AI detectors are imperfect and flag real human writing too. Here is what its AI detection does, how reliable it is, and why integrity matters more than the detector.
July 7, 2026 · 8 min read - AI & Studying
Is Using AI to Study Cheating? Where the Line Is
AI can deepen your understanding or quietly do your thinking for you. Here is the honest line between studying with AI and crossing into cheating.
July 7, 2026 · 8 min read - AI & Studying
Can ChatGPT Solve Math Problems? What to Know
ChatGPT can solve arithmetic, algebra, and many word problems, but because a plain model predicts text rather than calculating, it can be confidently wrong. Here is when to trust it and how to check.
July 24, 2026 · 9 min read - AI & Studying
AI Hallucinations: Why ChatGPT Invents Facts and Citations
Language models hallucinate because their training and their benchmarks reward confident guessing over admitting uncertainty. Here is the mechanism, what the research shows about how stubborn the problem is, and why generating study material from a source you uploaded closes off the most common route to an invented citation.
August 2, 2026 · 8 min read