AI & Studying By Shannon August 21, 2026 9 min read

Can AI Replace Tutors? What It Can and Cannot Do

Can AI replace tutors? It matches a tutor on explaining and drilling, and falls short on diagnosing why you are stuck and on making you show up.

Not completely. AI already covers the explaining half of tutoring well: it restates a concept as many times as you need, works through examples, and generates practice at any hour. What it does not do is diagnose why you keep getting something wrong, or make you show up. Those two jobs are what most tutors are really being paid for.

The question usually gets argued at the wrong altitude, as a verdict on AI in general. An hour with a tutor is several different jobs bundled into one rate, and AI has taken over some of them almost completely while barely touching others. Splitting the bundle is what makes the decision answerable, so what follows takes each half in turn, then checks both against what the research has actually measured.

What can AI already do as well as a tutor?

More than the sceptical version of this argument tends to admit. Four things in particular, and every one of them is something students currently pay for by the hour.

  • Explaining the same idea again, differently. A tutor explaining a concept for the fourth time is doing skilled work under two constraints: the hour is finite and they are tired. A model has neither problem. Ask for the same derivation three ways, with a fresh analogy each time, and the fourth attempt costs exactly what the first one did.
  • Being there at 11pm. Confusion arrives the night before the problem set is due. The tutor arrives on Thursday. That mismatch is one of the most common reasons a tutoring relationship fails to help with the specific thing you were stuck on, because by Thursday you have either worked around it or given up on it.
  • No social cost to asking. Plenty of students will not admit to a person that they never understood the material from two months ago. A text box removes the audience, and with it the incentive to bluff.
  • Generating practice. Turning a chapter into twenty questions used to be either tutor homework or an evening of your own. It is now close to free, which changes how much drilling is realistic in a week and makes active recall and spaced repetition far cheaper to actually run.

If you want the practical version of this rather than the argument about it, the technique is a role prompt that keeps the model asking questions instead of handing over answers, which we walk through step by step in our guide to using ChatGPT as a tutor. If you would rather use a tool built for tutoring than configure a general chatbot yourself, our roundup of the best AI tutor apps compares the purpose-built options.

One caveat belongs in this section rather than the next one, because it undercuts the strengths above instead of sitting alongside them. A model will produce a wrong formula, date, or citation in exactly the fluent register it uses for correct ones. A tutor who is unsure at least hesitates, and the hesitation is information. That failure mode has a recognisable shape, which we set out in why AI invents facts and citations.

What does a human tutor do that AI cannot?

Four things again, and none of them are about subject knowledge. This is where the replacement argument runs out of road.

  • Diagnosis. A good tutor spends much of the hour watching how you work, noticing the step where you slow down, and forming a theory about the single misconception producing a whole run of wrong answers. AI answers the question you thought to ask, which is a genuinely different job. Doing it on yourself is unreliable for a structural reason: the errors you can spot unaided are mostly the ones you have already fixed, and the feeling of understanding while you reread is a poor guide to whether you could produce the answer cold. That trap has its own anatomy, covered in why rereading feels like learning when it is not.
  • Accountability. A scheduled hour with a person who is expecting you creates an obligation that an open browser tab never generates. For a large share of students paying for tutoring, this is the actual product, whatever the invoice says it is. Nobody notices when you stop opening the chatbot.
  • Reading what you did not type. Hesitation, a guess delivered in a confident voice, the particular quality of silence that means you are lost: a tutor in the room reads all of it and adjusts before you have said anything. A text box sees only the words you chose to send, already filtered through whatever you judged worth saying.
  • Continuity across months. A tutor who has taught you since September knows which topic you quietly dropped in October, and brings it back in February without being asked. A fresh chat starts cold, holding only what you remembered to tell it, and what you remember to mention is filtered by exactly the blind spots you needed help finding.

What does the research actually show?

Both sides of this argument reach for the same number, and it is the wrong one. The claim that one-to-one tutoring lifts students by two standard deviations comes from Benjamin Bloom in 1984, and the evidence underneath it is thinner than the fame of the figure suggests. The Nickow, Oreopoulos and Quan meta-analysis of tutoring experiments notes that Bloom was presenting results from small-scale experiments run by two doctoral students. Three more recent findings are more sober and far more useful.

  • Human tutoring measures around 0.79, and software around 0.76. Kurt VanLehn reviewed the experiments comparing human tutoring, computer tutoring, and no tutoring for Educational Psychologist in 2011. The widely believed ladder ran 0.3 for answer-based systems, 1.0 for intelligent tutoring systems, and 2.0 for adult human tutors. His review did not confirm it. Human tutoring came in at d = 0.79 and intelligent tutoring systems at 0.76, which in his own words makes them nearly as effective as human tutoring. That gap was small well before generative AI existed.
  • Who tutors you matters, and so does whether the sessions are enforced. Nickow, Oreopoulos and Quan pooled dozens of randomised tutoring trials in preK-12 education and found an overall effect of 0.37 standard deviations. The interesting part is the variation. Effects ran stronger for teacher and paraprofessional tutors than for volunteer and parent tutors, and stronger for programmes held during school than for those held after it. Expertise and structure are doing real work in that number, which is a useful thing to know when you are pricing a replacement for either.
  • A well-built AI tutor can beat an active learning class. A randomised controlled trial published in Scientific Reports in June 2025 put 194 Harvard physics undergraduates through both conditions in a crossover design, one lesson on surface tension and one on fluid flow. Median learning gains in the AI condition were over double those in the classroom condition, with the effect size estimated at 0.63 by linear regression and between 0.73 and 1.3 standard deviations once the ceiling effect in the post-test was accounted for.

That 2025 trial is the strongest evidence available on the pro-AI side, which is exactly why its caveats deserve reading. The tutor was custom-built by physics instructors, with question-specific prompts, scaffolding, and guardrails designed to stop students shortcutting the thinking, so it is a poor proxy for a general chatbot you steer yourself. The comparison ran against an active learning classroom, and no private tutor appeared in the trial at all. And the authors write that an AI tutor “should not replace in-person teaching”, recommending instead that it bring every student up to a level where class time can be spent on harder work. The people with the best result in this field are the ones arguing hardest against reading it as a replacement.

AspectWhat a general AI tool coversWhat a human tutor adds
Explaining a concept againUnlimited attempts, new analogy each time, no fatigue and no impatienceVery little, once the fourth explanation is free
AvailabilityAny hour, including the night before the deadlineNothing, unless the scheduled slot is the point
Generating practiceA chapter becomes twenty questions in secondsQuestions pitched at your exam and your known weak spots
Checking your workingWill find the arithmetic slip and the wrong step reliablyNotices you got it right for the wrong reason
Finding the misconceptionAnswers the question you asked, and cannot see the pattern behind your errorsForms a theory about why a run of answers went wrong, then tests it
Making you show upNone. Nobody notices when you stop opening itA person expecting you at a fixed time, which is often the real product
Continuity across a termStarts cold, holding only what you remembered to paste inRemembers the topic you dropped in October and returns to it
CostFree to a few pounds or dollars a monthTypically an hourly rate, which is what makes the trade-off worth doing properly
Tutoring is several jobs sold as one hourly rate. The top four automate well and the bottom four barely automate at all, which is why the answer depends entirely on which rows you were buying.

So do you still need a tutor?

Look at what your hour is actually buying, because that decides it rather than any general view about the technology. Three cases cover almost everybody.

Case one: the sessions are mostly re-explanation and practice. AI covers a large share of that now, at a fraction of the cost, and you can test the claim cheaply before you cancel anything. Run a fortnight of your own revision through a chatbot configured as a tutor, keep a list of the moments it genuinely unstuck you, and compare that honestly against what the same fortnight of paid hours produced. Most people find this test decides it in one direction or the other within two weeks.

Case two: the sessions exist because you would otherwise not do the work. Keep the tutor. The standing appointment is the product, and nothing in software has replaced it yet. Every AI tool in this space is opt-in every single time, which is precisely the property that fails a student who cannot reliably opt in. Being honest with yourself about which case you are in is worth more here than any feature comparison.

Case three: your tutor keeps finding the thing you misunderstood. This is the expensive one to replace, and the one most often mistaken for case one. A tutor who surfaces a different specific misconception week after week is doing diagnostic work you cannot reliably do on yourself. If you can name three separate misunderstandings your tutor caught this term that you had no idea you held, you are paying for diagnosis, and that is a very different purchase from paying for explanation.

A hybrid is also allowed, and for most students it is the honest answer. Fewer paid hours, spent on the diagnostic work and the topics that keep breaking, with the explaining and the drilling moved onto tools that do it for nearly nothing. That reallocation is the real change AI has made to tutoring, and it is a much less dramatic claim than replacement.

Where a question generator fits into this

GeniusPal occupies one band of the table above, and naming which band is more useful here than a pitch would be. Hand it a file you already own, whether that is a chapter, a lecture deck, or the notes you took yourself, and what comes back is a set of questions about that material. Sit them as a quiz, flip them as flashcards, or take them in recall, which asks you to write the answer out from memory before it shows you the source to mark yourself against. That is the drilling row, and nothing above or below it. The reason a tool exists for just that row is that generating the practice is the piece of tutoring which automates cleanly, while being the piece that eats an evening when you do it by hand.

An article arguing for honest limits has to apply a few to itself, so here are the ones that matter against the rest of the table. Diagnosis is absent: the tool holds no theory about which misconception is generating your errors, only the material you gave it. Conversation is absent too, since one file goes in and one set of questions comes out, with nothing to reply to and no way to ask a follow-up. Only files carrying text can be read, so a hint your tutor said out loud reaches it solely by way of notes you typed. And it carries no picture of you across a term, which means nothing in it will resurface the topic you dropped in October. A new free account carries two study-set generations for the lifetime of that account rather than a monthly refill. Where diagnosis or accountability is the thing you are buying, a question generator will not replace it, and the sensible move is to keep those hours and stop paying anyone to write your practice questions.

Frequently asked questions

Can AI replace human tutors?

Not completely, and the reason is narrower than the debate around it suggests. AI already handles the explaining half of tutoring well: it restates a concept five different ways, works an example step by step, and generates practice questions at three in the morning without getting bored or impatient. The measured evidence backs that up. Kurt VanLehn reviewed the tutoring experiments in Educational Psychologist in 2011 and found computer tutoring systems reached an effect size of 0.76 against no tutoring, close to the 0.79 he measured for adult human tutors. What AI does not do is diagnose. A skilled tutor watches you work, notices the step you hesitate on, and forms a theory about the one misconception producing a whole pattern of errors. AI answers the question you thought to ask, which is a different job. It also supplies no accountability, since nobody is expecting you on Tuesday and nobody notices when you stop turning up.

Is an AI tutor as good as a human tutor?

On measured learning gains the gap is smaller than most people assume. VanLehn found human tutoring produced an effect size of 0.79 and computer tutoring 0.76, both a long way below the two standard deviations often quoted from Benjamin Bloom. A randomised controlled trial published in Scientific Reports in June 2025 went further, finding that 194 Harvard physics students learned more in less time from a custom AI tutor than from an active learning class, with a median learning gain over double that of the classroom group. Two caveats decide how much that result means for you. The tutor in that trial was built by physics instructors with expert-written prompts and guardrails, so it is a poor proxy for a general chatbot you steer yourself. And the comparison was against a classroom, never against one-to-one human tutoring. The authors state plainly that an AI tutor should not replace in-person teaching.

Should I stop paying for a tutor if I use AI?

Work out what the sessions are actually buying, because that decides it. If your hour is mostly re-explanation, worked examples, and practice questions, AI covers a large share of that at a fraction of the price, and you can test the claim cheaply by running a fortnight of revision through a chatbot before you cancel anything. If the sessions exist because you would otherwise not do the work at all, keep the tutor. The standing appointment is the product in that case, and no software has replaced it yet. The third case deserves the most weight. A tutor who keeps finding the specific thing you misunderstood, week after week, is doing diagnostic work you cannot reliably do on yourself, because the errors you are able to spot unaided are mostly the ones you have already fixed. Paying for diagnosis is very different from paying for explanation.

What can a human tutor do that AI cannot?

Four things, and none of them are about knowing more than the model does. First, diagnosis: a tutor watches how you solve a problem, spots the step where you slow down, and forms a theory about the misconception underneath a run of wrong answers rather than treating each one in isolation. Second, accountability: a scheduled hour with a person creates an obligation that an open browser tab never generates. Third, reading you: hesitation, a guess delivered in a confident voice, the particular silence that means you are lost, all of which are invisible to a text box that only sees what you chose to type. Fourth, memory across months: a tutor who has taught you since September knows which topic you quietly dropped in October and brings it back unprompted. A fresh chat starts cold, holding only what you remembered to tell it.

Try our study app free