Exam Prep By Shannon July 31, 2026 11 min read

How to Review a Practice Test: The Error Log Method

How to review a practice test: log every wrong answer by why you missed it, not just that you missed it. Careless slip, knowledge gap, misread, or timing.

Reviewing a practice test means logging every wrong answer by why you got it wrong, not just that you got it wrong. Sort each miss into a cause: careless slip, knowledge gap, misread question, timing pressure, or wrong approach. The pattern across those causes is what tells you what to study next. The score does not.

That record is called an error log, and it is the difference between a student whose practice scores climb and a student whose practice scores wander sideways for three months. The test is not the training. The test is the diagnostic, and a diagnostic you do not read is just an expensive afternoon.

Why most students waste a practice test

The standard review takes eleven minutes. You total the score, feel something about it, click through the explanations for the questions you missed, nod at each one, and close the tab. Nothing about that is lazy. It feels like reviewing. It is the thing almost everyone does, and it produces close to nothing, for three specific reasons.

Reading an explanation is not the same as being able to answer. A worked solution is written to be followable, so following it is easy and feels like understanding. You never find out whether you could have produced it, because the moment you read the explanation you lose the ability to test that. The information you most needed was destroyed by the act of looking.

You end up with a list of topics, not a list of causes. Ask a student what went wrong on their last test and you get an answer like "stoichiometry and inference questions". That is a description of where the misses happened, not why. Two students can miss the identical stoichiometry question, one because they cannot balance an equation and one because they read grams where the question said moles, and those two students need completely different weeks.

The score gets all the attention and carries the least information. A single practice score is noisy enough that a swing of a few points means almost nothing, and it is also the one number you cannot act on. You cannot study harder at a 1340. You can study the four question types that produced eleven of your nineteen misses.

The research on learning from errors makes the same point from the other direction. In a review of the field published in the Annual Review of Psychology, Janet Metcalfe found that errorful learning followed by corrective feedback is beneficial to learning, and specified what that feedback has to contain: corrective feedback, including analysis of the reasoning leading up to the mistake, is crucial. The mistake alone teaches nothing. The reasoning that produced it is the payload, and an error log is simply a structured way of writing that reasoning down before you forget it.

What is the error log method?

An error log is one row per missed question, recording what you answered, what the correct answer was, and, most importantly, a cause drawn from a fixed short list. Fixed categories are the whole trick, because categories can be counted and free-form notes cannot. Twelve misses that are all labelled "careless" is a finding. Twelve paragraphs of reflection is a diary.

1Score only

Mark right and wrong. Do not open a single explanation yet. You need the misses intact.

2Redo cold

Reattempt each miss with no notes and no key. This is the test that separates a slip from a gap.

3Read the key

Now read the explanation, and read it against the second attempt you just produced.

4Label the cause

One category per miss from the fixed list. If two apply, pick the one that came first.

5Count and act

Tally the causes across tests. The biggest pile is what next week is actually for.

The five-pass review: score without explanations, redo every miss cold, then label the cause before you let the answer key tell you anything.

Four fields per row is enough, and keeping it to four is what stops the log turning into a second homework assignment: the question reference, your answer, the correct answer, and the cause. A fifth optional field, one sentence in your own words describing the correction, is worth adding when the miss was a real gap. Write that sentence without looking at the explanation. If you cannot, the gap is still open.

Here are the six causes, and what each one is actually telling you to do.

CauseHow you recognise itWhat it tells you to do
Careless slipYou produce the correct answer cold in under a minute and can name the slipBuild a checking routine, not more content study
Knowledge gapYou still cannot do it after the explanation, or had never met the conceptGo back to the source material and relearn it properly
Misread questionYou solved a real problem correctly, just not the one that was askedDrill question parsing before you drill more content
Timing pressureYou would have had it with two more minutes, and you know itFix pacing, and find out where the earlier minutes went
Wrong approachYou knew the content but chose a method that could not get thereCollect the method that does work, indexed by question type
Lucky guessMarked correct, but you were not confident and you know itTreat it as a miss and redo it cold

Keep the list at roughly this size. Three categories is too coarse to act on and twelve turns every miss into a classification problem you will quietly stop doing by test three. If your exam has a distinctive failure mode, swap a category rather than adding one. A student sitting a heavily quantitative section might replace "wrong approach" with "arithmetic error", because those two need genuinely different fixes.

How do you tell a careless mistake from a real knowledge gap?

Redo the question cold, before you read anything, and time yourself. If you produce the correct answer in under a minute with no hints and can say in one sentence what you did wrong the first time, it was careless. If you cannot, it was a gap, whatever it felt like.

This test matters because "careless" is the most over-used label in the whole system, and the reason is comfortable rather than sinister. Careless means you already knew it, which means nothing needs to change. Gap means you have work to do. Everyone drifts toward the first explanation, and the drift is invisible from the inside. So make the label earn itself with the two-part rule above: you can produce the answer cold, and you can name the specific slip. "I rushed" is not naming the slip. "I solved for x when the question asked for 2x" is.

A second signal is free if you collect it during the test. Put a small dot next to any question you are not certain about, which costs about a second per question. In review, a miss you had flagged is almost always a gap: you felt the shakiness in real time. A miss you had not flagged, where you were confident and still wrong, is the more interesting case, because it means your sense of what you know is off in that area, which is a worse problem than not knowing something.

Those confident errors are also, encouragingly, the ones that fix best. Metcalfe's review reports that the benefit of corrective feedback is particularly salient when people strongly believe their error is correct, so that errors committed with high confidence are corrected more readily than low-confidence errors. The miss that stings the most, the one where you were sure, is the one most likely to stay fixed once you have looked at it properly.

There is one more rule that keeps the system honest over time. If the same "careless slip" appears three times across your log, relabel every instance as a process gap. A slip that recurs on schedule is not an accident, it is a missing habit, and habits get fixed by a procedure rather than by resolving to concentrate harder.

How often should you redo the questions you missed?

Categorise on the same day, then redo the misses cold about a week later, and retire a question only after two correct cold attempts on separate days. Redoing a question ten minutes after reading its explanation tests your short-term memory of the explanation and nothing else.

The gap between answering and finding out is doing real work there. Andrew Butler, Jeffrey Karpicke and Henry Roediger had participants take a multiple-choice test and receive feedback either immediately or after a delay, then measured retention with a later cued-recall test. Delayed feedback produced better final performance than immediate feedback, which they attributed to the spaced presentation of the information. The practical reading is small and specific: do not check answers as you go. Sit the whole section, then review it. The habit of peeking after each question feels efficient and quietly costs you retention.

A schedule that works for most people, and survives contact with a real week:

Same day, while it is fresh. Redo cold, read the key, label the cause, write the one-sentence correction. This is the expensive pass and it is the one that cannot be skipped.

Three to seven days later, cold. Reattempt the missed questions with no notes. Expect to get a real fraction of them wrong again. That is not failure, it is the first honest measurement, because the same-day version was contaminated by the explanation you had just read.

Two weeks out, only the survivors. Anything still wrong is a genuine gap that content study has not closed, and it deserves a different treatment: go back to the source, not to the question.

Retire a question after two consecutive correct cold attempts on different days. One correct answer is too easy to fake with recall of the answer choice rather than command of the material. This is the same logic that drives a proper spaced repetition schedule, applied to your own wrong answers instead of to a deck of cards, and it works for the same reason: spacing the attempts is what turns a recognised answer into a retrievable one.

Log the questions you got right but were not sure about

A lucky guess is a miss with a delay on it. On the real exam the same coin lands the other way about as often as not, so a log that only records wrong answers is systematically optimistic about where you stand.

This is also where the confidence dots pay for themselves twice. The same team behind the feedback-timing work found that feedback does more than repair wrong answers: feedback doubled the retention of correct low-confidence responses compared with no feedback at all. A shaky right answer that you check and confirm becomes a far more durable right answer. Skip it because the mark was already green and you leave that entirely on the table.

In practice this adds two or three items per section rather than doubling your review. Questions you answered quickly and confidently and got right need nothing at all, and most of a test falls into that bucket.

How the log changes what you study next

Once you have three tests logged, sort by cause and count. The largest pile is your study plan, and it very often is not the pile you expected, because the causes point at completely different kinds of work.

A pile of knowledge gaps means content study. This is the only cause that sends you back to the material, and it is the one students assume is behind every miss. Go to the source, relearn the concept, then build retrieval practice on it rather than rereading.

A pile of misreads means a parsing routine, not more content. Nothing about knowing more chemistry will stop you from answering the wrong question. What stops it is a mechanical habit: underline the actual ask, circle any qualifier such as "not", "least", or "except", and restate the question in four words before you start solving.

A pile of careless slips means a checking routine. Find the two or three slips you actually make, which are usually a small and embarrassingly stable set, and write a check for each. Sign errors, units, solving for the wrong variable, and answering with the intermediate value cover most of them.

A pile of timing misses means pacing work, and an audit. Running out of time at the end is a symptom, and the cause is nearly always earlier: two or three questions where you spent four minutes refusing to move on. Log where the minutes went, not just that you ran out.

A pile of wrong approaches means a method index. Keep a short page of question type mapped to the approach that works, built from your own log rather than copied from a book, and review it before your next test.

This is where the plateau usually lives. A student whose score has not moved in six weeks is typically doing content review every week while their log, if they had one, would be sixty per cent misreads and timing. More chemistry cannot fix a misread. For multiple-choice sections in particular, the fix is a set of habits at the question level, which is what our guide to multiple choice test taking strategies covers in detail.

Rank what you fix by frequency multiplied by cost, not by how annoying the miss felt. A misread that appears seven times and costs one mark each is worth more attention than a genuinely hard gap that appeared once.

Where this fits for your specific exam

The method is exam-agnostic, but what counts as a scarce resource is not. Where official practice material is strictly limited, the review discipline matters far more, because every badly reviewed test is a test you can never get back.

On the digital SAT, the Bluebook practice tests are limited and adaptive, so a wasted one is genuinely costly, and the error log is what turns a baseline test into a question-type priority list. Our guide to studying for the SAT covers how that baseline fits into a months-long plan. On AP exams, the log has to run across two very different formats, since a multiple-choice miss and an FRQ miss fail in different ways, and an FRQ cause is usually "did not answer the verb in the prompt" rather than a content gap at all. Our AP Biology study plan works through that split for a science subject. And on the MCAT, where full-length AAMC exams are the scarcest resource in the whole preparation, review time regularly exceeds test time by a wide margin, which is exactly the ratio our MCAT study guide assumes.

The same applies to every other timed exam with an official question bank behind it: the ACT, GRE, GMAT, LSAT, NCLEX, TEAS, ASVAB, GED, USMLE Step 1, and the rest of the AP subjects all have their own study guides on this site, and every one of them gets better with a log running underneath it. The categories do not change. Only the balance between them does.

Six ways an error log stops working

It is written in prose. Reflection paragraphs cannot be sorted or counted, so the log never produces a finding. Fixed categories in a column, always.

It records the topic instead of the cause. "Missed a percentages question" is where, not why. Where is the part you already knew.

Nobody ever reads it back. A log that is only ever written to is a journalling habit. Set a standing appointment to sort and count it after every second test, or the whole exercise is theatre.

The correction is copied from the explanation. Transcribing the official solution produces a neat entry you understand no better. Close the explanation and write the correction from memory, in your own words. What you cannot write is what you did not learn.

Everything gets labelled careless. The single most common failure, and the two-part rule above is the fix. If more than about a third of your misses are careless, you are almost certainly mislabelling.

It becomes the work instead of measuring it. Colour coding, elaborate spreadsheets, and a fifteen-field template all feel productive and none of them are studying. Four fields, six categories, one sentence. If the log takes longer than the fixing, cut the log.

How GeniusPal helps

The boundary first, because it is easy to overstate. GeniusPal does not grade your practice test, cannot read your answer sheet, and has no chat window to ask why an answer was wrong. Diagnosing your own miss is the entire skill this post is about, and handing it to software would leave you with a tidy label and none of the judgement that produced it.

What it does is the step after the diagnosis. Once your log has told you which topics are real gaps, upload the source material for those topics, a PDF, a set of lecture slides, or a plain text or Markdown file, and GeniusPal turns it into a question set you can work through three ways: as a multiple-choice quiz, as flashcards, or in active recall mode where you write the answer from memory before comparing it with the model answer. That is the retrieval practice a knowledge gap needs, aimed only at the topics your log actually flagged.

Two things it does line up unusually well with the method. When you go back through a finished quiz in answer review, each question carries a private notes field, so you can write the cause and your one-sentence correction on the question itself rather than in a separate document, which means the reason travels with the item instead of getting lost. And the app keeps a persistent weak pile per study set: a wrong answer sends a question back to the start, two correct answers in a row retire it, and a "review weak questions" option redrills only the ones still outstanding, in active recall. That is the two-cold-attempts retirement rule from the section above, running automatically across sessions.

The free plan covers 2 study sets for the lifetime of the account, which is enough to run one set of logged gaps through it before deciding whether it earns a place in your routine.

None of this costs anything to start. Take your last practice test, the one you scored and closed, and spend twenty minutes putting its misses into six categories. The count at the bottom will tell you something about the next month that the score at the top never could.

Frequently asked questions

What is the error log method?
The error log method is a running record of every question you get wrong, sorted by the reason you got it wrong rather than by the topic it covered. One row per miss, with four fields: the question, what you answered, what the correct answer was, and the cause. The cause field is the one that does the work, and it comes from a fixed short list: careless slip, knowledge gap, misread question, timing pressure, or wrong approach. Fixed categories matter because they make the log countable. After three practice tests you can sort by cause and discover that eleven of your thirty misses were misreads and only four were genuine gaps, which points to a completely different study plan from the one you would have written by feel. A log kept in free-form prose cannot be counted, so it cannot tell you anything.
How long should it take to review a practice test?
Plan for at least as long as the test itself took, and often longer. A three hour practice test deserves a three hour review, which sounds absurd until you notice that sitting the test teaches you almost nothing and reviewing it teaches you everything. A workable split is roughly five minutes per missed question: one to two minutes redoing it cold before you read any explanation, one minute reading the explanation, and one to two minutes writing the cause and a one sentence correction in your own words. Twenty misses at five minutes each comes to about a hundred minutes. If that feels like too much, take fewer practice tests rather than reviewing them faster. Two tests reviewed properly beat six tests scored and forgotten, because the second approach leaves you with a graph of your scores and no idea why any of them happened.
Should you retake the same practice test after reviewing it?
Retake the questions you missed, not the whole test. A full retake of a test you have already sat mostly measures how well you remember the answers, so the score comes back inflated and tells you nothing about whether the underlying weakness is fixed. Pull your missed questions into a separate set instead and redo them cold about a week later, with no notes and no explanation in front of you. Retire a question only once you get it right twice in a row on separate days, because a single correct answer can easily be recall of the answer rather than command of the material. Save full-length tests for fresh material and for pacing practice, where the point is stamina and timing rather than the individual items. Official practice tests are a limited resource, so spending one on a retake you cannot learn much from is expensive.
Do you need to log questions you got right?
Log the ones you were not sure about. A question you guessed and happened to get right is a miss that has not caught up with you yet, and on the real exam that same guess lands the other way roughly as often as not. The habit that makes this work is marking your confidence while you take the test: a small dot beside any question you are not certain of, which costs about a second. Then in review, treat every low-confidence correct answer as though it were wrong, redo it cold, and log it with a cause. Questions you answered quickly and confidently and got right need nothing at all, so this does not double your workload. It usually adds two or three items per section, and those items are the ones most likely to turn into real misses under exam pressure.
Try our study app free