Study Techniques By Shannon Loy September 20, 2026 10 min read

Levels of Processing: Depth, Memory and How to Study

Levels of processing explained: the 1975 experiment where a quick question about meaning beat a slower surface task, its critics, and how to study for meaning.

Levels of processing is the idea that memory depends on what you did with material while studying. On this view, thinking about what it means leaves a more durable memory than attending to its surface, such as how it is printed or how it sounds, and in a 1975 experiment that held even when the surface task took longer.

For a student, the working rule is to ask a quick question about meaning each time you meet a new idea: what kind of thing it is, what it does, and where it fits. That rule is our extrapolation from laboratory work with single words, and the sections below set out both the evidence and the distance between it and a textbook chapter.

What is levels of processing?

Levels of processing is a framework for studying memory. Fergus Craik and Robert Lockhart proposed it in a 1972 paper titled Levels of processing: A framework for memory research. Our summary of the framework comes from the abstract of a 1975 follow-up by Craik and Endel Tulving, Depth of processing and the retention of words in episodic memory, which set out to explore “the levels of processing framework for human memory research proposed by Craik and Lockhart (1972).”

That abstract states the basic notions in two parts. First, a memory trace “may be thought of as a rather automatic by-product of operations carried out by the cognitive system.” Second, the trace lasts longer as processing goes deeper, “where depth refers to greater degrees of semantic involvement.” Semantic means to do with meaning. So on this view, a memory is what your mind leaves behind after working on something, and work that engages with meaning leaves a longer-lasting trace.

On our reading, the by-product idea is the part a student can use straight away: two study sessions of equal length can leave very different memories, depending on the operations that filled them.

In the 1975 experiments, depth was set by the question each subject answered about a word. The abstract describes three levels: shallow encodings came from “asking questions about typescript”, intermediate levels from “asking questions about rhymes”, and deep levels from “asking whether the word would fit into a given category or sentence frame.”

AspectWhat the question is aboutAn example question for the word BEAR
ShallowHow the word is printed, its typescriptIs the word printed in capital letters?
IntermediateHow the word sounds, whether it rhymesDoes the word rhyme with CHAIR?
DeepWhat the word means, whether it fits a category or a sentenceIs the word a kind of animal?
The three kinds of encoding question described in the abstract of Craik and Tulving (1975). The example questions for the word BEAR are our own illustrations.

What did Craik and Tulving find in 1975?

The paper reports ten experiments. Subjects answered a question about each word in a list, and after that encoding phase, in the words of the abstract, “subjects were unexpectedly given a recall or recognition test for the words.” The surprise test points any difference in memory back to the questions.

The main result, again from the abstract: “In general, deeper encodings took longer to accomplish and were associated with higher levels of performance on the subsequent memory test.” Put plainly, meaning generally beat print and sound.

The abstract adds a detail about answers. Questions “leading to positive responses were associated with higher retention levels than questions leading to negative responses, at least at deeper levels of encoding.” A word that fitted its category or sentence was remembered better than a word that clashed with it. The abstract then adds a qualifier: “Negative responses were remembered as well as positive responses when the questions led to an equally elaborate encoding in the two cases.”

The authors also draw on a principle of congruity from a 1974 paper by Schulman, which the abstract says “appears necessary for a complete description of the effects obtained.” As the abstract puts it, “Memory performance is enhanced to the extent that the context, or encoding question, forms an integrated unit with the word presented.” They give two reasons, in the abstract’s words: “a more elaborate trace is laid down”, and “the structure of semantic memory can be utilized more effectively to facilitate retrieval.” A word that completes a sentence becomes part of one coherent idea, and both reasons point at that coherence.

Does deeper processing work because it takes longer?

In the original results, deeper encodings took longer, so time was an obvious rival explanation, and the authors checked whether encoding time accounted for the pattern. The abstract describes the test: “an experiment was designed in which a complex but shallow task took longer to carry out but yielded lower levels of recognition than an easy, deeper task.”

So a quick question about meaning beat a slower task about the word’s form. The abstract puts the general conclusion in one line: “retention depends critically on the qualitative nature of the encoding operations performed; a minimal semantic analysis is more beneficial than an extensive structural analysis.”

This is the result with the plainest message for students. In these experiments, the kind of work predicted memory better than the time it took. Our guide to the generation effect reaches a compatible conclusion from separate research, where unscrambling letters, an effortful task that never routes through meaning, left people very slightly worse off than reading.

Does it still work when you know a test is coming?

Students usually know a test is coming, so the surprise test raises a fair question. The abstract reports that the authors found “the typical pattern of results under intentional learning conditions”, and also “where each word was exposed for 6 sec in the initial phase.” Knowing about the test left the advantage for meaning in place, and so did a fixed viewing time for every word.

The abstract also records a shift in the authors’ own vocabulary. They suggest that how elaborate an encoding is, and how widely it spreads, may describe their results better than depth does. In their words, “spread and elaboration may indeed be better descriptive terms for the present findings”, and they hold that retention still depends critically on the kind of encoding performed.

Elaboration, in everyday terms, means connecting a new item to more of what you already know, and two methods covered elsewhere on this blog put it to work. Asking why a fact is true is the method in our guide to elaborative interrogation, and explaining each step of a worked example in your own words is self-explanation. Both ask a question about meaning, and both link the answer to prior knowledge.

Where was levels of processing criticised?

The framework drew criticism within six years. In 1978 Alan Baddeley published The trouble with levels: A reexamination of Craik and Lockhart’s framework for memory research in Psychological Review. Its ERIC abstract says the paper discusses problems in applying a levels-of-processing approach, “suggests that some of the basic assumptions are false”, and argues for information-processing models “devised to study working memory and reading, which aim to explore the processes underlying memory in detail.” He preferred that detailed approach to broad general principles. The abstract names no specific assumption, so this post names none either.

Craik returned to the framework himself. In Levels of processing: Past, present... and future?, published in the journal Memory in 2002, he writes that he will “briefly survey some enduring legacies” of the 1972 article and “address some common criticisms”, which the abstract leaves unnamed. A separate, later section of the article discusses topics that include “measurement of ‘depth’ in LOP” (LOP is levels of processing) and “encoding-retrieval interactions.”

The 2002 abstract gives no detail on measuring depth. One difficulty is plain enough on our own reading: ranking two tasks by depth needs some yardstick other than the memory result itself. Craik’s own discussion is in the full article, which we have not read.

Encoding-retrieval interactions are the subject of our post on transfer-appropriate processing, which covers a 1977 experiment where studying for rhyme beat studying for meaning on a rhyming test.

How do you use levels of processing when you study?

Everything in this section is our reading of the evidence above, and it involves a real leap. The 1975 experiments used single words, one question per word, in a laboratory. A textbook chapter is connected prose, and an exam asks for far more than recognising a word. The direction of the advice follows from the findings, and the size of any benefit for a chapter is something none of these abstracts measured.

  1. Ask one meaning question per idea. When you meet a new term or claim, ask what kind of thing it is, what it does, or what would change without it. A minimal semantic analysis beat an extensive structural one in 1975, so a quick question is worth asking even on a rushed evening.
  2. Fit the idea into something larger. The congruity result favoured questions that formed an integrated unit with the word. For a student, that means placing each idea inside a bigger picture: the process it belongs to, the example it explains, the cause it follows from.
  3. Answer in your own words before you check. Say or write the answer to your meaning question first, then compare it with the text and correct anything you got wrong.
  4. Judge a session by the questions it asked. Copying, restyling and colour-coding keep you working on the look of your notes, and on our reading they sit closer to the shallow tasks than their time cost suggests. Our post on whether rewriting notes works covers copying. After a session, count how many ideas you questioned for meaning.
  5. Match the depth to the exam. Meaning questions prepare you for questions about meaning. If your exam also asks you to spell terms, label a diagram or run a calculation, practise those actions as well, in the form the exam uses.

What does this research leave open?

Every quotation above comes from an abstract, and nothing here reports a number, a sample or a method detail beyond what an abstract states. The 1972 paper appears as the proposal that the 1975 paper names, and we attribute no argument to it.

The 1975 experiments used lists of single words, so any application to paragraphs, problem sets or essays is an inference. Baddeley’s abstract says some assumptions are false, Craik’s 2002 abstract says he addresses some common criticisms, and neither abstract says which ones. None of the three abstracts reports an effect size, and none describes students revising for a real exam.

Where GeniusPal fits

GeniusPal turns a document you upload into a set of practice questions. The instructions it gives the model ask for multiple choice questions that test whether an idea is understood, and they favour why, what-happens-if and applied questions over definitions. That aim lines up with the advice above. An instruction is a request, so read each question with one test in mind: does answering it require knowing what the idea means?

A free account covers 2 generations for the life of the account, at 10 questions per set, with 2 full runs of each set, and its quizzes are multiple choice throughout. Student and Genius add written answers inside the quiz, marked for you when you finish the round, plus flashcards, written recall and sets of up to 30 questions. Student is 14.99 dollars a month for 100 generations. Genius is 59.99 dollars a year, shown as 5.00 dollars a month, under a fair-use ceiling. The daily review is free on every plan: up to 20 questions from across your sets, with the ones you missed first, then the ones you have never answered, then the ones still short of mastery.

GeniusPal reads a PDF, a Word (.docx) file, a PowerPoint (.pptx) deck, plain text, Markdown or CSV up to 10 MB, or a pasted link. It has no optical character recognition, so a scan with no text layer is refused. It writes no summaries, draws no mind maps, and offers no chat window or tutoring dialogue.

Frequently asked questions

What is levels of processing?

Levels of processing is a framework for memory research proposed by Fergus Craik and Robert Lockhart in 1972. A 1975 follow-up by Craik and Endel Tulving summarises its basic notions in its abstract: a memory trace is a fairly automatic by-product of the operations the mind carries out on material, and the trace lasts longer when those operations go deeper, where depth means greater involvement with meaning. In the 1975 experiments, shallow questions asked about the typescript of a word, intermediate questions asked about its rhyme, and deep questions asked whether it fitted a category or a sentence. Words given the deeper questions were generally remembered better on the surprise memory test that followed. For a student, the working reading is to ask questions about what material means while studying, with the caveat that those experiments used single words in a laboratory.

What did the Craik and Tulving experiment show?

Craik and Tulving reported ten experiments, published in 1975. In the basic design, subjects answered a question about each word in a list and were then unexpectedly given a recall or recognition test. Questions about typescript produced shallow encoding, questions about rhyme an intermediate level, and questions about whether the word fitted a category or a sentence frame produced deep encoding. According to the abstract, deeper encodings generally took longer to carry out and were followed by better performance on the memory test. To check whether time explained the result, the authors designed a complex shallow task that took longer than an easy deep one, and the deep task produced better recognition. The typical pattern also appeared when subjects knew a test was coming. The abstract states the conclusion directly, saying that a minimal semantic analysis is more beneficial than an extensive structural analysis.

What are the criticisms of levels of processing?

The framework was challenged within six years of its 1972 proposal. In 1978 Alan Baddeley published The trouble with levels in Psychological Review. Its ERIC abstract says the paper discusses problems in applying a levels of processing approach, suggests that some of the basic assumptions are false, and argues for information-processing models of the kind devised to study working memory and reading, which aim to explore the processes underlying memory in detail. He preferred that detailed approach to broad general principles. The abstract names no specific assumption, so none is named here. In 2002 Fergus Craik returned to the framework in the journal Memory. His abstract says he addresses some common criticisms, which it leaves unnamed, and a later section discusses topics including the measurement of depth and encoding-retrieval interactions. That second topic is the ground covered by transfer appropriate processing, which asks whether the best way to study depends on the test.

How can students use levels of processing when studying?

Ask a short question about meaning whenever you meet a new idea: what kind of thing it is, what it does, why it is true, or where it fits in a larger process. In the 1975 experiments by Craik and Tulving, an easy task about meaning produced better recognition than a longer task about the form of a word, and meaning still won when subjects knew a test was coming, so the hours a session takes are a weak guide to what it did for your memory. Questions that tie an idea into a coherent whole are a good target, since the abstract reports better memory when the encoding question formed an integrated unit with the word. All of this is an extrapolation. Those experiments used single words in a laboratory, so pair meaning questions with self-testing on the tasks your exam will set, and judge the results by your own practice scores.

Try our study app free