How to Take Notes on a PDF (And Get Them Back Out)
Take notes on a PDF in two layers: highlights inside the file as an index, and the notes themselves in a separate document keyed by page number.
Take notes on a PDF in two layers: highlights inside the file, which act as an index, and the notes themselves in a separate document where every line starts with a page number. Before either layer, check that the PDF has a text layer, because a scan blocks selecting, searching, and exporting.
A PDF is not a book on a screen and it is not a document you can edit. It is a fixed picture of a set of pages, sometimes with machine-readable text sitting behind that picture and sometimes not. Almost every frustration in this topic comes from that one fact, and so does the single real advantage a PDF has over paper.
Check whether your PDF has a text layer first
There are two kinds of PDF and they behave nothing alike. A born-digital PDF was exported from a word processor, a slide deck, or a typesetting system, so the characters are still in the file. A scanned PDF is a photograph of a page. The two can look identical on screen and one of them is completely inert.
The test takes five seconds. Try to drag-select a sentence. If the text highlights the way it does on a web page, you have a text layer and everything below works. If nothing happens, or the whole page selects as a single block, you are looking at an image.
| Aspect | PDF with a text layer | Scanned PDF |
|---|---|---|
| Select a sentence | Drag across it, as on any web page. | Nothing selects. The page is one image. |
| Highlight | The colour attaches to the words themselves. | You draw a box over the area. It is a sticker, not a highlight. |
| Search the document | Find every occurrence of a term in one keystroke. | Search returns nothing, however legible the page looks. |
| Copy into your notes | Copy and paste, then repair the line breaks. | Retype it by hand, or run OCR first. |
| Export your highlights | The reader can list them all as text. | There is no text to list. |
Optical character recognition is the fix, and every major reader and most reference managers now ship it. It runs over the image, guesses which shapes are which characters, and writes an invisible text layer underneath. After that, selecting and searching usually work.
Treat that layer as a guess rather than a transcript, because OCR fails hardest on exactly what students scan. Handwriting, mathematical notation, chemical structures, two-column journal pages, and tables all degrade, and tables degrade in the worst way: the characters come out roughly right while the reading order comes out scrambled, so a row of results can end up welded to the wrong heading. That limitation follows the file wherever it goes, which is why whether ChatGPT can read a PDF depends far more on which of the two kinds you uploaded than on which model is running. Spot-check a page of OCR output against the original before you trust it.
How do you take notes on a PDF?
In two layers, because the file is good at holding one of them and bad at holding the other.
The first layer is marks. Highlights, plus a comment pin at the few places that need one. Its job is to make a sixty-page reading navigable, so keep it sparse enough that a colour still means something. A file with two thirds of its text shaded scrolls past at exactly the speed of an unmarked one.
The second layer is the notes, meaning the sentences where you work out what the reading is doing. Those do not belong in the file, and this is where PDF note-taking most often goes wrong.
A PDF has no margin. That sounds cosmetic and is not. A margin note on paper is visible while you flip past it, which is the whole reason a marked-up book is scannable. A PDF comment is folded into an icon and stays hidden until you click it, so a page carrying eight comments shows you eight identical little squares and tells you nothing. The comment layer is bad at precisely the job margins exist to do.
So put the words in a separate document, and start every line with the page number followed by the note. That number is the join between the two layers, and a line lacking one cannot be traced back to the page it came from. The document, rather than the marked-up file, is the thing you are really producing here. Where it should sit alongside everything else is a filing question rather than a PDF question, and organising your notes for studying only needs answering once per course.
What to write in that document is a separate skill, and it is the same one paper asks for, so the rules in annotating a textbook for studying transfer intact: name what a passage is doing, record the connection to something else, pin down the precise sentence that lost you. What changes on a PDF is only where those words live, never which words are worth writing.
Highlighting a PDF only pays if you get the highlights back out
On paper, a highlight is stuck to the page for good. In a PDF it is a record in a list, and that difference is the entire case for highlighting a digital reading.
Most readers can show you every annotation in a document in one panel and export the lot: look for an annotations or comments sidebar, then an option to summarise, export, or copy. Reference managers go further. Zotero will build a single note containing every annotation in a PDF, usually with a page reference attached, which is most of a notes document assembled for nothing.
Run that export once, early, because annotations are not always stored where you assume. Some readers write them into the PDF file itself, so they travel with the file. Others, Zotero among them, keep them in their own database, which means the copy you send a classmate arrives blank. Open one of your marked files in a different reader and see what survives. Whatever comes back is what you actually own.
None of this rescues bad highlighting. The evidence on whether highlighting is effective is not kind to it as a technique on its own, and a coloured passage still marks a spot without recording why the spot mattered. The export step is what changes the arithmetic. A highlight pulled into a document, with one line beside it saying why you pulled it, has been converted into a note. A highlight left in the file is a coloured rectangle in a file you will not reopen.
Where a PDF is genuinely worse than paper
Four constraints are worth knowing before you commit a course to this format.
- The layout is fixed. A PDF does not reflow, so a two-column journal article on a phone is a choice between four words to a line and panning sideways. An ebook adapts to the screen. A PDF cannot.
- Page anchoring is brittle. Annotations attach to a position in one specific file. Re-export the deck, move to the second edition, or download the same paper from a different repository, and the marks do not follow.
- The comment layer is invisible until clicked, which is the margin problem above and the reason the words belong somewhere else.
- A scan blocks all of it until OCR has run, and OCR is a guess.
Against that, the file gives you four things paper never will: search across the whole reading, copying that introduces no typos, an annotation list you can export in one action, and unlimited room for notes that are not squeezed into two centimetres of white space. All four are conveniences of access rather than gains in understanding, which is the honest way to hold them.
Should you print it instead?
Sometimes, and less often than the argument around this suggests.
In a 2018 meta-analysis in Educational Research Review, Pablo Delgado, Cristina Vargas, Rakefet Ackerman and Ladislao Salmeron pooled research published between 2000 and 2017: 38 between-participants and 16 within-participants studies, covering 171,055 participants. Both designs gave the same result, a small advantage for paper on reading comprehension, Hedges g of -0.21. Three moderators reached significance. The paper advantage grew when reading was timed rather than self-paced, held for informational texts and mixed sets but not for narrative-only ones, and increased across the publication years in the sample.
Read the size of that number before the direction of it. About a fifth of a standard deviation is real and small, and it has to be weighed against losing search, copying, export, and every note you would have made in the second layer. The narrow case for printing is a dense reading you will do once, under time pressure, and never come back to. Anything you will return to is worth keeping in a file.
Turning two layers into questions you have to answer
Both layers are still preparation. A file full of highlights and a document full of page-numbered notes give you an index and a summary, and neither one has made you answer anything from memory. That is the step that moves material into memory, and the step that gets dropped, usually because drafting questions from your own material is slow and unrewarding at the exact moment it needs doing.
GeniusPal takes a file and generates a study set from it: a quiz, flashcards, and a written recall drill built from that document. It accepts PDF, Word, PowerPoint, plain text, Markdown, and CSV files up to 10 MB. It writes no summaries and draws no mind maps, which matters here, because by this point the notes document is already the summary.
The useful move is to upload that notes document rather than the original PDF. The questions then come from the passages you judged worth keeping, instead of from all sixty pages weighted evenly by a model that has never seen your syllabus. Uploading the raw file is a fair shortcut when you have not read it yet, and making flashcards from a PDF directly covers that route. A scanned PDF is a poor upload for the same reason it is a poor file to highlight: there is nothing in it to read.
The Free plan covers two generations for the lifetime of the account rather than two a month, enough to run it once over a reading you have already marked up. Student is 14.99 dollars a month for 100 generations, and Genius is 59.99 dollars a year with a fair-use ceiling shown as unlimited.
If you keep one thing from this page, keep the test at the top. Drag-select a sentence before you do anything else, because everything downstream, the highlighting, the searching, the exporting, and the uploading, depends on whether anything selects at all.
Frequently asked questions
How do you take notes on a PDF?
Take notes on a PDF in two layers, because the file is good at holding one of them and bad at holding the other. The first layer lives inside the file: highlights, plus a comment pin at the few places that need one. Its job is to make a long reading navigable, so keep it thin enough that the marks still discriminate. The second layer is the notes themselves, and those belong in a separate document rather than in the comment pins, because a PDF has no margin. A margin note on paper is visible while you flip past it. A PDF comment is folded into an icon and stays hidden until you click it, so it is invisible at exactly the moment you are scanning for something. Start every line in that document with the page number, then the note. The page number is the join between the two layers, and a note without one is unfindable.
Can you highlight a scanned PDF?
Not the way you can highlight a typed one, unless you run OCR on it first. A scanned PDF is a photograph of a page, so there are no characters in the file for a highlighter to attach to. Most readers will still let you draw a coloured rectangle over the area, which looks like a highlight and behaves like a sticker: it cannot be searched, copied, or exported as text. Optical character recognition adds a text layer by guessing which shapes are which characters, and once it has run, selecting and searching usually work. Treat the result as a guess rather than a transcript. OCR degrades on exactly what students scan most: handwriting, mathematical notation, two-column journal pages, and tables, where the characters can come out roughly right while the reading order comes out scrambled. Spot-check a page against the original before you build on it.
How do you get your highlights out of a PDF?
Most PDF readers can list every annotation in a document and export that list, though the menu wording varies: look for an annotations or comments panel, then an option to summarise, export, or copy. Reference managers go further. Zotero will build a single note containing every annotation in a PDF, usually with a page reference attached, which is most of a notes document assembled for nothing. Run that export early, before you commit a term of reading to one app, because annotations are not always stored where you assume. Some readers write them into the PDF file itself, so they travel with the file. Others, Zotero among them, keep them in their own database, which means the copy you send a classmate arrives blank. Open one of your marked files in a different reader and see what survives. Whatever comes back is what you actually own.
Is it better to print a PDF or read it on screen?
On average paper reads slightly better, but the gap is small and printing costs you everything a PDF is good at. A 2018 meta-analysis in Educational Research Review by Pablo Delgado, Cristina Vargas, Rakefet Ackerman and Ladislao Salmeron pooled 38 between-participants and 16 within-participants studies covering 171,055 participants, and found a small advantage for paper on reading comprehension, Hedges g of -0.21. Three moderators mattered: the paper advantage grew when reading was timed rather than self-paced, held for informational texts but not for narrative-only ones, and increased across the publication years in the sample. Weigh the size of that number against what printing removes, which is search, copy, export, and any note you would have made in a second document. The narrow case for printing is a dense reading you will do once, under time pressure, and never return to.
Keep reading
- Study Techniques
How to Take Notes From a Textbook (Without Copying It Out)
Copying sentences out of a textbook feels productive but builds almost no memory. Here is a method for taking notes from a textbook that turns reading into real recall.
July 7, 2026 · 8 min read - Study Techniques
How to Take Cornell Notes (and Actually Use Them)
Most students take Cornell notes and never open them again. The layout is only half the method. Here is how to take Cornell notes and turn them into real retrieval practice.
July 6, 2026 · 8 min read - Study Techniques
How to Take Lecture Notes That Actually Help
You cannot listen closely and write down every word at once. Here is how to take lecture notes that capture structure and key ideas, and the after-class review that turns them into memory.
July 23, 2026 · 8 min read - Study Techniques
How to Study From Recorded Lectures Without Zoning Out
A live lecture has no rewind, and that constraint does more work than it ever gets credit for. Remove it and the promise that you can always watch the thing again is usually what turns a recording into an hour you sat through once and never opened. Szpunar, Khan and Schacter found that students watching an online lecture reported mind wandering on 41 percent of probes, and that interrupting the video with short tests cut that to 19 percent. Here is what actually changes when playback controls arrive, what the evidence says about watching at 2x speed, and a workflow that uses pause, timestamp and rewatch deliberately rather than passively.
August 7, 2026 · 9 min read