Ask AI to summarize a PDF: what survives and what gets dropped

A 300-page annual report goes in, four hundred fluent words come out, and every sentence in them is true. The problem is the sentence that is not there: the one about the covenant that becomes binding in the second quarter. Nobody reading the summary would know to ask. This is the specific reason experienced users still search for how to ai summarize pdf documents, long after they have stopped being impressed by the output.

Summarising is not compression of a document into a shorter document. It is a judgment about what matters, made by something that does not know why the document was opened. What follows is what that judgment reliably keeps, what it reliably discards, and how to ask in a way that changes the outcome.

What a summary optimises for

A summary is produced to read well and to cover the document's main thrust. Both of those goals work against the reader who has a specific reason for being there.

Reading well means the output is coherent prose, which favours the general over the particular. A sentence stating that the agreement includes standard termination provisions is smoother than one listing four notice periods with their conditions, and the smoother sentence is the one that gets written.

Covering the main thrust means frequency wins. A theme repeated across forty pages is central by any reasonable measure, so it survives. A single clause appearing once on page 212 is, statistically, a detail. It is also, quite often, the only part of the document with consequences.

The material a summary drops is exactly the material that is unusual, and unusual is what people read contracts and reports to find. That is not a flaw in any particular product. It is what the task rewards, and it is why a summary is a triage tool rather than a substitute for reading the part that matters.

What gets dropped, predictably

Six categories go missing often enough to check for by name every time.

Numbers with their units and their basis. A summary will report that fees increased without the figure, or report the figure without saying whether it is annual or monthly, gross or net, per seat or per organisation. The number alone is not the fact; the number with its basis is.

Negations and conditions. The difference between "the licence may be transferred" and "the licence may be transferred only with prior written consent" is one clause, and it is the clause most likely to be smoothed away. Conditional structures suffer the same way, because they are long and a summary is short.

Dates and their effect. Documents contain a signature date, an effective date, a review date and a termination date, and a summary frequently reports one of them as though it were the others.

Attribution. Who is obliged to do a thing is dropped more often than the thing itself, which turns an obligation on one party into an abstract requirement floating in the document.

Exceptions and carve-outs. These are almost always short, almost always late in the document, and almost always the operative text.

Anything that lives in a chart. Whether a chart is read at all depends on the route and the length. In the Claude apps, PDFs of 100 pages or fewer are analysed for both text and visual elements, while documents of 101 to 1,000 pages are processed as text only and their visual elements are not analysed. A 350-page report summarised through that path has had its graphs ignored, silently.

Length changes the answer, not just the speed

The limits matter here in a different way than they do for a single question about a document.

A short document is summarised in one pass, with the whole text available at once, so a clause on page 8 can be related to a clause on page 60. That cross-reference is the most valuable thing a summary can contain.

A long document is summarised in pieces, then the pieces are summarised again. The published guidance is explicit that splitting is the remedy for a document too dense or too large to process in one request, alongside the numbers: a maximum request size of 32 MB and 600 pages per request through the API, falling to 100 pages when the context window is under one million tokens, with the warning that dense pages can exhaust the context before the page limit is reached.

The consequence of that second pass is worth stating plainly. In a summary of summaries, nothing connects section three to section nine, because the pass that wrote the final text never saw either in full. Contradictions between distant parts of a document, which is one of the things a careful reader is looking for, cannot be found this way. When that is the goal, the sections have to be compared deliberately, two at a time, with both in front of the model.

Ask for a structure instead of a summary

The single change that improves output most is to stop asking for a summary and to ask for a specific shape. A structure forces the model to look for particular things and makes an omission visible as an empty row.

Request What it is for Why it works
One sentence stating what this document is and who it binds Triage of a stack Too short to hide vagueness
A table of every figure with its unit and page number Financial and pricing review Numbers cannot be paraphrased away
Obligations listed by party Contracts Forces attribution into the output
Every date with what it triggers Renewals and deadlines Separates the four kinds of date
Conditions and exceptions, quoted verbatim Risk review Quotation blocks smoothing
What changed against the previous version Redrafts Diffing beats summarising

Two habits sharpen all of these. Ask for page numbers on every row, which makes verification a matter of seconds instead of a re-read. And ask for the model to state when something was not found, rather than leaving a row out, because an absent row and a missing fact look identical otherwise.

The last row deserves emphasis. When a previous version of the document exists, comparing the two is a better use of the tool than summarising either, and it is a task where the output is checkable line by line.

Two kinds of summary, and which one to ask for

There is a useful distinction hiding behind the single word summary, and naming it changes what gets asked for.

One kind selects. It pulls sentences out of the document and presents them in order, unchanged. The output reads less smoothly and is far more trustworthy, because every line exists in the source and can be found there. For a contract, a policy or anything that will be relied on, this is the safer form, and it is requested by asking for quoted passages rather than a description of them.

The other kind rewrites. It produces new sentences that describe the document, which is what most tools do by default and what most people mean by the word. It is better for deciding whether a document is worth reading at all, and worse for anything that depends on precise wording, because the precision is what the rewriting removes.

The sensible pattern uses both in sequence. Start with a rewritten summary of a few sentences to decide whether this document is relevant, which is a cheap question to answer wrongly. Then, for the documents that survive that filter, ask for selected and quoted passages on the specific points in question. Using the second kind for triage and the first kind for the actual work avoids the common failure, which is making a decision from output that was optimised for readability.

A third form is worth knowing about for recurring documents. When a set of files shares a structure, such as monthly statements or intake forms, the request should be a fixed set of fields rather than any kind of summary at all. The output is then comparable across files, and a missing value is visible as a blank rather than as a sentence that quietly says nothing.

Verifying a summary in two minutes

A summary that has not been checked is a rumour about a document. The check does not have to be long.

Pick three numbers from the summary and find them in the source. Preview's search is enough. A number that cannot be located, or that appears with a different unit, is the early warning that the extraction was partial.

Search the document for the negation words the summary did not use. The words "not", "except", "unless", "only" and "subject to" each take a few seconds, and they land on the clauses most likely to have been smoothed away.

Read the first two and last two pages yourself. Purpose and parties are at the front; termination, governing law and the awkward schedules are at the back. Both ends are short, and both are where summaries are weakest.

Then spot-check one claim per section. Not every claim, one. A summary that survives four spot-checks is usually sound; a summary that fails one was produced from an incomplete reading, and the whole thing needs to be treated as unverified.

Summarising a folder rather than a document

One report is a task. Forty monthly statements, or a due diligence folder, is a different problem, and it is where the browser upload route ends. Chat uploads take up to 20 files in a single conversation, which is a hard stop rather than a suggestion, and re-uploading by hand for the next twenty is the point at which the process quietly gets abandoned.

Three things make a batch work. A consistent output shape, decided once, so that forty summaries can be compared instead of read individually. Output written next to the source, one file per document with the same base name, so the summary and the original never drift apart. And a record of which files have been done, because a batch always gets interrupted.

That arrangement wants a folder, a terminal in it, and an agent that can read the files and write its results back, which is a different tool from a chat window. Keeping the three in one window is what makes the loop survive an interruption, and the comparison with other file managers sets out where a dedicated PDF tool still does the job better. A long batch left running can be checked from an iPhone or iPad rather than requiring a seat at the desk.

What to change first

Replace the next request for a summary with a request for a table: every figure, every date, or every obligation, each with a page number. Then spend two minutes checking three rows against the source. If this is monthly work across a folder rather than a one-off report, Atriens is the part that removes the uploading and keeps each summary beside the document it came from.

Frequently asked questions

Why does a summary miss the one clause that mattered?

Because a rare clause is, by the measure a summary uses, a detail. Summaries favour what recurs and what reads smoothly, and a single conditional sentence late in a document is neither. Asking for conditions and exceptions quoted verbatim, as a list, is what surfaces them.

How long can a document be before the summary degrades?

Degradation starts well before any hard limit, at the point where the document no longer fits in one pass and has to be summarised in sections. In the Claude apps, charts and other visual elements stop being analysed above 100 pages, and through the API a request is capped at 32 MB and 600 pages, or 100 pages when the context window is under one million tokens. Dense pages can exhaust the context earlier than the page count suggests.

Is a summary of a scanned PDF reliable?

Only if character recognition has been applied. A scan is a set of images, so a text-based route returns nothing and a vision-based route reads the pages as pictures, which works better for prose than for tables. Selecting text in Preview is the quick test for which situation the file is in.

Does asking for page numbers make the summary less accurate?

No, and it usually makes it more useful, because a claim with a citation can be checked in seconds. It also changes the failure mode: a model that cannot locate a passage is more likely to say so than to produce a fluent paraphrase of something it did not read.

Is it better to summarise section by section?

For any long document where the sections have different purposes, yes. Section summaries keep more detail, they can be checked one at a time, and the parts that need no attention can be skipped. The one thing they lose is the connection between distant sections, which has to be asked for separately with both parts in view.

Back to all posts