You paste a 30-page PDF, hit send, and get "conversation too long" — or worse, a confident answer that clearly only read the first few pages. The model isn't necessarily out of room. Your PDF is spending its budget on junk. Here's where the tokens actually go, and how to reclaim them in about a minute.
Models count tokens, not pages
Large language models read text in tokens — word fragments of roughly four characters. Everything the model can pay attention to at once — your instructions, the document, the conversation history so far, and the answer it's writing — has to fit inside one context window, measured in tokens. A "10-page report" means nothing to the model; what matters is that the extracted text of those pages might run anywhere from three thousand to fifteen thousand tokens before it says anything useful.
The window is generous on paper and tight in practice: history accumulates as the conversation grows, long answers consume budget, and many consumer plans impose their own per-message or per-hour limits well below the model's theoretical maximum.
Where the tokens actually go
Extract the text of a typical business PDF and you'll find that a large share of it was never written for a reader. Depending on the document, 20–50% of the extracted text is overhead:
- Repeated headers and footers. A letterhead, a classification banner, or
Page 12 of 40stamped on every page becomes forty copies in the extracted text — each one costs tokens on every single page. - Boilerplate and disclaimers. The confidentiality footer that appears under every page of a legal document, standard preamble blocks, revision tables, "This document is intended solely for…".
- Whitespace and layout noise. Soft hyphens, runs of double spaces, and hard line breaks inserted to match the printed column width — line breaks and spaces are tokens too.
- Broken line wraps. PDF text extraction slices sentences at whatever width the page happened to be, leaving a paragraph shaped like confetti.
- Multi-column scrambles. Extract a two-column layout linearly and the two columns interleave: one sentence from the left column, half a sentence from the right, and so on down the page.
None of this carries signal. All of it consumes the same context window the answer needs.
It's not just size — it's fidelity
The waste would be bad enough if it only cost tokens. But damaged text also misreads:
- Hyphenated wraps break words.
coordi-at the end of a line andnationat the start of the next becomes "coordi nation" — and the model either stumbles or quietly invents the join. - Attention gets diluted. A model has finite attention to spread across the window. Key facts buried in page furniture are the classic "the answer was in the document" failure — it was in the document, under forty repetitions of a footer.
- The model can't tell signal from structure. Without cleanup, a repeated confidentiality banner looks like content. Models will occasionally quote your footer back to you as if it were a finding.
A real before-and-after, from a quarterly report extraction:
Q4 2026 — CONFIDENTIAL, INTERNAL USE ONLY Page 7 of 32
Q4 2026 — CONFIDENTIAL, INTERNAL USE ONLY
Revenue increased 4.2% quarter-over-quarter, driven
primarily by enterprise renewals. The coordi-
nation team's onboarding plan is attached in Appendix C.
Q4 2026 — CONFIDENTIAL, INTERNAL USE ONLY Page 8 of 32
After cleanup:
Revenue increased 4.2% quarter-over-quarter, driven primarily by
enterprise renewals. The coordination team's onboarding plan is
attached in Appendix C.
Same information — a fraction of the tokens, no broken words, and nothing for the model to trip over.
Fix it in a minute, on-device
LoyalCloak automates exactly this cleanup, and it runs entirely on your device — the document is never uploaded. The flow in the free web app (same engine in the desktop apps):
- Drop the PDF in. DOCX, spreadsheets, EPUB, HTML, and scanned documents (via on-device OCR) work the same way. Parsing happens locally in your browser.
- Header/footer stripping and whitespace repair run automatically. Repeated banners are detected and removed, wrapped lines are rejoined, soft hyphens are fixed, and runs of whitespace collapse.
- Watch the token counter. LoyalCloak counts exact BPE tokens using the same tokenizers your target model uses — cl100k_base and o200k_base for ChatGPT, plus a Claude estimator — so you see before/after counts, not a guess.
- Still too long? Auto-chunk it. LoyalCloak splits the document on section boundaries into segments labeled
<part_1_of_3>,<part_2_of_3>… that you paste in sequence, and can wrap everything in a structured<context>tag so the model knows what it's holding. - Paste into ChatGPT, Claude, DeepSeek, or NotebookLM.
Why not just upload the PDF to the chat?
Chat file uploads get preprocessed for you — and that's the problem. You don't control the extraction: the same headers, footers, and wrap damage come along, and long files are often truncated or summarized rather than read in full, without much indication of what was dropped. Cleaning the text first means you decide exactly what the model sees, you can read it yourself before sending, and you know the real token count up front.
Manual alternatives, ranked
- Copy-paste from the PDF viewer. Fastest for a paragraph, painful for pages — and it preserves every wrap artifact you were trying to escape.
- Export to Word or Markdown. Fixes some structure, but keeps the headers, footers, and boilerplate that are eating your window.
- "Free PDF to text" converter websites. They upload your document to someone else's server — you just traded a token problem for a privacy problem.
- On-device cleanup. The whole pipeline — parse, strip, repair, count, chunk — locally, in about a minute.
Frequently asked questions
How many tokens is a PDF page?
A clean business-text page is typically a few hundred to over a thousand tokens depending on density, tables, and footers — "per page" is exactly the wrong unit to reason with. LoyalCloak shows the real count using the actual tokenizer of your target model.
What does "prompt too long" actually mean?
Your instructions + document + conversation history + expected answer exceeded the model's context window. The durable fix is sending less text: strip the overhead, start a fresh chat, or split the document into chunks.
Does stripping headers and footers lose information?
They're page furniture — repeated on every page and carrying no payload. The actual content is untouched, and you can review the sanitized output at a glance before sending it.
How do I fit a 100-page PDF into ChatGPT?
Chunk it. Split on section boundaries (not arbitrary character counts), label the segments <part_1_of_N> so the model knows the order, and paste them in sequence — then ask your question against the relevant part. LoyalCloak automates the split.
Reclaim your context window
Drop a PDF in the browser and watch the token count fall — free, no sign-up, nothing uploaded.