You paste a 30-page PDF, hit send, and get "conversation too long" — or worse, a confident answer that clearly only read the first few pages. The model isn't necessarily out of room. Your PDF is spending its budget on junk. Here's where the tokens actually go, and how to reclaim them in about a minute.

Models count tokens, not pages

Large language models read text in tokens — word fragments of roughly four characters. Everything the model can pay attention to at once — your instructions, the document, the conversation history so far, and the answer it's writing — has to fit inside one context window, measured in tokens. A "10-page report" means nothing to the model; what matters is that the extracted text of those pages might run anywhere from three thousand to fifteen thousand tokens before it says anything useful.

The window is generous on paper and tight in practice: history accumulates as the conversation grows, long answers consume budget, and many consumer plans impose their own per-message or per-hour limits well below the model's theoretical maximum.

Where the tokens actually go

Extract the text of a typical business PDF and you'll find that a large share of it was never written for a reader. Depending on the document, 20–50% of the extracted text is overhead:

None of this carries signal. All of it consumes the same context window the answer needs.

It's not just size — it's fidelity

The waste would be bad enough if it only cost tokens. But damaged text also misreads:

A real before-and-after, from a quarterly report extraction:

Q4 2026 — CONFIDENTIAL, INTERNAL USE ONLY                Page 7 of 32
Q4 2026 — CONFIDENTIAL, INTERNAL USE ONLY
Revenue increased 4.2% quarter-over-quarter, driven

primarily by enterprise renewals. The coordi-

nation team's onboarding plan is attached in Appendix C.

Q4 2026 — CONFIDENTIAL, INTERNAL USE ONLY                Page 8 of 32

After cleanup:

Revenue increased 4.2% quarter-over-quarter, driven primarily by
enterprise renewals. The coordination team's onboarding plan is
attached in Appendix C.

Same information — a fraction of the tokens, no broken words, and nothing for the model to trip over.

Fix it in a minute, on-device

LoyalCloak automates exactly this cleanup, and it runs entirely on your device — the document is never uploaded. The flow in the free web app (same engine in the desktop apps):

  1. Drop the PDF in. DOCX, spreadsheets, EPUB, HTML, and scanned documents (via on-device OCR) work the same way. Parsing happens locally in your browser.
  2. Header/footer stripping and whitespace repair run automatically. Repeated banners are detected and removed, wrapped lines are rejoined, soft hyphens are fixed, and runs of whitespace collapse.
  3. Watch the token counter. LoyalCloak counts exact BPE tokens using the same tokenizers your target model uses — cl100k_base and o200k_base for ChatGPT, plus a Claude estimator — so you see before/after counts, not a guess.
  4. Still too long? Auto-chunk it. LoyalCloak splits the document on section boundaries into segments labeled <part_1_of_3>, <part_2_of_3>… that you paste in sequence, and can wrap everything in a structured <context> tag so the model knows what it's holding.
  5. Paste into ChatGPT, Claude, DeepSeek, or NotebookLM.
Bonus: the same pass can redact PII before the text leaves your device. If the document contains names, account numbers, or API keys, see our guide to pasting documents into AI chats safely.

Why not just upload the PDF to the chat?

Chat file uploads get preprocessed for you — and that's the problem. You don't control the extraction: the same headers, footers, and wrap damage come along, and long files are often truncated or summarized rather than read in full, without much indication of what was dropped. Cleaning the text first means you decide exactly what the model sees, you can read it yourself before sending, and you know the real token count up front.

Manual alternatives, ranked

Frequently asked questions

How many tokens is a PDF page?

A clean business-text page is typically a few hundred to over a thousand tokens depending on density, tables, and footers — "per page" is exactly the wrong unit to reason with. LoyalCloak shows the real count using the actual tokenizer of your target model.

What does "prompt too long" actually mean?

Your instructions + document + conversation history + expected answer exceeded the model's context window. The durable fix is sending less text: strip the overhead, start a fresh chat, or split the document into chunks.

Does stripping headers and footers lose information?

They're page furniture — repeated on every page and carrying no payload. The actual content is untouched, and you can review the sanitized output at a glance before sending it.

How do I fit a 100-page PDF into ChatGPT?

Chunk it. Split on section boundaries (not arbitrary character counts), label the segments <part_1_of_N> so the model knows the order, and paste them in sequence — then ask your question against the relevant part. LoyalCloak automates the split.

Reclaim your context window

Drop a PDF in the browser and watch the token count fall — free, no sign-up, nothing uploaded.