It takes two seconds to paste a contract, a medical report, or a client email into an AI chat β and zero seconds to decide it was a mistake. Once text leaves your device, you can't un-send it. This guide explains what actually happens to pasted documents, which details count as sensitive (it's more than names), why manual find-and-replace fails, and gives you a repeatable routine you can finish in about a minute with a free, on-device tool.
What happens to text after you paste it
When you paste into a consumer AI chat, the text is processed on the provider's servers. Depending on the provider, your plan, and your settings, conversations may be retained for a period of time, reviewed for abuse, and in some cases used to improve future models. Providers offer controls β training opt-outs, temporary chats, enterprise data agreements β and they're worth using. But they all share one property: you're trusting a policy that can change, not a technical guarantee.
And like any online service, AI providers are breach targets. What you pasted sits in someone else's infrastructure, in whatever form and for however long their retention rules say.
The point isn't that AI chats are unsafe β it's that your control ends at the paste button. Everything you keep on your side of it is the part you actually own.
What counts as sensitive (most people under-redact)
Names and email addresses are the obvious stuff. Documents built for humans are full of identifiers that are harder to spot but just as identifying:
- Direct identifiers β names, email addresses, phone numbers, messaging handles.
- Financial data β payment card numbers (valid ones pass the Luhn checksum), IBANs (ISO 13616 with mod-97 check digits), and account numbers.
- Government IDs β US Social Security Numbers, UK National Insurance numbers, Israeli Teudat Zehut. Many carry checksums, which is exactly how software can tell a real one from a random digit string.
- Health identifiers β medical record numbers (MRNs), NPI provider IDs, DEA registration numbers, NDC drug codes, ICD-10 diagnosis codes. A diagnosis code on its own can reveal more about a person than their name.
- Digital secrets β AWS access keys (
AKIAβ¦), GitHub tokens, JWTs, Slack and Stripe keys. Developers paste these inside stack traces and config dumps constantly, often without noticing. - Legal identifiers β court docket and case numbers. US federal filing rules (FRCP 5.2) even mandate that SSNs appear only in partial form β last four digits β precisely because these values shouldn't circulate.
An identifier doesn't need a name attached to be dangerous. A medical record number next to a diagnosis code is re-identifying in the wrong hands β no name required.
Why manual find-and-replace fails
The standard advice is "just redact it yourself first." In practice, manual scrubbing breaks down in four predictable ways:
- Variants.
+1 (555) 019-2831,+1-555-019-2831,555.019.2831β you'll miss every format you didn't think to search for, and international numbers multiply the permutations. - You can't tell real from fake. A 16-digit number that looks like a card might fail the Luhn check (test data β safe to keep). A human reading can't evaluate checksums; software can, in milliseconds.
- Fatigue. Page 30 of a 40-page document gets a fraction of the attention page 1 did. Leaks are usually the value you skimmed past.
- The dumb-bucket problem. Replacing everything with
[REDACTED]destroys the context the AI needs. Ask it to draft a reply to "Dear [REDACTED], regarding invoice [REDACTED]β¦" and you'll get a useless answer.
The routine: sanitize on-device, then paste
Instead of hunting identifiers by eye, run the document through a sanitizer that executes entirely on your own device β nothing is uploaded β before it ever touches a chat window. Here's the routine using LoyalCloak's free web app (the same flow works in the macOS, Windows, and Linux desktop apps):
- Drop in the document or paste the text. PDF, DOCX, spreadsheets, EPUB, HTML, or plain text all work. Parsing happens locally in your browser β files are never uploaded.
- Choose what to scrub. Emails and phone numbers are covered on the free tier. Pro adds Luhn-verified cards, SSNs, IBANs, IP addresses, API keys and secrets, person names via on-device AI, and a pattern library of checksum-validated health, legal, and financial IDs. There are one-tap profiles for legal work, development, and GDPR-style scrubbing.
- Review what was found. LoyalCloak shows what it detected and where, and you can allowlist specific values you know are harmless (your own support address, for example).
- Copy the sanitized output. You get clean Markdown β typically substantially smaller than the raw text, which also helps with prompt limits (see our guide to fixing "prompt too long" errors).
- Paste into ChatGPT, Claude, DeepSeek, or NotebookLM. The AI works with the structure and meaning of the document β not with your private values.
"But the AI's answer will reference the fake data"
This is the legitimate objection to redaction: if you blank out the client's name, the AI's drafted email can't address them. LoyalCloak Pro solves this with reversible pseudonymization. Instead of an empty token, each real value is swapped for a realistic stand-in drawn from reserved, can't-be-real ranges β example.com email domains (RFC 2606), 555-01xx phone numbers, documentation-range IP addresses (RFC 5737).
The AI sees a coherent, human-looking document and answers in full context. When the answer comes back, paste it into LoyalCloak and the stand-ins are swapped back to the original values β the mapping never left your device.
Original: Contact: Sarah Connor <[email protected]>, +1 (555) 019-2831
Token mode: Contact: [EMAIL_1], [PHONE_1]
Pseudonym: Contact: Emma Larson <[email protected]>, +1 (555) 014-2211
β realistic for the AI; restored to the original on paste-back
The 30-second pre-paste checklist
- Am I about to paste anything a stranger shouldn't own? If yes β sanitize first.
- Did I look beyond names: account numbers, national IDs, API keys, diagnosis codes, docket numbers?
- Would this document be a problem if it appeared in a breach or a screenshot?
- Do I need the real values in the AI's answer? Use reversible pseudonyms, not blank tokens.
- Is my sanitizer running on-device β so my sanitized copy doesn't become a second leak?
Frequently asked questions
Is it safe to paste confidential documents into ChatGPT?
It depends on the provider, your plan, and settings you can't audit β retention periods and training policies differ and change over time. The only part of the pipeline you fully control is what you send. Removing sensitive data before pasting makes the question moot: there's nothing sensitive left to mishandle.
If I turn off chat history and training, is my data safe?
Those controls reduce one specific use (model training). Conversations are still processed and retained per the provider's policies β for abuse monitoring, for example. Sanitizing first protects you even if policies change, you misconfigure a setting, or you're on a plan without those controls.
What's the fastest way to redact a PDF before using AI?
An on-device sanitizer. Print-to-PDF, blurring, and screenshot workarounds are slow and easy to get wrong, and online "PDF redactor" sites mean uploading the very document you're trying to protect. LoyalCloak parses PDFs locally, scrubs the PII, and hands you copy-paste-ready Markdown.
Will redacting data confuse the AI?
Blank tokens in the wrong places can β that's the dumb-bucket problem. Reversible pseudonymization keeps the document readable and coherent, so summaries, drafts, and analysis stay accurate while the real values never leave your device.
Does LoyalCloak upload my documents anywhere?
No. The web app runs entirely in your browser and the desktop apps run entirely on your machine β parsing, PII detection, and token counting all happen in local RAM. Nothing is stored on disk, and there are no analytics or telemetry in the apps.
Sanitize before you paste
Try it now in your browser β free, no sign-up, and your document never leaves the page.