In the "AI-safe tools" market, everyone says "anonymise your documents." Yet as soon as a product offers to reinject original names after analysis, it's no longer anonymization under GDPR - it's pseudonymization. The confusion isn't just vocabulary : it drives your legal risk level and the trust you can place in a vendor.
What GDPR says
GDPR art. 4(5) defines pseudonymization as processing personal data so they can no longer be attributed to a data subject without additional information, kept separately and subject to technical and organisational measures.
Anonymization aims for a state where re-identification is no longer reasonably possible - without an exploitable mapping table. Anonymized data generally falls outside GDPR scope, provided anonymity is real and durable.
A DPO or privacy lawyer will immediately challenge marketing that says "anonymization" while selling a "restore data" button.
Why the distinction matters for AI use
The most common business flow today:
- Protect the document (replace names, precise amounts, identifiers…)
- Send the cleaned version to ChatGPT, Claude, Gemini or Copilot
- Retrieve a summary, translation or analysis
- Reinject original values into the deliverable
That fourth step requires a mapping - correspondence file, encrypted table, consistent tokens across occurrences. Technically and legally, you're in reversible pseudonymization, not irreversible anonymization.
Both approaches are legitimate depending on use case. The problem is confusing them or selling under the wrong label.
Tools that say "anonymize" but reinject
Some vendors - including lawyer-oriented solutions - use "anonymization" in the UI while keeping a workspace, session history, or entity restoration chain. Then:
- Data isn't "anonymous" in the strict sense
- Re-identification risk depends on masking quality and mapping security
- Processing responsibility remains that of pseudonymized personal data
Safe-Doc uses terms precisely: pseudonymization when mapping allows de-pseudonymization; anonymization for outputs with no planned return (Clean mode without reversibility).
Which mode to choose?
Pseudonymization - recommended for contracts, due diligence, internal notes where you need real names back in the AI response. Consistent tokens ([PERSON_1], [ORG_2]), JSON mapping export, local de-pseudonymization.
Anonymization - when you don't need to re-identify: sector monitoring, communication drafts, external sharing without return. Residual re-identification risk by context remains : human review is still essential.
In both cases, masking reduces exposure to the AI model and vendor training policies - without alone eliminating all GDPR obligations or contextual leak risk.
Safe-Doc: the right term, zero AI transit
Safe-Doc doesn't replace your AI tool. You pseudonymize or anonymize before pasting text into your chosen interface. We don't see your AI query : no built-in chat, no logging of content sent to the model. Source documents aren't stored after processing; reversible mapping is encrypted and under your control.
That's a different positioning from "all-in-one" platforms that centralize documents, prompts and history - and sometimes call it anonymization.
Try on a real document - pseudonymization or anonymization in seconds.