
Processing sensitive documents without storing data means applying irreversible anonymization or secure local processing to avoid recording any personal data. This approach, referred to in GDPR compliance as processing without retention, addresses a growing requirement in the legal, financial, and human resources sectors. A lawyer who submits a contract to ChatGPT, an HR director who analyzes a personnel file via Claude, or a financial analyst who processes a data room-all expose personal data to uncontrolled third parties. Tools like Gemma 4, GLiNER, and Presidio now make it possible to process these documents locally, without any transfer of raw data to external servers.
Table of Contents
- What are the regulatory requirements for processing sensitive documents without storing data?
- Which tools should you use to pseudonymize or process sensitive documents locally?
- How do you implement secure processing without storage in practice?
- What are the common pitfalls and best practices to ensure compliance and security?
- Key takeaways
- What I learned working on processing workflows without storage
- Safe-doc: pseudonymization and compliance without storage
- Frequently asked questions
- Recommended resources
What are the regulatory requirements for processing sensitive documents without storing data?
The GDPR distinguishes two concepts that professionals regularly confuse: anonymization and pseudonymization. Irreversible anonymization is the only method that removes data from the scope of the GDPR. Pseudonymization, on the other hand, maintains the status of personal data because re-identification remains possible via a protected key.
Under the GDPR, pseudonymization does not eliminate legal obligations. Even if the correspondence key is restricted to a single administrator, pseudonymized data remains personal data. This distinction has direct consequences for security obligations, retention periods, and breach notification requirements.
Sensitive data under the GDPR includes health information, biometric data, ethnic origin, political opinions, and sexual orientation. Processing such data is prohibited except under strictly regulated exceptions. The legal, HR, and finance sectors may access it under conditions of enhanced security and documented justification.
Article 25 of the GDPR mandates minimization, access limitation, and retention limitation as fundamental principles. The data controller must configure systems to minimize the exposure surface of sensitive data. This means concretely: no unnecessary copies, no excessive logging, no open access.
Here are the key obligations to comply with for processing without storage:
- Minimization: collect only data strictly necessary for the declared purpose.
- Limited duration: delete data as soon as the purpose is achieved, with no residual retention.
- Restricted access: limit access rights to only those who need them.
- Minimal traceability: log access without retaining the content of processed data.
- Privacy by design: integrate protection from system design, not as an afterthought.
Pro tip: Before choosing a processing tool, verify whether the vendor can demonstrate contractually that no content is retained, reused for training, or transmitted to subprocessors. A simple commercial commitment is not enough: demand a Data Processing Agreement (DPA) compliant with Article 28 of the GDPR.

Which tools should you use to pseudonymize or process sensitive documents locally?
Local processing is the safest method to guarantee confidentiality without storage. Local anonymization with deletion of identifiers before any transfer to a third-party AI prevents data leaks and complies with the principle of privacy by design. The question is no longer whether it is technically possible, but which architecture to choose.

The local anonymization pipeline relies on three steps: text extraction via OCR, named entity recognition (NER) on a local machine, then replacement of sensitive data with generic markers called placeholders. This workflow ensures that no raw data leaves the controlled environment. Anonymization by named entity recognition is more reliable than simple masking for legal documents, as it understands semantic context.
| Tool | Type | Processing | Primary use |
|---|---|---|---|
| Gemma 4 | Local language model | Local, on-device | Analysis and generation on anonymized documents |
| GLiNER | Lightweight NER model | Local, on-device | Named entity detection in legal texts |
| Presidio | Open-source library | Local or private cloud | Detection and replacement of PII in documents |
| Emvista | Sovereign SaaS | French cloud | GDPR-compliant semantic analysis |
Gemma 4 is a language model designed to run entirely locally, with no calls to external APIs. GLiNER specializes in named entity detection with a reduced memory footprint, making it suitable for constrained environments. Presidio, developed by Microsoft, is an open-source library that automatically detects and replaces personally identifiable information (PII) in text documents.
Hosting or processing via US cloud APIs exposes sensitive data to the US Cloud Act. The CNIL recommends prioritizing local anonymization or locally deployed language models. For law firms, finance departments, and HR teams, this risk is not theoretical: a breach can result in sanctions and loss of client trust.
Pro tip: For highly structured documents like contracts or payslips, combine GLiNER for entity detection and Presidio for automated replacement. This combination covers both standard named entities (names, addresses) and structured identifiers (social security numbers, IBANs).
How do you implement secure processing without storage in practice?
Setting up a compliant workflow requires rigorous preliminary analysis before any technical deployment. Here are the concrete steps for professionals in the legal, financial, and HR sectors.
1. Map the scope of processed data. Identify which documents contain sensitive data: contracts with personal data, employee files, notarial deeds, nominative financial reports. This mapping determines the anonymization rules to apply.
2. *Design the architecture with privacy by design.* Privacy by design architecture requires strict access control, limited duration, and minimal traceability. The system must be designed never to retain data beyond the declared purpose.
3. Apply the extraction and anonymization pipeline. The document enters the system, the OCR engine extracts the text, the NER model identifies sensitive entities, and each entity is replaced with a generic placeholder. For legal documents, NER must be contextualized to precisely replace sensitive data without altering the meaning of the document.
4. Process the anonymized document, then delete immediately. The AI or analysis tool receives only the anonymized document. As soon as processing is complete, the temporary file is deleted. No copy remains in logs, caches, or shared workspaces.
5. Log access without retaining content. True processing without storage requires controlling logs, prompts, and any retraining mechanisms of the tools used. Log who accessed what and when, without recording the content of processed data.
6. Audit the workflow regularly. Verify that placeholders properly cover all categories of identified sensitive data. Quarterly audits detect drift before it becomes a breach.
Concrete example in HR: a dismissal file contains health data, union information, and personal financial data. Before any AI tool analysis, the pipeline automatically anonymizes these categories. The tool receives a document where "Mr. Dupont, diabetic, union representative" becomes "[PERSON], [HEALTH], [UNION_STATUS]". The legal meaning is preserved; personal data does not transit.
What are the common pitfalls and best practices to ensure compliance and security?
The risk of re-identification is the first pitfall to avoid. Improper use of pseudonymization, with easy reconstruction via quasi-identifiers like age, postal code, and profession combined, maintains the status of personal data. This risk must be rigorously assessed before considering a treatment as anonymized.
Best practices to apply systematically:
- Test re-identification: after anonymization, verify that a third party cannot recover an individual's identity from the remaining data.
- Delete unnecessary logs: temporary files, prompt histories, and caches constitute often-neglected leak vectors.
- Control access by role: only collaborators directly involved in processing access documents, even anonymized ones.
- Audit subprocessors: any vendor involved in the processing flow must sign a Data Processing Agreement compliant with Article 28 of the GDPR.
- Distinguish local LLM from cloud LLM: a model deployed locally transmits no data. A cloud LLM, even with contractual guarantees, presents residual risk.
The CNIL severely sanctions failures related to excessive retention of sensitive data and lack of access control. A poorly designed architecture can turn legitimate processing into a characterized GDPR violation.
Pseudonymization remains useful for internal workflows where controlled reversibility is necessary, for example to retrieve a file after processing. But it does not exempt organizations from GDPR obligations. For transfers to third-party AI tools, only irreversible anonymization offers real protection.
Key takeaways
Processing sensitive documents without storing data requires irreversible anonymization, a privacy by design architecture, and strict control of logs, access, and subprocessors.
| Point | Details |
|---|---|
| Anonymization vs pseudonymization | Only irreversible anonymization removes data from the scope of the GDPR. |
| Recommended local pipeline | Combine OCR, local NER (GLiNER, Presidio), and immediate deletion after processing. |
| Cloud Act risk | US cloud APIs expose sensitive data; prioritize locally deployed models. |
| Privacy by design mandatory | Article 25 of the GDPR requires minimization, limited duration, and access control from design. |
| Re-identification to be tested | Systematically verify that no quasi-identifiers allow re-identification after anonymization. |
What I learned working on processing workflows without storage
The biggest obstacle is not technical. Tools like GLiNER, Presidio, or Gemma 4 are accessible and well documented. The real problem is internal culture: legal and HR teams have entrenched work habits, and the temptation to copy-paste a contract into ChatGPT remains strong when the anonymized pipeline requires two additional steps.
I have seen legal departments invest in a perfectly designed local architecture, then bypass it within six months because no one had trained the staff. Technology without human support protects nothing. Compliance for legal departments depends as much on training as on the tool.
The other point often underestimated: logs. Processing "without storage" can easily retain traces in prompt histories, system temporary files, or automatic workstation backups. Controlling logs is as important as controlling the anonymization pipeline itself.
Finally, the client benefit is concrete and measurable. A law firm or finance department that can demonstrate to its clients that their data is not durably stored or transmitted to third parties gains a real competitive advantage in trust. This is not a marketing argument. It is a contractual guarantee that few competitors can offer today.
- Jacques
Safe-doc: pseudonymization and compliance without storage
Safe-doc precisely addresses the needs identified in this article. The platform pseudonymizes sensitive documents in real time before any AI processing, without durably storing the original content. Legal, financial, and HR teams continue to use ChatGPT or Claude, but with a protective layer that guarantees GDPR compliance.

Safe-doc handles GDPR-compliant pseudonymization with a zero-storage-by-design architecture. For DPOs and compliance officers, the platform also offers audit and traceability features tailored to Article 25 requirements. Processing remains local, data does not transit, and compliance is documented at every step.
Frequently asked questions
What is the difference between anonymization and pseudonymization under the GDPR?
Irreversible anonymization removes data from the scope of the GDPR. Pseudonymization maintains the status of personal data because re-identification remains possible via a protected key.
Can you use ChatGPT or Claude with sensitive documents?
Yes, provided you anonymize documents before sending. A tool like Safe-doc pseudonymizes content in real time, without storing data, before transmission to the AI.
Which tools enable local anonymization without data transfer?
GLiNER and Presidio enable detection and replacement of sensitive entities locally. Gemma 4 then processes anonymized documents without calling an external API.
Does the US Cloud Act concern French companies?
Yes. Any data hosted or processed via a US vendor's API may be subject to the Cloud Act. The CNIL recommends local anonymization or locally deployed models for sensitive data.
When is pseudonymization sufficient for GDPR compliance?
Pseudonymization suffices for internal workflows where controlled reversibility is necessary. For any transfer to a third-party tool, only irreversible anonymization guarantees real protection of personal data.