Blog

Pseudonymization of legal documents: GDPR 2026 guide

What exactly does the GDPR say about the pseudonymization of legal documents

Pseudonymization is defined in Article 4, paragraph 5 of the GDPR as the processing of personal data in such a manner that they can no longer be attributed to a natural person without the use of additional information, provided that this information is kept separately and is subject to appropriate technical and organizational measures. In other words, a contract whose names of the parties have been replaced by codes remains personal data as long as the correspondence table exists somewhere. The technique does not take the processing outside the scope of the GDPR. It reduces the risk, without eliminating it.

For legal professionals, this framework has direct consequences. A pseudonymized judgment transmitted to an analysis service provider, a customer file processed by an artificial intelligence tool, archived correspondence with substituted identifiers: all these documents remain subject to the obligations of the GDPR as long as reidentification remains possible by reasonable means.

Here are the key points of the applicable legal framework:

  • Article 4, paragraph 5 of the GDPR: legal definition of pseudonymization, based on the separation of additional information.
  • Recital 26 of the GDPR: specifies that the identifiable nature of a person must be assessed with regard to all the means reasonably likely to be used, taking into account the cost, time and available technologies.
  • Article 25 of the GDPR: pseudonymization is one of the data protection measures by design (privacy by design).
  • Article 32 of the GDPR: pseudonymization is cited as an appropriate technical security measure to protect the processed data.
  • CJEU case law, judgment of March 7, 2024 (C-479/22): pseudonymized data remains personal for the data controller. For a third party, it is only if re-identification is physically impossible, that is to say it would involve a disproportionate effort in terms of time, cost and human resources.
  • Decision of the Council of State n°498628 of February 13, 2026: confirms the CNIL sanctions in the case of the “Thin” and “Gers Études client” databases, recalling that pseudonymization only constitutes anonymization if the risk of reidentification is insignificant.
  • Position of the CNIL: pseudonymization is a technical security measure, not a legal basis for processing. Consent or another legal basis remains necessary.

The CJEU also confirmed in September 2025 that pseudonymization does not exempt the data controller from its transparency and documentation obligations. Processing a legal document with pseudonyms without informing the persons concerned or documenting the measure exposes you to sanctions.

Pseudonymization or anonymization: what are the concrete legal differences?

A legal expert reviewing legal files

Confusion between these two notions is common, and it is costly. The distinction made by the CNIL is however clear: pseudonymization is reversible, anonymization is not. An anonymized document falls outside the scope of the GDPR; a pseudonymized document remains fully submitted.

CriterionPseudonymizationAnonymization
ReversibilityYes, with the lookup tableNo, irreversible by definition
Scope of GDPRMaintained for the data controllerExcluded if the risk of re-identification is insignificant
Risk of re-identificationResidual, to be assessed contextuallyTheoretically null
Legal basis requiredYesNo (outside GDPR scope)
Rights of data subjectsMaintained (access, rectification, erasure)Not applicable
Typical usage in lawTransmission to third parties, use of AI, archivingPublication of court decisions, research

Infographic: the key differences between pseudonymization and anonymization

The practical consequences for legal documents are significant. A lawyer who pseudonymizes a file before submitting it to an AI tool must always respect professional secrecy and the rights of the persons concerned. Pseudonymization facilitates processing, it does not exempt it.

Two situations clearly illustrate where pseudonymization is not enough to go beyond the scope of the GDPR:

  • Cross-referenced data: in the case judged by the Council of State in 2026, patient codes combined with data on age, pathology, date of consultation and prescriber identifiers made it possible to reconstruct individualized care pathways with a simple spreadsheet. Pseudonymization was considered insufficient.
  • Dense legal context: a contract whose names are replaced by codes but which contains references to addresses, file numbers or specific amounts can remain identifiable by cross-checking.

Professional secrecy adds an additional layer of requirement. For lawyers and notaries, the pseudonymization of documents transmitted to external service providers is not only a good GDPR practice: it is also part of the ethical obligation to protect the information entrusted by clients.

Which pseudonymization techniques are suitable for legal documents?

Not all methods are equal depending on the context. The techniques recommended for legal documents generally combine several approaches, depending on the sensitivity of the data and the purpose of the processing.

  • Substitution of identifiers: replacement of last names, first names and file numbers with randomly generated alphanumeric codes or pseudonyms. Most common method in contracts and correspondence.
  • Encryption: identifying data is encrypted with a key kept separately. The encrypted data remains in the document; the key is stored in an isolated environment. Suitable for long-term archives.
  • Hashing: irreversible transformation of an identifier into a digital fingerprint. Useful for verifications without re-identification, but be careful: a hash without salt (salt) can be reversed by dictionary attack on predictable data like social security numbers.
  • Tokenization: replacement of the real identifier by a token (token) managed by a secure third-party system. Common in document management platforms for legal departments.
  • Generalization: replacing a specific value with a range or category (for example, replacing an exact date of birth with an age range). Useful for statistical analyzes on files.

Strict separation of the lookup table is non-negotiable. The CNIL regularly audits the robustness of what it calls the “pseudonymization vault”: all of the information allowing pseudonymization to be lifted must be stored in a separate environment, with rigorous access controls and complete traceability of consultations.

The legal qualification of the data does not depend on the technique used. It is based solely on the objective assessment of the actual risk of reidentification, with regard to the means reasonably available to achieve this. Council of State, decision no. 498628, February 13, 2026.

To choose the appropriate technique for a legal document, three questions arise: should reidentification remain possible internally? Will the document be transmitted to third parties? What is the expected shelf life? A contract transmitted to an AI provider for analysis calls for a substitution of identifiers coupled with encryption of the table. A judgment intended for publication may, depending on the level of generalization achieved, tend towards anonymization.

How to guarantee GDPR-compliant pseudonymization in your processes?

Compliance is more than just applying a technique. It presupposes an organizational and documentary architecture that the CJEU and the CNIL examine closely during controls.

  • Register of processing activities: each pseudonymized processing must appear there with the description of the technique used, the categories of data concerned, the recipients and the associated security measures. The absence of this precise documentation is one of the first reasons for sanction during a CNIL inspection.
  • Information notices: data subjects must be informed that their data is subject to pseudonymized processing, even if they do not see the pseudonyms. The legal basis for the underlying processing must be clearly identified.
  • Access control to the correspondence table: only strictly authorized people can access the pseudonymization vault. Each access must be logged.
  • Reidentification risk assessment: it must be documented and updated, particularly when new sources of public data become accessible or when technologies evolve. The Council of State reminded in 2026 that a simple spreadsheet can be enough to remove a poorly designed pseudonymity.
  • Regular audit: technical and organizational measures must be tested periodically. The CNIL may request to examine access logs and key management procedures.
  • Impact analysis (AIPD): for high-risk processing, in particular those involving health data or judicial profiles, an impact analysis relating to data protection remains mandatory even if the data is pseudonymized.

Pro tip: The pseudonymization key should never be stored in the same environment as the pseudonymized data. Use a dedicated secrets manager (enterprise digital vault type) with multi-factor authentication for all access to the lookup table. Document each access with operator identity, timestamp and justification.

Secure key management is the element that supervisory authorities examine first. The robustness of the system relies less on the existence of the key than on the conditions in which it is protected and used.

Pseudonymization in practice: three concrete cases in legal documents

Contracts and notarial deeds transmitted to an AI service provider

A law firm wants to use a generative AI tool to analyze M&A contracts. Before any transmission, the names of the parties, their contact details, SIREN numbers and amounts are replaced by randomly generated codes. The correspondence table remains in the firm's internal environment, encrypted and accessible only to the partners responsible for the file. The AI ​​tool receives a functional document but without direct identifiers. After analysis, the results are reintegrated into the original file via the table. This flow allows use AI on sensitive files without exposing customer data.

Judgments and court decisions for legal research

A legal department of a large company archives court decisions to provide an internal case law database. The names of the parties, witnesses and experts are substituted by identifiers of type “Party_A”, “Expert_02”. The precise dates are generalized to the year. The resulting database can be queried by in-house lawyers without exposing identities. However, if very specific contextual elements (atypical amounts, rare sector of activity, precise location) remain, the risk of reidentification remains to be assessed.

Pseudonymization is a risk management tool, not a legal purpose in itself. Residual risks must always be assessed, documented and reduced as much as possible. CNIL, recommendations on the pseudonymization of personal data.

Correspondence and HR files transmitted to third parties

A data protection officer supervises the transmission of employee files to a social audit service provider. Names, social security numbers and addresses are pseudonymised before transmission. The rights of the persons concerned remain fully applicable: if an employee exercises his right of access, the data controller must be able to find his data via the correspondence table and communicate to him all the information concerning him. Pseudonymization does not suspend GDPR rights, it simply makes them more complex to exercise in practice, hence the importance of a clear re-identification procedure on request.

DPO in charge of managing anonymized HR files

The main limit of pseudonymization in these contexts remains the risk of reidentification by crossing. The less precise the residual data, the stronger the protection. But generalizing too much can make the document unusable for the intended analysis. Finding the right balance is an exercise in professional judgment, not a simple mechanical application of rules.

Safe-doc: pseudonymize your legal documents without changing your tools

The main obstacle to compliant pseudonymization in legal firms and departments is not lack of will. This is operational friction. Manually pseudonymizing a hundred-page contract before submitting it to ChatGPT or Claude takes time, generates errors and discourages teams, who end up using AI tools directly on the original documents. It is precisely this phenomenon that we call Shadow AI: the uncontrolled use of consumer tools on sensitive data, outside of any compliance framework.

Safe-doc addresses this problem with an automatic pseudonymization layer that is inserted between your documents and the AI ​​tools you already use. The pseudonymization is carried out locally, without the documents passing through third-party servers or being stored on the platform. Identifying data is detected and substituted in real time, before transmission to the AI ​​tool. The lookup table remains in your environment.

For data protection managers, Safe-doc directly meets the requirements of Article 32 of the GDPR in terms of technical security measures, while simplifying the compliance and audit of processing involving AI. Processing logs, data separation and the absence of external storage are concrete arguments during a CNIL inspection.

https://safe-doc.ai

Pro Tip: Before deploying Safe-doc in your firm or legal department, map the document flows that are already passing through unsecured AI tools. This mapping of existing Shadow AI will allow you to prioritize the use cases to secure first and demonstrate a proactive approach during an audit.

Pseudonymization of legal documents with Safe-doc also makes it possible to maintain the full functionality of AI tools for analysis, assisted drafting and documentary research, without compromising professional secrecy or exposing client data. For teams dealing with mergers and acquisitions, litigation files or social audits, it is an operational response to the requirements of the GDPR.

Key points

The pseudonymization of legal documents remains subject to the GDPR as long as re-identification is possible by reasonable means, which requires precise technical, organizational and documentary measures at each stage of processing.

PointDetails
Strict legal definitionArticle 4, paragraph 5 of the GDPR requires the separation of additional information under appropriate security measures.
Pseudonymization ≠ anonymizationPseudonymization is reversible and keeps the processing within the scope of the GDPR; anonymization excludes it if the risk of reidentification is insignificant.
Reidentification risk to be assessedThe Council of State confirmed in 2026 that a simple spreadsheet can be enough to remove a poorly designed pseudonymity.
Mandatory documentationThe processing register, information notices and access logs to the correspondence table are audited by the CNIL.
Safe-doc for legal teamsSafe-doc pseudonymizes locally, without external storage, and integrates with existing AI tools for operational GDPR compliance.

Recommendation