How it works

How Safe-Doc pseudonymizes or anonymizes documents.

Safe-Doc secures external AI usage with a simple process: pseudonymize (or anonymize), analyze, de-anonymize.

Safe-Doc flow diagram: input (PDF, DOCX, TXT, Data Room), protection (anonymization, pseudonymization, levels N1-N2), outputs (pseudonymized text and optional JSON mapping).

Interactive demo

Follow the guided demo: anonymize a document, send it to an AI, get the response back in clear text. Simulated exchange with no real data sent anywhere.

1) Import: File, Paste, or Data Room

Import options

  • Upload a file: drag & drop or select a document (PDF/DOCX/TXT).
  • Paste text: great for an email, an excerpt, a note.
  • Data Room (multi-docs): import multiple documents at once for multi-document analysis.
File Paste Data Room

Why Data Room helps

With pseudonymization, the same entity remains the same pseudonym across documents.

Doc 1
Doc 2
Carrefour
Carrefour
[ORG_2]
[ORG_2]

Great for diligence, contracts, multi-piece case files.

2) Choose a mode: Anonymize, Pseudonymize, or Fake data

Anonymized

Generic tokens: [PERSON], [LOCATION]

When you don't need to keep links between occurrences.

Pseudonymized

Numbered, consistent tokens: [PERSON_1], [ORG_2].

Best to preserve narrative consistency across a document set.

Fake data

Readable replacements (invented names/addresses) for a "natural" text.

Useful for review and presentation while masking real values.

3) Increase protection: N1-N2, N3 (roadmap)

Principle

Higher levels reduce re-identification through context (dates, amounts, locations, writing style).

N1-N2 Standard · Advanced

Direct identifiers (PII) and risky context: tokenization, cleanup and generalization (dates, amounts, locations, references).

N3 High security

Stylistic fingerprint and weak signals. Roadmap Q3 2026.

Visual example

"-17.3M - Feb 12, 2026 - Rouen"
N1-N2: contextual reduction · N3 (roadmap)
"mid-teen millions - Q1 2026 - North France"

Goal: keep useful meaning while reducing identifiability.

4) Review: detected entities and residual scan

What you control

  • Detected entities: people, orgs, locations, emails, phones, IBAN, amounts…
  • User choice: uncheck what you don't want to mask.
  • Comparisons: views by source (e.g., model vs regex) depending on UI.

Why human review matters

Automated detection can produce false positives and false negatives. Indicators (residual scan / leakage score) help assess risk, but don't replace a final review.

Ready to try it on a real document?

Pseudonymize or anonymize in seconds.