Blog/Security & compliance
PII redaction in LLM pipelines: where to redact, how, and what GDPR says
Where to redact PII in an LLM pipeline, reversible tokens vs masking, Presidio and cloud DLP, German and Hungarian gaps, GDPR on pseudonymised data, and tests.
Balázs Csorba··12 min read
- PII redaction
- GDPR
- Microsoft Presidio
- LLM security
- Pseudonymisation

Key takeaways
- Redact at every boundary, not once: ingest, prompt, output, logs and traces each leak differently, and traces and logs are the boundary teams forget most often.
- Reversible tokens (a vault that maps placeholders back to real values) keep answers useful, but the vault itself becomes a personal-data store that needs keys, access control and a short retention.
- No detector finds everything. Presidio says so itself, and coverage for German is far better than for Hungarian, so measure recall per language on your own data instead of trusting a vendor list.
- Under the GDPR, pseudonymised data stays personal data for whoever holds the key. The CJEU ruling of 4 September 2025 adds that it may not be personal data for a recipient who cannot re-identify, but that has to be shown, not assumed.
- Treat redaction as defence in depth next to contracts, EU hosting and access control, and test it like any other feature: a labelled multilingual set, recall targets and a regression gate in CI.
Most teams add PII handling to an LLM feature the same way: one regex for e-mail addresses in front of the API call, and a note in the backlog. It works in the demo and fails in production, because personal data does not enter an LLM system at one point. It arrives in the user message, in the documents you index, in the tool results an agent reads, and then it is copied into the model output, the application log, the trace backend and the evaluation set.
This article is how I would design redaction for a European company: where the gates go, which technique to use at each one, what the tools can and cannot do (including for German and Hungarian), how to read the current GDPR position on pseudonymised data, and how to test the whole thing. It is engineering advice, not legal advice, and it complements the hosting and contract questions in GDPR and LLM APIs: EU data residency.
Where to redact: five boundaries
Think of the pipeline as five boundaries where text crosses into a system you do not fully control or that lives longer than the request. Each one needs its own decision.
- Ingest. Redact documents before chunking and embedding. A vector store full of raw personal data is hard to delete from and widens every later leak. Decide per source whether you need the real value at all. See the chunking and indexing trade-offs in RAG pipeline: chunking, hybrid search, reranking.
- Prompt. The user message, retrieved context and tool results all go through the gate right before the API call. This is the most important boundary, because it is the one that controls what the provider receives.
- Output. The model can repeat, infer or invent personal data. Check the answer before it is shown or stored, and restore tokens only where the viewer is allowed to see the real value.
- Logs. Application and gateway logs are the classic leak. Log the redacted prompt, or a hash and a size, never the raw message.
- Traces and evals. Observability tools store full prompts and completions by design. Langfuse, for example, offers masking hooks that run before data is exported, and plain OpenTelemetry setups can mask in the application or in a collector. Evaluation datasets built from production traffic inherit whatever the traces contain.
A gateway such as LiteLLM can host the prompt-side gate. Its Presidio guardrail can run before the call, after the response, or only for logging, and it can parse model output to replace masked tokens with the original values. That is a convenient place to start, but note the limits: it handles the request and response, not your ingest job or your trace backend.
Redaction, masking, tokenisation: choosing the technique
Detection finds the spans; the technique decides what replaces them. The choice is a trade-off between utility for the model, reversibility and the damage if the output leaks. The table is my assessment, built on the operators that Presidio and Google Cloud Sensitive Data Protection document.
| Technique | Reversible | Model utility | Main risk | Good for |
|---|---|---|---|---|
| Removal (empty or REDACTED) | No | Low: sentence structure breaks | Information loss | Logs, analytics, anything that never needs the value |
| Typed placeholder (PERSON_1) | Only with a vault | High: the model still sees roles and relations | Vault becomes a data store | Prompts, summaries, support tickets |
| Character masking (**1234) | No | Low to medium | Leaks partial values and length | Display of card or phone tails |
| Salted or keyed hash | No | Medium: stable joins, unreadable text | Guessable for low-entropy values | Deduplication, join keys in analytics |
| Deterministic or format-preserving encryption | Yes, with the key | Medium to high: same value gives same token | Key management; equality leaks | Structured fields, cross-document consistency |
| Realistic surrogate (fake name) | Only with a vault | High, reads naturally | Fake value may collide with a real person | Demos, test data, evals |
Presidio ships operators for replace, redact, hash, mask, encrypt and custom functions, with decrypt as the built-in reverse. Google documents deterministic encryption (AES-SIV), format-preserving encryption (FPE-FFX) and HMAC-SHA-256 hashing, the first two reversible, and recommends keys wrapped by Cloud KMS. My default for prompts is typed, numbered placeholders: the model can still reason that PERSON_1 wrote to PERSON_2, and you decide at the output boundary who may see what.
Reversible tokenisation: useful, with a vault attached
Reversible tokens solve the usability problem: the user asks for a reply to a customer, the model drafts it around PERSON_1, and your gate restores the real name before display. Three design rules keep it safe.
- Scope tokens to a session or request. A fresh mapping per conversation avoids a global lookup table and stops tokens from becoming cross-conversation identifiers.
- Encrypt and expire the vault. Keep it in your own infrastructure, encrypt it with a managed key, and delete mappings when the conversation ends or after a short TTL. It is personal data and needs the same deletion path as the rest.
- Restore only at the edge. Do the substitution in the layer that renders to an authorised user, not inside the agent loop. Otherwise a tool call can carry real values back to places the model should not reach.
Tools: Presidio, cloud services and NER models
There is no single answer; there are three families, and many production setups combine them.
- Microsoft Presidio (open source, self-hosted). It combines named-entity recognition, regular expressions, rule-based logic, checksums and context words. It is the usual starting point because you control where the text goes and you can add recognizers. Its documentation is candid: because detection is automated, there is no guarantee that it finds all sensitive information.
- Cloud services. Google Cloud Sensitive Data Protection offers a long list of infoTypes with a location per type, including German ones such as passport, identity card, driver's licence, taxpayer ID and SCHUFA ID. Azure Language lists German and Hungarian for text PII, and its conversation PII is documented for English, French, German and Spanish only. Amazon Comprehend documents PII detection for English or Spanish text. These are managed and easy to start with, but sending raw text to a third-party detector is itself a transfer you must justify.
- NER models and hybrids. Fine-tuned transformers can be run locally. One research paper on hybrid detection (regular expressions plus LLMs, tested on 13 low-resource languages) reports clearly better weighted F1 than fine-tuned NER models and zero-shot LLMs; treat it as a pointer to combine deterministic patterns with context-aware models, not as a ready product.
My rule: deterministic recognizers with validation (IBAN checksums, tax-ID formats) for structured identifiers, an NER model for names and places, and an LLM-based pass only where recall matters more than cost and the model runs inside your boundary.
German and Hungarian: the coverage gap
Most detectors are strongest in English. For a company in Austria or Hungary that is the practical risk, because the text is German or Hungarian, often mixed with English.
- German. Presidio documents German recognizers for tax IDs, passports, national ID cards, health insurance numbers and vehicle plates, spaCy ships trained German pipelines, and LiteLLM lists German as a supported guardrail language. Names also have to be told apart from the many capitalised common nouns, so test name recall on German text separately.
- Hungarian. I found no Hungarian-specific recognizers in Presidio's entity list or in Google's infoType reference, and spaCy shows no trained Hungarian pipeline. Azure lists Hungarian for text PII. Open models exist, for example a huBERT-based NER model fine-tuned on the NerKor corpus with PER, ORG, LOC and MISC labels, but it is GPL-licensed and limited to 448 tokens of input, which matters for long documents.
- Language-specific context. Presidio recognizers support one language each, and its documentation notes that while patterns such as regular expressions are language agnostic, the context words that raise confidence are not. A German recognizer needs words like "Steuernummer"; a Hungarian one needs "adószám" or "TAJ-szám", and Hungarian suffixes make names change form (Péter, Péternek, Péterrel).
Hungarian-specific identifiers such as the tax number or the social security (TAJ) number are easy to add as custom pattern recognizers with checksums, and that is where I would start. Names are the hard part and need a model plus your own evaluation.
False negatives: the failure that matters
A false positive replaces a harmless word and costs a little quality. A false negative sends a real name to a provider and writes it to a log. Optimise for recall on the entities that matter, and accept noisy precision.
- Names in free text: nicknames, lower-case typing in chat, names that are also common words, and inflected forms.
- Context-dependent identifiers: a job title plus a small town plus a date can identify a person without any single obvious entity.
- Format variants: phone numbers with odd spacing, IBANs split across lines, identifiers inside URLs or code blocks.
- Non-text inputs: OCR output from scanned documents, tool results in JSON, and file names.
Because of these gaps, do not rely on redaction alone. Add provider-side controls (EU region, no training on data, zero retention where offered), least-privilege retrieval, and a rule that special-category data (health, for example) is not sent to a general model at all unless a documented basis exists.
The GDPR view: pseudonymised is not anonymous
Article 4(5) GDPR defines pseudonymisation as processing so that data can no longer be attributed to a specific person without additional information, provided that information is kept separately under technical and organisational measures. It is a safeguard, not an exit from the regulation. The EDPB adopted its Guidelines 01/2025 on pseudonymisation on 16 January 2025 and consulted on them in early 2025; I could not confirm a final version, and the EDPB held a stakeholder event on the topic in December 2025, so treat the guidelines as still evolving.
The CJEU added a nuance on 4 September 2025 in EDPS v SRB (C-413/23 P). The data in that case had been pseudonymised by the Single Resolution Board, which kept the key, before being sent to Deloitte. The court confirmed that such data can be personal data for the original controller but not necessarily for a recipient who has no reasonable means to re-identify the people, and that the controller's duty to inform data subjects exists independently of the recipient's view. Commentators advise documenting why a recipient cannot re-identify and reassessing when technology or datasets change.
What this means for an LLM pipeline, in my reading: your own systems that hold the vault or key still process personal data. Whether the model provider receives personal data depends on whether it has reasonable means to re-identify, which is a factual question about the text you send. Free-text prompts that survive redaction with rare combinations of details are weak evidence. Document the assessment, keep your privacy notice accurate about the transfer, and do not call placeholder text anonymous.
Testing redaction like a feature
Redaction is a classifier, so test it like one, and wire it into the same practices as other LLM features (see LLM evals for product features).
- Build a labelled set per language you serve, with realistic noise: typos, lower-case names, mixed German or Hungarian and English, tables and code blocks.
- Report recall and precision per entity type, and set recall targets for names, contact data and identifiers separately. Track the rate of missed entities, not only the average.
- Add canary values (fake but valid-looking names, IBANs and tax IDs) to test traffic and assert in CI that they never appear in provider requests, logs, traces or caches.
- Test the round trip: tokenise, call the model, restore. Check that tokens survive paraphrasing, that unknown tokens are not restored, and that restoring is impossible without the right session.
- Re-run the suite whenever the NLP model, a recognizer or the language configuration changes, and sample live traffic with human review to find drift.
What I would do first
Start with a prompt-side gate using Presidio or an equivalent, typed placeholders with a per-session vault, redaction before indexing, and masking in your trace backend. Add custom recognizers for German and Hungarian identifiers, measure recall on your own text, and pair all of it with EU hosting and contracts. The goal is not perfect detection, which no tool promises, but a pipeline in which one missed name does not end up in five different systems.
Sources
- Microsoft Presidio: documentation (limitations, methods)
- Presidio: supported entities and country-specific recognizers
- Presidio: supporting additional languages
- Presidio: anonymizer operators
- Google Cloud: infoTypes reference
- Google Cloud: pseudonymization in Sensitive Data Protection
- Microsoft Learn: Azure Language PII detection language support
- AWS: Detecting PII entities with Amazon Comprehend
- spaCy: models and languages
- Hugging Face: novakat/nerkor-hubert (Hungarian NER)
- arXiv: An Evaluation Study of Hybrid Methods for Multilingual PII Detection
- LiteLLM: Presidio PII masking guardrail
- Langfuse: masking
- GDPR Article 4: definitions (pseudonymisation, 4(5))
- EDPB: Guidelines 01/2025 on Pseudonymisation
- IAPP: EDPB publishes draft guidelines on pseudonymization
- Taylor Wessing: Analysis of the CJEU judgment in C-413/23 P (EDPS v SRB)
- Jones Day: CJEU clarifies scope of personal data in EDPS v SRB
- IAPP: leaked Council Digital Omnibus compromise drops the revised personal data definition
- Law Health Tech: Pseudonymisation under the GDPR and the Digital Omnibus (May 2026)
Frequently asked questions
How do I remove PII before sending data to an LLM?
Put a redaction gate between your application and the model API. Detect entities with a mix of patterns, checksums and an NER model (for example Microsoft Presidio), replace each one with a typed placeholder such as PERSON_1, send the redacted text, and optionally map the placeholders back in the answer. Apply the same gate to documents before indexing, and to logs and traces.
What is the difference between redaction, masking and tokenisation?
Redaction removes the value, masking replaces characters with a symbol, and tokenisation replaces the value with a stand-in that can be mapped back through a separate vault. Only tokenisation (or encryption) is reversible. Hashing is one-way but can be guessed for low-entropy values such as phone numbers.
Is pseudonymised data personal data under the GDPR?
For the party that holds the additional information needed to re-identify people, yes. Article 4(5) defines pseudonymisation as a safeguard, not as anonymisation. The CJEU held on 4 September 2025 (C-413/23 P) that for a recipient who has no reasonable means to re-identify, the same data may not be personal data, so the assessment depends on perspective.
Does Microsoft Presidio support German and Hungarian?
Presidio can run in other languages through its NLP engine configuration, and its documentation lists German-specific recognizers such as tax IDs and ID cards. I found no Hungarian-specific recognizers in the list, and spaCy has no trained Hungarian pipeline, so for Hungarian you need your own recognizers or a transformer model and your own evaluation.
Can PII detection guarantee that nothing leaks?
No. Presidio states that because it uses automated detection there is no guarantee that it finds all sensitive information. Names, free text, typos and context-dependent identifiers produce false negatives, so redaction should be one layer next to access control, contracts and EU data residency.
How do I test a PII redaction pipeline?
Build a labelled set in every language you serve, with realistic noise, and measure recall per entity type, because a missed name matters more than a false alarm. Add canary values that must never appear in logs or provider requests, run the suite in CI, and re-run it whenever the model, the language models or the recognizers change.