In one multi-turn conversation benchmark, Microsoft Presidio covered 0.79 of the gold PII characters in English conversations but only 0.42 in Russian ones, according to the PII-TRACE benchmark paper. That gap, not any headline F1 score, is the practical answer to the rules-versus-NER question: Presidio-style recognizer lists remain a reasonable choice for structured, well-specified patterns in English-heavy traffic, but the moment your product serves non-English users, regex-style coverage degrades quietly, and the cost of finding out shifts onto evaluation sets you have to build yourself.
The freshest evidence on multilingual PII extraction makes the same point with a large asterisk. Meddies-PII, a multilingual clinical de-identification framework from a September 2026 arXiv preprint, reports a mean F1 of 0.827 across fifteen external benchmarks against 0.658 for the strongest baseline. Those numbers are author-reported, from an unreplicated preprint, trained on an entirely synthetic dataset, and scoped to clinical text. They are a useful test case for how multilingual PII extraction is being built, not a certification of anything for your logs or prompts. Everything in this article carries some version of that caveat.
Read the provenance before the numbers
Every result cited here comes from an academic benchmark, not production traffic. That matters more than usual for PII redaction, because the failure that actually hurts you is a missed name in a customer’s support message, not a benchmark aggregate.
Meddies-PII is the clearest case. The framework trains a BIOES token classifier on a synthetic multilingual clinical dataset, and the authors are unusually direct about the limits: the dataset “is entirely synthetic and may not capture all linguistic variation, documentation practices, and annotation ambiguity found in real-world clinical records,” and its deterministic validation gates “do not directly measure the semantic naturalness of every generated document,” per the Meddies-PII paper. So the 0.827 mean F1 tells you the approach trains well on data with known structure; it does not tell you how the model behaves on your analytics events or LLM prompt logs, and no cited study ran there.
The same discipline applies in the other direction. The independent evaluation of OpenAI’s Privacy Filter covered 32 benchmarks across 14 languages and 5 domains, which is genuinely broad, but it is still benchmark-scoped, and none of these papers measures serving cost or compliance outcomes at a given deployment’s scale. Latency and compute do appear in one place: PRvL reports inference latency and training GPU hours, but only for the LLM redactors it evaluates, and nothing in these papers prices Presidio against a trained detector or a hybrid in a like-for-like deployment. What the evidence can settle is directional: where rule-based coverage breaks, where trained models win and lose, and what a hybrid buys. That is enough to make a decision, if you treat the final validation step as yours.
Where recognizer lists break first: the language edge
Rule-based redaction fails across languages for a structural reason, not an implementation one. The PRvL study of LLM redaction capabilities puts it plainly: rule-based methods “rely on deterministic pattern matching or regular expression matching. While fast and interpretable, their use case is extremely limited as PII redaction requires language and context understanding. Patterns rarely generalize across languages or domains. For e.g., U.S. phone numbers are completely different from those in the UK.” Every locale-specific pattern you ship is a small piece of maintenance debt denominated in a language your team may not read.
PII-TRACE measures the consequence directly. In multi-turn LLM conversations, Presidio’s gold-character coverage fell from 0.79 on English to 0.42 on Russian. A small trained model did not escape either: OpenMed’s 434M-parameter detector saw recall drop from 0.94 on English to 0.69 on Korean. OpenAI’s Privacy Filter, by contrast, “maintained recall across different scripts” in the same benchmark. Two readings follow. First, cross-lingual robustness is a property you have to buy deliberately, whether through training data or architecture; a small clinical-domain detector does not get it for free. Second, the degradation is quiet. Nothing throws an error when a Russian surname slides past a recognizer tuned for Anglo names. The redaction pipeline reports success, the log ships, and the PII is in your analytics warehouse.
This is the strongest argument in the whole comparison, because it is a measurement of the thing developers actually default to, failing at the thing products actually do as they grow.
The quieter failure: precision collapse at perfect recall
Coverage is not the only axis where rules misbehave. A robustness study of PII detection systems dissects Presidio’s date recognizer and finds it “produces excessive false positives by aggressively matching any date like numeric pattern regardless of semantic context, tagging standalone numbers like “6574” and “79327” as dates” (Mind the Gap). The over-matching “causes DATE_OF_BIRTH precision to drop to 0.492 on Set B despite maintaining perfect recall.”
Perfect recall with 0.492 precision sounds acceptable until you price it. Every false positive in a redaction pipeline is text destroyed or distorted before it reaches a log, a prompt, or an analyst. If your redactor tags order IDs, timestamps, and version strings as dates of birth, downstream consumers stop trusting the data, and someone files a ticket asking for the redaction to be relaxed, which is how real PII leaks get reintroduced. Precision failures create organizational pressure that recall failures never do, precisely because recall failures are invisible.
The decision implication: when you evaluate any approach, rules or model, score precision and recall separately per entity type and per language. An aggregate F1 hides exactly the asymmetry that will define your operational experience.
What a trained multilingual detector buys, and where Presidio still wins
The strongest head-to-head evidence is the independent 32-benchmark evaluation of OpenAI’s Privacy Filter (OPF), a 1.5B-parameter model that converts an autoregressive language model into a bidirectional PII detector. In zero-shot evaluation, “OPF outperforms Presidio on all PII-annotated datasets (F1 gains: +0.42 AI4Privacy, +0.12 Nemotron, +0.17 Gretel/Kiji, +0.19/+0.17 SPY medical/legal) and also on CoNLL-2003 newswire NER (+0.18).”
That looks conclusive until the same paper reports the reversal: “On MultiCoNER general NER, Presidio (0.497) surpasses OPF (0.397) because its spaCy PERSON recognizer covers the NER class without domain adaptation.” Two lessons hide in that result. The obvious one is that “trained model beats rules” is benchmark-dependent; on general person-name detection, the tuned pipeline still won, and 0.497 is nobody’s idea of a finished solution either way. The subtler one is that Presidio is not purely regex. Part of its strength comes from a bundled statistical recognizer, so the honest framing was never “rules versus neural” but “which recognizers, tuned for which entity types and languages.”
So what does the trained detector actually buy you? Cross-script recall that holds (PII-TRACE), broad zero-shot gains on PII-annotated text (the 32-benchmark study), and freedom from hand-maintaining per-locale pattern lists. What it costs: a 1.5B-parameter serving footprint instead of a regex library, and a model whose behavior on your domain you cannot read from a pattern file. Interpretability is the rule-based approach’s real remaining advantage, and the PRvL paper names it explicitly alongside speed. When a regulator or a customer asks why a specific string was redacted, “this deterministic pattern matched” is a complete answer. Model outputs require a different kind of audit story.
The third option: regex plus LLM, measured
The RECAP study of hybrid multilingual PII detection quantifies the combination directly: a hybrid of regex patterns with prompt-based LLMs “outperforms the state-of-the-art baselines, NER by 82% and LLM by 17% (weighted F1-score)” (RECAP). The low-resource result is the one worth remembering: in Polish (pl_PL), RECAP reached F1 0.60 against 0.26 for the NER baseline, a 130.77% relative improvement, and 0.49 for the zero-shot LLM (RECAP).
The mechanism is intuitive once the failure modes above are clear. Regex catches the structured, high-precision patterns (ID formats, phone numbers, card numbers) cheaply and interpretably. The LLM handles contextual, free-text PII where language understanding is the bottleneck. Each component covers the other’s weakest region.
But the LLM-in-the-loop path has its own ceiling, and it lands exactly on the multilingual axis. A study on scalable multilingual PII annotation reports that “current large language models exhibit substantial limitations in detecting locale-specific PII, achieving a recall of no more than 65% in the case of Chinese” (multilingual annotation study), with missed detections and incorrect type classifications driven by limited digital representation of those entities. So the hybrid is not “add an LLM, solve languages.” It is “add an LLM, move the bottleneck,” and the bottleneck still has to be measured locale by locale.
Mapping the decision
The evidence supports a three-way choice, gated on four variables you can measure about your own traffic.
| Your situation | Evidence-backed choice | Why |
|---|---|---|
| English-heavy traffic, structured PII (IDs, phones, dates) | Keep Presidio-style recognizers | Fast, interpretable, deterministic; the generalization failure has not been triggered yet (PRvL) |
| Person names and free-text PII in English | Test, do not assume: Presidio beat a 1.5B model on MultiCoNER (0.497 vs 0.397) | The bundled spaCy recognizer already covers the class (OPF evaluation) |
| Growing non-English traffic (Cyrillic, CJK, other scripts) | Add a trained multilingual detector or hybrid before launch | Presidio coverage fell 0.79 → 0.42 English→Russian; a 434M model fell 0.94 → 0.69 English→Korean (PII-TRACE) |
| High missed-PII risk, tolerance for false positives | Model or hybrid, then watch precision per entity type | Rules kept perfect date recall but precision collapsed to 0.492; models have their own error shapes (Mind the Gap) |
| Low-resource languages, budget for LLM calls | Hybrid regex+LLM | RECAP: F1 0.60 vs 0.26 NER and 0.49 zero-shot LLM on Polish (RECAP) |
| Locale-specific PII in underrepresented locales (e.g., Chinese) | Any approach, plus mandatory per-language validation | LLM recall capped at no more than 65% on locale-specific Chinese PII (multilingual annotation study) |
Deployment cost is the row the table cannot fill in honestly. Presidio is cheap and CPU-friendly; OPF is a 1.5B-parameter model; RECAP pays per LLM call. No cited paper measures serving cost or latency for these approaches against each other in a like-for-like deployment: PRvL reports inference latency and training GPU hours, but only for its own LLM redactors. Measure that axis in your own stack before committing.
The step nobody can skip: build the per-language evaluation set
Every source here converges on the same operational requirement, and it is the least glamorous one. The gains from switching approaches are real but benchmark-scoped; the failures are locale-specific; and the locale-specific ceiling applies to rules, small models, and LLMs alike. So the gating artifact for any switch is a labeled evaluation set per language you serve, built from text that resembles what your pipeline actually sees.
This does not have to be enormous. It has to be honest: sampled from real traffic patterns (with appropriate care, since it contains the PII you are trying to detect), covering your actual entity mix, and scored per entity type with precision and recall separated. Without it, you are importing someone else’s benchmark conditions. With it, the decision table above becomes checkable: run Presidio, run your candidate model or hybrid, and compare on your data. I would treat any redaction switch that skips this step as a rename, not an improvement, because the evidence says the failure modes differ by language and entity type in ways no aggregate benchmark predicts for your corpus.
Also budget for maintenance either way. Recognizer lists drift as locales and formats accumulate; models drift as traffic composition changes. The choice is which maintenance burden your team is equipped to carry.
The verdict, and its limits
Keep Presidio-style recognizers where your traffic is English-heavy and your PII is structured; they are fast, interpretable, and in at least one measured case (MultiCoNER person names, 0.497 vs 0.397) still ahead of a much larger trained model. Before serving non-English users, add a trained multilingual detector or a regex-plus-LLM hybrid, because the measured coverage drop, 0.79 to 0.42 from English to Russian in multi-turn conversations, is the kind of failure that produces no alerts. Gate the switch on per-language evaluation sets you build yourself, and remember the LLM ceiling: no more than 65% recall on locale-specific Chinese PII (multilingual annotation study) means “we added a model” is not a compliance answer.
What this evidence cannot establish is transfer. The headline multilingual number, Meddies-PII’s 0.827 mean F1, is an author-reported preprint result from clinical text trained on synthetic data whose own authors caution it may not capture real-world linguistic variation. The OPF comparison is third-party and benchmark-scoped. No cited study evaluated any of these approaches on general product text such as logs, prompts, or analytics events, and none compares serving cost or latency across Presidio, a trained detector, and a hybrid under matching conditions. The decision framework is durable; every specific number in it is a 2025–2026 preprint result subject to revision, and none of them stand in for a measurement on your own data.
Frequently Asked Questions
How does Presidio’s PII detection performance change when moving from English to Russian text?
In one multi-turn conversation benchmark, Microsoft Presidio covered 0.79 of the gold PII characters in English conversations but only 0.42 in Russian ones, according to the PII-TRACE benchmark paper.
What is the main operational requirement before switching PII redaction approaches?
So the gating artifact for any switch is a labeled evaluation set per language you serve, built from text that resembles what your pipeline actually sees.
How does a hybrid regex and LLM approach perform on low-resource languages like Polish?
in Polish (pl_PL), RECAP reached F1 0.60 against 0.26 for the NER baseline, a 130.77% relative improvement, and 0.49 for the zero-shot LLM (RECAP).

Join the discussion
Share a useful perspective or ask a question about this article.