How Hexagone AI performs.
Every anonymization vendor claims high accuracy. Very few say on which dataset, against which metric, or compared to what. This page collects the evaluations we have run and published, with the methodology behind each and the papers and reports to read in full. If you find a mistake, tell us and we will correct the page.
Headline results
- of annotated personal values removed on PII-INDEXBENCH
- Measured over the whole dataset rather than a sample of it: 12,079 documents and 226,428 annotated values. The next system, an open-weights Qwen3-14B prompted as a detector, found 91.7%, and precision on the same run was 93.5%.
- of annotated personal values removed, across six languages
- 47,566 documents in English, French, German, Spanish, Italian and Dutch, and the best recall of the five systems compared. The spread between the strongest and weakest language is 1.3 points.
- re-identification rate at level 1, the lowest of the tools compared
- Azure scores 22% and a GPT-4.1 purifier prompt 26% on the same texts. At level 2, where identifiers are obfuscated: 27% against 33% and 58%.
- direct identifiers recovered from the anonymized texts
- No name, phone number or social security number was recoverable at either difficulty level. Azure and the GPT-4.1 prompt both leaked some.
Measured on: PII-INDEXBENCH v1.2, every split · September 2026
Measured on: AI4Privacy PII-Masking-300k, full validation split · September 2026
Measured on: RAT-Bench test set · Imperial College London
Measured on: RAT-Bench test set · Imperial College London
These figures describe performance on the public test sets named above, and they are not a guarantee about your own documents. A different corpus gives different numbers, which is why every result here names the corpus it was measured on.
The studies
- Full-dataset evaluation
PII-INDEXBENCH: 12,079 documents built to be hard
Every record through the production API, scored against the benchmark's own annotations alongside five other systems, one of which is Qwen3-14B prompted as a detector. The precision figure is taken apart against those same annotations further down the page.
Run by Hexagone AI, September 2026. Dataset PII-INDEXBENCH v1.2, CC BY 4.0.
- 97.9%
- recall, highest of six systems
- 93.5%
- precision
- Full-dataset evaluation
AI4Privacy: 47,566 documents in six languages
The same protocol as PII-INDEXBENCH, applied to the most cited public PII corpus and scored per language and per kind of personal data. It also comes with the caveats that follow from a corpus this widely distributed.
Run by Hexagone AI, September 2026. Public dataset, reproducible.
- 95.3%
- recall, highest of five systems
- 6
- languages, 1.3 points apart
- Independent benchmark
RAT-Bench: can an attacker still find the person?
A re-identification benchmark built at Imperial College London, and a better question than most evaluations ask. We ran our pipeline against it alongside Azure and a GPT-4.1 masking prompt, using the authors' own scoring criteria unchanged.
Benchmark by Krčo, Yao, Meeus and de Montjoye, Imperial College London, 2026
- 14%
- re-identified at level 1, the lowest compared
- 0
- direct identifiers recovered
- 0.83
- BLEU: how much text survives
- In-house evaluation, 2024
Detection rate against Microsoft Presidio
Our first published comparison, from November 2024, and the one that comes with a ten-page methodology PDF. Superseded in scope by the two full-dataset runs above, and kept here because its numbers still stand.
Run by Hexagone AI, November 2024. Public dataset, reproducible.
- 96.95%
- of entities detected, against 75.30%
- 93.34%
- on degraded text, against 65.89%
How to read any vendor benchmark. Ask three questions. Which dataset, and is it public? Which metric, and does it measure risk or just recall? Which systems were compared, and were they configured fairly? A number without those three answers, ours included, is marketing. Every study above carries all three.
Run it on your own documents.
A benchmark is someone else's data. One week free, no credit card, and the files never leave your machine, so there is nothing to risk in checking for yourself.