47,566 synthetic emails, forms, chats and JSON payloads in six languages: English, French, German, Spanish, Italian and Dutch. It is the most widely cited public PII corpus, and that popularity cuts both ways, because it has been public long enough that any large model has probably read it during training.
Same protocol as the PII-INDEXBENCH run, same five systems, same scorer, over the whole validation split: 47,566 documents and 293,712 annotated values.
95.3%
- Recall
- Share of the annotated personal values the output removed, the highest of the five systems here as well, with the OpenAI Privacy Filter next at 82.9%.
80.2%
- Precision
- Lower than on PII-INDEXBENCH, for a reason worth reading: three quarters of what costs us here is never annotated as personal data anywhere in the corpus. The section below takes the figure apart.
1.3 pts
- Spread between languages
- Between the best language (French, 95.9%) and the weakest (Dutch, 94.7%) on recall.
Hexagone AI, language by language
Hexagone AI results per language on AI4Privacy| Language | Documents | Recallshare of personal data found | Precision |
|---|
| English | 7,942 | 95.9% | 79.9% |
|---|
| French | 8,390 | 95.9% | 81.9% |
|---|
| German | 8,100 | 95.2% | 78.9% |
|---|
| Spanish | 7,803 | 95.6% | 81.1% |
|---|
| Italian | 7,878 | 94.6% | 80.1% |
|---|
| Dutch | 7,453 | 94.7% | 79.4% |
|---|
French comes out ahead of English here, which is what you would expect from a system built in France, and which is the opposite of the usual pattern for tools in this category.
Detection coverage by kind of personal data on AI4Privacy| Kind of personal data | Spans | Hexagone AI | OpenMed | OPF | Presidio | GLiNER |
|---|
| NamesPERSON | 52,298 | 97.8% | 93.7% | 95.2% | 36.3% | 83.4% |
|---|
| Street addressesADDRESS | 41,024 | 95.4% | 70.8% | 97.4% | no label | 56.9% |
|---|
| PlacesLOCATION | 36,318 | 91.9% | 94.0% | 40.2% | 37.3% | 41.5% |
|---|
| Phone numbersPHONE_NUMBER | 12,601 | 97.0% | 85.2% | 99.4% | 64.1% | 60.5% |
|---|
| Email addressesEMAIL_ADDRESS | 15,883 | 99.8% | 99.5% | 99.6% | 99.7% | 58.3% |
|---|
| Social security numbersSSN | 15,791 | 98.0% | 86.8% | 99.7% | 71.2% | 23.6% |
|---|
| Dates and timesDATE_TIME | 49,285 | 91.2% | 92.8% | 53.1% | 50.1% | 62.5% |
|---|
| National ID numbersNATIONAL_ID | 16,693 | 96.4% | no label | 95.5% | no label | no label |
|---|
| Passport numbersPASSPORT_NUMBER | 15,570 | 96.0% | no label | 95.9% | 46.6% | 29.7% |
|---|
| Driving licencesDRIVER_LICENSE | 15,207 | 97.4% | no label | 99.0% | 74.7% | 18.9% |
|---|
| IP addressesIP_ADDRESS | 14,110 | 99.0% | 99.4% | 99.1% | 99.6% | no label |
|---|
| Passwords, keys and tokensPASSWORD | 10,303 | 90.3% | 74.6% | 97.7% | no label | no label |
|---|
Best in each row is bold. « No label » means the system has no category for that kind of data at all, so it is neither scored nor penalised on that row, which is the generous reading of a missing vocabulary. The shading of each cell follows the value in it.