What was measured
For each of 46 languages the corpus holds a letter and a running-prose document with one person's data — name, street, house number, postcode, city, date of birth, e-mail, phone and, where the country has them, an IBAN and a national identifier (16–20 values per language) — plus one document with no personal data at all. The values are generated and pass their own check digits; they belong to no real person. A value counts as found only if it is covered completely. The second value set ("holdout") keeps the documents and replaces the values, so a repair written for one set of strings cannot pass as a general one.
Measured on 2026-09-26 with the product's own detection path, on the development build (desktop commit
7bae21d). That build is not yet released; the version you can download today was not
measured for this page.
Every language
Values fully found at 85%. 15 languages reach every value in both sets. With 16 to 20 values per language, a single value moves a language by about six points, so read the spread, not a single row.
| Language | Values (set 1) | Values (set 2) |
|---|---|---|
| Arabic | 13/18 | 14/18 |
| Bengali | 10/16 | 10/16 |
| Bulgarian | 20/20 | 20/20 |
| Chinese | 16/18 | 14/18 |
| Croatian | 18/20 | 17/20 |
| Czech | 20/20 | 20/20 |
| Danish | 16/18 | 16/18 |
| Dutch | 18/20 | 18/20 |
| English | 17/20 | 15/20 |
| Estonian | 20/20 | 19/20 |
| Finnish | 18/20 | 18/20 |
| French | 20/20 | 20/20 |
| German | 20/20 | 20/20 |
| Greek | 17/18 | 17/18 |
| Hebrew | 18/18 | 17/18 |
| Hindi | 12/16 | 12/16 |
| Hungarian | 17/18 | 18/18 |
| Icelandic | 20/20 | 20/20 |
| Indonesian | 16/16 | 15/16 |
| Irish | 18/18 | 18/18 |
| Italian | 19/20 | 20/20 |
| Japanese | 14/16 | 15/16 |
| Korean | 15/16 | 15/16 |
| Latvian | 20/20 | 20/20 |
| Lithuanian | 20/20 | 20/20 |
| Macedonian | 16/18 | 16/18 |
| Malay | 16/16 | 14/16 |
| Maltese | 17/18 | 18/18 |
| Norwegian | 18/18 | 18/18 |
| Persian | 16/16 | 16/16 |
| Polish | 20/20 | 18/20 |
| Portuguese | 20/20 | 20/20 |
| Romanian | 20/20 | 19/20 |
| Russian | 15/16 | 16/16 |
| Serbian | 19/20 | 19/20 |
| Slovak | 20/20 | 16/20 |
| Slovenian | 19/20 | 20/20 |
| Spanish | 20/20 | 20/20 |
| Swahili | 14/16 | 13/16 |
| Swedish | 20/20 | 20/20 |
| Tagalog | 12/16 | 10/16 |
| Thai | 11/16 | 11/16 |
| Turkish | 19/20 | 20/20 |
| Ukrainian | 18/18 | 18/18 |
| Urdu | 13/16 | 13/16 |
| Vietnamese | 16/16 | 16/16 |
The threshold matters
The threshold is a setting in every preset. Detection does not fall off gradually above 85% — it drops in one step, because many recognizers score exactly 0.85. A higher threshold redacts less text by mistake and misses more. The app shows a note when a preset is set above 85%.
| Threshold | Set 1 | Set 2 | Characters redacted in the 46 documents without personal data |
|---|---|---|---|
| 85% | 791/846 (93.5%) | 779/846 (92.1%) | 1631 |
| 86% | 421/846 (49.8%) | 419/846 (49.5%) | 768 |
| 90% | 410/846 (48.5%) | 411/846 (48.6%) | 752 |
| 95% | 370/846 (43.7%) | 375/846 (44.3%) | 670 |
The last column is the cost of caution: in documents that contain no personal data at all, 1631 characters were redacted anyway at 85%.
Limits of these figures
- They are measured on test texts written for this purpose, not on customer documents. Your documents, layouts and names can score higher or lower.
- They are measured on typed text. Pages read by text recognition (scanned documents, photos) are less accurate and are not part of these figures.
- For Indonesian, recognition of place names was measured only on real place names; invented or very rare names were not recognised in our test.
- Whether the models recognise invented names of people and places as well as real ones is not established by this measurement, and no figure for it is published here.
- Per language the sample is small (16–20 values). The overall figure is the more stable one.