Detection accuracy, measured

At the default threshold of 85%, the engine fully found 93.5% of the personal data in our test documents (791 of 846 values in 46 languages), and 92.1% (779 of 846) on a second set of values. By language the figure ranges from 62.5% to 100%. These are test texts, not customer documents.

What was measured

For each of 46 languages the corpus holds a letter and a running-prose document with one person's data — name, street, house number, postcode, city, date of birth, e-mail, phone and, where the country has them, an IBAN and a national identifier (16–20 values per language) — plus one document with no personal data at all. The values are generated and pass their own check digits; they belong to no real person. A value counts as found only if it is covered completely. The second value set ("holdout") keeps the documents and replaces the values, so a repair written for one set of strings cannot pass as a general one.

Measured on 2026-09-26 with the product's own detection path, on the development build (desktop commit 7bae21d). That build is not yet released; the version you can download today was not measured for this page.

Every language

Values fully found at 85%. 15 languages reach every value in both sets. With 16 to 20 values per language, a single value moves a language by about six points, so read the spread, not a single row.

LanguageValues (set 1)Values (set 2)
Arabic13/1814/18
Bengali10/1610/16
Bulgarian20/2020/20
Chinese16/1814/18
Croatian18/2017/20
Czech20/2020/20
Danish16/1816/18
Dutch18/2018/20
English17/2015/20
Estonian20/2019/20
Finnish18/2018/20
French20/2020/20
German20/2020/20
Greek17/1817/18
Hebrew18/1817/18
Hindi12/1612/16
Hungarian17/1818/18
Icelandic20/2020/20
Indonesian16/1615/16
Irish18/1818/18
Italian19/2020/20
Japanese14/1615/16
Korean15/1615/16
Latvian20/2020/20
Lithuanian20/2020/20
Macedonian16/1816/18
Malay16/1614/16
Maltese17/1818/18
Norwegian18/1818/18
Persian16/1616/16
Polish20/2018/20
Portuguese20/2020/20
Romanian20/2019/20
Russian15/1616/16
Serbian19/2019/20
Slovak20/2016/20
Slovenian19/2020/20
Spanish20/2020/20
Swahili14/1613/16
Swedish20/2020/20
Tagalog12/1610/16
Thai11/1611/16
Turkish19/2020/20
Ukrainian18/1818/18
Urdu13/1613/16
Vietnamese16/1616/16

The threshold matters

The threshold is a setting in every preset. Detection does not fall off gradually above 85% — it drops in one step, because many recognizers score exactly 0.85. A higher threshold redacts less text by mistake and misses more. The app shows a note when a preset is set above 85%.

ThresholdSet 1Set 2Characters redacted in the 46 documents without personal data
85%791/846 (93.5%)779/846 (92.1%)1631
86%421/846 (49.8%)419/846 (49.5%)768
90%410/846 (48.5%)411/846 (48.6%)752
95%370/846 (43.7%)375/846 (44.3%)670

The last column is the cost of caution: in documents that contain no personal data at all, 1631 characters were redacted anyway at 85%.

Limits of these figures

  • They are measured on test texts written for this purpose, not on customer documents. Your documents, layouts and names can score higher or lower.
  • They are measured on typed text. Pages read by text recognition (scanned documents, photos) are less accurate and are not part of these figures.
  • For Indonesian, recognition of place names was measured only on real place names; invented or very rare names were not recognised in our test.
  • Whether the models recognise invented names of people and places as well as real ones is not established by this measurement, and no figure for it is published here.
  • Per language the sample is small (16–20 values). The overall figure is the more stable one.