TAR Review Set Anonymization with anonym.plus

Clear PII before a machine-review model trains, all on your own device.

In simple terms, PII redaction is the on-device process of finding and masking personally identifiable information in a document before it is shared.

TAR-set anonymization is the removal of personal data from documents fed to technology-assisted review. FRCP 26(b)(1) makes discovery proportional, which is the reason ranking is used at all. anonym.plus keeps the corpus on your own device.

When this applies

A ranking model reads every item in the collection, privileged ones included, before any lawyer does. Send that corpus to a hosted service and you have disclosed the whole collection. Proportionality under Rule 26(b)(1) is why you rank.

How anonym.plus handles it

  1. Point anonym.plus at the document set on your device.
  2. Local OCR reads any scanned items in the set.
  3. The tool flags names, contacts, and IDs across files.
  4. Use steady labels so relevance signals survive.
  5. Replace or mask each confirmed value.
  6. Save the clean set for the ranking workflow.

What you need to provide

PII entity types detected

Categoryanonym.plus entity typeExample
NamesPERSONcustodian name → [PERSON_n]
ContactEMAIL_ADDRESSsender email → [EMAIL]
DatesDATE_TIMEdoc date → [DATE]
IdentifiersUS_SSNSSN → [SSN]
LocationLOCATIONaddress → [ADDRESS]
AccountUS_BANK_NUMBERaccount no. → [ACCOUNT]

Compliance achieved

Anonymize TAR document sets offline — see plans & start free →

Limitations & cautions

Anonymizing before machine ranking can shift how a model reads context. Steady labels keep most signals, but test recall on a control set first. Free-text clues that survive redaction still need a human pass on responsive items.

Frequently asked questions

Do courts accept technology-assisted review?

Yes. Da Silva Moore v. Publicis Groupe (S.D.N.Y. 2012) was the first opinion to approve predictive coding, and later rulings followed. Courts look at the validation protocol, not the brand of software.

Does anonymizing first hurt recall?

It can shift context. Steady labels preserve most of the signal, since one custodian keeps one token throughout. Validate on a control set and measure elusion before you rely on the ranking.

Why clear the data before the model trains?

A hosted model ingests the whole collection at once, privileged items included. Clearing first, on your own hardware, keeps that exposure at zero.