Clinical Letter Batch Anonymisation with anonym.plus

Clean a whole folder of letters in one local run with steady labels.

In simple terms, PII redaction is the on-device process of finding and masking personally identifiable information in a document before it is shared.

Batch anonymisation removes personal data from an entire folder of clinical letters in one run, with a shared alias map so the same patient maps to the same label everywhere. That is the shape a research or teaching corpus needs. UK GDPR Art. 89(1) requires safeguards, including data minimisation and pseudonymisation, for processing for scientific research or statistical purposes, and DPA 2018 s.19 sets the UK conditions. anonym.plus does the whole run on your device.

When this applies

Building a corpus means cleaning hundreds of letters at once, with the same patients and clinicians recurring across them. Note the trade-off: a retained shared map makes the corpus pseudonymous under UK GDPR Art. 4(5), so it is still personal data. Drop the map and Recital 26 applies instead.

How anonym.plus handles it

  1. Point anonym.plus at the folder on your machine.
  2. It scans each letter for the full set of direct and indirect identifiers.
  3. A shared map keeps recurring people steady across the run.
  4. Review the summary and fix any low-confidence flags.
  5. Save the clean corpus on your device.

What you need to provide

Patient data entity types detected

Categoryanonym.plus entity typeExample
NamesPERSONpatient across files → [PATIENT_1]
NamesPERSONrecurring clinician → [PROVIDER_1]
DatesDATE_TIMEletter dates → [DATE]
Record IDsMEDICAL_RECORD_NUMBERMRNs → [MRN_n]
LocationLOCATIONaddresses → [ADDRESS]
ContactEMAIL_ADDRESSemails → [EMAIL]

Compliance achieved

Anonymise clinical letters offline — see plans & start free →

Limitations & cautions

Mixed folders (some scanned, some native) depend on OCR for the image files, so review low-confidence flags from scans. A shared map keeps results consistent, but it is also a re-identification key: while you hold it, UK GDPR Art. 4(5) means the corpus is still personal data and every duty still applies.

Frequently asked questions

How does batch mode keep one patient steady?

A shared map records each detected person once, so the same patient or clinician receives the same label in every file across the folder. That is what makes a corpus analysable — you can still count how often one person recurs without knowing who they are.

Does the shared map keep the corpus in UK GDPR scope?

Yes, while you hold it. UK GDPR Art. 4(5) defines pseudonymisation as processing where data can no longer be attributed to a person without additional information kept separately — and pseudonymous data remains personal data. Destroy the map and the Recital 26 anonymity test becomes available.

What governance does a research corpus need?

That depends on the study, not the tool. UK GDPR Art. 89(1) and DPA 2018 s.19 set the safeguards, Schedule 1, Part 1, paragraph 4 supplies the Art. 9 condition, and the UK Policy Framework for Health and Social Care Research sets out the HRA approvals route. Anonymising the corpus reduces the burden but does not replace the approvals.