Patient History De-Identification with anonym.plus

Clear identifiers from medical, social, and family history while the story stays.

In simple terms, PII redaction is the on-device process of finding and masking personally identifiable information in a document before it is shared.

Patient-history de-identification removes identifiers from the medical, social, and family sections of a record. Those last two sections carry the hardest problems. Family history describes inherited characteristics. That makes it genetic data under UK GDPR Art. 4(13), and special category under Art. 9(1). Social history lists occupation, locality, and habits. UK GDPR Art. 4(1) treats each of those as an identifier in its own right. anonym.plus runs locally and gives both sections extra attention.

When this applies

A full history is invaluable for teaching and research. It is also the densest source of indirect identifiers in the whole record. A named relative with a rare inherited condition identifies a family, not just a patient.

How anonym.plus handles it

  1. Load the file into anonym.plus on your device.
  2. It scans the medical, social, and family sections.
  3. Named relatives, occupations, and places get flagged as indirect identifiers.
  4. Swap them so the clinical story still flows.
  5. Save the clean history on your device.

What you need to provide

Patient data entity types detected

Categoryanonym.plus entity typeExample
NamesPERSONpatient & relatives → [NAME]
LocationLOCATIONminer, Barnsley → [LOCATION]
DatesDATE_TIMEsmoker since 1998 → [DATE]
RelativesPERSONfather, dx 2010 → [RELATIVE]
Record IDsMEDICAL_RECORD_NUMBERMRN → [MRN]
NHS NumberUK_NHSNHS 943 476 5919 → [NHS_NO]

Compliance achieved

Anonymise patient histories offline — see plans & start free →

Limitations & cautions

The social and family sections are the hardest to clear fully. A rare occupation in a small town can re-identify someone after every direct identifier is gone. So can a named relative's rare inherited condition. Genetic information also identifies relatives who never consented to anything. Review these sections and apply the ICO motivated-intruder test for unusual cases.

Frequently asked questions

Why are the social and family sections riskier?

Because they are built from exactly the material UK GDPR Art. 4(1) calls out. That means factors specific to a person's economic, cultural, and social identity. An occupation, a town, a household, and a relative's diagnosis can point to one individual. No name or number has to be present at all.

Is family history really genetic data?

Where it describes inherited characteristics, yes. UK GDPR Art. 4(13) defines genetic data as data on inherited or acquired genetic traits. Those traits give unique information about a person's health, and Art. 9(1) makes them special category. It also means the text concerns relatives who are not your patient.

Are relatives' details removed too?

Yes. A named relative in the history is a data subject in their own right, so their name, their diagnosis, and their age at diagnosis are all flagged and swapped. The clinical pattern — that a first-degree relative was affected early — can be preserved without naming anyone.