Skip to main content
Knowledge · Data breach 7 min read Updated June 2026

PII & PHI data breach review

A breach review is a narrow, high-recall form of document review with a statutory clock attached. Here is what it asks of a review team — and the three ways it goes wrong.

A PII/PHI breach review is a kind of managed document review, but a narrow one. Where a litigation review asks whether a document is relevant or privileged, a breach review asks a smaller, more exacting question of every record: does this contain protected information, and if so, whose? The answer feeds a notification list, and the list is the entire point.

What counts as a reportable breach

Not every incident triggers a duty. The working definition most counsel use: unauthorized access to records containing PII or PHI, where the data can’t be confirmed encrypted or otherwise unreadable. The precise trigger varies by regulator and jurisdiction, which is itself part of the review’s job — the same record can be reportable in one state and not in another, and the analysis tracks each affected person back to their jurisdiction.

The review may also surface information that doesn’t fit the PII/PHI definition but still matters to the client — trade secrets, contractually protected data, material someone would rather not see leaked. Good protocols flag it; the client decides what to do with it.

What the review team actually does

The team extracts entities and the sensitive data points attached to them — names, contact details, account numbers, dates, health and financial attributes — from a corpus that is rarely clean. Breach data arrives as partially OCR’d scans, multilingual records, and image-based HR and healthcare forms that extraction tools can’t read. Much of the work is the part software hands back: a reviewer reading a scanned intake form because the OCR dropped the one field that makes the record reportable.

A project manager runs the protocol around them:

  • Calibrating the team to the client’s definition of in-scope data, and re-calibrating as edge cases surface.
  • Running keyword and pattern searches to catch PII the first pass may have missed — recall insurance, not a substitute for review.
  • Building tags for each category of sensitive data so the output is structured, not a pile of flagged documents.
  • Running QC continuously, during the project and after, against a recall target rather than a precision one.

The deliverable is a list, not a production

The output of a breach review is a deduplicated list of affected individuals and businesses, and for each one: full name as extracted, contact information, the specific protected data points exposed, and the jurisdiction that governs their notification. The client uses that list to discharge its notice obligations — in most jurisdictions, written notice to each affected person.

Everything about the workflow serves that deliverable. Which is why the failure modes are specific.

The three ways it goes wrong

  • Recall failure. A person who was in the data and never reached the list. This is the cardinal error — it’s the one with regulatory and reputational teeth — and it’s why breach QC is built to favor recall over precision at every layer.
  • Entity-resolution failure. The same individual appears across a dozen documents under inconsistent spellings. Resolve them wrong and you either notify one person three times or miss them entirely. The tooling helps; the judgment call on a genuine ambiguity stays with a reviewer.
  • Calendar failure. The most common of the three, and it has nothing to do with the substance. Most breach reviews fail because staffing started on day 35 of a 60-day clock. The fastest reviewers alive don’t recover two weeks you’ve already spent.

Where judgment stays in the loop

Extraction and pattern matching surface candidates well, and that share of breach work is increasingly automated — appropriately so. What doesn’t automate is the determination layered on top: whether a flagged record is reportable, whether two entries are one person, whether an image-based form holds something the model never saw. An attorney resolves those. The documents that reach a human in a breach review are, by construction, the ones the software couldn’t close — so the human work concentrates into judgment, and the team you put on it matters more, not less.

For how breach review compares to litigation review across pace, protocol, and tooling, see comparing data breach review with litigation document review. For the broader service model these reviews run inside, see managed document review.

Talk to our delivery team about a live breach matter — we scope the corpus and the clock before anything is committed.

Have a review that needs an attorney in the loop?

We scope, staff, and run document review end to end — calibrated teams, deployed under your brand or ours. Tell us about the matter and we'll walk the bench and the plan with you.