Pass HIPAA Audits: Remove 18 Safe Harbor Identifiers On Device

· 19 min read

Pass HIPAA Audits: Remove 18 Safe Harbor Identifiers On Device

Isometric privacy gate for clinical data

Safe Harbor de-identification means removing 18 specific identifier categories under 45 CFR 164.514(b)(2) and confirming you have no actual knowledge the remaining data could still identify someone. Do both correctly, and the dataset stops being protected health information under the Privacy Rule. Skip the knowledge test, or mishandle even one identifier category, and you have not achieved Safe Harbor no matter how many fields you scrubbed. HHS guidance treats both requirements as inseparable.


TL;DR:

  • Removing all 18 HIPAA identifiers and having no actual knowledge of identifyability are both required for Safe Harbor to be valid; failing either invalidates the de-identification.
  • For geographic data, ZIP codes below 20,000 population must be replaced with “000,” and dates more precise than the year must be stripped, with ages over 89 aggregated to 90 or older.
  • Unstructured clinical notes require NLP and pattern matching tools with confidence scoring to detect hidden identifiers, especially in images or pasted texts, and manual review is essential for low-confidence matches.
  • The catch-all category can include internal codes, unique characteristics, or rare diagnoses, which may function as identifiers if specific enough to trace back to an individual, necessitating controlled, separate re-identification keys.
  • When dataset needs detailed dates or geography, or involves small populations, Expert Determination is preferable to Safe Harbor, requiring documented statistical analysis by a qualified expert.

Medscrub
medscrub.ai
De-Identify Patient Data On Device
MedScrub transforms complex patient data into automated insights while anonymizing sensitive information on your machine.
Explore MedScrub

Table of Contents

The 18 Safe Harbor identifiers: the complete list and exact removal rules

The list in 45 CFR 164.514(b)(2)(i) is specific, and the specificity is the point. Vague instructions like “remove identifying information” invite inconsistent judgment calls. This list does not. Each category has a defined removal rule, and several have exceptions that trip up teams who assume “remove” always means “delete entirely.”

The scope matters as much as the list itself. Safe Harbor requires removing these identifiers not just for the patient, but for relatives, employers, and household members named in the record. A progress note that mentions “patient’s employer, Springfield General Hospital, where her husband John also works as a surgeon” carries three separate identifier problems, none of which belong to the patient directly.

Here is the full list with the operational rule for each:

  • Names. Full names, nicknames, and initials tied to an individual, their relatives, employers, or household members must go. Watch for names embedded in narrative text, not just dedicated name fields.
  • Geographic subdivisions smaller than a state. Street addresses, cities, counties, precincts, and most ZIP codes must be removed, with a specific carve-out for 3-digit ZIP prefixes covered separately below.
  • All elements of dates directly related to an individual, except year. Birth dates, admission dates, discharge dates, death dates, and any date tied to a specific service must lose the month and day. The year alone can stay, with a special rule for ages over 89.
  • Telephone numbers. Any phone number linked to the individual or their household, including numbers appearing in call logs or referral notes.
  • Fax numbers. Same treatment as phone numbers.
  • Email addresses. Personal or work email addresses tied to the individual.
  • Social Security numbers. No partial masking that leaves a memorable fragment.
  • Medical record numbers. Internal MRNs must be removed or replaced with a code that meets the re-identification safeguards discussed later.
  • Health plan beneficiary numbers. Insurance ID numbers, including subscriber and dependent identifiers.
  • Account numbers. Billing account numbers and financial account identifiers tied to care.
  • Certificate or license numbers. Professional license numbers when they identify the patient (rare, but it happens with provider-patients).
  • Vehicle identifiers and serial numbers, including license plates. Relevant in trauma, EMS, and workers’ compensation records.
  • Device identifiers and serial numbers. Implant serial numbers, pacemaker IDs, and durable medical equipment tags.
  • Web URLs. Any URL that could route back to an individual’s record or portal account.
  • IP addresses. Increasingly common in telehealth and patient portal logs.
  • Biometric identifiers, including fingerprints and voiceprints. Retinal scans and voice recordings used for identity verification fall here too.
  • Full-face photographs and comparable images. Photos that could allow visual identification, including some clinical photography.
  • Any other unique identifying number, characteristic, or code. The catch-all category, covered in detail later in this guide.

Two implementation notes save teams from the most common mistakes. First, the “characteristic” language in the last category is broader than most compliance teams initially assume. It is not limited to numbers. Second, removal does not always mean deletion. Several categories permit structured transformation instead of blank redaction, which is exactly where the ZIP code and date rules come in.

What does the ‘actual knowledge’ test require?

Removing all 18 identifiers is necessary but not sufficient. HHS guidance requires that the covered entity also have no actual knowledge that the remaining information could be used, alone or combined with other reasonably available information, to identify the individual. This is a subjective, fact-specific test, and it is where a surprising number of technically compliant datasets fail.

Consider a dataset that removes every one of the 18 identifiers but includes a case note reading “the only pediatric liver transplant patient treated at this facility in 2025.” No name, no MRN, no ZIP code. Anyone with local knowledge of the facility’s transplant volume could identify that patient in seconds. That is actual knowledge territory, and it is exactly the trap that industry guidance flags as a recurring failure mode, particularly with small-population datasets or rare diagnoses.

A practical residual-risk review should check for:

  • Rare diagnoses or procedures with very small local incidence
  • Small geographic or demographic subpopulations (a single-employer health plan, a rural clinic with few patients matching an age and sex combination)
  • High-profile or publicly known patients whose treatment episodes made local news
  • Combinations of quasi-identifiers, such as occupation plus age plus treatment date, that narrow a population to one person
  • Free-text narrative that describes unique circumstances, even without naming anyone directly

Pro Tip: Run a “small cell” check on any demographic combination in your dataset. If fewer than roughly 5 to 10 people in the source population share a given combination of age, sex, ZIP prefix, and diagnosis, treat that row as a residual-risk flag requiring manual review before release.

Document the review itself, not just its conclusion. Auditors want to see who ran the residual-risk assessment, what criteria they checked, what they found, and what they did about it. A one-line sign-off saying “reviewed, no issues found” carries far less weight than a log showing the specific small-population and rare-diagnosis checks performed, the tool or method used, and the reviewer’s name and date. This governance layer is arguably as important to a defensible Safe Harbor claim as the redaction work itself.

How do you apply the ZIP code, date, and age rules?

These three rules generate more implementation questions than any other part of Safe Harbor, mostly because they are the identifiers most likely to carry real analytic value.

The 3-digit ZIP rule allows you to retain the first three digits of a ZIP code, but only if the combined population of all ZIP codes sharing that 3-digit prefix exceeds 20,000 according to the most recent Census data. If the combined population falls at or below that threshold, the entire ZIP code, including the 3-digit prefix, must be changed to “000.” HHS publishes the list of low-population 3-digit ZIP prefixes that fail this test, and it changes only when Census figures are updated, so check it against the current release rather than an old copy sitting in a compliance folder.

Dates follow a simpler but stricter rule: remove every element more precise than the year. Birth dates become birth years. Admission and discharge dates lose month and day. A death date recorded as “March 14, 2025” must become “2025.” The one specific numeric exception in the entire Safe Harbor list is ages over 89, which must be aggregated into a single “90 or older” category rather than reported as a specific age like 92 or 95, per implementation guidance built around the HHS rule.

A few concrete examples clarify where teams go wrong:

  • Allowed: “Patient born in 1958, admitted in 2025, ZIP prefix 902 (population over 20,000).”
  • Not allowed: “Patient born 03/14/1958, admitted 06/02/2025, ZIP 90210.”
  • Allowed: “Age 90+.”
  • Not allowed: “Age 93,” even without a name attached.

When a research protocol genuinely requires specific dates or fine-grained geography, for example calculating exact time-to-event intervals or mapping disease clusters at the ZIP+4 level, Safe Harbor is the wrong tool. That is a signal to move to Expert Determination, which is covered in detail further down.

How do you detect identifiers in unstructured clinical notes?

Structured fields are the easy part of Safe Harbor. A Social Security number field, a phone number field, a date-of-birth field: these are deterministic, pattern-matchable, and a regular expression will catch nearly all of them reliably. Free-text clinical notes are where de-identification programs actually fail, because identifiers hide inside narrative sentences with no consistent format.

A defensible detection pipeline generally runs in four stages:

  1. Deterministic pattern matching for high-structure identifiers: Social Security numbers, phone numbers, email addresses, IP addresses, and URLs all follow recognizable formats that regular expressions catch with high reliability.
  2. NLP-based entity detection for names, places, and organizations embedded in narrative text, since a clinician’s free-text note might read “referred by Dr. Alvarez at Riverside Family Practice” with no structured field marking any of those as identifiers.
  3. Confidence scoring on every flagged entity, so the pipeline can distinguish a near-certain match from an ambiguous one rather than treating all detections as equally reliable.
  4. Human review for low-confidence flags, routing anything under a conservative threshold to a person rather than auto-redacting or, worse, auto-ignoring it.

That confidence-threshold step deserves specific attention. Detection guidance from practitioner projects recommends tagging every automated match with a confidence score and capturing both the model version and the threshold chosen directly in the audit log. That way, if a name slips through six months later, you can reconstruct exactly what the pipeline saw, what score it assigned, and why it fell above or below the review line.

Common failure modes cluster in predictable places. Physician signatures at the bottom of notes often escape both regex and NLP detection because they sit outside the narrative body. Scanned documents and embedded images, including handwritten notes or photographed insurance cards, defeat text-based detection entirely and need separate image-review workflows. Free-text fields where a clinician pastes in a referral letter or an email thread carry a much higher identifier density than the structured chart around them, and they are exactly where automated tools under-perform.

Clinical note detection and manual review process

Pro Tip: Treat any scanned PDF, embedded image, or copy-pasted external document inside a chart as an automatic manual-review flag. No automated text pipeline reliably catches identifiers inside image content, and assuming it does is the single most common gap in otherwise solid de-identification programs.

The output of this pipeline should be a redaction log, not just a redacted document: what was flagged, at what confidence level, what a human reviewer decided, and when.

What counts as the catch-all identifier category?

The 18th category, “any other unique identifying number, characteristic, or code,” is deliberately broad, and it is where compliance teams most often underestimate their exposure. This is not a narrow legal formality tacked onto the end of a list. It exists specifically to catch identifiers nobody thought to name in 1996 when the rule was written.

Common items that fall under the catch-all include internal case numbers, clinical trial participant IDs, UUIDs generated by an EHR system, custom patient-tracking codes used in a research registry, and even distinctive characteristics described in narrative text, such as a rare genetic condition affecting only a handful of known patients nationally. Any of these can function as a unique identifier even without a name attached, if the code or characteristic is specific enough to trace back to one person.

There is one important, narrow exception. 45 CFR 164.514© permits a covered entity to assign a code or other means of record identification to allow de-identified information to be re-identified later, provided the code is not derived from or related to information about the individual and the covered entity does not disclose the key or the derivation method. In practice, this means you can maintain a secure, reversible mapping between de-identified records and source patients for legitimate internal purposes, as long as that mapping lives in a separate, access-controlled system and the code itself carries no information that could be reverse-engineered.

Getting this exception right requires documenting a few specific controls:

  • The re-identification key or crosswalk table is stored separately from the de-identified dataset, with distinct access permissions
  • The code assigned is generated independently (a random UUID, for example) rather than derived from any PHI element like a birth date or Social Security number fragment
  • Access to the key is logged and restricted to a defined, documented list of authorized personnel
  • The disclosure policy for the key itself is written down, not assumed

Safe Harbor, Expert Determination, or Limited Data Set: which one fits?

Safe Harbor is the right tool when your analytics needs tolerate the loss of dates, fine geography, and ages above 89, and when the dataset does not represent a small or identifiable population. It is checklist-based, requires no statistician, and scales well across large datasets with routine structured fields. That predictability is exactly why it is the default choice for most operational reporting and general research extracts.

Expert Determination, the other de-identification path recognized under the Privacy Rule, works differently. A qualified statistical expert applies accepted methods to determine that the risk of re-identification is very small, then documents the analysis, methods, and assumptions used to reach that conclusion. This path suits datasets where dates, precise ages, or granular geography carry real analytic value, since preserving those fields typically gives the best balance between data utility and re-identification risk, as long as the expert’s report backs it with defensible methodology.

A Limited Data Set under a Data Use Agreement is the third path, and it occupies a middle ground. It permits retention of dates and certain geographic detail (though not full addresses) for specific purposes like research, public health, or health care operations, but it requires a signed DUA with the recipient restricting further use and disclosure. It remains PHI under HIPAA, just PHI that can move under a defined, restricted, agreement rather than needing full de-identification.

A few decision rules simplify the choice:

  • If the dataset needs exact dates or address-level geography, Safe Harbor is not viable. Choose Expert Determination or a Limited Data Set with a DUA.
  • If the population is small enough that even properly de-identified data raises actual-knowledge concerns, Expert Determination’s statistical rigor is usually the safer path.
  • If the data will move to an external research partner for a defined project with clear use restrictions, a Limited Data Set with a signed DUA is often faster than either de-identification method.
  • If the dataset is large, routine, and does not require precise temporal or geographic detail, Safe Harbor’s checklist approach is the most efficient choice.

An Expert Determination report should document the statistical methods applied, the specific re-identification risk threshold used, the assumptions about what data an adversary might reasonably access, and the expert’s named qualifications. Without that documentation, the determination itself is difficult to defend later.

What documentation do auditors expect to see?

A Safe Harbor claim that exists only in someone’s memory does not survive an audit. The process needs to leave a paper trail at every stage, and the artifacts matter as much as the redaction work itself.

The core workflow runs in six steps:

  1. Scope the dataset. Identify every field, note type, and attachment that could contain one of the 18 identifier categories.
  2. Run automated detection. Apply deterministic pattern matching and NLP-based entity detection across structured fields and free text.
  3. Route low-confidence flags to human review. Anyone below your confidence threshold gets a manual look before the record moves forward.
  4. Apply the ZIP, date, and age transformations. Convert 3-digit ZIP prefixes below the population threshold to “000,” strip date elements finer than year, and aggregate ages over 89.
  5. Conduct the residual-risk review. Check for small-population flags, rare diagnoses, and any actual-knowledge concerns before finalizing.
  6. Obtain sign-off. A named reviewer confirms the dataset meets both Safe Harbor requirements and logs the date of that confirmation.

The artifacts an auditor will look for include:

  • Removal and redaction logs showing what was flagged and what action was taken
  • Reviewer notes from the manual review stage, including reasoning for any edge-case decisions
  • Expert Determination reports, when that path was used instead of or alongside Safe Harbor
  • Signed Data Use Agreements for any Limited Data Set disclosures
  • A change-control history showing when de-identification rules or tooling versions were updated

At the field level, a useful audit log captures who performed the review, when it happened, which method or tool was used, the software or model version, and the confidence threshold applied to automated detections. That level of specificity is what separates a defensible compliance record from a vague assurance that “the data was de-identified.”

Practitioner perspective: fitting de-identification into real clinical workflows

The tension nobody resolves cleanly is analytic fidelity versus re-identification risk. Every date you strip, every ZIP code you zero out, is a data point a researcher or care team might have wanted. Safe Harbor’s rigidity is also its usefulness: it removes the guesswork, at the cost of granularity that sometimes genuinely mattered.

The practical fix is piloting on a representative sample before rolling a pipeline out fully. Run it against real clinical notes, not clean synthetic examples, and measure both false positives (over-redaction that destroys useful context) and false negatives (missed identifiers, which carry the compliance risk). Tune the confidence threshold from there rather than guessing at it up front.

On-device processing changes the risk calculus in a way that manual review alone cannot. When de-identification happens locally rather than through a cloud upload, the exposure window shrinks and the audit trail becomes a byproduct of the workflow instead of an extra task layered on top of it. That is the difference between de-identification as a compliance chore and de-identification as something that just happens while a clinician works.

— Clint

Automating Safe Harbor redaction without moving PHI off-site

This tool offers an alternative to manual chart-by-chart redaction by syncing with EMR systems and performing anonymization locally on the clinician’s machine rather than routing patient data through an external server.

Medscrub

Every recommendation in this guide, the confidence-scored detection, the human review step for low-confidence flags, the audit log capturing who reviewed what and when, is exactly the kind of infrastructure that gets skipped when teams rely on manual chart review alone. This workflow includes automated chart preparation, care-gap tracking, and on-device PHI de-identification that produces the documentation trail auditors expect, without a single record leaving the clinician’s device. Developers building on top of chart data can also work with a reversibly de-identified FHIR proxy, which mirrors the secure re-identification-code pattern permitted under §164.514© without exposing raw PHI. See how one practice put this to work in the eSpiral case study, or visit the clinician product page to start a trial.

This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.

Sources

FAQ

What are the 18 identifiers defined by HIPAA?

The 18 identifiers cover names, geographic subdivisions smaller than a state, dates (except year), phone and fax numbers, email addresses, Social Security numbers, medical record and health plan numbers, account numbers, license numbers, vehicle and device identifiers, URLs, IP addresses, biometric identifiers, full-face photographs, and any other unique identifying number or characteristic under 45 CFR 164.514(b)(2)(i). These apply to the individual as well as their relatives, employers, and household members.

How do you de-identify PHI under HIPAA?

You can use either the Safe Harbor method, removing all 18 identifiers and confirming no actual knowledge of residual identifiability, or Expert Determination, where a qualified expert statistically certifies the re-identification risk is very small. HHS guidance treats both as equally valid paths to the same legal outcome.

Does PHI include all 18 identifiers?

Protected health information includes any of the 18 identifiers when linked to health data, but data stops being PHI once all 18 are properly removed or transformed and the actual-knowledge test is satisfied. A single remaining identifier, even an internal case number under the catch-all category, keeps the dataset classified as PHI.

What are examples of de-identified data?

A dataset showing “Patient born 1958, admitted 2025, ZIP prefix 902, diagnosis code X” with no name, contact information, or exact dates qualifies as de-identified under Safe Harbor, provided the 902 ZIP prefix covers a population over 20,000. A tool like Medscrub applies these transformations automatically as part of its on-device processing, generating the redaction log alongside the output.

When should you use Expert Determination instead of Safe Harbor?

Choose Expert Determination when your analysis genuinely needs exact dates, fine-grained geography, or specific ages above 89, since Safe Harbor strips all three by default. It also fits small or identifiable populations where even properly de-identified data raises actual-knowledge concerns that Safe Harbor’s rigid checklist cannot resolve on its own.

Related articles