Statistical De-identification
Statistical de-identification is one of the ways health information can be stripped of its connection to a specific person so that it is no longer considered protected health information under the HIPAA Privacy Rule. Under this approach, a person with appropriate expertise uses statistical and scientific methods to determine that the risk of re-identifying an individual from the data is very small. It differs from simply removing a fixed list of identifiers, because it relies on expert analysis of the specific data set. Readers should verify the specific requirements against current HHS guidance.
Statistical de-identification refers, in the HIPAA context, to the Expert Determination method of de-identification recognized under the HIPAA Privacy Rule, as distinguished from the Safe Harbor method that removes an enumerated set of identifiers. De-identification is generally defined as any process of removing the association between a set of identifying data and the data subject. Under the statistical/expert determination approach, a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable applies such methods to determine that the risk is very small that the information could be used, alone or in combination with other reasonably available information, to identify an individual, and documents that determination. Practitioners typically must account for quasi-identifiers, data elements that, while not direct identifiers, may enable re-identification in combination. Note that HHS OCR guidance describes the methods and approaches to achieve de-identification in accordance with the Privacy Rule, and the precise standards, documentation expectations, and acceptable techniques should be confirmed against the current HHS guidance and applicable regulatory text. This entry addresses HIPAA de-identification only; other frameworks (for example, NIST publications on de-identifying datasets) use the general term de-identification and may apply different risk criteria, and state law may impose additional requirements.
Why it matters
De-identification matters because it is one of the primary mechanisms under the HIPAA Privacy Rule for allowing health information to be used and shared for research, analytics, public health, and other purposes without triggering the full set of Privacy Rule obligations. Once data is properly de-identified in accordance with the Privacy Rule, it is generally no longer considered protected health information, which changes what a covered entity or business associate may do with it. The Expert Determination (statistical) method offers an alternative to the Safe Harbor method's fixed list of identifiers, and it can be valuable when an organization needs to retain more analytically useful detail than Safe Harbor would permit.
The statistical method carries meaningful responsibility because the determination hinges on expert judgment about re-identification risk rather than a mechanical checklist. A key concern in practice is quasi-identifiers, data elements that are not direct identifiers but that, in combination with other reasonably available information, may allow an individual to be re-identified. As one industry resource put it, with quasi-identifiers 'the devil is in the details.' If the risk analysis is flawed or the data environment changes, information believed to be de-identified could carry more re-identification risk than intended.
Because the precise standards, acceptable techniques, and documentation expectations are matters of HHS guidance and regulatory text that are updated over time, organizations should not treat any single de-identification exercise as a permanent guarantee. Statistical de-identification reduces re-identification risk to a level the Privacy Rule describes as very small; it does not eliminate that risk entirely, and readers should confirm current requirements against HHS guidance before relying on a particular approach.
Who it's relevant to
Inside Statistical De-identification
Common questions
Answers to the questions practitioners most commonly ask about Statistical De-identification.