Skip to main content
Category: De-identification and PHI Types

Statistical De-identification

Also known as: Expert Determination Method, Statistical Method of De-identification
Simply put

Statistical de-identification is one of the ways health information can be stripped of its connection to a specific person so that it is no longer considered protected health information under the HIPAA Privacy Rule. Under this approach, a person with appropriate expertise uses statistical and scientific methods to determine that the risk of re-identifying an individual from the data is very small. It differs from simply removing a fixed list of identifiers, because it relies on expert analysis of the specific data set. Readers should verify the specific requirements against current HHS guidance.

Formal definition

Statistical de-identification refers, in the HIPAA context, to the Expert Determination method of de-identification recognized under the HIPAA Privacy Rule, as distinguished from the Safe Harbor method that removes an enumerated set of identifiers. De-identification is generally defined as any process of removing the association between a set of identifying data and the data subject. Under the statistical/expert determination approach, a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable applies such methods to determine that the risk is very small that the information could be used, alone or in combination with other reasonably available information, to identify an individual, and documents that determination. Practitioners typically must account for quasi-identifiers, data elements that, while not direct identifiers, may enable re-identification in combination. Note that HHS OCR guidance describes the methods and approaches to achieve de-identification in accordance with the Privacy Rule, and the precise standards, documentation expectations, and acceptable techniques should be confirmed against the current HHS guidance and applicable regulatory text. This entry addresses HIPAA de-identification only; other frameworks (for example, NIST publications on de-identifying datasets) use the general term de-identification and may apply different risk criteria, and state law may impose additional requirements.

Why it matters

De-identification matters because it is one of the primary mechanisms under the HIPAA Privacy Rule for allowing health information to be used and shared for research, analytics, public health, and other purposes without triggering the full set of Privacy Rule obligations. Once data is properly de-identified in accordance with the Privacy Rule, it is generally no longer considered protected health information, which changes what a covered entity or business associate may do with it. The Expert Determination (statistical) method offers an alternative to the Safe Harbor method's fixed list of identifiers, and it can be valuable when an organization needs to retain more analytically useful detail than Safe Harbor would permit.

The statistical method carries meaningful responsibility because the determination hinges on expert judgment about re-identification risk rather than a mechanical checklist. A key concern in practice is quasi-identifiers, data elements that are not direct identifiers but that, in combination with other reasonably available information, may allow an individual to be re-identified. As one industry resource put it, with quasi-identifiers 'the devil is in the details.' If the risk analysis is flawed or the data environment changes, information believed to be de-identified could carry more re-identification risk than intended.

Because the precise standards, acceptable techniques, and documentation expectations are matters of HHS guidance and regulatory text that are updated over time, organizations should not treat any single de-identification exercise as a permanent guarantee. Statistical de-identification reduces re-identification risk to a level the Privacy Rule describes as very small; it does not eliminate that risk entirely, and readers should confirm current requirements against HHS guidance before relying on a particular approach.

Who it's relevant to

Privacy Officers and Compliance Teams
Privacy officers evaluating whether to release data for research or analytics need to understand when the Expert Determination method is appropriate versus Safe Harbor, and what documentation the organization must retain to support a very small risk determination. They are typically responsible for confirming that the chosen approach aligns with current HHS guidance.
Statisticians and Data Scientists
The statistical method depends on a qualified expert applying generally accepted statistical and scientific principles. These professionals assess re-identification risk, address quasi-identifiers, and produce the analysis and determination that supports treating data as de-identified under the Privacy Rule.
Researchers and Data Analytics Users
Researchers who want more analytically useful data than Safe Harbor allows often rely on the statistical method. They should understand that de-identified data is no longer PHI under the Privacy Rule, but that this status depends on a sound expert determination and may be affected if the data environment changes.
Business Associates and Vendors Handling Data
Vendors that process or analyze health information on behalf of covered entities may be involved in de-identification workflows. Their obligations attach through business associate agreements, and they should be clear on whether they are permitted to de-identify data and how the resulting data may be used, noting that state law or other frameworks may impose additional requirements.

Inside Statistical De-identification

Expert Determination Method
One of the two de-identification methods recognized under the HIPAA Privacy Rule, sometimes referred to as the statistical method. It relies on a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable.
Very Small Risk Standard
Under the expert determination approach, the qualified expert must determine that the risk is very small that the information could be used, alone or in combination with other reasonably available information, to identify an individual who is a subject of the information. This is a risk-based standard rather than an absolute guarantee of anonymity.
Documentation of Methods and Results
The expert is generally expected to document the methods and results of the analysis that justify the determination. This documentation supports the covered entity's or business associate's ability to demonstrate that de-identification was performed appropriately.
Qualified Expert
The individual applying appropriate statistical or scientific principles. The Privacy Rule does not prescribe a specific degree or certification, but the person must have appropriate knowledge and experience for the analysis to be defensible.
Relationship to De-identified Data Status
Information that meets the de-identification standard is generally no longer considered PHI and falls outside many Privacy Rule restrictions on use and disclosure. This status applies only so long as the de-identification determination remains valid and the data has not been re-linked to identifiers.

Common questions

Answers to the questions practitioners most commonly ask about Statistical De-identification.

Does statistical de-identification mean the data is completely anonymous and can never be re-identified?
No. Statistical de-identification, sometimes called the expert determination method, generally reduces the risk of re-identification to a level a qualified expert determines to be very small, but it does not claim to make re-identification impossible. The standard is about small risk, not zero risk. Because residual risk can change as external data sources and re-identification techniques evolve, an expert determination is typically valid only in the context and conditions under which it was made, and readers should verify the applicable requirements against the current Privacy Rule text.
Is statistical de-identification the same thing as removing the identifiers listed under the Safe Harbor method?
No. These are two distinct methods for de-identifying PHI under the Privacy Rule. The Safe Harbor method involves removing a specified set of identifiers and having no actual knowledge that the remaining information could identify an individual. The statistical, or expert determination, method instead relies on a person with appropriate knowledge and experience applying generally accepted statistical and scientific principles to conclude that the re-identification risk is very small. They are alternative pathways, not the same process, and satisfying one does not require satisfying the other.
Who qualifies as an expert for the expert determination method?
The Privacy Rule generally describes a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable. It does not prescribe a specific certification or license. In most cases, organizations engage individuals with relevant statistical, mathematical, or data science expertise. Because the rule focuses on demonstrable knowledge and experience rather than a formal credential, you should confirm the current regulatory language and any applicable guidance when selecting an expert.
How should the expert's analysis and conclusions be documented?
As a practical matter, organizations typically retain documentation describing the methods and results of the analysis that justify the determination that the re-identification risk is very small. This documentation generally supports the ability to demonstrate compliance if questioned. The specific form and retention expectations should be confirmed against the current Privacy Rule text and any HHS OCR guidance, and organizations should note that state law may impose additional recordkeeping requirements.
Does an expert determination expire or need to be revisited?
The determination is generally tied to the specific data set, intended recipients, and conditions evaluated at the time. Because the risk of re-identification can change as new external data sources or techniques become available, many organizations treat an expert determination as time- and context-bound and reassess it when circumstances change. There is no single fixed expiration built into the concept, so you should confirm current guidance and consider re-evaluation when the data, its uses, or the surrounding data environment change materially.
Once data is de-identified through expert determination, is it still subject to HIPAA?
Information that has been properly de-identified under the Privacy Rule is generally no longer considered PHI, and the Privacy Rule's restrictions typically do not apply to it. However, this outcome depends on the de-identification being valid and maintained. Organizations should be cautious about re-linking data or combining it with other information in ways that could reintroduce identifiability, and should be aware that other frameworks, contractual obligations, or state laws may still impose requirements on the resulting data set beyond HIPAA.

Common misconceptions

Statistical (expert determination) de-identification guarantees that data can never be re-identified.
The standard requires the expert to conclude the risk of re-identification is very small, not zero. It is a risk-based determination, and no method can guarantee that all re-identification is prevented, particularly as external data sources change over time.
Expert determination and Safe Harbor are interchangeable ways to reach the same result.
They are two distinct methods under the Privacy Rule. Safe Harbor requires removal of a specified list of identifiers and no actual knowledge of remaining identifiability, while expert determination relies on statistical or scientific analysis and documented judgment. They have different requirements and are not simply substitutes for one another.
Once data is de-identified by an expert, the determination is permanent.
A determination reflects conditions and reasonably available information at the time it is made. Changes in available data or use of the information could affect re-identification risk, so covered entities and business associates should treat determinations as bounded by their stated assumptions and periodically reassess.

Best practices

Engage an individual with appropriate knowledge and experience in generally accepted statistical and scientific de-identification principles rather than assuming any staff member can perform the analysis.
Ensure the expert documents the methods and results supporting the very small risk determination, and retain that documentation to demonstrate compliance if questioned.
Define and record the assumptions underlying the determination, including what reasonably available external data was considered, since the conclusion is bounded by those conditions.
Reassess de-identification determinations periodically or when data uses, disclosures, or external data availability change materially, because a determination is not automatically permanent.
Do not treat expert determination as a substitute for Safe Harbor without understanding that each is a separate method with distinct requirements; select the approach that fits the specific dataset and use case.
Verify the current regulatory text and any applicable HHS OCR guidance, and account for additional obligations that state law or other frameworks may impose beyond the HIPAA Privacy Rule.