Skip to main content
Category: De-identification and PHI Types

Re-identification

Also known as: De-anonymization, Data re-identification
Simply put

Re-identification is the process of matching previously de-identified or anonymized data back to the specific person it describes, often by combining it with other available information. It effectively reverses de-identification, undermining the privacy protection that de-identification was intended to provide. In the HIPAA context, this generally matters because de-identified health information is no longer treated as protected health information (PHI) unless it can be re-identified.

Formal definition

Re-identification is a process by which information is attributed to de-identified data in order to identify the individual to whom that data relates, typically by matching de-identified or anonymized records against other available or publicly accessible data sources. Under the HIPAA Privacy Rule, data that has been properly de-identified (through either the Expert Determination method or the Safe Harbor method) is generally not considered PHI and falls outside many Privacy Rule restrictions; however, the risk that such data could be re-identified is central to whether a given de-identification approach is adequate. Note that the Privacy Rule addresses re-identification codes and mechanisms as part of its de-identification standard, and a covered entity that itself re-identifies data it has de-identified may again be subject to Privacy Rule obligations. This entry describes the general concept; readers should verify specific de-identification and re-identification requirements against the current text of the HIPAA Privacy Rule, and note that the HIPAA Security Rule governs only ePHI and does not itself define de-identification. The term also carries distinct, unrelated meanings in fields such as computer vision (object re-identification), which are out of scope here.

Why it matters

Re-identification is central to how the HIPAA Privacy Rule treats de-identified health information. When data is properly de-identified, it generally falls outside many Privacy Rule restrictions and is no longer treated as PHI. That regulatory relief depends entirely on the assumption that the data cannot readily be matched back to an individual. Re-identification directly challenges that assumption: if de-identified records can be recombined with other available or publicly accessible information to pin down specific people, the privacy protection de-identification was meant to provide is undermined.

For compliance officers and privacy officers, this makes re-identification risk a practical question rather than an abstract one. The adequacy of a de-identification approach, whether pursued through the Expert Determination method or the Safe Harbor method, hinges on how vulnerable the resulting data is to being re-identified. Organizations that share, sell, or publish de-identified datasets should treat re-identification risk as an ongoing consideration, because the availability of external data sources that could be used for matching changes over time.

It is also important to recognize that re-identification can pull data back within the scope of the Privacy Rule. A covered entity that itself re-identifies data it previously de-identified may again be subject to Privacy Rule obligations with respect to that information. Because specific de-identification and re-identification requirements are set out in the current text of the Privacy Rule, readers should verify the precise standards and any applicable code or mechanism provisions against that text, and should note that state law or other frameworks may impose additional requirements.

Who it's relevant to

Privacy Officers and Compliance Officers
Those responsible for HIPAA privacy compliance need to understand re-identification risk when deciding whether a de-identification approach is adequate and whether resulting data can be treated as outside PHI. They should also account for the possibility that re-identifying previously de-identified data can bring it back within Privacy Rule obligations, and verify specific requirements against the current text of the Rule.
Data Analytics and Research Teams
Teams that work with de-identified datasets for analysis, research, or secondary use rely on the assumption that the data cannot be readily matched back to individuals. They should treat re-identification risk, particularly the availability of external data sources that could be used for matching, as a practical factor in how datasets are prepared and shared.
Organizations Sharing or Disclosing De-identified Data
Covered entities and their business associates that share, publish, or otherwise disclose de-identified health information should weigh re-identification risk, because that risk determines whether de-identification continues to hold up. State law or other frameworks may impose additional requirements beyond HIPAA, and specific de-identification standards should be confirmed against current guidance.

Inside Re-identification

Re-identification Defined
Re-identification is the process of matching de-identified data back to a specific individual, either by reversing the de-identification method or by combining the data with other available information. It is the counterpart concept to de-identification under the HIPAA Privacy Rule.
Relationship to De-identification Standards
The HIPAA Privacy Rule generally recognizes two de-identification methods: the Expert Determination method and the Safe Harbor method. Re-identification risk is the central concern these methods are designed to minimize, and the durability of de-identification depends on how well re-identification is guarded against.
Re-identification Code Provision
The Privacy Rule generally permits a covered entity to assign a code or other means of record identification to de-identified data so that it may be re-identified later, provided the code is not derived from or related to information about the individual and cannot itself be translated to identify the individual. The mechanism for such re-identification is typically restricted.
Scope Under HIPAA
Data that is properly de-identified generally falls outside the definition of protected health information and is therefore not subject to the Privacy Rule while it remains de-identified. Re-identification can return the data to PHI status, restoring the associated obligations. Readers should verify specific provisions against the current regulatory text.
Sources of Re-identification Risk
Risk can arise from combining de-identified datasets with external or publicly available information, from insufficient generalization or masking of quasi-identifiers, and from small population sizes where unique combinations of attributes point to individuals.

Common questions

Answers to the questions practitioners most commonly ask about Re-identification.

Does de-identified data automatically stay de-identified if it can never be linked back to an individual?
Not necessarily. De-identification reduces, but does not always permanently eliminate, the risk that data could be linked back to an individual. Re-identification refers to the process by which information previously stripped of identifiers is matched back to a specific person, often by combining the dataset with other available information. Because external data sources and analytical techniques evolve, the risk that a given dataset could be re-identified may change over time. Under HIPAA, this is why the two recognized de-identification approaches (the Expert Determination method and the Safe Harbor method) address re-identification risk in different ways, and why re-identification remains a relevant ongoing consideration rather than a one-time determination. You should verify the specific standards against the current text of the HIPAA Privacy Rule.
Is re-identifying data always a HIPAA violation?
Not in every case. Re-identification is a technical process, and whether it raises HIPAA concerns depends on who performs it, under what conditions, and how the resulting information is used. The HIPAA Privacy Rule contemplates that a covered entity may, in certain circumstances, assign a code or mechanism to enable information to be re-identified by the covered entity itself, subject to conditions in the regulatory text. Conversely, unauthorized re-identification by a recipient of de-identified data, or use of information contrary to applicable agreements, may implicate HIPAA, contractual, or other legal obligations. The characterization is fact-specific, and state law or other frameworks may impose additional requirements. Confirm the applicable conditions against the current regulation.
If we release data using the Safe Harbor method, how should we think about residual re-identification risk?
Safe Harbor is one of the two HIPAA de-identification methods and generally involves removing the categories of identifiers specified in the Privacy Rule and having no actual knowledge that the remaining information could be used to identify an individual. Even after meeting Safe Harbor requirements, some residual re-identification risk can theoretically remain, particularly as external datasets and techniques change. In most cases, organizations pair the method with sound data governance and periodic review rather than treating the initial determination as permanent. You should verify the current identifier list and the actual-knowledge standard against the applicable regulatory text.
How can we manage re-identification risk when sharing data with a recipient?
A common practice is to use data use agreements or other contractual terms that restrict recipients from attempting re-identification, from contacting individuals, or from combining the data with other sources for that purpose. These controls are typically administrative in nature and complement the technical de-identification itself. Because such measures are not part of the de-identification standard alone and do not by themselves guarantee that re-identification cannot occur, they are generally most effective as one layer within a broader governance program. The specific enforceability of such agreements may also depend on state law and other frameworks.
Who within an organization should evaluate re-identification risk?
Evaluating re-identification risk often involves privacy and security personnel, and where the Expert Determination method is used, an individual with appropriate knowledge of and experience applying generally accepted statistical and scientific principles for rendering information not individually identifiable. The precise qualifications and documentation expectations are set by the applicable regulatory standard, so organizations typically document the basis for their determinations and retain that documentation. Confirm the current requirements for expert qualifications and documentation against the applicable regulatory text.
How does re-identification risk relate to the HITRUST CSF or other control frameworks?
Control frameworks such as the HITRUST CSF may include controls addressing data handling, de-identification practices, and access restrictions that can support an organization's management of re-identification risk. However, HITRUST is a private organization and the HITRUST CSF is a certifiable control framework; certification is not a legal requirement and does not by itself establish HIPAA compliance or determine whether data is properly de-identified under the Privacy Rule. Organizations should treat such frameworks as supporting tools and verify their de-identification approach against the current HIPAA regulatory standards and, where applicable, the current HITRUST CSF version.

Common misconceptions

De-identified data can never be re-identified, so it carries no privacy risk.
De-identification reduces but does not necessarily eliminate re-identification risk. Data de-identified under Safe Harbor or Expert Determination can, in some circumstances, be re-identified by combining it with other available information. Expert Determination in particular addresses a very small, not zero, risk of re-identification.
Once data is de-identified it is permanently outside HIPAA regardless of what happens later.
While properly de-identified data generally is not PHI under the Privacy Rule, re-identifying that data can restore its PHI status and the corresponding HIPAA obligations. The exemption applies only while the data remains de-identified.
Any code used to link de-identified data back to individuals is acceptable.
The Privacy Rule generally permits a re-identification code, but the code must not be derived from or related to information about the individual and must not be independently translatable to identify the person. Codes that fail these conditions can undermine the de-identification itself.

Best practices

Treat de-identification as a risk-reduction measure rather than a guarantee, and periodically reassess re-identification risk as external data sources and analytic techniques evolve.
When using the Expert Determination method, retain documentation of the expert analysis and the basis for concluding that re-identification risk is very small; verify current requirements against the applicable regulatory text.
If assigning a re-identification code, ensure it is not derived from or related to individual information and cannot be translated to identify the person, and restrict access to any re-identification key.
Establish contractual and access controls so that recipients of de-identified data do not attempt re-identification, and consider whether state law or other frameworks impose additional obligations beyond HIPAA.
Recognize that re-identifying de-identified data can return it to PHI status, and apply appropriate Privacy Rule and, for electronic data, Security Rule safeguards once that occurs.
Confirm de-identification methods and re-identification provisions against the current HIPAA regulatory text and relevant HHS OCR guidance, since specific requirements may be updated over time.