Re-identification
Re-identification is the process of matching previously de-identified or anonymized data back to the specific person it describes, often by combining it with other available information. It effectively reverses de-identification, undermining the privacy protection that de-identification was intended to provide. In the HIPAA context, this generally matters because de-identified health information is no longer treated as protected health information (PHI) unless it can be re-identified.
Re-identification is a process by which information is attributed to de-identified data in order to identify the individual to whom that data relates, typically by matching de-identified or anonymized records against other available or publicly accessible data sources. Under the HIPAA Privacy Rule, data that has been properly de-identified (through either the Expert Determination method or the Safe Harbor method) is generally not considered PHI and falls outside many Privacy Rule restrictions; however, the risk that such data could be re-identified is central to whether a given de-identification approach is adequate. Note that the Privacy Rule addresses re-identification codes and mechanisms as part of its de-identification standard, and a covered entity that itself re-identifies data it has de-identified may again be subject to Privacy Rule obligations. This entry describes the general concept; readers should verify specific de-identification and re-identification requirements against the current text of the HIPAA Privacy Rule, and note that the HIPAA Security Rule governs only ePHI and does not itself define de-identification. The term also carries distinct, unrelated meanings in fields such as computer vision (object re-identification), which are out of scope here.
Why it matters
Re-identification is central to how the HIPAA Privacy Rule treats de-identified health information. When data is properly de-identified, it generally falls outside many Privacy Rule restrictions and is no longer treated as PHI. That regulatory relief depends entirely on the assumption that the data cannot readily be matched back to an individual. Re-identification directly challenges that assumption: if de-identified records can be recombined with other available or publicly accessible information to pin down specific people, the privacy protection de-identification was meant to provide is undermined.
For compliance officers and privacy officers, this makes re-identification risk a practical question rather than an abstract one. The adequacy of a de-identification approach, whether pursued through the Expert Determination method or the Safe Harbor method, hinges on how vulnerable the resulting data is to being re-identified. Organizations that share, sell, or publish de-identified datasets should treat re-identification risk as an ongoing consideration, because the availability of external data sources that could be used for matching changes over time.
It is also important to recognize that re-identification can pull data back within the scope of the Privacy Rule. A covered entity that itself re-identifies data it previously de-identified may again be subject to Privacy Rule obligations with respect to that information. Because specific de-identification and re-identification requirements are set out in the current text of the Privacy Rule, readers should verify the precise standards and any applicable code or mechanism provisions against that text, and should note that state law or other frameworks may impose additional requirements.
Who it's relevant to
Inside Re-identification
Common questions
Answers to the questions practitioners most commonly ask about Re-identification.