Skip to main content
Category: De-identification and PHI Types

De-identified Information

Also known as: De-identified Data, De-identification
Simply put

De-identified information is health information that has had identifiers removed or altered so it can no longer reasonably be linked back to a specific individual. Because it can no longer identify a person, it is generally treated differently under HIPAA than protected health information. However, in some cases de-identified information can be re-identified if a code, algorithm, or pseudonym is used to link it back to the original records.

Formal definition

Under the HIPAA Privacy Rule, de-identified information is health information from which identifiers have been removed or manipulated such that there is no reasonable basis to believe it can be used to identify an individual. HHS guidance generally describes two recognized approaches to de-identification: an Expert Determination method and a Safe Harbor method involving removal of specified identifiers (readers should verify the specific identifiers and requirements against the current HHS de-identification guidance and the applicable CFR text). Information that meets the Privacy Rule's de-identification standard is generally not treated as PHI and therefore falls outside many Privacy Rule use and disclosure restrictions, though this is a specific regulatory standard that differs from general or colloquial uses of the term 'de-identification.' Note that HHS acknowledges de-identified information may be re-identified through a code, algorithm, or pseudonym; the conditions under which a re-identification code may be assigned and retained are governed by the Privacy Rule and should be confirmed against current guidance. This scope is limited to HIPAA; state law, the HITECH Act, or other frameworks may impose additional or differing requirements. Related but distinct concepts, such as a 'limited data set,' are not synonymous with de-identified information.

Why it matters

De-identification is one of the primary mechanisms that allows health information to be used and shared for purposes such as research, analytics, and product development without triggering the full set of use and disclosure restrictions the HIPAA Privacy Rule imposes on protected health information (PHI). Once information meets the Privacy Rule's de-identification standard, it is generally no longer treated as PHI, which meaningfully expands what covered entities and business associates can do with it. This makes correct de-identification a high-stakes compliance decision: information that is believed to be de-identified but does not actually meet the standard remains PHI and continues to carry the Privacy Rule's obligations.

The stakes are heightened by the fact that de-identification is a specific regulatory standard, not a colloquial one. Simply stripping obvious identifiers such as names or addresses does not necessarily satisfy the HIPAA de-identification standard, and organizations that treat any partially masked dataset as 'de-identified' may be misclassifying PHI. HHS recognizes only defined approaches to de-identification, and readers should confirm the specific methods and identifier lists against current HHS de-identification guidance and the applicable CFR text.

Re-identification risk further complicates the picture. HHS acknowledges that de-identified information can be re-identified using a code, algorithm, or pseudonym assigned to link data back to the original records. The conditions under which such a re-identification code may be assigned and retained are governed by the Privacy Rule, and mishandling those codes can undermine the de-identified status of the data. Because state law, the HITECH Act, or other frameworks may impose additional or differing requirements, organizations should not assume that meeting the HIPAA standard resolves all obligations.

Who it's relevant to

Privacy Officers
Privacy officers are typically responsible for determining whether data qualifies as de-identified under the Privacy Rule and therefore falls outside many use and disclosure restrictions. They should ensure the organization applies a recognized method (Expert Determination or Safe Harbor) correctly and does not treat merely masked data as de-identified. Verifying the specific method requirements against current HHS guidance is essential.
Researchers and Data Analytics Teams
Teams that use health information for research or analytics often rely on de-identified information to work with data outside the full scope of Privacy Rule restrictions. They should confirm that datasets genuinely meet the de-identification standard and understand that de-identified information is not the same as a limited data set, which carries different requirements.
Covered Entities and Business Associates
Covered entities and business associates that share data for secondary purposes need to correctly classify information, since data that fails to meet the de-identification standard remains PHI and retains its associated obligations. They should also manage any re-identification code, algorithm, or pseudonym in accordance with the Privacy Rule and confirm handling conditions against current guidance.
Compliance and Legal Professionals
Compliance and legal professionals should recognize that the HIPAA de-identification standard is a specific regulatory concept that differs from colloquial uses of the term. They should also account for the possibility that state law, the HITECH Act, or other frameworks may impose additional or differing requirements beyond HIPAA.

Inside De-identified Information

Health Information Stripped of Identifiers
De-identified information is health information from which identifiers that could reasonably be used to identify an individual have been removed or manipulated, such that it is no longer considered protected health information (PHI) under the HIPAA Privacy Rule and is generally not subject to its restrictions on use and disclosure.
Safe Harbor Method
One of two de-identification approaches recognized under the Privacy Rule. It generally involves removing a specified set of identifiers relating to the individual and their relatives, employers, and household members, and requires that the covered entity has no actual knowledge that the remaining information could be used to identify the individual. Practitioners should verify the current list of identifiers against the applicable regulatory text.
Expert Determination Method
The alternative de-identification approach, under which a person with appropriate knowledge and experience applying generally accepted statistical and scientific principles and methods determines that the risk of identifying an individual is very small, and documents the methods and results of that analysis.
Relationship to the Privacy Rule
De-identification is a concept defined under the HIPAA Privacy Rule, which governs PHI in all forms. Because de-identified data is no longer PHI, it typically falls outside the scope of the Privacy Rule, the Security Rule, and the Breach Notification Rule, though other laws or contractual terms may still apply.
Re-identification Considerations
The Privacy Rule addresses re-identification codes or mechanisms; a covered entity may, in certain circumstances, assign a code to allow later re-identification, provided the code is not derived from or related to information about the individual and is not otherwise disclosed.

Common questions

Answers to the questions practitioners most commonly ask about De-identified Information.

Is de-identified information still considered protected health information (PHI) under HIPAA?
No. Information that has been properly de-identified in accordance with the HIPAA Privacy Rule's standards is no longer considered PHI, and the Privacy Rule's restrictions on use and disclosure generally no longer apply to it. However, this only holds when the de-identification is done correctly under one of the recognized methods; information that is merely partially masked or informally stripped of some identifiers may still qualify as PHI. Readers should verify their approach against the current regulatory text and applicable HHS OCR guidance.
Does removing patient names and other obvious identifiers automatically make data de-identified?
Not necessarily. Simply removing names or a few obvious identifiers does not meet the HIPAA de-identification standard. The Privacy Rule recognizes specific methods for achieving de-identification, and each has its own requirements. Data that retains a combination of remaining elements may still allow an individual to be identified, in which case it would generally continue to be treated as PHI. State law or other frameworks may also impose additional considerations.
What methods does the HIPAA Privacy Rule recognize for de-identifying information?
The Privacy Rule generally recognizes two approaches: a method relying on a qualified expert's determination that the risk of re-identification is very small, and a method based on removing a specified set of identifiers relating to the individual and their relatives, employers, or household members, combined with having no actual knowledge that the remaining information could identify the individual. Organizations should confirm the specific requirements of each method against the current regulatory text before relying on either.
Who can perform an expert determination for de-identification?
The expert-determination approach generally relies on a person with appropriate knowledge of and experience with accepted statistical and scientific principles and methods for rendering information not individually identifiable. This person applies such methods to determine that the risk of re-identification is very small and documents that determination. The specific qualifications and documentation expectations should be confirmed against current HHS OCR guidance.
Can de-identified data be re-linked to individuals later?
The Privacy Rule generally permits assigning a code or other means of record identification to allow information to be re-identified, provided the code is not derived from or related to the individual's information and cannot otherwise be translated to identify the individual, and provided the covered entity does not use or disclose the code for other purposes or disclose the mechanism for re-identification. Organizations should verify the precise conditions against the current regulatory text.
How does de-identified information differ from a limited data set?
These are distinct concepts under the Privacy Rule and should not be conflated. De-identified information, when properly created, is generally no longer PHI and falls outside the Privacy Rule's use and disclosure restrictions. A limited data set, by contrast, still retains certain identifiers, remains subject to the Privacy Rule, and typically requires a data use agreement before it may be shared for permitted purposes. Readers should confirm the specific requirements for each against the current regulation.

Common misconceptions

De-identified information is the same as a limited data set.
These are distinct concepts under the Privacy Rule. A limited data set still contains certain identifiers (such as some dates and geographic elements) and remains PHI subject to the Privacy Rule, generally requiring a data use agreement. De-identified information, by contrast, is no longer PHI when properly de-identified.
Simply removing names and Social Security numbers makes data de-identified.
Proper de-identification under the Privacy Rule requires either the full Safe Harbor process (removal of all specified identifiers plus no actual knowledge of residual identifiability) or a documented Expert Determination. Removing only a few obvious identifiers typically does not meet either standard, and the remaining data may still be PHI.
Once data is de-identified, there are no further legal or contractual obligations of any kind.
While properly de-identified information is generally outside the scope of the HIPAA Privacy and Security Rules, state law, the HITECH Act, other frameworks, or contractual terms may impose additional restrictions. Re-identification safeguards and organizational policies may also continue to apply.

Best practices

Select and document which de-identification method you are relying on, Safe Harbor or Expert Determination, and ensure the process is fully carried out and recorded, since documentation supports defensibility.
When using Safe Harbor, verify the current list of required identifiers against the applicable regulatory text and confirm you have no actual knowledge that the residual information could identify an individual.
When using Expert Determination, engage a qualified individual with appropriate statistical and scientific expertise and retain their documented methods and conclusions regarding re-identification risk.
Manage any re-identification codes carefully, ensuring codes are not derived from information about the individual and are not disclosed in a way that undermines de-identification.
Distinguish de-identified information from a limited data set in your policies, since a limited data set remains PHI and typically requires a data use agreement.
Confirm whether state law, the HITECH Act, other frameworks, or contractual terms impose obligations beyond HIPAA before treating de-identified data as fully unrestricted, and periodically reassess re-identification risk.