If you de-identify health data under HIPAA, that does not mean it is anonymous under GDPR. That is the main point. In cross-border healthcare work, the same dataset can be outside HIPAA but still inside GDPR.

Here’s the short version:

  • HIPAA gives you defined paths: Safe Harbor or Expert Determination
  • GDPR asks a harder question: can someone still be identified by means reasonably likely to be used?
  • Pseudonymized data is still regulated under GDPR
  • Limited data sets are still PHI under HIPAA
  • For U.S.-EU data sharing, I would usually treat HIPAA de-identified data as still regulated in the EU unless a separate anonymization review says otherwise

That gap matters for:

HIPAA vs GDPR Anonymization Standards: Key Differences at a Glance

HIPAA vs GDPR Anonymization Standards: Key Differences at a Glance

Learn Data Anonymization Techniques for GDPR & HIPAA Compliance

Quick Comparison

Criteria HIPAA GDPR
Main test Remove 18 identifiers or show very low re-identification risk Ask whether a person can still be identified by any means reasonably likely to be used
Data outside the law De-identified data Anonymous data only
Pseudonymization May still leave data regulated, depending on key and context Still personal data
Limited data set Still PHI Usually still personal data
Cross-border effect May allow U.S. use with fewer HIPAA limits May still require lawful basis, transfer tool, and security steps

In other words: HIPAA is more rule-based, while GDPR is more context-based. If you work with U.S. and EU health data, I’d build to the stricter GDPR standard, document the re-identification risk, and review the dataset again when the data, recipient, or outside matching sources change.

HIPAA de-identification: Safe Harbor, Expert Determination, and limited data sets

Under the HIPAA Privacy Rule, PHI counts as de-identified when it no longer identifies a person and there is no reasonable basis to think it could be used to identify them.[2] Once data is de-identified through either allowed method, it is no longer PHI and is exempt from most HIPAA Privacy Rule requirements, including authorization and minimum necessary.[2] That choice matters. The method you use changes what data you can keep and how you can share it.

Safe Harbor: removing 18 specified identifiers

Safe Harbor, codified at 45 C.F.R. §164.514(b)(2), works like a checklist. If you remove all 18 specified identifiers - for the person, their relatives, employers, and household members - the data is treated as de-identified, as long as the covered entity has no actual knowledge that the remaining data could be used, by itself or with other reasonably available information, to identify the person.[1][9][17]

The list includes the obvious items, like names, Social Security numbers, and medical record numbers. But it also reaches items people often miss, such as IP addresses, device serial numbers, URLs, biometric identifiers like fingerprints and voiceprints, and full-face photographs.[19] It also requires removal of all date elements tied to a person except the year, and ages 90 and above must be grouped into a single 90-or-older category.

That "no actual knowledge" standard is not just a box to check. It calls for a situational look, especially for small clinics, rare disease groups, and ZIP codes with small populations.[1][10] Safe Harbor is used a lot in U.S. healthcare research because the rule-based list is clear and repeatable, which works well for IRB review and large data pipelines.[1][9]

If that checklist cuts too much out of the data, HIPAA gives you another path.

Expert Determination: documenting a very low re-identification risk

Expert Determination, under 45 C.F.R. §164.514(b)(1), takes a different route. Instead of following a fixed list, a qualified expert uses accepted statistical and scientific principles to review the dataset and decide that the risk of re-identification is very low for expected recipients, including likely linkage with other reasonably available information.[1][12][14]

In practice, the expert usually tests linkage risk against realistic outside data sources and may use methods such as data generalization, suppression, noise addition, and aggregation.[1][6][7] HIPAA also requires written documentation of the full determination, including the methods used, assumptions about who will receive the data and what outside datasets they could access, and a clear statement of residual risk.[1][11][12] That record matters for audits and for cross-border data transfer assessments.

Expert Determination makes more sense when a study needs finer-grained data, like full date ranges or detailed geographic detail, that Safe Harbor would remove.[7][1] It takes specialized skill and heavier documentation, but it lets teams keep more analytic usefulness and can fit harder use cases better.

Limited data sets are not the same as anonymized data

A limited data set is reduced PHI, not anonymized data. A limited data set sits in the middle: most direct identifiers are removed, but some elements can stay, including dates of admission, discharge, and service; dates of birth and death; and city, state, and ZIP code.[13][15][7] That leftover detail is why limited data sets are often used for outcomes research and quality improvement, where date fields and limited geography can make all the difference.

A limited data set is still PHI.[8][18] It is not de-identified under HIPAA and should not be treated as anonymous. HIPAA also requires a data use agreement that limits use, sets safeguards, and bars re-identification.[8][16] Treating a limited data set as anonymous would break both HIPAA and GDPR analysis.

That line matters even more under GDPR, where identifiability is judged more strictly.

GDPR anonymization: identifiability, Recital 26, and pseudonymization

Under GDPR, personal data means any information tied to an identified or identifiable natural person. Health data stays personal data unless the person can no longer be identified. In plain English: data falls outside GDPR only when identifying someone is no longer reasonably likely for the controller or for another party.[3][27][29] That line is what separates anonymization from pseudonymization.

GDPR sets a higher anonymization threshold than HIPAA

Recital 26 is the key text here. It says identifiability must be judged against all means reasonably likely to be used by the controller or any other person to identify someone, whether directly or indirectly.[3][20][24] A practical way to read that is to look at three questions: can a person be singled out, linked, or inferred?[21][22][25]

That test sets a tougher bar than HIPAA. Under HIPAA, removing 18 identifiers or showing a very small re-identification risk through Expert Determination may be enough. Under GDPR, that does not automatically clear Recital 26. For complex health datasets, GDPR anonymization often calls for more aggressive stripping, aggregation, or deletion than HIPAA de-identification.[20][22][25]

That’s why pseudonymization matters so much. It lowers risk, but it does not make data anonymous.

Pseudonymization reduces risk but does not remove GDPR obligations

Article 4(5) defines pseudonymization as processing personal data so it cannot be tied to a person without separate, protected additional information.[23][27][28] In practice, that usually means direct identifiers are swapped out for codes, while the key is stored separately.

Recital 26 is clear on the next point: pseudonymized data is still personal data because the underlying medical data can still be linked back through the key or by other means that are reasonably likely to be used.[3][26][29] So the legal duties stay in place. Pseudonymized health data still needs a lawful basis for processing, still falls under special category protections, and still needs a valid transfer tool for cross-border sharing.[27][29][30]

The European Data Protection Board (EDPB) Guidelines 01/2025 push this point further. They explain that pseudonymized data remains personal data even when a different entity holds the re-identification key, if re-identification is still reasonably possible.[23][31]

For healthcare organizations and vendors, pseudonymization is a useful safeguard within a robust third-party risk management framework recognized in Article 32 and Article 89(1) for research. But it does not turn regulated health data into unregulated data. That difference matters in the cross-border comparison below.

This gap matters because data that falls outside HIPAA after de-identification may still count as personal data under GDPR. And once that happens, lawful basis, security, and transfer rules can still apply, taking the risk out of healthcare data management. That becomes a big deal when the same dataset is sent to an EU researcher, vendor, or model-training workflow.

The table below shows how the same dataset can shift in legal status under each regime.

Topic HIPAA GDPR Practical impact
Legal scope De-identified data is outside HIPAA once Safe Harbor or Expert Determination is met Only truly anonymous data is outside GDPR scope A dataset may be outside HIPAA once de-identified but still regulated under GDPR
Threshold for anonymization or de-identification HIPAA can be satisfied by a defined method; GDPR needs a context-specific anonymity showing GDPR uses a broader contextual test GDPR often requires stronger technical and contextual risk reduction
Treatment of pseudonymized data HIPAA does not treat pseudonymization as a full regulatory off-ramp; coded data may still be PHI depending on the key and context Pseudonymized data remains personal data Coding alone is usually not enough for cross-border compliance
Re-identification standard HIPAA relies on the defined de-identification method GDPR asks whether identification is reasonably likely by any means GDPR analysis is broader and more contextual
Typical healthcare research use cases Common for U.S. secondary use, analytics, and some research workflows Often requires more controls, a lawful basis, and transfer analysis unless data is truly anonymous Cross-border studies need dataset-by-dataset review and governance

These differences hit hardest when data leaves the country where it was de-identified.

Why HIPAA de-identified data may still be personal data under GDPR

HIPAA de-identification does not automatically meet GDPR anonymization. A dataset that clears HIPAA Safe Harbor can still be identifiable under GDPR if linkage, inference, or outside datasets make identification reasonably likely.[32][3][5]

That can happen, for example, in a rare disease cohort or a small geographic area. In those cases, Recital 26 may keep the data inside GDPR's scope.[32][3][5] In practice, EU regulators often treat that kind of dataset as pseudonymized instead of anonymous, which means GDPR duties still apply.[5][4]

The plain-English takeaway is simple: in cross-border work, teams should usually assume HIPAA de-identified data is still regulated personal data in the EU unless a context-specific anonymization review shows that identification is no longer reasonably possible.[5][32][3]

Cross-border healthcare research and data transfer requirements

When data moves between the U.S. and EU, its legal status under each regime shapes the safeguards required. If data is truly anonymized under GDPR, it can move across borders with fewer transfer requirements because it sits outside GDPR altogether.[3][5][32]

Dataset type U.S.–EU multicenter clinical studies AI model training with healthcare data Practical compliance point
HIPAA-de-identified data May be acceptable for U.S. use without HIPAA restrictions if properly de-identified May be used for analytics or model development in the U.S. under HIPAA de-identification rules Still may be personal data under GDPR if re-identification remains reasonably possible
GDPR-anonymized data More likely to move across borders without GDPR transfer rules because it is outside GDPR scope Lower transfer requirements for model training if anonymization is defensible Utility may decrease as anonymization strength increases
Pseudonymized data Remains regulated and usually requires a lawful basis, security controls, and transfer safeguards Common in research and AI workflows but still subject to GDPR obligations Strong governance, contracts, and documented risk assessment remain necessary

Pseudonymized data stays personal data under GDPR, so transfer safeguards still apply. For transfers from the EU to the U.S., organizations need a valid transfer mechanism. After Schrems II, Standard Contractual Clauses are the tool most groups use. But they can't stand alone. They also need a Transfer Impact Assessment and extra technical measures, such as encryption and access controls, to provide comparable protection.[33][34][35][36]

And there's another twist. If a U.S. entity keeps the key, the dataset may still be PHI under HIPAA, which adds one more layer of compliance work.[1][10][6]

Practical alignment strategies and conclusion

How to align anonymization practices across U.S. and EU requirements

This gets practical fast: how do teams build one dataset that can work under both U.S. and EU rules? The safest move is to design to the stricter standard. Start with HIPAA de-identification, then add GDPR-level generalization for indirect identifiers. That usually means generalizing dates, using broader age bands, coarsening geography, and suppressing rare combinations.[5][39][40]

After that reduction, the next step is simple but often skipped: write down why the remaining re-identification risk is acceptable. The final dataset should be checked through a documented risk assessment, not just a checklist.[38] That record should include the anonymization method, the fields removed or generalized, the intended recipients, and any external datasets that could realistically make linkage possible.[1][2]

Using risk management in healthcare to support ongoing oversight

Anonymization is not a one-and-done technical task. Risk can change over time as external datasets become more detailed and re-identification methods improve. That is especially important for longitudinal research datasets. A better way to handle it is to treat anonymization as a living control: reassess risk when the dataset changes, when recipients change, or when new external linkage sources appear.[37][38]

Conclusion: the key difference is risk threshold, not just terminology

In day-to-day practice, the takeaway is straightforward. HIPAA and GDPR differ mostly in the level of risk they allow, so cross-border teams should build to the stricter standard and document that decision.[41]

FAQs

Does HIPAA de-identification satisfy GDPR?

No. HIPAA de-identification does not automatically satisfy GDPR requirements.

Under U.S. law, data that has been de-identified under HIPAA may no longer count as PHI. But GDPR sets a tougher bar. Under GDPR, data is anonymous only when re-identification is not reasonably likely.

That’s the key difference.

In practice, HIPAA de-identified data is often viewed as pseudonymized under GDPR, which means it still falls within GDPR’s scope and remains subject to its rules.

Is pseudonymized health data still regulated?

Yes. Under the GDPR, pseudonymized health data is still regulated because it can be linked back to a person through extra information or re-identification keys. That means it still counts as personal data.

So organizations still need to follow all GDPR rules that apply. Even if data has been de-identified under HIPAA, the GDPR may still treat it as pseudonymized personal data if any re-identification risk is still on the table.

When is a limited data set still PHI?

Under HIPAA, a limited data set is still PHI. Some direct identifiers are stripped out, but the data can still include city, state, ZIP code, and full dates.

That matters because those details can still point, directly or indirectly, to a specific person. So the data stays under HIPAA rules. Organizations need a Data Use Agreement and should put the right safeguards in place to lower re-identification risk.

Related Blog Posts