Removing names doesn’t make healthcare data de-identified. I start by choosing one of HIPAA’s 2 methods: Safe Harbor, which removes 18 identifier categories and requires no actual knowledge of remaining identification risk, or Expert Determination, which documents a finding that identification risk is very small.
Here’s the workflow I follow - and what you should check before sharing:
- Define the purpose and inventory the data. Review legal limits, contracts, patient commitments, and who will receive the release. Pseudonymized data may still be PHI; limited data sets remain PHI.
- Choose and document the HIPAA method. Match privacy protections to the detail your research, analytics, or AI task needs.
- Transform and test. Check tables, clinical notes, images, device records, and AI outputs for identification risk and loss of clinical usefulness. Tokens, shifted dates, and synthetic records don’t prove compliance by themselves.
- Approve and control sharing. Record the dataset version, test results, approvers, access rules, key controls, recipient restrictions, and any required agreements.
- Review when conditions change. New recipients, outside data, model updates, or suspected re-identification can call for new testing or a pause in sharing.
My rule: <u>approve the tested release - not every future use</u>. Keep the evidence, name the owners, and set review dates. Third-party vendor risk management tools can help track reviews, but they don’t replace a HIPAA de-identification determination.
Healthcare Data De-Identification: 5-Step Workflow
The Ten HIPAA De-identification Myths, Corrected on the Spot
sbb-itb-535baee
Choose and Document a HIPAA De-Identification Method
Choose Safe Harbor or Expert Determination under 45 C.F.R. § 164.514. Treat only final, effective requirements as binding.[4][2][8] Use the inventory and legal review to select the HIPAA method that discloses the least information while still supporting the intended use.
Safe Harbor: Remove 18 Identifier Categories
Apply Safe Harbor to the patient and the patient’s relatives, employers, and household members. Review structured fields, free text, filenames, image metadata, and derived variables - not just columns labeled “identifier.”[7][2]
| Identifier category | What to remove or transform |
|---|---|
| Names | Names, including those in narrative text. |
| Geography | Subdivisions smaller than a state, subject to the ZIP exception below. |
| Dates and ages | Restricted date elements and ages, subject to the rules below. |
| Telephone numbers | All telephone numbers. |
| Fax numbers | Fax numbers. |
| Email addresses | Email addresses. |
| Social Security numbers | Social Security numbers. |
| Medical record numbers | Medical record and patient-chart identifiers. |
| Health plan beneficiary numbers | Member, subscriber, and beneficiary numbers. |
| Account numbers | Billing, financial, and other account numbers. |
| Certificate or license numbers | Certificate and license numbers. |
| Vehicle identifiers and serial numbers | Vehicle identifiers, license plates, and serial numbers. |
| Device identifiers and serial numbers | Device identifiers and serial numbers. |
| URLs | URLs. |
| IP addresses | Internet Protocol addresses. |
| Biometric identifiers | Fingerprints, voiceprints, and comparable biometric identifiers. |
| Full-face photographs and comparable images | Full-face photographs and similarly identifying images. |
| Other unique identifiers, except a permitted re-identification code | Unique identifying numbers, characteristics, or codes, except those permitted under § 164.514(c). |
Keep a 3-digit ZIP code only when its prefix covers more than 20,000 people, based on current publicly available Census data. Otherwise, use 000.
Keep only year-level dates, and group ages over 89 as 90 or older. Apply these rules to derived ages and narrative references, too. Remove date elements - including years - that reveal an age over 89. Before release, confirm there is no actual knowledge that the remaining data could identify a person, either alone or combined with other information.[2][7]
Expert Determination: Assess Identification Risk
A qualified expert must use generally accepted statistical and scientific methods to determine that identification risk is very small for the anticipated recipient.[4][9]
The assessment should examine whether attributes recur, appear in outside sources, or make records distinguishable. It must also account for recipient capabilities, intended uses, external linkage sources, and release controls. HIPAA sets no universal numerical risk threshold.[2][10]
Document the dataset version, expert qualifications, methods and results, assumptions, external sources, safeguards, limitations, rationale, conclusion, approval, and re-review triggers.[4][9]
Compare the 2 HIPAA Methods
| Consideration | Safe Harbor | Expert Determination |
|---|---|---|
| Compliance requirements | Remove 18 categories and meet the no-actual-knowledge condition. | Document an expert finding of very small identification risk. |
| Dates and geography | Year-only dates and qualifying 3-digit ZIP codes, subject to age restrictions. | Retain detail supported by the risk analysis. |
| Analytical utility | Standardized removals reduce temporal and geographic detail. | Tailored transformations preserve needed detail. |
| Expert involvement | No expert required. | Qualified expert required. |
| Documentation | Record field rules, transformations, validation, and actual-knowledge review. | Record methods, results, assumptions, recipient conditions, and rationale. |
| Research and AI use | Use when finer detail is unnecessary. | Use when finer detail is justified by the analysis. |
| Remaining linkage risk | Address known linkage paths after removal. | Assess linkage against reasonably available information and release conditions. |
Choose the method that fits the task and record the rationale in the release record. If Safe Harbor removes dates, geography, or longitudinal detail needed for the intended use, assess Expert Determination before release.
Suppression, generalization, aggregation, perturbation, and role-based access controls can work together, but they do not create a third HIPAA method. Document which recognized method the resulting dataset satisfies. Neither method automatically validates every downstream research or AI use.[2][4][8]
After selecting the method, map each field to a transformation and test the highest-risk fields and uses.
Transform Data and Test High-Risk Uses
Compare Privacy Techniques and Analytical Trade-Offs
Once you’ve selected the HIPAA method, match each transformation to the data, recipient, outside data, and required precision. Choose the lightest approach that supports the intended use. Transformations change the data; privacy models set the protection target. Test edge cases and check disclosure risk under the conditions in which the data will be used.[2][8]
| Technique or model | Purpose | Utility trade-off | Main limitation | Suitable uses |
|---|---|---|---|---|
| Suppression - transformation | Remove records, fields, or rare values | Shrinks the sample; may introduce selection bias | Rare or clinically important cases may disappear | Public summaries, small-cell control |
| Generalization - transformation | Replace precise values with broader categories | Reduces geographic, temporal, or clinical precision | Attribute combinations may still be linkable | Cohort reporting, population analysis |
| Masking - transformation | Hide part of a value | Retains display value but little analytic value | Values may still identify people or allow linkage | Operational display, limited troubleshooting |
| Tokenization - transformation | Replace identifiers with tokens | Keeps consistent longitudinal linkage | The mapping remains sensitive; tokens are not de-identified by default | Controlled research, internal analytics |
| Aggregation - transformation | Report counts, rates, or summaries | Removes record-level analysis | Small cells, outliers, and repeated queries may reveal individuals | Dashboards, public reporting |
| Date shifting - transformation | Move dates by an offset | Keeps some intervals or sequences; disrupts cross-dataset alignment | Shifted dates remain sensitive and do not satisfy Safe Harbor alone | Controlled longitudinal analysis |
| Top- and bottom-coding - transformation | Replace extreme values with thresholds | Retains most distributions but loses extreme-case detail | Thresholds may still identify rare cases | Claims, age, cost, length-of-stay data |
| Perturbation - transformation | Add noise or alter values | Keeps approximate distributions; may distort individual relationships | Weak noise may be reversible or leave values linkable | Statistical reporting, exploratory analysis |
| k-anonymity - privacy model | Require each quasi-identifier pattern to appear in at least k records | May require extensive generalization or suppression | Misses sensitive-value similarity and some linkage risks | Structured data with known quasi-identifiers |
| l-diversity - privacy model | Require diverse sensitive values within each k-anonymous group | Reduces utility; makes grouping harder | Similar or skewed values may allow inference | Sensitive-attribute protection |
| t-closeness - privacy model | Limit differences between group and overall sensitive-value distributions | Requires more suppression or generalization | Complex to configure; does not eliminate risk | High-risk tabular releases |
| Differential privacy - privacy framework | Add calibrated randomness to queries or model outputs | Privacy loss accumulates and can reduce accuracy | Requires sound implementation, sensitivity controls, and privacy-budget management | Repeated queries, public releases, some synthetic-data workflows |
| Synthetic data - generation approach | Generate artificial records modeled on source data | Keeps some correlations; may weaken rare-event fidelity | Models can memorize source data; disclosure risk and utility need testing | Development, testing, controlled research |
For differential privacy, document the mechanism, sensitivity assumptions, privacy parameters, and cumulative budget across releases.[11][12] For synthetic data, test for near-duplicates, memorized text, rare-record reproduction, and disclosure risk.
Before release, compare counts, missingness, clinical relationships, temporal ordering, and subgroup model performance against a trusted reference. Set acceptance thresholds before seeing the results.
Check Clinical Text, Images, and Medical Device Data
High-risk formats need checks beyond those used for tabular data. Pair automated scans with targeted human review. In notes, check narrative identifiers, rare diagnoses, family details, occupations, and unusual events. In images, inspect DICOM metadata, burned-in text, annotations, filenames, embedded files, and recognizable facial anatomy. Removing metadata does not remove identifying image content.
For devices, inspect serial numbers, timestamps, waveforms, locations, alarm and event sequences, firmware, configurations, maintenance records, and connected-service copies. Test whether these details link to admissions, procedures, or service logs, with deliberate attention to rare sequences.
Use timestamp coarsening, location generalization, and protected token maps where appropriate. After transformation, check waveform fidelity and clinical usefulness. These checks address re-identification through the content itself - not just individual fields.
Control Disclosure Throughout the AI Lifecycle
Maintain data lineage and purpose limits from curation and labeling through training, validation, monitoring, debugging, and synthetic-data evaluation. Restrict access, check labels for sensitive details, redact debugging artifacts, and keep human review for high-risk outputs.
Test memorization, membership inference, model inversion, prompt leakage, output leakage, and retrieval risks based on exposure and the threat model. Models can still reproduce de-identified training data. Inspect logs and telemetry, too.
Record the test scope, results, fixes, and retest triggers. A passing result applies only to the tested conditions. Treat every model update as a new disclosure event that requires review, and carry the test results into release approval and third-party review.
Govern Data Releases and Third-Party Sharing
Assign Data Owners and Release Roles
Approve releases only after validation. Inventory source PHI, derived data, keys, and downstream outputs. Apply separate access controls to source data, released copies, and re-identification keys. Never include keys in the transfer package.
Each release record should document the requestor and recipient, purpose, legal and policy basis, dataset version, data elements, population and time range, de-identification method, and validation results. Also record approved users, contract status, retention and deletion dates, the residual-risk decision, approvers, release date, review date, and incident contacts.
Use a responsibility matrix to keep decision-making separate from execution. Treat the matrix as a template, and name the actual owners in each release record. Require documented privacy and legal sign-off before sharing, plus separate authorization for source-PHI transfers, key access, and downstream model or dataset releases.
| Activity | Privacy | Security | Legal | Clinical/Research | Data Governance | AI/Technical | Executive/Release Owner |
|---|---|---|---|---|---|---|---|
| Approve purpose and scope | A/R | C | C | R | R | C | A |
| Select de-identification method | R | C | C | C | R | R | A |
| Validate transformation | R | R | C | C | R | R | A |
| Approve recipient and contract | R | R | A/R | C | R | C | A |
| Manage access and keys | C | A/R | C | C | R | R | A |
| Review residual risk | A/R | R | C | C | R | R | A |
| Monitor downstream use | R | R | C | C | A/R | R | A |
A = accountable; R = responsible; C = consulted.
Adjust assignments to match the organization’s policies and applicable state, federal, contractual, and research requirements. Then manage third-party risk and contract terms.
Define Recipient and Contract Controls
Assess the recipient’s identity, role, purpose, users, hosting location, security program, linkage capability, external data access, subcontractors, international access, retention, and incident history. Determine whether the recipient can access the re-identification key, create derived datasets or models, or share those outputs with others.
Review the risk of combining the release with public records, claims databases, geolocation data, social-media content, genomic databases, or other datasets. The same dataset can pose different risks for different recipients.
Contracts should specify the purpose, fields, users, security controls, re-identification bans, linkage limits, subcontractor approval, and onward-disclosure limits. Include retention terms, deletion or return requirements, incident notice deadlines, audit rights, logging, data-location limits, and change-control procedures.
Address outputs that could retain or reveal source information, too: derived datasets, dashboards, synthetic data, embeddings, model weights, prompts, fine-tuning artifacts, and other outputs. Require review before introducing new linkage sources, subcontractors, production uses, or materially different AI uses.
A vendor that creates, receives, maintains, or transmits PHI generally needs a written BAA.[13][14] If the vendor performs de-identification work, the BAA should expressly authorize it.[1][17] Use a DUA when sharing a limited data set.[15][7] A BAA and DUA may be combined when both apply.[13][16]
Neither agreement is required solely to share properly de-identified information, though other laws, policies, and contracts may impose obligations.[2] Limiting the scope of releases of de-identified information is a governance practice - not an additional HIPAA requirement.
Track vendor risk findings and release approval in the same review record.
Coordinate Risk Reviews with Censinet RiskOps™
Use Censinet RiskOps™ to coordinate third-party and enterprise risk reviews. Keep those findings separate from the de-identification determination.
Reassess Identification Risk as Conditions Change
De-identification holds only as long as the assumptions behind it remain valid.
Even an unchanged dataset can become easier to identify. New public records, better linkage tools, or a recipient with more outside information can increase the risk. Expert Determination considers the recipient and reasonably available outside information - not just the dataset. Material changes to those assumptions should prompt updated expert review.[1][2] Changes in recipients, linkage sources, or uses may also mean the approved HIPAA method no longer supports the release.
Set a documented review schedule: check high-risk data before each release and quarterly, moderate-risk data every six months, and lower-risk static data annually. Review sooner if a trigger occurs. These intervals are governance recommendations, not HIPAA deadlines.
Use a Repeat Risk Review Checklist
Whenever the recipient, data, or model changes, use this checklist to confirm that the release still matches its approved de-identification basis:
- Confirm the release terms. Recheck the approved purpose, recipients, permitted uses, retention period, and contractual restrictions.
- Review the data and outside sources. Check identifiers, dates, geography, text, images, device data, and model outputs. Review possible linkage sources, including public registries, social media, voter files, property records, data brokers, breach datasets, and disease repositories. Test rare combinations and longitudinal patterns - not just individual columns.
- Check for missed information and AI risks. Sample clinical text, scanned documents, images, audio, and metadata for missed redactions. Review access and deletion records, and retest relevant AI memorization, prompt-based extraction, and retrieval risks.
For Safe Harbor, recheck all 18 identifier categories.[8] For Expert Determination, obtain updated expert analysis when changes undermine its assumptions or findings.
Record test results, limitations, remediation owners, approvals, and the next review date. A failed test calls for mitigation or escalation, not just a checked box.
Set Review Actions for Data and Use Changes
Use these triggers to decide whether sharing can continue. The approval roles below are governance assignments; adjust them to match your organization’s policies. Pause affected transfers or uses when the approved basis no longer applies.
| Change or trigger | Analysis and default action | Accountable approvers |
|---|---|---|
| New recipient or subcontractor | Review linkage capabilities, access, and contract changes. Pause release until approval; update Expert Determination if recipient assumptions change. | Data owner, privacy officer, security or legal reviewers |
| More precise geography or dates | Retest uniqueness and reconstruction risk. Generalize or suppress fields; rerun Safe Harbor review or obtain updated expert review. | Data owner, privacy officer, qualified expert when applicable |
| New AI use or model provider | Test memorization, prompt extraction, retrieval exposure, and retention. Pause ingestion or transfer until controls pass validation. | AI governance owner, privacy officer, security lead |
| New device or telemetry integration | Check serial numbers, MAC addresses, IP addresses, Bluetooth identifiers, timestamps, room or bed location, and device fingerprints. Minimize fields before integration or release. | Clinical engineering or device owner, data owner, and security reviewer |
| Major expansion, merge, or new fields | Reinventory identifiers and rerun combination and linkage tests. Reapply Safe Harbor checks or update Expert Determination before distribution. | Data owner, privacy officer, qualified expert when applicable |
| Suspected re-identification or security incident | Preserve evidence, contain access, and assess whether PHI was exposed. Pause sharing; investigate and notify as required. | Privacy officer, incident-response lead, legal counsel |
| New public source | Assess whether the source is reasonably available to the anticipated recipient and whether it enables matching. Restrict access or transform fields; update expert analysis when assumptions change. | Privacy officer, data owner, qualified expert when applicable |
Conclusion: Document, Validate, and Review
After validation, confirm that the release meets HIPAA’s de-identification standard. Pseudonymization can make people harder to identify, but only Safe Harbor or Expert Determination satisfies that standard.[2][8]
Check how identifiable and useful the data remains - not just whether a transformation ran. Assess remaining disclosure risk across text, image metadata, device data, and AI outputs. Keep the test results and document any unresolved limitations.[6][18]
If the data passes testing, move from analysis to release controls. Approval applies to that specific release. Put recipient restrictions, retention limits, onward-sharing rules, and deletion requirements in writing. Keep the approved dataset version, validation evidence, and approver names together in one release record.
Follow this sequence: inventory the data, choose and document a HIPAA method, validate, approve the release, and schedule re-reviews. Use that documentation as the baseline for each later re-review.[2][5][6]
FAQs
How do I choose a qualified de-identification expert?
Look for proven experience applying generally accepted statistical and scientific principles, along with expertise in re-identification risk analysis and Statistical Disclosure Control [1][2]. HIPAA experts typically have backgrounds in statistics, mathematics, or related fields [1].
Ask for a written report that documents their methods, assumptions, and findings [3][1][4]. Give priority to experts who assess threats from external data and can issue time-limited certifications that account for changing risks [3][1].
How can I preserve rare cases without identifying patients?
Rare diagnoses or unusual encounter patterns can identify patients even after HIPAA identifiers are removed. Standard de-identification alone may not be enough. Suppress high-risk records or fields, check patterns to flag outliers, and analyze cell sizes to confirm that rare combinations meet a minimum threshold.
For higher-risk datasets, group rare codes into broader categories or consider synthetic alternatives that retain analytical value without exposing individuals.
What evidence should I request before using de-identified AI data?
To check compliance and risk mitigation, request documentation of the de-identification method: either a HIPAA Safe Harbor checklist or an Expert Determination report.
For Expert Determination, get the expert’s credentials, statistical methodology, risk threshold, and signed attestation that the risk of re-identification is very small.
Also request a detailed data inventory, results from linkage and cell-size testing, and a Data Use Agreement (DUA) that prohibits attempts to re-identify individuals.