If an AI vendor wants patient data, I’d treat that as a stop-point until six things are clear in writing: what data they touch, how long they keep it, whether they train on it, which controls protect it, who their outside providers are, and who pays and acts if something goes wrong.

This is the short version: a “HIPAA compliant” claim is not enough. I’d want a signed BAA, a written no-training term, a current SOC 2 Type II report, a retention schedule, a subprocessor list, and breach terms with deadlines like 24 to 72 hours. If any of that is missing, I’d hold the deal.

Here’s the full checklist I’d use for managing third-party AI risk before go-live:

  • Data access: Ask for a field-by-field list, not “we need the whole chart.”
  • Retention: Get deletion dates for prompts, logs, backups, and third-party copies.
  • Training use: Ban use of PHI, prompts, outputs, metadata, and logs for model improvement.
  • Security: Check encryption, MFA, audit logs, pen-test records, and incident response steps.
  • Subcontractors: Get the named list of cloud, model, speech, and support vendors.
  • Ownership: Put breach notice, remediation, and audit support into the contract.

A simple rule works here: paper beats promises. If the vendor will not put these points into the BAA and main contract, I would not move ahead.

6 Questions to Ask AI Vendors Before Sharing Patient Data

6 Questions to Ask AI Vendors Before Sharing Patient Data

How to Vet Healthcare AI Vendors

Why AI Vendors Require Stricter Vetting Than Standard SaaS

AI tools handle more than saved data. They often take in prompts, file attachments, transcripts, images, and records, then send that data to outside models for inference, evaluation, or product improvement. That changes the risk picture in a big way.

With standard SaaS, data usually stays inside a more predictable path. AI systems can be different. They may use data for fine-tuning, abuse monitoring, or internal product work. Those behind-the-scenes paths can create risk far beyond the final output a user sees.

Take a clinician using an ambient-scribing tool. Just by summarizing a visit, that person may send a patient’s name, diagnosis, medications, and voice recording to the vendor. And even if the vendor doesn’t keep the final note, the input may still show up in logs, error reports, test environments, or human review queues.

The bigger problem is fourth-party risk. Your AI vendor may depend on a model provider, cloud platform, speech-to-text service, or vector database that your team never picked and never negotiated with directly. HHS states that a cloud service provider can be a business associate even when it cannot view encrypted ePHI or does not possess the decryption key.[2] So before any PHI moves downstream, ask for:

  • A current subcontractor list
  • Named model providers
  • Flow-down privacy and security terms[1][4]

That’s why the first diligence question is so simple: what data will the vendor access, and why?

State privacy laws can add more duties on top of HIPAA, including rules tied to biometric data, consumer-health data, deletion, minimization, and breach notification and response. HIPAA is the floor, not the full legal review. The questions that follow turn those risks into a vendor review process.

1. What Data Will You Access, and Why?

Start with scope. Ask for a field-level inventory of every data element the vendor will touch, including inputs, outputs, metadata, logs, backups, and support copies. That can include names, dates of birth, diagnoses, clinical notes, audio recordings, and insurance information. This inventory sets the minimum access the vendor should have.

Classify each item as PHI/ePHI, de-identified data, prompts, outputs, metadata, or logs. That matters because a vendor can still handle ePHI even if it can't decrypt it.[5]

Then go field by field and ask why the vendor needs it. Tie each field to a specific product function, like transcription, clinical summarization, coding assistance, or workflow automation. A scheduling assistant may need appointment type and availability, but not a full progress note. A document-classification tool may need referral text, but not a patient's Social Security number or home address. Under HIPAA's minimum-necessary rule, access should stay limited to what's reasonably needed for the stated purpose.[2][3] If a field isn't needed for that function, leave it out.

Good answers are specific. They spell out the workflow, the exact fields involved, and any stored copies.

Before you approve any access scope, ask for a data-flow diagram and API documentation that shows the available scopes.

Use the answers below to tell the difference between narrow, purpose-based access and requests that go too far.

Vendor answer What it signals What to do
We need the entire chart for all features. May exceed minimum necessary for a specific workflow. Request feature-by-feature justification; test a restricted integration.
We need only selected fields and encounter documents. Fits purpose limitation and least-privilege access. Verify the exact fields, access mechanism, and logging.

2. How Long Will You Retain Our Data?

Once access scope is clear, the next step is retention.

A vendor can handle PHI the right way during active use and still keep it far too long after the relationship ends. That’s where problems start. So after you define access, ask how long the vendor keeps your data.

The question to ask is direct and blunt:

"Can you contractually commit to deleting or returning our data within a defined timeframe after request or termination, including from backups and all subcontractors, and what proof of deletion will you provide?"

A BAA gives you a starting point for HIPAA-compliant vendor risk management. But it does not set retention limits, spell out backup deletion, or require proof that data was destroyed. In plain English, deletion is a contract matter, not something you should assume from a policy page.

Your contract should state:

  • the deletion deadline
  • that backups are included
  • that inference logs are included
  • that third-party copies are included
  • that the vendor must give written proof of deletion

This matters even more if PHI is used to train or fine-tune a model. In that case, deletion gets much harder to confirm. Put simply, PHI used for training or fine-tuning weakens deletion assurance.

A few vendor answers should make you pause. Watch for no post-termination deletion plan at all, fuzzy wording like "data may be used to improve our models," or deletion terms that cover only primary storage while ignoring inference logs, backups, or third-party model providers. The table below helps you read those answers for what they are.

Vendor answer What it signals What to do
"We delete or return all data within the contractually defined timeframe, including backups, inference logs, and subcontractors, and provide written proof." Clear, enforceable, and complete. Verify scope in the contract and confirm subcontractor coverage.
"We retain data under our standard policy." Vague - no defined timeframe or scope. Request the written policy; push for specific contract language.
"Data may be used to improve our models." PHI could be used for training; deletion may be harder to verify. Treat as a red flag; require written confirmation that PHI will not be used to fine-tune base models.
No post-termination deletion plan. Significant gap in vendor governance. Do not proceed without negotiating explicit deletion terms.

If retention is unclear, stop there. If it’s clear, move on to whether the vendor uses your data to train or improve models.

3. Will You Use Our Data to Train or Improve Models?

Put this question in front of every AI vendor:

"Will you use any PHI, prompts, outputs, metadata, or logs to train, fine-tune, validate, or improve a model - and if not, will that prohibition be written into the BAA and master contract?" [6]

That question gets to the point fast. A signed BAA is a legal floor, not proof that everything is handled the right way. That’s why the master contract also needs plain language that bars PHI, prompts, outputs, metadata, and log data from any model-improvement pipeline.

Ask the vendor directly whether PHI or related data can enter any training or fine-tuning process. Then get the answer in writing and tie it to the contract. If a vendor can’t give you a clear answer, they should not get PHI.

Vendor answer What it signals What to do
"No - PHI, prompts, outputs, metadata, and log data are never used for training or fine-tuning, and that's stated in the contract." Clear and enforceable. Check that the prohibition appears in both the BAA and master contract.
"We use interactions to improve our models." Your data may be used in model improvement. Treat this as a red flag. Ask for written clarification and a contract ban.
No written position on training use. Major governance gap. Do not move forward without direct contract language.

If training use is banned in writing, the next step is to look at how the vendor protects the data it still handles.

4. Which Security Controls Protect the Data?

If a vendor says it won’t use your data for training, that’s only the first hurdle. The next one is simpler: can it protect the data it still touches? Before any PHI moves, ask for proof, not promises.

A BAA is paper. It does not prove that the vendor has encryption, logging, incident response, or systems that can hold up under stress. You need records that show those controls exist and work.

Ask for proof of:

  • independent security controls
  • recent penetration testing
  • incident-response procedures
  • PHI access logs

A self-attestation or a badge on a website doesn’t count as proof here.

Under the May 2026 HIPAA Security Rule updates, encryption and MFA should be written into the contract.

Use the checklist below to separate actual controls from sales language.

Security Element What to Require Red Flag
Encryption & MFA Both required in the contract Listed as optional or "planned"
Audit Logs PHI access tracked and available for buyer review No logging or log access for the buyer
Data Isolation PHI in an isolated environment Shared model environment with other clients
Breach Notification Specific SLA (e.g., 24–72 hours) in the contract Legal minimum only
Third-Party Audits SOC 2 Type II or HITRUST CSF report on request Self-attestation only

Under HIPAA, a breach at a vendor is generally reported as a breach at the covered entity [7]. That means the vendor’s incident-response records and audit logs become part of your risk picture, not just theirs.

If these controls look weak, the next step is to check who else can get to the data. Those controls should match the vendor’s subcontractor and model-provider map.

5. Who Are Your Subcontractors and Model Providers?

Most AI vendors depend on other companies behind the scenes. You need a full list of every subprocessor, hosting provider, and outside model provider that may store, process, transmit, or access PHI, along with the country or region where that work happens. This is the downstream side of the data path you mapped earlier.

If the answer is vague, treat that as a warning sign. “HIPAA-ready” means very little if the vendor won’t name its subprocessors. And while a signed BAA matters, it doesn’t solve the whole problem. Every downstream provider also needs to be named and covered.

Question for Vendor What It Tests Red Flag
Who are your subprocessors? Transparency and downstream compliance Vague "HIPAA-ready" claims without a named list
Does PHI stay in a dedicated instance? Data isolation and security architecture PHI processed in shared or multi-tenant environments
What is the data residency? Jurisdictional and regulatory compliance "Data is handled globally" without specifics
Are model providers used? Third-party risk and BAA chain Consumer-grade AI APIs without BAAs
Is PHI pseudonymized before any external model provider receives it? Data minimization and exposure control No pseudonymization or tokenization in place

After you map the downstream chain, set clear ownership for compliance, incident response, and remediation across every party in that chain.

6. Who Owns Compliance and Remediation?

Once you know who handles the data after it leaves your hands, the next step is simple: name the party on the hook when something goes wrong.

That means no fuzzy “shared responsibility” language. If there’s a breach, a regulator asks questions, a subcontractor slips up, or fixes cost money, the contract should say exactly who handles it. Don’t leave that for later. Lock it in before go-live.

Your BAA and master contract should spell out who owns:

  • training restrictions
  • logging
  • prompt-injection response

If you’re dealing with agentic systems, push for more than general promises. Require traceable decisions, drift detection, and automatic performance alerts. The contract should also lay out exact breach notification SLAs and a named process for telling you when regulations change.

That ownership also covers regulatory updates. The vendor should have a defined timeline for notifying you about HIPAA and other rule changes. Your team shouldn’t have to monitor vendor-side compliance shifts on its own.

And if the vendor is in charge of remediation, don’t hand over full control. Keep the right to bring in your own team, or a trusted third party, for fixes and audits. The contract should pin down who pays, who takes action, and what proof they need to provide.

Contract Category Required Terms / Evidence
Remediation Specific breach notification SLAs; drift detection and performance alerts
Regulatory Cooperation Named process for regulatory updates; defined notification timeline
Accountability BAA assigning ownership for logging and training restrictions; subcontractor enforcement responsibility
Remediation Expense Vendor responsible for fixing findings at its own expense; independent remediation rights and audit access
Compliance Evidence Third-party audit reports on request

How to Evaluate Each Question in Practice

Use the same test for every vendor answer: written proof, contract language, and a named control. That simple filter helps you tell the difference between a polished sales pitch and something you can actually review.

Question 1 - Data access and purpose. Start by mapping the data path. Then check whether the vendor can justify each field under minimum necessary access. Ask for a data-flow diagram, sample API payloads with sensitive fields clearly labeled, a data-classification matrix, and a written purpose limit. The point is simple: if a vendor says it needs data, it should be able to show which data and why.
Red flag: fuzzy language about using data to improve services, or collection of full records when a smaller de-identified, tokenized, or synthetic data set would do the job.

Question 2 - Retention. Ask for separate retention rules for source files, prompts, outputs, backups, logs, and support data. Then ask for written deletion proof and whether legal holds can delay deletion. This matters because “deleted” can mean very different things depending on where the data lives.
Red flag: indefinite retention or deletion that only removes data from the user interface.

Question 3 - Model training. Check whether the vendor can limit training use across prompts, outputs, metadata, logs, and third-party model providers. Ask directly whether prompts, clinical notes, imaging files, outputs, feedback, and logs are used for fine-tuning, reinforcement learning, evaluation, or human review. Then confirm that any third-party foundation model provider is bound by the same limits.
Red flag: opt-out-only controls, or answers that don’t match between the sales proposal and the privacy policy.

Question 4 - Security controls. Map safeguards to the actual workflow, not just a badge on a security page. Ask how clinical notes are protected during API transit, whether support engineers can access production data, and whether audit logs cover prompts, exports, and administrative actions. If any safeguard depends on another company, verify that the downstream provider is covered before you move to Question 5.
Red flag: a nice-looking certification that doesn’t apply to the AI or model environment you’re reviewing.

Question 5 - Subcontractors and model providers. Ask for a named, current list that includes cloud infrastructure providers, foundation-model APIs, annotation firms, support platforms, and any offshore personnel with data access. Then ask what data each party receives, where that data is processed, and whether that party can retain or train on it on its own.
Red flag: unnamed partners or a model provider whose own terms allow unrestricted training.

Question 6 - Compliance and remediation. Assign one accountable party for each task. Ask the vendor to name who owns risk-analysis inputs, access approvals, breach investigation, regulator and patient notifications, audit support, and corrective action. A solid answer names accountable executives, sets response-time commitments, and explains how the vendor will preserve evidence and support the investigation.
Red flag: no contractual breach-notification timeline, refusal to support audits, or remediation commitments that can’t be measured.

Use the table below as a fast screen after reviewing each answer.

Question Key Evidence to Request Primary Red Flag
Data access and purpose Data-flow diagram, field-level inventory, sample API payloads Broad rights to use data for any business purpose
Retention Retention schedule by data type, deletion SOP, backup policy, written deletion proof Indefinite retention; UI-only deletion
Model training Data-use policy, model-provider terms, opt-out configuration Opt-out-only controls; answer differs from privacy policy
Security controls SOC 2 Type II, pen-test summary, access-control matrix, IR plan Certification excludes AI/model environment
Subcontractors Subprocessor register, downstream BAAs, geographic processing map Unnamed partners; no subprocessor change notice
Compliance ownership BAA, responsibility matrix, breach-notification SLA, audit reports Customer solely responsible; no contractual remedy

How to Turn Vendor Answers Into a Risk Decision

After you’ve gathered responses to all six questions, score them as a set, not one by one. The goal is to turn the full picture into one of three calls: approve, conditionally approve, or reject.

That matters because a single decent answer doesn’t tell you much on its own. What you want is a pattern. Written proof and clear controls lean toward approval. Vague policy wording or verbal promises push things toward rejection.

Use the decision table below to make a clear approve, conditionally approve, or reject call.

Some gaps should lead to an automatic reject. Reject vendors that:

  • refuse a BAA
  • give vague training answers
  • offer only a SOC 2 Type I report

A conditional approval should happen only when the vendor agrees, in writing, to fix specific gaps before go-live. That could mean putting AES-256 encryption at rest in place or setting a defined retention schedule with automated deletion. If the gap can be fixed, move it into conditional approval and track the remediation.

Keep each answer, evidence item, gap, owner, and due date in a single risk register. Include the supporting evidence for every answer. Assign one accountable owner and a due date to every open item, then review the documentation at least annually to confirm the vendor’s controls and disclosures haven’t changed [8].

Track compensating controls in that same record too. Common examples include RBAC, audit logs, breach response, and pseudonymization [8].

Decision Table for the Six Questions

Use this table to turn each vendor answer into a plain approve, conditional, or reject decision [8].

Question Acceptable Answer Evidence Required Risk Implication if Missing Fix Required Owner / approver Approval Status
What data will you access, and why? PHI access is limited to what’s needed for operations; identifiers are pseudonymized before outside processing Data processing policy; BAA template Unauthorized PHI exposure; HIPAA violation Reject the vendor if it refuses a BAA; require written pseudonymization before go-live Legal / Compliance Officer; Practice Owner / CEO Reject if BAA is absent; Conditional if scope is unclear and narrowed in writing
How long will you retain our data? Set retention windows with automated deletion; written zero-retention for audio when audio is not needed Written retention schedule; automated deletion process PHI stays on vendor servers with no end date; breach and subpoena risk Set strict, automated deletion windows for transcripts Privacy Officer; Compliance Lead Conditional until the schedule is documented
Will you use our data to train or fine-tune models? Clear contract language banning PHI use for training or fine-tuning Contract clause; architecture diagram showing data segregation Patient data could be used to improve outside models Add a contract clause that clearly bans training use; verify logical data siloing Data Governance / Privacy Officer; Chief Medical Officer (CMO) Reject if there is no written prohibition
Which security controls protect the data? TLS 1.2 or 1.3 in transit; AES-256 at rest; a recent SOC 2 Type II report covering the live service SOC 2 Type II report; encryption configuration docs; RBAC and audit log records Security posture is unverified; only marketing claims back it up Ask for a SOC 2 Type II report; verify encryption protocols; confirm RBAC and audit log controls IT / Security Manager; CISO Reject until a Type II report and control evidence are provided
Who are your subcontractors and model providers? Full list of third-party model providers; each covered by a BAA or equivalent Subprocessor list; evidence of downstream BAAs Downstream risk; PHI could be shared with uncovered vendors Require a full subprocessor list; verify each provider meets HIPAA rules Procurement / Security; Legal Counsel Conditional until the list is verified
Who owns compliance and remediation? Vendor sets breach notification timelines and accepts remediation responsibility in writing BAA breach-notification clause; written remediation commitments Unclear ownership slows breach response; regulatory exposure Define notification timelines and remediation duties inside the BAA Compliance Officer / Legal Counsel Conditional until the BAA clause is confirmed

A conditional approval shouldn’t stay informal. If a vendor says the right thing on a call but the contract says nothing, that’s a problem waiting to happen. Before go-live, every conditional item should show up in signed contract language.

What to Put in the Contract Before Go-Live

After you score the six answers, move forward to contract review only with vendors that still pass. The six-question review is your filter. It tells you whether a vendor even deserves contract language in the first place. From there, the contract turns every “yes, if...” into something you can enforce.

Treat the BAA as a starting point, not a green light. At this stage, the issue isn't whether the vendor sounds okay. It’s whether the contract makes those answers binding. That means adding AI-specific terms for data use, retention, and security.

The contract should clearly cover:

  • Restrictions on model training
  • Logging of prompts and outputs
  • Protection against prompt-manipulation attacks
  • Data deletion at the end of the relationship
  • What happens to any data used for fine-tuning or model improvement

Data use and deletion are only part of the picture. The contract also needs to lock in security commitments and audit rights. Require a current SOC 2 Type II report or HITRUST CSF evidence before signing. You should also require audit rights, proof that controls are working, and a set timeline for fixing gaps. Security proof has to live in the contract. A vendor’s verbal or written assurance on its own is not enough to support a go-live decision. The agreement should also name the process the vendor will use to notify your organization about any changes that affect your compliance duties.

Incident response and remediation need the same level of detail. Spell them out. Assign a named operational owner to each requirement so there’s no confusion about who is on the hook. Set breach-notification timing in hours, not vague language. Define what remediation looks like if output drift or degradation shows up. If any of these terms are missing, the vendor is not ready for approval.

Conclusion

After working through these six questions, the call gets pretty clear: ask for written proof or walk away.

These questions shift the conversation from sales talk to evidence you can check. That matters because the deal-breakers tend to show up in the same places every time: retention, training, security, subcontractors, and accountability. Paper beats pitch.

You don’t need a perfect vendor. You need one that will put the key terms in writing before go-live:

  • a signed BAA
  • a written no-training clause
  • a current SOC 2 Type II report

If a vendor drags its feet or won’t commit on paper, stop there. Do not onboard them.

FAQs

What should we ask for before signing a BAA?

Before you sign a Business Associate Agreement (BAA) with an AI vendor, check that it covers AI-specific risk, not just broad legal boilerplate.

One point matters a lot: the agreement should clearly ban the vendor from using PHI for model training, fine-tuning, or service improvement unless you give direct, written consent. If that language is vague, that's a red flag.

You’ll also want the BAA to spell out:

  • permitted uses and disclosures
  • breach notification timelines
  • the same HIPAA duties for subcontractors
  • audit rights for AI workflows and security controls
  • secure PHI return or destruction
  • coverage across the full AI stack, including inputs, outputs, and logs

That last part is easy to miss. An AI tool doesn’t just touch the data you enter. It can also create outputs, store logs, and pass data through other layers behind the scenes. If the BAA only covers one slice of that flow, you may have a gap.

How can we verify an AI vendor is not training on PHI?

Don’t settle for broad promises. The contract should clearly ban the use of PHI for model training, fine-tuning, or service improvement unless you’ve given prior written approval.

You should also ask for:

  • A data-flow diagram that shows where data goes, where it’s stored, and who can access it
  • Proof of technical guardrails and audit logs
  • Written confirmation that the BAA covers AI inputs, outputs, and logs
  • Terms for the return or destruction of data at the end of the contract, including any fine-tuning artifacts

That last point matters more than it may seem. If PHI touches a custom model, prompt history, or logging system, you need clear language on what happens when the relationship ends. Otherwise, data can linger in places no one meant to leave it.

What red flags should make us reject an AI vendor?

Reject an AI vendor if they can’t support their claims with hard proof, like validation metrics, model cards, or architecture diagrams.

Other red flags are just as serious:

  • Unclear training data provenance
  • Missing subprocessor or fourth-party details
  • Refusal to sign a BAA that covers AI features
  • Vague terms around PHI retention, deletion, or model training
  • No clear plan for human oversight, incident response, or model update notifications

If a vendor gets fuzzy on any of this, that’s a bad sign. In healthcare, vague answers aren’t good enough.

Related Blog Posts