The model is the easy part. Every major LLM provider signs a BAA now. In healthcare, agent products are won or lost on three questions: who the agent acts as, where PHI goes, and whether you can prove both in a security review.
The BAA chain has a gap at the tool layer. Anthropic's BAA excludes most MCP and third-party integrations, and OpenAI's excludes live web search. Tool calls that carry PHI need to run through infrastructure that's covered by your BAA chain, or inside your customer's boundary.
Service accounts are the most expensive shortcut in healthcare. They demo well, then fail the audit-controls and minimum-necessary questions on every health-system questionnaire. Retrofitting per-user identity later means a rewrite.
Design for the security review from day one: delegated access, metadata-only audit logs, approval gates on writes, a written PHI data path, and a deployment option that runs in the customer's cloud.
Use the buyer's checklist at the end to check your architecture before the first enterprise questionnaire arrives.
Why Healthcare Is Different for AI Agents
Healthcare is the fastest-adopting AI market and the hardest to ship into. Menlo Ventures counted $1.4B of healthcare AI spend in 2025, nearly 3x 2024, with health systems adopting at 27%. Bessemer reports AI took 55% of health tech venture funding in 2025.
The spend is moving from documentation to action. Ambient scribes ($600M) came first. Prior authorization grew about 10x, and coding and billing reached $450M. Actions mean integrations, and integrations in healthcare come with four constraints most SaaS agents never face:
Every system that touches PHI is in scope for HIPAA, including your tool-calling layer, logs and vector stores.
Every action must be attributable to a person. HIPAA's audit controls (45 CFR 164.312(b)) and minimum-necessary rules assume actors have roles.
The buyer runs a formal security review before any PHI flows, and increasingly one specific to AI.
Incumbents are building agents too. Epic has Art, Emmie and Penny, plus Agent Factory (previewed at HIMSS26). Oracle Health has its Clinical AI Agent. Your agent has to work alongside the EHR's own agents, through the certified APIs.
Step 1: Map Your BAA Chain Before You Write Code
A business associate (BA) is any vendor that creates, receives, maintains or transmits PHI for a covered entity. BAs must flow the same obligations down to every subcontractor that touches PHI.
For an AI agent product, the chain typically looks like this: the health system (covered entity) signs a BAA with you (the AI agent vendor), and you sign BAAs with your LLM provider, your tool-calling/auth layer, your logging and observability stack, and your cloud host. Your tool-calling layer then signs BAAs with its own subprocessors.
Two points trip teams up:
Encryption doesn't remove BA status. HHS FAQ 2076 says a cloud provider holding only encrypted ePHI, without the key, is still a business associate. Any layer that sees tool payloads or stores PHI in logs is in the chain.
The conduit exception is narrow. It covers transmission-only services like ISPs. A gateway that caches, logs or retries requests almost certainly doesn't qualify.
Our view: the cleanest chain has the fewest links that persist PHI. Choose infrastructure that processes payloads in memory, logs metadata only, and can publish a short subprocessor list for BAA accounts.
Step 2: Choose Your LLM Provider for Its BAA Coverage, Not Just Its Model
Every major provider signs a BAA as of 2026, so the differentiator is which features the BAA covers.
Provider
BAA path
Notable exclusions or conditions
OpenAI API
Request via baa@openai.com; case-by-case
Web search with live internet access isn't HIPAA eligible. Default abuse-monitoring retention is up to 30 days unless zero data retention (ZDR) or modified retention is approved.
Anthropic (Claude API / Enterprise)
BAA for commercial customers; Claude for Healthcare launched January 2026
Excludes Batch, Files, Code Execution, Computer Use, Web Fetch, and most integrations that send data to third parties, including MCP connectors.
AWS Bedrock / AgentCore
AWS BAA
Both are HIPAA eligible; confirm model availability by region.
Check the covered products list for the exact service names.
What the table implies: if your agent reaches EHRs, payer portals and SaaS tools through a model provider's hosted connectors, those calls may fall outside the model provider's BAA. Run tool execution through your own BA-covered layer instead, or inside the customer's cloud, and give the model only the tool results it needs.
Step 3: Decide Who Your Agent Acts As
This is the architectural decision that's hardest to reverse.
Service account
Delegated (per-user) access
How it works
One credential for all users and all workflows
Each user authorizes once; the agent acts with their scoped, short-lived tokens
EHR audit log shows
svc-agent-prod
Dr. Chen via delegated authorization, patient/MedicationRequest.write
Minimum necessary
The union of every permission any workflow needs
Only what that user's role allows
Offboarding
Access survives every departure
Tokens die with the user's account
Security review
Stalls
Answers the attribution question by construction
For EHRs, delegated access means SMART on FHIR user-context authorization. The fhirUser claim ties the token to the actual clinician or patient. For SaaS tools (email, CRM, ticketing), it means per-user OAuth. The same principle applies across both.
Keep system-level credentials for jobs where no person is acting, such as bulk FHIR export for analytics, and isolate them from the interactive agent.
Most healthcare agent workflows touch three to six systems, each with its own auth model and BAA.
Use case
Systems touched
Integration and auth reality
Prior authorization
EHR (FHIR), payer portals, fax, clearinghouse (X12 278), CMS coverage data
Portal logins with MFA per payer; FHIR Prior Authorization APIs due from payers by Jan 1, 2027 (CMS-0057-F). Legacy and FHIR paths coexist until then.
RCM, claims, eligibility
Clearinghouses (Availity, Optum, Stedi), practice management, X12 270/271, 837, 835
API keys, SFTP, enrollment per clearinghouse
Scribe write-back
Epic, Oracle Health, athenahealth, eClinicalWorks
User-context SMART on FHIR; per-health-system app activation; narrow write scopes
Scheduling and patient access
EHR scheduling APIs, telephony, CRM, SMS
Patient identity verification; AI disclosure laws when content is clinical
Care coordination
EHR, care management, HIE, email, fax
Cross-organization sharing, consent
Step 5: Write Down the PHI Data Path, Hop by Hop
"We don't store PHI" is the most common claim in healthcare security reviews, and the most commonly wrong. Define it for every hop:
Model provider: ZDR or modified retention enabled? Which endpoints?
Tool-calling layer: are request and response payloads processed in memory only, or written to disk, queues or caches?
Logs and traces: do they include payloads or prompts? Is redaction applied before storage?
Agent memory and vector stores: namespaced per patient and session? What TTL?
Retries and dead-letter queues: where do failed payloads go?
Credentials: OAuth tokens and API keys aren't PHI, but they grant access to it. Where are they encrypted, and who holds the key?
Publish this as a one-page diagram. Security teams will ask for it, and having it ready shortens the review.
Step 6: Scope Tools to Minimum Necessary
HIPAA's minimum-necessary standard applies to business associates too. For agents, the engineering control is least-privilege tool exposure:
Give each agent (or each role) only the tools its job requires. An inbox summarizer doesn't need delete_mail, and a chart-prep agent doesn't need write tools.
Enforce scope before the call reaches the upstream API, not in the prompt.
Build the tool list from the scopes the user actually granted. SMART servers can grant less than you requested.
Separate read and write tools, and expose write tools only to configurations that need them.
OWASP's 2025 LLM Top 10 lists Excessive Agency as a core risk, and its agentic Top 10 (December 2025) names Identity & Privilege Abuse. Least-privilege tool scoping addresses both. More on this in our guide to agent tool calling auth patterns and anti-patterns.
Step 7: Put Approval Gates on Writes
OAuth scope answers "can this agent call this tool?" It can't answer "should this call, with these arguments, run now?" In healthcare, that second question matters for every write: a note filed to the chart, an allergy updated, a prior-auth submitted.
Pause before high-impact writes and route them to a clinician or reviewer.
Re-validate credentials when the task resumes, since the token valid at pause time may not be valid later.
Log the approval as part of the audit record.
Epic builds this into its API: medication orders from third parties arrive as unsigned orders via CDS Hooks that a clinician must sign. Human oversight on agentic writes is not optional—it's a design requirement in regulated environments.
Step 8: Build Audit Logs That Prove Control Without Storing PHI
Health-system reviewers ask whether you can attribute every access to an individual and produce the record for a specific incident. Design the log for that:
Log metadata: who authorized (the human), which agent, on whose behalf, which tool and connector, which scope, resource identifiers, policy decision (allowed, denied, approved), timestamp, status, latency.
Don't log payloads by default. If audit entries carry request and response bodies, your monitoring pipeline becomes another PHI store with its own BAA and breach surface.
Stream to the customer's SIEM (Datadog, Splunk, Sentinel), since health systems increasingly require agent activity in their own monitoring.
Retention: HIPAA requires six-year retention for required documentation (164.316(b)(2)). It doesn't set an explicit period for raw logs. Many buyers ask for six years anyway. Make retention configurable.
Step 9: Offer a Deployment Model That Fits the Customer's Boundary
Large health systems and payers increasingly ask for private deployment. There are several drivers:
Data residency law. Texas SB 1188 requires electronic health records of Texas patients to be physically stored in the US, including by cloud vendors, from January 1, 2026, with penalties of up to $250,000 per violation.
A shorter subprocessor chain.
Supply-chain caution after the Change Healthcare breach, which affected about 192.7M people.
Plan for at least: managed cloud with a BAA, a VPC option, and ideally customer-infrastructure deployment, all with the same API. The component that most needs to move inside the boundary is the one holding credentials and brokering PHI: your tool-calling and auth layer.
Step 10: Prepare for the Security Review
What health-system reviews ask AI vendors in 2026, based on common healthcare third-party risk questionnaires and the Health Sector Coordinating Council's third-party AI risk guide (April 2026):
A signed BAA and a subprocessor list with flow-down BAAs
SOC 2 Type II dated within 12 months. Larger systems and payers may ask for HITRUST (r2 is the most stringent tier).
Encryption at rest and in transit, key management, BYOK
No training on customer data without explicit consent
MFA and SSO for your admin access
Per-user attribution and minimum-necessary access
Traceable, exportable audit logs
Breach notification within 24–72 hours
Data residency and deployment options
AI-specific disclosures: intended use, known risks, human oversight
Build to the proposed HIPAA Security Rule now. HHS's December 2024 proposal (MFA, encryption, asset inventory and network maps, 72-hour restore, annual audits and pen tests, 24-hour BA notice of contingency-plan activation) wasn't final as of September 2026. It's on the long-term agenda with final action estimated for July 2027. Buyers already put most of it in their questionnaires.
Governance frameworks are arriving too. The Joint Commission and CHAI published Responsible Use of AI in Healthcare guidance in September 2025, with a voluntary certification planned. Expect health systems to pass its data-security and monitoring elements through to vendor contracts.
In force. The Security Rule update was proposed; final action estimated July 2027.
21st Century Cures Act (information blocking)
EHR vendors can't block access through certified APIs; up to $1M per violation. HTI-5 (proposed) would add automated access, including autonomous AI systems, to the definitions; ONC FAQs cover agentic AI.
In force; HTI-5 proposed
CMS-0057-F
Payer FHIR Prior Authorization APIs
Decision timeframes since Jan 1, 2026; APIs due Jan 1, 2027
Section 1557 §92.210
Covered entities must identify and mitigate discrimination risk in patient-care decision support tools
In force since May 1, 2025
FDA CDS guidance
Clarifies when clinical decision support is a device; relaxed in January 2026
In force
California AB 3030
Generative AI patient communications about clinical information need a disclaimer and a path to a human
In force since Jan 1, 2025
Texas TRAIGA and SB 1188
AI-use disclosure to patients; US-only storage of Texas patients' EHRs
TRAIGA and SB 1188's data localization since Jan 1, 2026 (other SB 1188 provisions since Sept 1, 2025)
Illinois HB 1806
AI can't make independent therapeutic decisions in behavioral health
In force since August 2025 (effective on signing)
Colorado SB 26-189
Notice and disclosure for consequential decisions (HIPAA-covered entities largely exempt for non-employment uses)
Effective Jan 1, 2027
Five Mistakes We See Healthcare Agent Teams Make
Shipping on a service account "for now." The rewrite arrives with the first enterprise questionnaire.
Assuming the model provider's BAA covers the tools. Check the feature exclusions.
Logging full payloads for debugging and then discovering the observability vendor is now a BA.
Hardcoding one EHR's configuration. Every health system is a separate tenant with its own base URL, activation and scopes.
Treating self-hosting as a later problem. The first large health system or payer often asks for it, and a separate self-hosted codebase is a second product.
How Scalekit Fits
Scalekit AgentKit is the authorization and tool-calling layer for healthcare agents. It's the part of this guide that most teams would otherwise build and maintain themselves.
Delegated access everywhere: SMART on FHIR connectors for any FHIR server (plus pre-built AdvancedMD), and 500+ connectors with 20K+ actions across SaaS and healthcare tools. Every call runs as the authorizing user.
PHI stays out of Scalekit: payloads are processed in memory and never written to disk or logs. For BAA accounts, the only subprocessors are GCP and CockroachDB, both under BAAs.
Scoped tool surfaces per agent and role with Virtual MCP Servers, enforced before the upstream call.
Metadata-only audit logs streamed to your SIEM.
Deployment: our cloud (US or EU), your VPC, your customer's infrastructure, or air-gapped, with the same SDKs. Bring your own key from your own KMS.
Compliance: SOC 2 Type II, ISO 27001, HIPAA BAA on Enterprise. We don't train models on your data.
Talk to us about your healthcare agent, or start with the SMART on FHIR connector. Understanding access control for multi-tenant AI agents is a natural next step as you design your tenant isolation model.
The Healthcare AI Agent Buyer's Checklist
Use this to audit your own architecture, or give it to the vendors you're evaluating.
Contracts and compliance
BAA with the customer, and flow-down BAAs with every subprocessor that touches PHI
US data residency (Texas SB 1188), EU where needed
VPC and customer-infrastructure deployment with the same API
BYOK for stored credentials; revocation tested
AI governance
Intended-use and known-risk documentation (CHAI-style model card)
Patient-facing AI disclosures where state law requires them (CA AB 3030, TX TRAIGA)
Prompt-injection controls on untrusted inputs (faxes, portal pages, patient messages)
Frequently Asked Questions
Do AI agent vendors need to sign a BAA?
Yes, if the agent creates, receives, maintains or transmits PHI for a covered entity or another business associate. The vendor is then a business associate and must sign a BAA with its customer, then flow the same obligations down to subprocessors that touch PHI, including the LLM provider and the tool-calling layer.
Does OpenAI or Anthropic sign a BAA?
Both do as of 2026, with feature exclusions. OpenAI's API BAA excludes web search with live internet access. Anthropic's BAA excludes Batch, Files, Code Execution, Computer Use, Web Fetch and most third-party integrations, including MCP connectors. AWS Bedrock, Azure OpenAI and Google Cloud's Gemini services are covered under their cloud BAAs.
Can I use MCP servers with PHI?
Yes, if every MCP server and gateway in the path is covered by your BAA chain and the data path is controlled. Hosted MCP connectors offered through a model provider may be excluded from that provider's BAA. Running MCP tool calls through your own BA-covered layer, or inside the customer's cloud, keeps them in scope.
Is zero data retention required for HIPAA?
No. HIPAA requires a BAA and appropriate safeguards, not zero retention. ZDR reduces risk by limiting where PHI persists, which simplifies the data path and the security review. Many health-system buyers now ask for it for model providers and tool layers.
How long must HIPAA audit logs be retained?
HIPAA requires six-year retention of required documentation (45 CFR 164.316(b)(2)). It doesn't set an explicit retention period for raw system logs. Many healthcare buyers still ask for six years, so make audit-log retention configurable and keep logs metadata-only so long retention doesn't mean long PHI storage.
Do hospitals require HITRUST from AI vendors?
Some do, especially large health systems and payers for vendors hosting PHI at scale. Many mid-market buyers accept SOC 2 Type II plus a BAA and a HIPAA-aligned control set. HITRUST has three levels (e1, i1, r2), and r2 is the most stringent. Ask your target buyers early, because an r2 assessment takes months.
What is the biggest mistake teams make building healthcare AI agents?
Using a shared service account for EHR and tool access. It works in a demo, but it can't answer the audit questions health systems ask: who performed each action, whether access matched the user's role, and what happens when a user leaves. Moving to per-user delegation after launch is usually a rewrite. For a broader view of why this matters, see our analysis of API access patterns for AI agents and identity architecture.