TL;DR
- Health systems ask for private deployment for concrete reasons: a shorter BAA chain, data-residency law (Texas SB 1188's US data-localization requirement took effect January 1, 2026), and supply-chain caution after Change Healthcare. "Can it run in our VPC?" usually comes up in the first security call.
- Self-hosting the model isn't the whole answer. Credentials and PHI concentrate in the tool-calling and auth layer: it holds OAuth tokens for every EHR and SaaS app, and it brokers every request and response. It's the component that most needs to sit inside the boundary.
- "PHI never persists" has to be true at every hop: model provider, tool layer, logs, traces, queues, memory and vector stores. Define it in writing.
- Audit logs should prove control without becoming a PHI store. Log who, which agent, which tool, which scope, and the outcome. Don't log payloads.
- Self-hosting only works if it's the same product. A separate on-prem fork is a second codebase. Look for identical SDKs and APIs across managed, VPC, customer-infrastructure and air-gapped deployments.
Why healthcare buyers ask for self-hosted deployment
1. The BAA chain gets shorter
Every hosted service that touches PHI is a subcontractor business associate. That includes services storing only encrypted PHI: HHS FAQ 2076 says a cloud provider holding encrypted ePHI without the key is still a business associate. Running components inside the customer's cloud removes links from the chain, or at least moves PHI storage inside the customer's control. (A vendor with remote admin access may still be a BA. Have counsel review the structure.)
2. Data residency is now law, not preference
Texas SB 1188 requires electronic health records of Texas patients to be physically maintained in the US, explicitly including by third-party and cloud vendors, from January 1, 2026 (its other provisions took effect September 1, 2025). Penalties run from $5,000 to $250,000 per violation (Hunton). Health systems serving many states are standardizing on "US-only, in our account."
3. Supply-chain risk
The Change Healthcare breach affected about 192.7 million people (HHS). The Health Sector Coordinating Council's April 2026 third-party AI risk guide tells health systems to surface hidden AI dependencies and monitor vendors continuously. Every hosted hop is a dependency to justify.
4. The proposed HIPAA Security Rule favors what you can evidence
HHS's proposed update (December 2024; final action estimated July 2027) would require:
- a technology asset inventory and a network map showing ePHI movement
- restoring certain systems within 72 hours
- annual compliance audits
- business associates verifying their safeguards every year
All of these are easier to evidence when the system runs in the customer's environment (HHS fact sheet).
5. Model-provider BAAs have gaps at the tool layer
Anthropic's BAA excludes most MCP and third-party integrations. OpenAI's excludes live web search. If tool calls carrying PHI run through the model provider's hosted connectors, they may be outside coverage. Running tool execution inside the boundary keeps it in scope.
The deployment spectrum
PHI crosses into the vendor's cloud (ideally in memory only)
Startups, practices, faster procurement
Dedicated / single-tenant cloud
Isolated stack, often with BYOK
Mid-market systems that want isolation without operating it
Your or the buyer's AWS, GCP or Azure account
Shared: vendor ships, you run
PHI and credentials stay in the account
Large health systems, payers, Texas residency
Your end customer's cluster, single tenant
Agent vendor or customer holds keys and controls upgrades
Fully inside the customer
Agent vendors selling to IDNs and payers with strict boundaries
Isolated network, no outbound internet
Government health (VA, DoD), research enclaves
Which parts of the agent stack belong inside the boundary
Most self-hosting conversations start and end with the model. That's the wrong unit. A healthcare agent stack has six components, and they don't carry equal risk.
Prompts and tool results (PHI) in transit
Covered by the provider's BAA and retention mode
Cloud-provider models over private endpoints with ZDR or modified retention; open-weight models only if air-gapped
Conversation state, plans
Your application boundary
Tool-calling and auth layer
OAuth tokens and API keys for every EHR and app; every request and response
Highest: a credential store plus a PHI broker
Inside the boundary, or a vendor with in-memory payloads, per-customer keys and BYOK
Long-lived refresh tokens (for example, three-month EHR refresh tokens)
High: standing access to PHI
Encrypted with a customer-managed KMS key
Whatever you choose to log
High if payloads are logged
Metadata only; stream to the customer's SIEM
Inside the boundary, namespaced per patient and session, with a TTL
Our view: the tool-calling and auth layer is the most underestimated component in healthcare agent security. It's the only piece that holds standing credentials to every clinical and business system and sees every payload. It needs the strongest placement.
Credentials: the vault and bring-your-own-key
An agent that works across Epic, a payer portal, Gmail and Salesforce holds a refresh token per user per system. Under ONC's (g)(10) rules, certified EHRs must issue refresh tokens valid for at least three months to eligible apps. That's standing access to PHI.
What to require:
- Per-customer encryption keys. One compromised key shouldn't expose every tenant.
- Envelope encryption, with the master key in a separate KMS rather than next to the database.
- BYOK. The customer registers a key from their own KMS. The vendor uses the KMS API for every encrypt and decrypt, and never holds the key material. If the customer revokes the key, the vendor can't decrypt stored credentials. That gives the customer a working kill switch.
- Tokens never enter the model context, application code or logs. Credentials are injected at call time.
- Revocation on the next call. When a user or admin revokes access, the next tool call fails closed.
More in secure token management for AI agents and the token vault guide.
"PHI never persists": define it hop by hop
It's the claim every vendor makes. Security reviewers will ask you to prove it for each component:
- Model provider: ZDR or modified retention configured for the endpoints you use.
- Tool-calling layer: requests and responses processed in memory, never written to disk, queues or caches.
- Logs and traces: payloads excluded or redacted before storage.
- Retries and dead-letter queues: no payload persistence, or short TTLs inside the boundary.
- Agent memory and vector stores: inside the boundary, namespaced, with expiry.
- Backups: covered by the same policy.
Document it as a one-page data-flow diagram. That diagram covers the proposed Security Rule's "network map" requirement too.
Audit logging: prove control without creating a PHI store
HIPAA's audit controls standard (45 CFR 164.312(b)) requires mechanisms that "record and examine activity" in systems containing ePHI. For agents, reviewers want to know whose authority was behind each action.
The log schema we recommend:
med-rec-agent / clinician-readonly
smart_fhir / medication_request_search
patient/MedicationRequest.rs
FHIR resource type and ID (not contents)
allowed / denied / approved-by
What not to log: request and response bodies. If audit entries carry payloads, the monitoring pipeline becomes another PHI system, with its own BAA and breach surface.
Retention: HIPAA's six-year rule (164.316(b)(2)) applies to required documentation, and many buyers extend it to audit logs. Metadata-only logs make long retention cheap and safe. Stream to the customer's SIEM (Datadog, Splunk, Sentinel) so their security team monitors agent activity alongside everything else.
Deeper dive: audit trails for agent auth and agent tool observability.
Certifications: SOC 2, ISO 27001, HITRUST, and what self-hosting changes
What healthcare buyers expect
Controls operated effectively over 3–12 months
Baseline; report dated within 12 months
Certified information security management system
Common for international buyers and larger systems
Healthcare-oriented control framework; r2 is risk-based and the most stringent
Often requested by large health systems and payers for vendors hosting PHI at scale. Timelines run months; costs vary widely (secondary estimates put all-in r2 costs in the tens to low hundreds of thousands of dollars).
There's no official HIPAA certification; compliance means a BAA plus safeguards
Signed BAA, subprocessor list, risk analysis evidence
Emerging differentiator in AI-specific vendor reviews
What self-hosting changes: when the infrastructure runs in the customer's account, many infrastructure controls (network, physical, hosting) fall under the customer's certifications. The vendor's certifications still cover the software development lifecycle, release integrity, vulnerability management and support access. Buyers will ask about both.
Agent-specific threats a self-hosted stack still has to handle
Moving infrastructure inside the boundary doesn't fix agent-layer risks.
- Prompt injection leading to PHI exfiltration. Instructions hidden in a faxed referral, a payer portal page or a patient message steer an agent with EHR read access into sending data to an attacker's tool. Mitigate with least-privilege tool surfaces (the agent reading faxes shouldn't have outbound email) and approval gates on sensitive actions.
- Confused deputy. An agent holding a broad token acts for a user who lacks access to that patient. Per-user delegated tokens make the EHR enforce the user's own permissions.
- Token leakage into model context, logs or agent memory. Inject credentials at call time and keep them out of the LLM entirely.
- Cross-patient memory contamination. Namespace memory per patient and session.
OWASP's Top 10 for Agentic Applications (December 2025) lists Identity & Privilege Abuse and Tool Misuse among the top agent risks. Both are addressed at the auth and tool layer, not the model.
Operating a self-hosted deployment: what to plan for
- Parity. Same SDKs, APIs and features as the managed service, or you'll maintain two integrations.
- Upgrades. Who controls the upgrade cadence, and how are releases signed and verified?
- Licensing in isolated networks. Air-gapped installs need offline license validation and images pulled from the customer's own registry.
- Connector updates. Upstream APIs change. How do connector fixes reach a self-hosted install?
- Observability. Metrics and audit events have to flow to the customer's stack, not the vendor's.
- Support access. Define whether the vendor has any access to the environment, and how it's granted and logged.
A decision framework
- Selling to practices and small groups → multi-tenant cloud with a BAA, in-memory payload handling, and metadata-only logs is usually enough.
- Selling to large health systems or payers → offer VPC or customer-infrastructure deployment with BYOK. Expect HITRUST questions.
- Texas patients, or buyers with US-only mandates → US residency at minimum; many will prefer the data in their own account.
- Government health or research enclaves → air-gapped, with locally hosted models.
- You're an agent vendor → pick infrastructure that deploys in your customer's environment with the same API, so a private-deployment requirement doesn't become a fork.
How Scalekit deploys
Scalekit AgentKit is the auth and tool-calling layer for healthcare agents, and it deploys across the whole spectrum without changing your integration:
- Managed cloud: fully managed, with US (Los Angeles) or EU (Frankfurt) data residency.
- Your VPC: deployed inside your AWS, GCP or Azure account. Your network boundary, our software.
- Your customer's infrastructure: single-tenant in your customer's cluster. You hold the keys and control upgrades.
- Air-gapped: no outbound calls to Scalekit, offline license validation, images from your own registry.
In every mode:
- No PHI stored by Scalekit. In cloud, request and response payloads are processed in memory and never written to disk. On-prem, Scalekit the company has no access to the data path; the software runs entirely in your environment.
- Per-customer encryption (AES-256 at rest, TLS 1.3 in transit), with keys wrapped by a KMS master key. BYOK is available: revoke the key and the stored data becomes unusable to us.
- Metadata-only audit logs, streamed in real time to your SIEM, with 90-day default retention (configurable).
- Identical SDKs and APIs across deployment models.
- HIPAA: both Cloud Enterprise and On-Premise Enterprise are HIPAA eligible with a signed BAA. For BAA accounts in our cloud, only GCP and CockroachDB process customer data, both under BAAs. All other subprocessors are disabled.
- Certifications: SOC 2 Type II, ISO 27001, GDPR and CCPA. Reports are available under NDA via our trust center.
- What it connects to: SMART on FHIR for any EHR, a pre-built AdvancedMD connector, and 500+ connectors with 20K+ actions across SaaS and healthcare tools. Every call runs as the authorizing user.
Read the self-hosted deployment announcement, or talk to us about scoping a private deployment.
For the full build and security-review checklist, see the healthcare agent tool-calling auth patterns and anti-patterns. For credential management at scale, see who holds the token across agent tool-calling patterns.
Frequently asked questions
Can agents be self-hosted for HIPAA compliance?
Yes. Self-hosting doesn't make an agent HIPAA compliant on its own, but it shortens the BAA chain and keeps PHI and credentials inside the customer's boundary. You still need safeguards: per-user access, minimum-necessary tool scoping, audit controls and encryption. Model calls typically go to a cloud provider under its BAA over private endpoints.
What's the difference between BYOC and on-prem?
BYOC (bring your own cloud) runs the vendor's software inside your AWS, GCP or Azure account, often with a vendor-managed control plane. On-prem or customer-infrastructure deployment runs entirely in the customer's environment, and they control upgrades and keys. Air-gapped is on-prem with no outbound network access.
Which part of an agent stack should be self-hosted first?
The tool-calling and auth layer. It holds OAuth tokens for every EHR and SaaS app the agent uses, and it brokers every request and response, so it concentrates both credentials and PHI. The LLM can usually stay with a cloud provider under a BAA, with zero or modified data retention over private endpoints.
Should audit logs contain PHI?
No. Log metadata: who authorized, which agent, which tool, which scope, the resource reference, the decision, status and latency. Logging request and response payloads turns your monitoring pipeline into another PHI store, with its own BAA and breach obligations. Metadata-only logs satisfy audit-control requirements and are safe to retain for years.
Is HITRUST required for healthcare AI vendors?
Not by law. HITRUST is a private certification that some large health systems and payers require from vendors that host PHI at scale. Many buyers accept SOC 2 Type II plus a BAA. HITRUST has three levels (e1, i1, r2). r2 is the most stringent and typically takes six months or more.
Does Texas SB 1188 require on-prem hosting?
No. It requires electronic health records of Texas patients to be physically stored in the United States, including by cloud and third-party vendors, from January 1, 2026. US-region cloud hosting satisfies residency, but many health systems use it as a reason to require deployment in their own cloud account.
What does BYOK protect against?
Bring-your-own-key lets the customer hold the master key in their own KMS. The vendor uses it for every encrypt and decrypt operation without holding the key material. If the customer revokes the key, the vendor can no longer decrypt stored data such as OAuth tokens. That gives the customer a kill switch and control over rotation.