Announcing CIMD support for MCP Client registration
Learn more

Self-Hosted Agent Integrations: VPC, On-Prem, and Air-Gapped, Explained

Self-hosted agent integrations: VPC, on-prem and air-gapped. Scalekit, the token vault, the root key and your agent inside your infrastructure boundary; connected apps reached across the line; Scalekit cloud outside with no route in.
Vishal Dhawani
Founding Architect @ Scalekit

TL;DR

  • When a customer asks whether your agent can run inside their boundary, they are asking about one layer of your stack: the one that holds their users' credentials and executes every action in their systems. That is the agent integration layer, and it is the piece that has to move.
  • There are three ways to run it on your side of the line. In your own VPC. As one isolated instance inside each customer's infrastructure. Or air-gapped, on a network with no route to the vendor at all.
  • The modes differ in what still crosses the boundary. In a VPC, license validation and update channels may call out. Air-gapped, the instance makes no calls to the vendor, and the only outbound traffic is whatever your connectors need to reach the apps they call.
  • Two things change the moment you self-host, and teams usually find them late: you register your own OAuth app with each provider, on your own domain, and every connector needs a network path to its provider.
  • The SDK has to stay the same. If the self-hosted edition needs different calls from the cloud edition, a customer's deployment requirement has just given you a second integration to maintain.

The call that starts this

It usually happens in the first security review. Your agent product has passed the demo, the champion is sold, and the customer's security architect asks the question you knew was coming: “Can this run in our VPC?” Sometimes it is phrased as “we cannot have credentials for our systems sitting in a third-party database.” Sometimes it is “our network has no outbound internet.” The question is the same each time. They want the part of your product that touches their systems to live where they can see it.

This post is about that part. It explains what the agent integration layer does, why it is the piece customers want inside their boundary, the three deployment modes you can offer, what crosses the boundary in each, and the two operational changes that catch teams by surprise. It uses Scalekit's self-hosted deployment as the worked example because that is what we build, and it compares how other vendors approach the same problem near the end.

What the agent integration layer does

Four things have to happen around every action an agent takes in another app. A user authorizes the agent through the app's OAuth flow. The access and refresh tokens that come back are stored, encrypted and refreshed before they expire. When the agent picks a tool, a service looks up that user's credential, attaches it to the request and calls the app's API. And something records who did what.

People call this layer different names depending on where they sit: AI agent integration middleware, a tool-calling gateway, an auth layer for agents. Whatever the name, it has three parts:

  • A token vault. Encrypted storage for each user's OAuth tokens and API keys.
  • A tool-calling gateway. Receives the agent's tool call, injects the right credential at call time and calls the app.
  • An audit trail. A metadata record of each call: which user authorized it, which agent made it, which tool ran and what happened.

The model, the orchestrator and the agent's memory sit elsewhere. They come up in a deployment review too, but they are usually already inside the customer's application boundary or covered by a model provider's contract. The integration layer is different. It is the piece most often bought as a third-party service, and it holds standing access to everything the agent connects to. That combination is why a review of AI agent integration architecture starts here.

Why customers want it inside their boundary

The request comes from one of four places, and it helps to know which one you are dealing with because each points to a different deployment mode.

  • Credential custody. The security team does not want refresh tokens for their CRM, mail and file storage in a vendor's database. Healthcare buyers add a regulatory version of the same concern: every hosted service that touches PHI is another business associate in the chain.
  • Data residency. Financial services and public-sector contracts name where data may be stored, and some name the customer's own cloud account.
  • Network policy. Defense networks and some government networks have no outbound internet, so any software they use has to run entirely inside.
  • Vendor risk. Each hosted hop is a dependency the customer has to assess, monitor and list in their own records, so every hop you remove shortens their review.

Many buyers who start with one of these are satisfied by a managed region in the right jurisdiction, and self-hosting would be more than they need. The EU-hosted agent integrations post in this series covers that case. This post is for the buyers a region does not satisfy.

The three deployment modes

Scalekit's self-hosted deployment runs in three modes, and they map directly onto the four reasons above. You can run different modes for different customers from the same codebase.

Three self-hosted deployment modes for Scalekit AgentKit: your VPC, your customer's infrastructure, and air-gapped. Each shows the tool-calling gateway, token vault, root key and audit events inside the boundary, and what, if anything, calls Scalekit.

Mode 1: your VPC

The integration layer runs inside a VPC in your own cloud account on AWS, GCP or Azure. Your team operates it, and the token vault lives in your Postgres. This is the usual answer when your own company has the requirement, or when your customers will accept a single-tenant deployment inside your account instead of a shared multi-tenant service.

What still crosses the boundary here: calls from the gateway to each connected app's API and OAuth endpoints, which is the whole point of the layer, and, depending on the vendor, license validation and software updates. In Scalekit's VPC mode, license validation and update channels may call out, both are configurable, and both can be routed through your egress controls so your network team sees exactly what leaves.

Mode 2: your customer's infrastructure

You deploy one isolated instance into each end customer's environment. The customer gets a single-tenant stack inside their own boundary. You, the agent vendor, hold the license keys and control upgrades. This is the mode for selling to banks, health systems and agencies that will not accept any multi-tenant component in the path, and it is the one most vendors have never had to think about until a large deal depends on it.

The cost scales with customers. Fifty customers in this mode means fifty deployments, each with its own upgrade schedule, its own OAuth app registrations and its own license key, because Scalekit licenses per deployment. A customer instance can also be air-gapped. That is the usual pattern for defense and other sealed networks: one isolated instance per customer, with no route out.

Mode 3: air-gapped

The network has no route to the internet, or none to the vendor. Three things have to be true before any software can run there.

  1. License validation works offline. A license check that calls home at boot fails in a sealed network.
  2. Images come from an internal registry. Kubernetes pulls every container image from a registry inside the network.
  3. Nothing phones home at runtime. No telemetry, no remote logging, no update checks, and no static assets fetched from a CDN.

Scalekit's air-gapped mode is built to those three conditions. It makes no outbound calls to Scalekit, validates the license offline, and pulls images from your own registry. Once installed, the instance has no connection to Scalekit's infrastructure: credentials, tool calls and logs stay in your network. Installation, configuration and upgrade guides come with the license.

One consequence of an air gap is easy to miss until the first connector fails: a connector only works if the cluster can reach the app it calls. In a fully sealed network that means internal apps and the custom connectors you build for internal APIs. A connector to a public SaaS app needs a route to that app. A deployment that is air-gapped from the vendor but still reaches Salesforce has a narrow, controlled egress to Salesforce, and the customer's review will treat that differently from a network with no egress at all. Describe it that way from the start.

What crosses the boundary in each mode

VPC
Customer infrastructure
Air-gapped
Token vault
Your Postgres
The customer's Postgres
Postgres inside the sealed network
Tool-calling gateway
Your Kubernetes cluster
The customer's cluster
A cluster inside the sealed network
Audit events
Your logging stack
The customer's logging stack
Logging inside the sealed network
Calls to connected apps
Outbound HTTPS to each provider
Outbound HTTPS, under the customer's egress policy
Internal apps only, or explicit allowlisted routes
Calls to Scalekit
License validation and update channels, configurable
Same as VPC, or none if air-gapped
None

In every self-hosted mode, user records, sessions, OAuth tokens and access keys live in your Postgres, not on Scalekit's infrastructure. There is no default outbound metrics channel, and logs reach Scalekit for support only if you decide to share them; by default, Scalekit has no log access.

Two changes teams miss when they self-host

Your OAuth apps move to your domain

In a managed cloud, many integration vendors offer their own OAuth app for popular providers so you can test without registering anything. Scalekit's managed cloud does this too: pick “Use Scalekit credentials” when you create a connection and you are testing against Gmail in a few minutes. That app belongs to Scalekit's cloud, and its redirect URI points at Scalekit's domain.

A self-hosted instance cannot use it. You create an OAuth app with each provider and register the redirect URI your instance shows in the connection form, which is on your own domain. Plan for this early, because some providers review apps that request sensitive scopes before external users can authorize them, and that review can take weeks. In customer-infrastructure mode, decide up front whether each customer registers their own OAuth apps, which puts the customer's name on the consent screen, or whether you register one set per instance.

Every connector needs a network path

A self-hosted gateway calls each connector's API and OAuth token endpoint directly. The cluster needs outbound HTTPS to every provider your agents use: Google, Slack, Salesforce, whichever apps your product connects to. Separately, your users' browsers need to reach each provider's consent screen when they authorize a connected account. That page loads in the user's browser, so a locked-down cluster does not block authorization as long as users can reach the provider.

Connectors for apps inside your network, including the ones you define for internal APIs, need no internet access at all. In an air-gapped deployment, those are the connectors that work without any exception being written.

Why the SDK has to stay the same

Teams build against the managed cloud first, often for months, before a customer's security review asks for a private deployment. If the self-hosted edition has a different API surface or a smaller feature set, that review becomes a rewrite, and you maintain two integrations for the life of the contract.

Scalekit's SDKs and REST API work the same way in the managed cloud and in a private deployment. Moving an app to a self-hosted instance is a configuration change: point it at the instance's environment URL and use credentials from that instance's own dashboard. The code you built against the cloud runs in your cluster instead of ours. The self-hosted AgentKit docs walk through the switch.

Before a deal depends on it, run your integration tests against the self-hosted instance as well as the cloud. You want any surprise in your own test run, not the customer's deployment.

Keys, retention and audit inside the boundary

Moving the gateway inside the boundary answers the “where does it run” question. A security architect will ask three more.

  • Who holds the root key. The vault encrypts each credential with a data key, and the data keys are wrapped by a root key held in a key management service. A vendor-managed root key in a KMS is a sound default. When you want the extra control, the root key can be yours: you register a key from your own Google Cloud KMS, rotation follows your policy, each wrap and unwrap of the data key shows up in your Cloud KMS audit logs (with Data Access logs enabled), and revoking Scalekit's access makes stored credentials unreadable to Scalekit until you restore it. For an air-gapped install, agree how the root key is managed with Scalekit when you scope the deployment. The bring-your-own-key post in this series goes through the key hierarchy and what revocation does in practice.
  • What is kept after a call. Tool calls return business data, and the question for a review is retention. Scalekit does not retain tool-call requests or responses. They are processed in memory for the duration of the call and are not written to Scalekit's database or logs. Logs keep call metadata.
  • Where the audit trail goes. Each call is recorded as metadata: who authorized it, which agent, which tool, status and latency. Successful responses are not written to that log. The app's error response is kept only if you turn on Store connector error details. Self-hosted, those records stay inside the customer's network with everything else.

What a self-hosted deployment needs

You provide the infrastructure; the platform ships as a Helm chart.

Dependency
Requirement
Kubernetes
1.27 or later, managed or self-managed, with Helm 3.12 or later
Ingress
Kubernetes Gateway API or the nginx ingress controller
PostgreSQL
15 or later (CockroachDB is also supported)
Redis
6.2 or later
SMTP
Any provider, for team invitations and sign-in email
Domain
A domain and TLS certificate for your instance

Kubernetes with Helm is the recommended production setup. Docker Compose packaging exists for development, testing and small deployments. Installation, configuration and upgrade guides come with the Enterprise license.

How other vendors approach self-hosting

The pattern above is Scalekit's. Here is how the other platforms in this category describe their own self-hosting, from their own pages as of October 2026. For the full side-by-side, including key custody and compliance, see Best self-hosted agent integration platforms.

  • Paragon offers three self-hosted models: unmanaged, where you install with its Terraform installer and Helm charts; managed, where Paragon runs it in your account; and forward-deployed, where Paragon deploys into your customer's VPC. Data lives in your VPC's Postgres, an S3-compatible object store and Redis. Billing, license verification and anonymized usage analytics reach Paragon's cloud, with the analytics disableable, and Paragon says it has arrangements for a complete air gap.
  • Nango publishes a free self-hosted edition that deploys with docker-compose and covers managed auth and the API proxy. The full platform, with Helm deployment on AWS, GCP or Azure, is on Enterprise.
  • Composio offers self-hosting on its Enterprise tier for teams whose infrastructure, network or residency requirements put the action layer inside their own environment.
  • Arcade deploys into your cloud through the Azure Marketplace, an AWS private offer or Helm on your own cluster, and lists a fully air-gapped option on Enterprise. Its hybrid MCP servers run in your environment but connect back to Arcade Cloud, so a sealed network needs the full deployment.

Choosing a mode

  • Your own company has the requirement, or your customers accept a single-tenant deployment in your account → VPC.
  • You sell an agent product to customers who require the stack inside their own environment → customer infrastructure, one instance per customer. Budget for per-customer OAuth apps, licenses and upgrade windows.
  • The customer's network has no outbound internet, or policy forbids any vendor connection → air-gapped. Confirm early which connectors the agent needs and whether any of them require an allowlisted route out.
  • The customer's real requirement is that data stays in a region → a managed region may be enough. The EU-hosted agent integrations post covers what that means hop by hop.

Where Scalekit fits

Scalekit AgentKit runs as a managed cloud with US (Los Angeles) and EU (Frankfurt) regions, and as a self-hosted deployment in all three modes above on the Enterprise plan. The full stack self-hosts: auth, sessions, the token vault and the tool-calling gateway, with the same SDKs and APIs you built against in the cloud. The catalog covers 500+ connectors and 20,000+ actions, plus custom connectors for internal APIs that never leave the network. Scalekit holds SOC 2 Type II and ISO 27001 and signs a HIPAA BAA on Enterprise, and in a self-hosted deployment Scalekit is not a sub-processor, because customer data stays in your deployment. If a customer has asked you the question this post opened with, Scalekit self-hosted is where to start scoping the answer.

The deployment questions here are one of three groups a regulated buyer's review covers. The security questionnaire for AI agent vendors has all three, with Scalekit's answers.

Frequently asked questions

What is the difference between a VPC deployment and an on-prem deployment for agent integrations?

A VPC deployment runs the integration layer in a cloud account you control, usually the agent vendor's own AWS, GCP or Azure account. An on-prem or customer-infrastructure deployment runs a separate, isolated instance inside each end customer's environment. With Scalekit, the agent vendor holds the license keys and controls upgrades for each instance. Air-gapped is an on-prem deployment on a network with no outbound route to the vendor.

Can AI agent integrations run in an air-gapped environment?

Yes, if the license validates offline, images come from an internal registry and nothing calls home at runtime. Scalekit supports this mode. Connectors still need a network path to the apps they call, so a fully sealed deployment works with internal apps, and SaaS connectors need an allowlisted route.

Which part of the agent stack should be self-hosted first?

The integration layer: the token vault and the tool-calling gateway. It holds OAuth tokens for every connected app and executes every action. The model can usually stay with a cloud provider under that provider's contract.

Do self-hosted agent integrations still need internet access?

The gateway needs to reach the APIs and OAuth endpoints of the apps it calls, and users' browsers need to reach each provider's consent screen. Calls to the vendor depend on the mode. In Scalekit's VPC mode, license validation and update channels may call out and can be routed through your egress controls. In air-gapped mode, nothing calls Scalekit.

Does self-hosting change the SDK or API?

It should not. With Scalekit, the SDKs and REST API work the same way in every deployment mode. You point your app at the self-hosted instance's environment URL and use credentials from that instance.

Is the integration vendor a sub-processor when the layer is self-hosted?

For Scalekit, no. In an on-prem deployment customer data stays in your deployment, so Scalekit is not listed as a sub-processor. A HIPAA BAA is still available on Enterprise if your compliance team wants one in place.

No items found.
Agent
Auth Quickstart
On this page
Share this article
Agent
Auth Quickstart

Acquire enterprise customers with
‍zero upfront cost.

Every feature unlocked. No hidden fees.