
Your agent needs to produce ad creatives, product shots, or short-form video. Higgsfield ships two official ways in: a hosted MCP server that any MCP client can sign into, and a REST API with Python and TypeScript SDKs. In a demo, they look interchangeable. They are not.
One path acts as your user and spends their credits. The other acts as your company and spends yours. One exposes 82 tools, including tools that publish to TikTok and edit website secrets. The other exposes models. Here is how to pick, and how to wire either one into a production agent.
Both paths reach Higgsfield's image, video, and audio models. What differs is who authenticates, what surface the agent sees, and how results come back.
Higgsfield runs a hosted, remote Model Context Protocol (MCP) server, launched in spring 2026 and documented on Higgsfield's official MCP page. It works with Claude, Claude Code, Cursor, OpenClaw, Hermes Agent, and any other MCP-compatible client. There is no local package to install.
Authentication is account sign-in through a browser OAuth flow. Higgsfield's own MCP page states that no API key is required, and Scalekit's connector lists the auth type as OAuth 2.1 with Dynamic Client Registration (DCR). Higgsfield's MCP guide for marketers says an active subscription is required, with a trial offered to eligible new users, and generations draw on the user's plan credits. Every generation lands in the user's Assets library on Higgsfield, tagged with the MCP source.
The Higgsfield API, documented in Higgsfield's official API reference, is an asynchronous REST API. You POST to a model-specific endpoint, receive a request_id with status and cancel links, then poll /requests/{request_id}/status or receive a webhook when the request reaches a terminal state.
Authentication is a server-side credential pair created in the Higgsfield Console, sent as Authorization: Key {key_id}:{key_secret}. There is no OAuth flow and no end-user consent. Billing is pay-per-generation against your account balance, with a dedicated /estimate/... endpoint for pricing a request before you submit it. Official SDKs exist for Python (higgsfield-client) and Node.js (@higgsfield/client).
Higgsfield also ships a CLI (@higgsfield/cli) that coding agents install themselves and authenticate with higgsfield auth login. It targets local developer agents, not multi-tenant backends, so this post treats it as out of scope.
The four dimensions below are the ones that change your architecture: capability coverage, auth model, operational surface area, and fit by use case.
The table compares documented capabilities. "Not documented" means the capability does not appear in Higgsfield's public API reference, not that it is impossible.
The most consequential gap for a production pipeline is event delivery. The API posts terminal results to your hf_webhook endpoint, retries network failures and 5xx responses for up to two hours, and expects you to deduplicate on request_id plus status. The MCP server has no event surface. An MCP agent generating video must keep polling, and jobs_wait caps each long-poll at 15 seconds, while Higgsfield's own tool description puts typical video jobs at 60 to 180 seconds.
That polling happens inside the model loop. Every wait is another tool call, another round of context, and more tokens. The second gap is cancellation. The API refunds a request canceled before processing starts; the MCP exposes no cancel tool at all.
The MCP gap runs in the other direction too: it exposes far more than a creative agent should hold. Launch-era coverage described five core tools. Scalekit's Higgsfield MCP catalog page now lists 82, and the descriptions are long. By a rough character count of that published tool list, the definitions and parameters come to more than 20,000 tokens before your agent does any work.
Several of those tools carry real side effects. select_workspace changes which workspace all later operations bill against, and the selection persists across sessions and clients. website_secrets writes environment variables. sandbox_exec runs shell commands in a remote Linux sandbox. tiktok_publish posts to a user's profile. A product-shot agent needs none of them.
The two paths do not just use different credential formats. They authenticate different principals, and that choice shapes everything downstream.
The MCP path is user-delegated. Each end user completes a browser OAuth consent, and the agent then acts as that user, inside that user's workspace, against that user's credits. That is the right model when your customers bring their own Higgsfield accounts.
It also means a background agent cannot bootstrap itself. Someone has to complete consent at least once, and Higgsfield's help center notes that expired MCP authorization requires a disconnect and reconnect in the client.
The API path is account-delegated. One key pair identifies your Higgsfield account, and Higgsfield's docs warn it must never ship in browser or mobile code. There is no notion of an end user at the Higgsfield layer.
A request ID created with your key returns 404 to any other account, but inside your account, every tenant's jobs sit side by side.
Recommended reading: Single-tenant vs multi-tenant tool calling
This is the decision that most Higgsfield comparisons skip. On MCP, generation cost lands on each user's plan, and the tool descriptions tell the model to confirm before expensive actions and never to spend unlimited or trial generations on its own initiative. On the API, generation cost lands on your balance, and you decide how to meter and resell it.
Neither is wrong. They describe two different products. A creative copilot for agencies that already pay for Higgsfield fits MCP. A SaaS feature that generates listing images for your customers, billed by you, fits the API.
Recommended reading: Tool calling authentication for AI agents
Higgsfield maintains the tool schemas, the model routing defaults, and the workflow bundles behind get_workflow_instructions. When Higgsfield adds a model, your agent can use it without a redeploy.
You still own token storage, refresh, revocation handling, and per-user isolation. You also own tool-surface churn: MCP tool definitions are unversioned, and a server that grew from five tools to 82 will keep changing. Some tools assume a chat UI. media_upload_widget and create_voice open Apps UI widgets that a headless backend cannot render, so a server-side agent has to use media_upload or media_import_url instead.
You own everything the MCP server abstracts: endpoint selection per model, request schemas, polling with backoff, webhook envelope validation and deduplication, and retries. Concurrency is the primary rate limit, set per account, and exceeding it returns 400 with no Retry-After header, so you need a worker pool or semaphore in front of submissions.
You also own storage. Higgsfield keeps generated output available for at least seven days, then may remove it, so completed files must be copied to your own storage. Credits expire one year after purchase. In exchange, you get stable, documented request contracts and model-specific schemas you can validate before a job ever runs.
Pick either path and a credential problem follows you into production. Its shape differs by path. The infrastructure it demands does not.
Two hundred customers means two hundred Higgsfield OAuth grants. Each one must be encrypted at rest, isolated by tenant, refreshed before it expires, and detected as revoked when a user disconnects Higgsfield. Without that, an agent discovers the problem mid-generation, after it has already planned a ten-shot storyboard.
Recommended reading: How to handle token refresh for AI agents
A single platform key is simpler until it isn't. Every tenant shares one concurrency pool, so one tenant's 500-image import can starve everyone else. A leaked key exposes your entire balance. If you instead collect each customer's own key pair, you are back to storing and rotating N secrets.
On the MCP path, Scalekit's Higgsfield MCP connector handles the OAuth flow, token storage, and refresh per user, so the credential lifecycle stops being code you maintain. The API path can move into the same vault through a custom connector, covered below.
Recommended reading: Credential ownership patterns for agent tool calling
The rest of this post builds a headless creative agent in Python: connect a user's Higgsfield account, restrict the agent to the tools its role needs, run a Claude tool-calling loop, then serve the same surface through a Virtual MCP server to a LangChain agent. Python is used throughout because the Scalekit Node.js SDK does not yet mint Virtual MCP session tokens.
Create a Higgsfield MCP connection under AgentKit > Connections in the Scalekit dashboard, then set SCALEKIT_ENVIRONMENT_URL, SCALEKIT_CLIENT_ID, and SCALEKIT_CLIENT_SECRET in your .env. The code below uses higgsfieldmcp as the connection name. That string must match the connection name in your dashboard exactly; a mismatch is the most common integration error.
Each user signs in to Higgsfield once. Scalekit stores the resulting token against your user's identifier and refreshes it, so the agent runtime never sees a Higgsfield credential.
list_scoped_tools does not return a flat connector catalog. It returns the tools this user's connected account is authorized to call. The agent role then narrows that further: a product-shot agent needs generation, polling, model lookup, and a balance check, not website deploys or TikTok publishing.
Every tool call goes through execute_tool, which injects the user's Higgsfield token server-side. Errors return to the model as is_error tool results, so a rejected parameter or an exhausted balance becomes something the model can reason about instead of a crash.
Client-side filtering works for one agent. It does not survive five agent roles, three frameworks, and a second connector. A Virtual MCP server moves the allowlist server-side: you define it once per agent role, and every MCP client that connects sees only those tools. This one pairs Higgsfield with Slack so the agent can post finished creatives for review.
Before each run, confirm the user's Higgsfield and Slack accounts are still active, then mint a short-lived session token bound to that user. The endpoint is static; the identity is per run. Session tokens default to about one hour and cannot be refreshed in place, so set expiry above your longest expected run.
Recommended reading: LangChain tool calling with per-user auth
For contrast, the direct API path through Higgsfield's official Python SDK needs no OAuth at all. The credential is your account's key pair, read from HF_KEY as key_id:key_secret, and every generation bills your balance.
That is the whole integration for one model. Everything around it, including per-tenant metering, concurrency control, webhook handling, and copying output before the seven-day retention window closes, is yours to build.
The code above is short because the hard parts moved into infrastructure. Three of them matter specifically for Higgsfield agents.
Media generation fails in ways that look identical from the agent's side: an expired grant, an exhausted balance, an NSFW rejection, a rate limit. Scalekit records tool execution logs per connected account, visible under AgentKit > Connected Accounts alongside token status and refresh history. When a call returns 429, the error_code tells you whether Scalekit (RATE_LIMITED) or Higgsfield (TOOL_ERROR) rejected it, and the tool call log carries the provider's message.
That turns "the agent stopped generating for one customer" into a lookup by identifier, not a reproduction exercise.
Recommended reading: Agent tool observability
The 82-tool problem is a least-privilege problem and a token-cost problem at once. A Virtual MCP server fixes both: the agent sees five Higgsfield tools and one Slack tool, not the full catalog, and the tools that switch workspaces, deploy websites, or publish to TikTok are simply absent. Scoping a 40-tool server to 5 to 10 tools cuts tool-definition overhead by roughly 80%; against Higgsfield's much larger surface, the savings are larger.
One definition serves every tenant: one server per agent role, one short-lived session token per user per run. Adding Google Drive later is another connection mapping, not another auth system.
Recommended reading: Launching scoped MCP
Scalekit's catalog ships Higgsfield as an MCP connector today. If your product needs the API path, for example to use webhooks or bill generation to your own balance, Scalekit's bring-your-own-connector model supports API key auth patterns with header overrides and value prefixes, and proxies calls through actions.request(). Higgsfield's Authorization: Key format fits that pattern; confirm the exact connector payload with the Scalekit team before production.
If your agent runs interactively and your users already pay for Higgsfield, build against the MCP server. You get the full platform, including Soul characters and Marketing Studio, without maintaining a single model schema, and generation cost stays on the user's plan.
If your agent runs headlessly, needs webhooks and cancellation, or generation is a feature you bill for, build against the API. The narrower surface is the point: versioned contracts, explicit concurrency, and cost estimates before submission.
Plenty of products will run both: MCP for the creative copilot, the API for the background catalog pipeline. Either way, the credential lifecycle and the tool surface are infrastructure problems, and that is where a production Higgsfield agent succeeds or fails.
Building a Higgsfield agent for multiple users or tenants? Talk to the Scalekit team for immediate help with connection setup, Virtual MCP design, or bringing the API path into the same vault.