Announcing CIMD support for MCP Client registration
Learn more

Higgsfield MCP vs Higgsfield API for AI Agents

Nishant Choudhary
Tech Evangelist

TL;DR

  • Higgsfield MCP exposes 82 tools: generation across 30+ models, plus billing, workspace, website, TikTok, and sandbox tools. The API exposes generation endpoints, status, cancellation, cost estimates, and webhooks.
  • MCP signs in each end user with OAuth and spends their plan credits. The API authenticates your account with a key pair and bills your balance.
  • Only the API offers webhooks and queued-job cancellation. MCP agents poll through jobs_wait and job_status.
  • Neither path solves credential lifecycle: MCP means N revocable OAuth tokens; the API means one shared key, concurrency pool, and blast radius.
  • Scalekit's Higgsfield MCP connector vaults and refreshes each user's token, scopes the 82 tools per agent role through Virtual MCP servers, and logs every tool execution.

Two official paths into Higgsfield

Your agent needs to produce ad creatives, product shots, or short-form video. Higgsfield ships two official ways in: a hosted MCP server that any MCP client can sign into, and a REST API with Python and TypeScript SDKs. In a demo, they look interchangeable. They are not.

One path acts as your user and spends their credits. The other acts as your company and spends yours. One exposes 82 tools, including tools that publish to TikTok and edit website secrets. The other exposes models. Here is how to pick, and how to wire either one into a production agent.

What Higgsfield MCP and Higgsfield API actually are

Both paths reach Higgsfield's image, video, and audio models. What differs is who authenticates, what surface the agent sees, and how results come back.

Higgsfield MCP

Higgsfield runs a hosted, remote Model Context Protocol (MCP) server, launched in spring 2026 and documented on Higgsfield's official MCP page. It works with Claude, Claude Code, Cursor, OpenClaw, Hermes Agent, and any other MCP-compatible client. There is no local package to install.

Authentication is account sign-in through a browser OAuth flow. Higgsfield's own MCP page states that no API key is required, and Scalekit's connector lists the auth type as OAuth 2.1 with Dynamic Client Registration (DCR). Higgsfield's MCP guide for marketers says an active subscription is required, with a trial offered to eligible new users, and generations draw on the user's plan credits. Every generation lands in the user's Assets library on Higgsfield, tagged with the MCP source.

Higgsfield API

The Higgsfield API, documented in Higgsfield's official API reference, is an asynchronous REST API. You POST to a model-specific endpoint, receive a request_id with status and cancel links, then poll /requests/{request_id}/status or receive a webhook when the request reaches a terminal state.

Authentication is a server-side credential pair created in the Higgsfield Console, sent as Authorization: Key {key_id}:{key_secret}. There is no OAuth flow and no end-user consent. Billing is pay-per-generation against your account balance, with a dedicated /estimate/... endpoint for pricing a request before you submit it. Official SDKs exist for Python (higgsfield-client) and Node.js (@higgsfield/client).

A third official path: the Higgsfield CLI

Higgsfield also ships a CLI (@higgsfield/cli) that coding agents install themselves and authenticate with higgsfield auth login. It targets local developer agents, not multi-tenant backends, so this post treats it as out of scope.

Comparing them where it matters for agents

The four dimensions below are the ones that change your architecture: capability coverage, auth model, operational surface area, and fit by use case.

What your agent can actually do

The table compares documented capabilities. "Not documented" means the capability does not appear in Higgsfield's public API reference, not that it is impossible.

Capability
Higgsfield MCP
Higgsfield API
Image and video generation
Yes (generate_image, generate_video)
Yes (per-model endpoints)
Parallel headless batches
Yes, 1 to 12 jobs (*_batch tools)
Yes, bounded by account concurrency
Speech and voice cloning
Yes (generate_audio, create_voice)
Audio models; cloning not documented
Cost preflight
Yes (get_cost: true)
Yes (/estimate/...)
Completion webhooks
No, polling only
Yes (hf_webhook)
Cancel a queued job
No tool exposed
Yes (cancel_url)
Soul character training
Yes (show_characters)
Not documented
Credits, transactions, workspaces
Yes
Not documented
TikTok publishing
Yes (tiktok_publish)
No
Website builder and deploy
Yes (create_website, deploy_website)
No
Remote shell sandbox
Yes (sandbox_exec)
No

Where the capability gaps bite

The most consequential gap for a production pipeline is event delivery. The API posts terminal results to your hf_webhook endpoint, retries network failures and 5xx responses for up to two hours, and expects you to deduplicate on request_id plus status. The MCP server has no event surface. An MCP agent generating video must keep polling, and jobs_wait caps each long-poll at 15 seconds, while Higgsfield's own tool description puts typical video jobs at 60 to 180 seconds.

That polling happens inside the model loop. Every wait is another tool call, another round of context, and more tokens. The second gap is cancellation. The API refunds a request canceled before processing starts; the MCP exposes no cancel tool at all.

Why 82 tools is its own problem

The MCP gap runs in the other direction too: it exposes far more than a creative agent should hold. Launch-era coverage described five core tools. Scalekit's Higgsfield MCP catalog page now lists 82, and the descriptions are long. By a rough character count of that published tool list, the definitions and parameters come to more than 20,000 tokens before your agent does any work.

Several of those tools carry real side effects. select_workspace changes which workspace all later operations bill against, and the selection persists across sessions and clients. website_secrets writes environment variables. sandbox_exec runs shell commands in a remote Linux sandbox. tiktok_publish posts to a user's profile. A product-shot agent needs none of them.

The auth path each one puts you on

The two paths do not just use different credential formats. They authenticate different principals, and that choice shapes everything downstream.

MCP authenticates your end user

The MCP path is user-delegated. Each end user completes a browser OAuth consent, and the agent then acts as that user, inside that user's workspace, against that user's credits. That is the right model when your customers bring their own Higgsfield accounts.

It also means a background agent cannot bootstrap itself. Someone has to complete consent at least once, and Higgsfield's help center notes that expired MCP authorization requires a disconnect and reconnect in the client.

The API authenticates your company

The API path is account-delegated. One key pair identifies your Higgsfield account, and Higgsfield's docs warn it must never ship in browser or mobile code. There is no notion of an end user at the Higgsfield layer.

A request ID created with your key returns 404 to any other account, but inside your account, every tenant's jobs sit side by side.

Recommended reading: Single-tenant vs multi-tenant tool calling

Whose credits the agent spends

This is the decision that most Higgsfield comparisons skip. On MCP, generation cost lands on each user's plan, and the tool descriptions tell the model to confirm before expensive actions and never to spend unlimited or trial generations on its own initiative. On the API, generation cost lands on your balance, and you decide how to meter and resell it.

Neither is wrong. They describe two different products. A creative copilot for agencies that already pay for Higgsfield fits MCP. A SaaS feature that generates listing images for your customers, billed by you, fits the API.

Recommended reading: Tool calling authentication for AI agents

What you own in production: the MCP path

Higgsfield maintains the tool schemas, the model routing defaults, and the workflow bundles behind get_workflow_instructions. When Higgsfield adds a model, your agent can use it without a redeploy.

You still own token storage, refresh, revocation handling, and per-user isolation. You also own tool-surface churn: MCP tool definitions are unversioned, and a server that grew from five tools to 82 will keep changing. Some tools assume a chat UI. media_upload_widget and create_voice open Apps UI widgets that a headless backend cannot render, so a server-side agent has to use media_upload or media_import_url instead.

What you own in production: the API path

You own everything the MCP server abstracts: endpoint selection per model, request schemas, polling with backoff, webhook envelope validation and deduplication, and retries. Concurrency is the primary rate limit, set per account, and exceeding it returns 400 with no Retry-After header, so you need a worker pool or semaphore in front of submissions.

You also own storage. Higgsfield keeps generated output available for at least seven days, then may remove it, so completed files must be copied to your own storage. Credits expire one year after purchase. In exchange, you get stable, documented request contracts and model-specific schemas you can validate before a job ever runs.

When Higgsfield MCP wins

  • Interactive creative agents in Claude, Claude Code, or Cursor, where the user is present for OAuth consent and wants results in their own Higgsfield Assets library.
  • Products where customers bring their own Higgsfield subscription and credits, and your agent should never spend on your balance.
  • Agents that need platform features the API does not document, such as Soul character training, Marketing Studio flows, dubbing, or TikTok publishing.
  • Prototypes that need to try many models quickly, since models_explore recommends a model per task without you writing per-model schemas.

When the Higgsfield API wins

  • Headless, event-driven pipelines, such as generating catalog images on every product import, where webhooks replace polling loops.
  • SaaS features where you own billing, meter usage per tenant, and resell generation inside your own pricing.
  • High-volume batch work that needs explicit concurrency control, cost estimates before submission, and cancellation of queued jobs.
  • Deterministic pipelines that call one or two known models, where a documented, per-model request contract matters more than tool discovery.

The credential problem that exists on both paths

Pick either path and a credential problem follows you into production. Its shape differs by path. The infrastructure it demands does not.

On the MCP path: N expiring OAuth tokens

Two hundred customers means two hundred Higgsfield OAuth grants. Each one must be encrypted at rest, isolated by tenant, refreshed before it expires, and detected as revoked when a user disconnects Higgsfield. Without that, an agent discovers the problem mid-generation, after it has already planned a ten-shot storyboard.

Recommended reading: How to handle token refresh for AI agents

On the API path: one key, shared blast radius

A single platform key is simpler until it isn't. Every tenant shares one concurrency pool, so one tenant's 500-image import can starve everyone else. A leaked key exposes your entire balance. If you instead collect each customer's own key pair, you are back to storing and rotating N secrets.

On the MCP path, Scalekit's Higgsfield MCP connector handles the OAuth flow, token storage, and refresh per user, so the credential lifecycle stops being code you maintain. The API path can move into the same vault through a custom connector, covered below.

Recommended reading: Credential ownership patterns for agent tool calling

Building a Higgsfield agent with Scalekit

The rest of this post builds a headless creative agent in Python: connect a user's Higgsfield account, restrict the agent to the tools its role needs, run a Claude tool-calling loop, then serve the same surface through a Virtual MCP server to a LangChain agent. Python is used throughout because the Scalekit Node.js SDK does not yet mint Virtual MCP session tokens.

Prerequisites and the connection name

Create a Higgsfield MCP connection under AgentKit > Connections in the Scalekit dashboard, then set SCALEKIT_ENVIRONMENT_URL, SCALEKIT_CLIENT_ID, and SCALEKIT_CLIENT_SECRET in your .env. The code below uses higgsfieldmcp as the connection name. That string must match the connection name in your dashboard exactly; a mismatch is the most common integration error.

pip install scalekit-sdk-python anthropic python-dotenv

Connect each user's Higgsfield account

Each user signs in to Higgsfield once. Scalekit stores the resulting token against your user's identifier and refreshes it, so the agent runtime never sees a Higgsfield credential.

import os from dotenv import load_dotenv from scalekit import ScalekitClient load_dotenv() scalekit_client = ScalekitClient( env_url=os.getenv("SCALEKIT_ENVIRONMENT_URL"), client_id=os.getenv("SCALEKIT_CLIENT_ID"), client_secret=os.getenv("SCALEKIT_CLIENT_SECRET"), ) actions = scalekit_client.actions CONNECTION_NAME = "higgsfieldmcp" # must match the connection name in your Scalekit dashboard USER_ID = "user_123" # your app's unique identifier for this user response = actions.get_or_create_connected_account( connection_name=CONNECTION_NAME, identifier=USER_ID ) if response.connected_account.status != "ACTIVE": link = actions.get_authorization_link( connection_name=CONNECTION_NAME, identifier=USER_ID ) print("Authorize Higgsfield:", link.link) input("Press Enter after authorizing...") response = actions.get_or_create_connected_account( connection_name=CONNECTION_NAME, identifier=USER_ID ) if response.connected_account.status != "ACTIVE": raise RuntimeError(f"Higgsfield is {response.connected_account.status}, not ACTIVE")

Scope the tool surface before the model sees it

list_scoped_tools does not return a flat connector catalog. It returns the tools this user's connected account is authorized to call. The agent role then narrows that further: a product-shot agent needs generation, polling, model lookup, and a balance check, not website deploys or TikTok publishing.

Tool
Purpose
higgsfieldmcp_balance
Check remaining credits before planning work
higgsfieldmcp_models_explore
Recommend a model for the brief
higgsfieldmcp_generate_image_batch
Submit up to 12 image jobs without a widget
higgsfieldmcp_generate_video_batch
Submit up to 12 video jobs without a widget
higgsfieldmcp_jobs_wait
Long-poll a group of jobs to completion
higgsfieldmcp_job_status
Check a single job
from google.protobuf.json_format import MessageToDict ALLOWED_TOOLS = { "higgsfieldmcp_balance", "higgsfieldmcp_models_explore", "higgsfieldmcp_generate_image_batch", "higgsfieldmcp_generate_video_batch", "higgsfieldmcp_jobs_wait", "higgsfieldmcp_job_status", } scoped_response, _ = actions.tools.list_scoped_tools( identifier=USER_ID, filter={"connection_names": [CONNECTION_NAME]}, page_size=100, # Higgsfield exposes 82 tools; fetch them all before narrowing ) llm_tools = [] for scoped_tool in scoped_response.tools: definition = MessageToDict(scoped_tool.tool).get("definition", {}) if definition.get("name") in ALLOWED_TOOLS: llm_tools.append({ "name": definition["name"], "description": definition.get("description", ""), "input_schema": definition.get("input_schema", {}), }) print(f"Passing {len(llm_tools)} of {len(scoped_response.tools)} tools to the model")

Run the Claude tool-calling loop

Every tool call goes through execute_tool, which injects the user's Higgsfield token server-side. Errors return to the model as is_error tool results, so a rejected parameter or an exhausted balance becomes something the model can reason about instead of a crash.

import anthropic client = anthropic.Anthropic() messages = [{ "role": "user", "content": ( "Check my credit balance first. Then generate three 4:5 product hero images " "of a matte black water bottle on wet slate, wait for them to finish, " "and return the image URLs." ), }] while True: response = client.messages.create( model="claude-sonnet-5", max_tokens=2048, tools=llm_tools, messages=messages, ) messages.append({"role": "assistant", "content": response.content}) if response.stop_reason != "tool_use": print("".join(block.text for block in response.content if block.type == "text")) break tool_results = [] for block in response.content: if block.type != "tool_use": continue try: result = actions.execute_tool( tool_name=block.name, identifier=USER_ID, connection_name=CONNECTION_NAME, tool_input=block.input, ) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": str(result.data), }) except Exception as exc: tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": f"Tool call failed: {exc}", "is_error": True, }) messages.append({"role": "user", "content": tool_results})

Serve the same tools through a Virtual MCP server

Client-side filtering works for one agent. It does not survive five agent roles, three frameworks, and a second connector. A Virtual MCP server moves the allowlist server-side: you define it once per agent role, and every MCP client that connects sees only those tools. This one pairs Higgsfield with Slack so the agent can post finished creatives for review.

from scalekit.actions.models.mcp_config import McpConfigConnectionToolMapping vmcp_response = scalekit_client.actions.mcp.create_config( name="creative-review-agent", connection_tool_mappings=[ McpConfigConnectionToolMapping( connection_name="higgsfieldmcp", tools=[ "higgsfieldmcp_balance", "higgsfieldmcp_models_explore", "higgsfieldmcp_generate_image_batch", "higgsfieldmcp_jobs_wait", "higgsfieldmcp_job_status", ], ), McpConfigConnectionToolMapping( connection_name="slack", # must match your Slack connection name in the dashboard tools=["slack_send_message"], ), ], ) config_id = vmcp_response.config.id mcp_server_url = vmcp_response.config.mcp_server_url # static; reuse for every user

Mint a per-user session token and connect LangChain

Before each run, confirm the user's Higgsfield and Slack accounts are still active, then mint a short-lived session token bound to that user. The endpoint is static; the identity is per run. Session tokens default to about one hour and cannot be refreshed in place, so set expiry above your longest expected run.

pip install langchain-anthropic "langchain-mcp-adapters>=0.3,<1"
import asyncio from datetime import timedelta from langchain_anthropic import ChatAnthropic from langchain_core.messages import HumanMessage, ToolMessage from langchain_mcp_adapters.client import MultiServerMCPClient accounts = scalekit_client.actions.mcp.list_mcp_connected_accounts( config_id=config_id, identifier=USER_ID, include_auth_link=True ) inactive = [a for a in accounts.connected_accounts if a.connected_account_status != "ACTIVE"] if inactive: for account in inactive: print(f"{account.connection_name} needs auth: {account.authentication_link}") raise SystemExit("Re-authorize the connections above, then run again.") session_token = scalekit_client.actions.mcp.create_session_token( mcp_config_id=config_id, identifier=USER_ID, expiry=timedelta(minutes=45), # video jobs can run minutes each ).token async def run(): mcp_client = MultiServerMCPClient({ "scalekit": { "transport": "streamable_http", "url": mcp_server_url, "headers": {"Authorization": f"Bearer {session_token}"}, } }) tools = await mcp_client.get_tools() tool_map = {t.name: t for t in tools} llm = ChatAnthropic(model="claude-sonnet-5", max_tokens=2048).bind_tools(tools) messages = [HumanMessage( "Generate two 1:1 lifestyle images of a ceramic pour-over set, wait for them, " "then post the URLs to #creative-review in Slack." )] while True: response = await llm.ainvoke(messages) messages.append(response) if not response.tool_calls: print(response.content) break for tc in response.tool_calls: result = await tool_map[tc["name"]].ainvoke(tc["args"]) messages.append(ToolMessage(content=str(result), tool_call_id=tc["id"])) asyncio.run(run())

Recommended reading: LangChain tool calling with per-user auth

What the API path looks like for comparison

For contrast, the direct API path through Higgsfield's official Python SDK needs no OAuth at all. The credential is your account's key pair, read from HF_KEY as key_id:key_secret, and every generation bills your balance.

import higgsfield_client result = higgsfield_client.subscribe( "higgsfield-ai/soul/v2/standard", arguments={"prompt": "Editorial portrait in soft daylight"}, ) print(result["images"][0]["url"])

That is the whole integration for one model. Everything around it, including per-tenant metering, concurrency control, webhook handling, and copying output before the seven-day retention window closes, is yours to build.

Why the Scalekit path holds up in production

The code above is short because the hard parts moved into infrastructure. Three of them matter specifically for Higgsfield agents.

Tool-call logs you can debug from

Media generation fails in ways that look identical from the agent's side: an expired grant, an exhausted balance, an NSFW rejection, a rate limit. Scalekit records tool execution logs per connected account, visible under AgentKit > Connected Accounts alongside token status and refresh history. When a call returns 429, the error_code tells you whether Scalekit (RATE_LIMITED) or Higgsfield (TOOL_ERROR) rejected it, and the tool call log carries the provider's message.

That turns "the agent stopped generating for one customer" into a lookup by identifier, not a reproduction exercise.

Recommended reading: Agent tool observability

One Virtual MCP server for multi-tool, multi-tenant agents

The 82-tool problem is a least-privilege problem and a token-cost problem at once. A Virtual MCP server fixes both: the agent sees five Higgsfield tools and one Slack tool, not the full catalog, and the tools that switch workspaces, deploy websites, or publish to TikTok are simply absent. Scoping a 40-tool server to 5 to 10 tools cuts tool-definition overhead by roughly 80%; against Higgsfield's much larger surface, the savings are larger.

One definition serves every tenant: one server per agent role, one short-lived session token per user per run. Adding Google Drive later is another connection mapping, not another auth system.

Recommended reading: Launching scoped MCP

Bringing the API path under the same vault

Scalekit's catalog ships Higgsfield as an MCP connector today. If your product needs the API path, for example to use webhooks or bill generation to your own balance, Scalekit's bring-your-own-connector model supports API key auth patterns with header overrides and value prefixes, and proxies calls through actions.request(). Higgsfield's Authorization: Key format fits that pattern; confirm the exact connector payload with the Scalekit team before production.

Which one to build against

If your agent runs interactively and your users already pay for Higgsfield, build against the MCP server. You get the full platform, including Soul characters and Marketing Studio, without maintaining a single model schema, and generation cost stays on the user's plan.

If your agent runs headlessly, needs webhooks and cancellation, or generation is a feature you bill for, build against the API. The narrower surface is the point: versioned contracts, explicit concurrency, and cost estimates before submission.

Plenty of products will run both: MCP for the creative copilot, the API for the background catalog pipeline. Either way, the credential lifecycle and the tool surface are infrastructure problems, and that is where a production Higgsfield agent succeeds or fails.

Get help building your Higgsfield agent

Building a Higgsfield agent for multiple users or tenants? Talk to the Scalekit team for immediate help with connection setup, Virtual MCP design, or bringing the API path into the same vault.

No items found.
Agent
Auth Quickstart
On this page
Share this article
Agent
Auth Quickstart

Acquire enterprise customers with
‍zero upfront cost.

Every feature unlocked. No hidden fees.