Announcing CIMD support for MCP Client registration
Learn more

Wispr Flow MCP vs Wispr Flow API for AI Agents

TL;DR

  • Wispr Flow's Model Context Protocol (MCP) server and its Voice Interface API do not overlap. MCP reads Notetaker meetings, transcripts, calendar events, and Scratchpad notes; the API only turns speech into text.
  • MCP is read-only and OAuth-only. Each user grants account-wide access in a browser, signing in with Google, Apple, Microsoft, or SSO.
  • The API requires Wispr's approval and an org-level key that mints short-lived per-user JWTs. Your end users never sign in to Wispr.
  • Multi-tenant agents hold N OAuth grants on the MCP path and one high-value org key on the API path. Neither path stores, refreshes, or revokes them for you.
  • Scalekit's Wispr Flow MCP connector handles OAuth, token refresh, and tool-call logs; Virtual MCP servers scope it per user.

Two Wispr Flow surfaces, two different jobs

Your agent needs Wispr Flow. Maybe it should brief a rep before a renewal call using what the customer said last month. Maybe your users would rather talk to your agent than type at it. Wispr ships a hosted MCP server and a developer API, and the natural assumption is that they are two routes to the same data.

They are not.

The MCP server reads what Wispr Notetaker captured. The API converts audio into formatted text. Choosing between them means deciding which job your agent has, and then solving a different credential problem for each. Here's the framework.

What Wispr Flow MCP and the Wispr Flow API actually are

Both surfaces are official and production-facing. They sit on opposite sides of your agent: one feeds it context, the other feeds it input.

Wispr Flow MCP

Wispr Flow MCP is a remote server hosted by Wispr at the /connect/mcp path on Wispr's API host. It is the agent interface to Wispr Notetaker, which launched on Mac on August 5, 2026 and on Windows on September 15, 2026. Wispr includes MCP access in its existing plans, including Free.

The server is read-only. It searches and reads meetings, transcripts, summaries, attendees, calendar events, upcoming-meeting pre-reads, recurring series, Scratchpad notes, and shared-note links. It cannot read dictation history or Notetaker chat history. Authentication is browser-based OAuth against the user's Wispr account, with no API key option. Wispr documents the setup in its help center article on connecting an MCP client to Wispr Flow.

Wispr Flow Voice Interface API

The Voice Interface API is Wispr's speech-to-text product for developers. It converts audio into formatted text across 100+ languages, with the same kind of cleanup the Flow apps apply: filler words removed, self-corrections resolved, and names fixed from context you pass in.

It offers a streaming WebSocket endpoint, which Wispr recommends, and a slower REST endpoint capped at 25MB or 6 minutes of 16kHz WAV audio per request. Access is by approval only. Approved orgs get an fl- prefixed API key that either authenticates backend calls directly or mints short-lived client JWTs. The published reference has no endpoints for meetings, transcripts, or notes. Wispr documents all of this in its Voice Interface API documentation.

Comparing capability coverage for agents

For most tools in this series, the capability table shows a large API surface and a smaller MCP subset. Wispr Flow's table shows two disjoint surfaces.

What your agent can read

Tool names are Wispr's; Scalekit's connector prefixes each with wisprflowmcp_.

Capability
Wispr Flow MCP
Voice Interface API
Search meetings by keyword, date, or attendee
Yes (search_meetings)
No
Read a meeting's summary, notes, and transcript
Yes, in bounded ranges
No
List past occurrences of a recurring meeting
Yes (list_meeting_series)
No
Read calendar events and pre-reads
Yes
No
Read Scratchpad notes
Yes
No
Resolve pasted share or calendar links
Yes
No
Create, edit, or delete meetings or notes
No
No
Read dictation history
No
No

What your agent can hear and react to

The voice side of the table flips, and the event row stays empty on both.

Capability
Wispr Flow MCP
Voice Interface API
Transcribe an audio file
No
Yes, REST
Stream speech-to-text with partial results
No
Yes, WebSocket
Format output by app type and dictionary
No
Yes
Webhooks or change events
No
No

Where the MCP ceiling is

Read-only is the headline constraint. An agent can find the action items from Tuesday's sync, but it cannot mark one done, write a note back, or rename a meeting.

The subtler limits show up in retrieval. search_meetings matches titles, summaries, and notes, not transcripts, so a phrase spoken only in conversation will not match. Notes and transcripts come back in character-bounded ranges with continuation offsets, and list tools page with cursors up to a 1,000-result cap. Attendee filters only match recordings linked to a calendar event. Pre-reads exist only when the desktop app generated one; otherwise the field is null.

Where the API ceiling is

The Voice Interface API has no memory of meetings. It does not read what Notetaker recorded; it is an input layer, audio in and text out. Usage is billed on tokens, and the developer platform has no configurable API limits, so spend control lives in your code.

The auth path each one puts you on

This is where the two surfaces diverge hardest. MCP delegates a user's Wispr account to your agent. The API delegates your org's Wispr contract to your users.

MCP: one browser consent per user, account-wide

Every user completes a browser OAuth flow against their own Wispr account. Scalekit's connector page classifies the flow as OAuth 2.1 with Dynamic Client Registration (DCR). Three details matter for agent builders:

  • Sign-in method: email-and-password Wispr accounts cannot complete MCP authorization. Users must sign in with Google, Apple, Microsoft, or enterprise SSO, and finish within 5 minutes.
  • Grant scope: authorization is account-wide and includes existing content. There is no per-meeting consent.
  • Revocation: Wispr has no disconnect control. Access ends when the client removes the connection, and policy changes can take a few minutes to apply.

After consent, the user's Flow app does not need to be open, because the server reads from the Wispr account itself.

API: an org key that mints per-user JWTs

The API credential belongs to your organization, not your user. Your backend holds the fl- key and calls /generate_access_token with a client_id, a duration_secs lifetime, and optional metadata. The returned JWT goes to the client as Authorization: Bearer <JWT> on /client_api, or as the client_key query parameter on /client_ws.

You can revoke a single token or every token for a client_id. Expired, revoked, and mismatched tokens all return 401 with a specific detail message. The API never asks your end users to sign in to Wispr; they are identified only by the client_id you choose.

What this means for multi-tenant agents

Both paths require per-user credential isolation in a multi-tenant B2B agent. They put it in different places.

On the MCP path, 500 users means 500 OAuth grants, each governed by the policies of that user's own Wispr organization. An admin can switch MCP off, lock Cloud Sync, or sign a HIPAA BAA that blocks meeting and calendar reads. An enterprise IP allowlist can reject calls from outside approved networks.

On the API path, isolation is entirely your job. Each JWT's client_id must be your tenant-scoped user ID, and a leaked org key can run up usage with no dashboard cap to stop it. Neither path solves storage, rotation, or revocation. For the broader pattern, see how tool calling auth changes from single-tenant to multi-tenant.

What you own in production

Wispr hosts both surfaces. What lands on your side of the line differs.

On the MCP path

Wispr maintains the server, the tool schemas, and the tool descriptions. You own the per-user token lifecycle and the error cases that come from Wispr's policy model: Cloud Sync disabled, HIPAA-covered accounts, org-disabled MCP or Scratchpad access, IP-allowlist denials, and too-many-requests responses, where Wispr advises waiting about a minute.

Two data-lifecycle gaps also reach your agent. If a user's connected calendar needs reauthorization, that provider's events drop out of MCP results until they reconnect. Retention policies delete transcripts and audio but keep summaries, so an agent can find a summary whose transcript no longer exists. There is no version pin for the tool contract; it changes when Wispr updates the server.

The token cost of Wispr's tool descriptions

Wispr writes long, prescriptive tool descriptions. They tell the model to convert UTC timestamps to local time, prefer transcripts over summaries, and keep paging until has_more is false. That guidance improves answers. It also costs context.

Across the 14 tools on Scalekit's connector page, descriptions total roughly 2,200 words before any parameter schema. Three tools — search_meetings, get_meeting, and get_meeting_by_calendar_id — run past 300 words each. An agent that loads the full server pays that on every request. Scoping the surface, covered below, is the fix; see MCP token costs and why they matter for the math.

On the API path

You own the audio pipeline. The API expects base64-encoded 16kHz, 16-bit mono PCM WAV; browsers record WebM by default, so conversion is on you. WebSocket sessions need equal-duration chunks, ideally 1 second each, a running position counter, and a final commit message with total_packets. Partial transcripts arrive roughly every 30 seconds.

You also own key custody, JWT minting and revocation, the warm-up call if latency matters, and spend control, since the platform has no usage caps.

When to use MCP, when to use the API

Because the surfaces do not overlap, this is less a choice than a job description. Some agents need both.

Use Wispr Flow MCP when

  • Your agent preps a rep for a call by pulling what the customer said in past Wispr-recorded meetings, filtered by attendee email.
  • Your agent turns action items from a meeting's summary and transcript into tickets, follow-up emails, or CRM updates in other tools.
  • Your agent catches a user up before today's instance of a recurring sync, using prior occurrences of the series.
  • Your users already record with Wispr Notetaker and want to ask your product what they committed to.

Use the Voice Interface API when

  • You are adding voice input to your agent's chat box or copilot, and want output formatted as an AI prompt rather than a raw transcript.
  • Your voice-first agent UI needs streaming partial transcripts over WebSocket.
  • A pipeline transcribes short voice memos, each under 6 minutes, with dictionary context for names and jargon.
  • Your end users do not have Wispr accounts and you want usage billed to your org.

Use both when

A voice-driven meeting assistant is the natural combination. The user speaks a question into your app through the Voice API, and the agent answers it from Notetaker data over MCP. Keep one user identifier across both: the Scalekit identifier on the MCP side and the Wispr client_id on the API side. If your voice layer is a realtime agent rather than dictation, the sales call prep agent guide covers authenticated tool calling for that shape.

The credential problem that exists on both paths

Neither Wispr surface gives you a vault, a refresh scheduler, or a revocation flow. The shape of the problem differs by path.

MCP path: N grants to keep alive

Fifty customers with twenty users each is 1,000 Wispr OAuth grants. Each needs encrypted, tenant-isolated storage, proactive refresh, and a clean path back to consent when a grant dies. The mechanics are in how to handle token refresh for AI agents in production.

Because Wispr offers no disconnect control, your product is also where users expect to revoke access. If your app does not delete the stored grant, nothing does.

API path: one key you cannot leak

The API path has one credential that matters: the org key. Keep it in a backend secret manager, never in a client bundle. Keep JWT lifetimes short with duration_secs, and wire /revoke_client_tokens into offboarding so a departed user's tokens die with their account.

Where Scalekit fits

Scalekit's Wispr Flow MCP connector handles the MCP side. Because the connector uses DCR, Scalekit handles OAuth client registration with Wispr; prefilled client fields in the connection form confirm it. Each user becomes a connected account: Scalekit runs the consent flow, stores and refreshes tokens, and delete_connected_account revokes stored credentials when a user disconnects in your app. Credentials never touch the agent runtime.

Scalekit does not ship a Voice Interface API connector. The org key stays in your backend, and the pattern is to reuse one user identifier on both sides.

Recommended Reading: Granola MCP vs Granola API for AI Agents (2026) works through the same decision for another meeting notetaker.

Build a Wispr Flow agent with Scalekit and the Claude SDK

The example uses Python and the Anthropic Python SDK. Python matters here: Scalekit's Node SDK does not mint Virtual MCP session tokens yet, and the multi-tool section below needs them. The sequence is the one that keeps per-user agents correct: authorize, retrieve the scoped tools, then execute. The Scalekit Anthropic example shows the same pattern for other connectors.

Prerequisites

  • A Scalekit environment, with SCALEKIT_ENVIRONMENT_URL, SCALEKIT_CLIENT_ID, and SCALEKIT_CLIENT_SECRET from Developers > API Credentials.
  • A Wispr Flow MCP connection created in AgentKit > Connections.
  • ANTHROPIC_API_KEY set, then pip install scalekit-sdk-python anthropic python-dotenv requests.

The connection name in your code must match the connection name in the dashboard exactly. This example assumes wisprflowmcp. A mismatch is the most common integration error.

Authorize the user's Wispr Flow account

Each user connects once. The helper looks up the user's connected account and, if it is not ACTIVE, generates an authorization link. In production, render that link in your UI. Tell users to sign in to Wispr with Google, Apple, Microsoft, or SSO, since email-and-password accounts cannot finish Wispr's MCP consent.

import json import os from datetime import datetime, timezone import anthropic from dotenv import load_dotenv from google.protobuf.json_format import MessageToDict from scalekit import ScalekitClient load_dotenv() scalekit_client = ScalekitClient( env_url=os.environ["SCALEKIT_ENVIRONMENT_URL"], client_id=os.environ["SCALEKIT_CLIENT_ID"], client_secret=os.environ["SCALEKIT_CLIENT_SECRET"], ) actions = scalekit_client.actions claude = anthropic.Anthropic() # reads ANTHROPIC_API_KEY # Must match the connection name in AgentKit > Connections exactly CONNECTION_NAME = "wisprflowmcp" USER_ID = "user_123" # your app's opaque user ID, not an email def ensure_wispr_connected(identifier: str) -> None: account = actions.get_or_create_connected_account( connection_name=CONNECTION_NAME, identifier=identifier ).connected_account if account.status == "ACTIVE": return link = actions.get_authorization_link( connection_name=CONNECTION_NAME, identifier=identifier ) # In production, show this link in your UI instead of the terminal print("Connect Wispr Flow (Google, Apple, Microsoft, or SSO sign-in):", link.link) input("Press Enter after authorizing...") account = actions.get_or_create_connected_account( connection_name=CONNECTION_NAME, identifier=identifier ).connected_account if account.status != "ACTIVE": raise RuntimeError(f"Wispr Flow is {account.status}, not ACTIVE") ensure_wispr_connected(USER_ID)

Retrieve the tools this user's connected account can call

The agent is not loading a connector catalog here. list_scoped_tools returns the tools bound to this user's connected account: the exact surface this user authorized, and nothing more. The filter is required; the API rejects a call without one.

scoped_response, _ = actions.tools.list_scoped_tools( identifier=USER_ID, filter={"connection_names": [CONNECTION_NAME]}, # required; empty filters are rejected page_size=100, ) claude_tools = [] for scoped_tool in scoped_response.tools: definition = MessageToDict(scoped_tool.tool).get("definition", {}) claude_tools.append({ "name": definition.get("name"), "description": definition.get("description", ""), "input_schema": definition.get("input_schema", {"type": "object", "properties": {}}), }) print(f"{len(claude_tools)} tools authorized for {USER_ID}")

Run the agent loop with execute_tool

execute_tool runs each call Claude requests as this user, through Scalekit, with the stored Wispr token. The system prompt carries today's date, because Wispr's search tools take ISO timestamps and the model cannot otherwise turn "the last 7 days" into a since value.

SYSTEM_PROMPT = ( f"Today is {datetime.now(timezone.utc).date().isoformat()} (UTC). " "You answer questions about the user's Wispr Flow meetings. " "Call wisprflowmcp_get_account_info first for first-person questions. " "For commitments and action items, read the transcript, not just the summary." ) messages = [{ "role": "user", "content": "What did I commit to in meetings with alex@acme.com in the last 7 days?", }] while True: response = claude.messages.create( model="claude-sonnet-5", max_tokens=4096, system=SYSTEM_PROMPT, tools=claude_tools, messages=messages, ) messages.append({"role": "assistant", "content": response.content}) if response.stop_reason != "tool_use": print("".join(block.text for block in response.content if block.type == "text")) break tool_results = [] for block in response.content: if block.type != "tool_use": continue try: result = actions.execute_tool( tool_name=block.name, tool_input=block.input, identifier=USER_ID, connection_name=CONNECTION_NAME, ) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": json.dumps(result.data), }) except Exception as exc: # Wispr policy blocks, rate limits, expired grants tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": f"Tool call failed: {exc}", "is_error": True, }) messages.append({"role": "user", "content": tool_results})

Tool failures go back to Claude as is_error results instead of crashing the loop, so a Cloud Sync or HIPAA block becomes an answer the user can act on.

Add voice input with the Voice Interface API

The Voice API path runs outside Scalekit, but it should share the same user identity. Your backend mints a short-lived Wispr JWT using the Scalekit identifier as the Wispr client_id. Set WISPR_API_BASE to the /api/v1/dash base path from Wispr's API reference.

import os import requests def mint_wispr_voice_token(identifier: str, ttl_secs: int = 900) -> str: """Mint a short-lived Voice API JWT. The org-level fl- key never leaves the backend.""" response = requests.post( f"{os.environ['WISPR_API_BASE']}/generate_access_token", headers={"Authorization": f"Bearer {os.environ['WISPR_ORG_API_KEY']}"}, json={"client_id": identifier, "duration_secs": ttl_secs}, timeout=10, ) response.raise_for_status() return response.json()["access_token"]

The browser opens /client_ws with that token in the client_key query parameter, streams audio, and hands the final text to the agent loop above as the user's message.

Multi-tool, multi-tenant agents with a Virtual MCP server

Most Wispr agents do something with meeting context. A follow-up agent reads Wispr and files issues in Linear. Connected naively, that is 14 Wispr tools plus 62 tools from the Linear connector: 76 tool definitions in context, including linear_issue_delete, before the agent does any work.

A Virtual MCP server fixes the bloat and the overreach at once. One server definition declares exactly which tools the agent sees, and each run gets a short-lived session token scoped to one user's connected accounts.

Create the server once per agent role

This runs once, not once per user. It exposes six tools: three from Wispr and three from Linear. The Wispr descriptions alone drop from about 2,200 words to about 700.

from scalekit.actions.models.mcp_config import McpConfigConnectionToolMapping # Run once per agent role, then store config_id and mcp_server_url vmcp = actions.mcp.create_config( name="meeting-followup-agent", description="Reads Wispr Flow meetings and files action items in Linear", connection_tool_mappings=[ McpConfigConnectionToolMapping( connection_name="wisprflowmcp", # must match the dashboard connection name tools=[ "wisprflowmcp_get_account_info", "wisprflowmcp_search_meetings", "wisprflowmcp_get_meeting", ], ), McpConfigConnectionToolMapping( connection_name="linear", # must match the dashboard connection name tools=["linear_teams_list", "linear_issue_search", "linear_issue_create"], ), ], ) config_id = vmcp.config.id mcp_server_url = vmcp.config.mcp_server_url

Mint a session token per user, per run

Before each run, confirm that the user's Wispr and Linear accounts are both active, then mint a token for that user only. Claude's MCP connector calls the Virtual MCP server from Anthropic's side, so there is no client-side tool loop. Full setup details are in the Virtual MCP server documentation.

from datetime import timedelta def run_followup_agent(identifier: str, prompt: str) -> str: # 1. Every connection on the server must be ACTIVE for this user state = actions.mcp.list_mcp_connected_accounts( config_id=config_id, identifier=identifier, include_auth_link=True ) pending = [a for a in state.connected_accounts if a.connected_account_status != "ACTIVE"] if pending: links = "; ".join(f"{a.connection_name}: {a.authentication_link}" for a in pending) raise RuntimeError(f"Authorization required before this run: {links}") # 2. Mint a fresh, user-scoped session token for this run only session = actions.mcp.create_session_token( mcp_config_id=config_id, identifier=identifier, expiry=timedelta(minutes=30), ) # 3. Claude's MCP connector calls the Virtual MCP server directly response = claude.beta.messages.create( model="claude-sonnet-5", max_tokens=4096, system=( f"Today is {datetime.now(timezone.utc).date().isoformat()} (UTC). " "Search Linear for an existing issue before creating a new one." ), messages=[{"role": "user", "content": prompt}], mcp_servers=[{ "type": "url", "url": mcp_server_url, "name": "scalekit", "authorization_token": session.token, }], tools=[{"type": "mcp_toolset", "mcp_server_name": "scalekit"}], betas=["mcp-client-2025-11-20"], ) return "".join(block.text for block in response.content if block.type == "text") print(run_followup_agent( USER_ID, "File a Linear issue for each action item I own from yesterday's platform sync.", ))

What this changes for multi-tenant agents

  • One endpoint, every user: the server URL is static; identity arrives with the session token.
  • Least privilege by construction: the agent cannot delete a Linear issue, because that tool does not exist on its server.
  • Smaller context: six tool definitions instead of 76. The fix is not better prompting. It is surface reduction.
  • Isolation per run: tokens default to about an hour, cap at 24 hours, and are minted fresh for every run.

Recommended Reading: What Is a Virtual MCP Server? Scoped Tools and Per-User Auth

Observability: logs for every downstream tool call

A meeting agent fails quietly. The user asks what they owe a customer, the agent returns nothing, and nobody knows whether the transcript was deleted, the grant expired, or the customer's admin switched MCP off.

What the tool call log records

AgentKit logs tool calls across all connectors, including Wispr Flow MCP. The overview shows total calls, success rate, connector errors, and API errors over windows from 1 hour to 30 days, filterable to one connector. Each log record carries the timestamp, connection, tool, user identifier, and latency, plus the error code and full message for failures. The launch post on agent tool observability walks through the dashboard.

Why the error split matters for Wispr

Scalekit separates two failure classes. An API error means the call never left Scalekit: invalid parameters, an expired token, a misconfigured tool. A connector error means the call reached Wispr and Wispr rejected it, which is where Cloud Sync, HIPAA, org-policy, and rate-limit responses fall.

Because every record carries the user identifier, "which customers lost meeting access this morning" becomes a filter, not an investigation.

Which one to build against

If your agent needs to know what was said in a meeting, build against Wispr Flow MCP; there is no API alternative for Notetaker data. If your product needs voice input, build against the Voice Interface API; the MCP server cannot hear anything. If your agent does both, run both under one user identifier.

The harder decision is where credentials live. MCP leaves you holding an OAuth grant for every user, each subject to someone else's admin policy. The API leaves you holding one org key with no spending cap. Both are infrastructure problems. The Wispr surface you pick decides which one you inherit, not whether you have one. For a deep look at credential ownership across agent tool-calling patterns, see our dedicated guide.

Get help shipping your Wispr Flow agent

Building a meeting-aware or voice-driven agent on Wispr Flow? Talk to us for immediate help with connector setup, Virtual MCP design, and multi-tenant auth.

No items found.
Agent
Auth Quickstart
On this page
Share this article
Agent
Auth Quickstart

Acquire enterprise customers with
‍zero upfront cost.

Every feature unlocked. No hidden fees.