
Your agent needs Wispr Flow. Maybe it should brief a rep before a renewal call using what the customer said last month. Maybe your users would rather talk to your agent than type at it. Wispr ships a hosted MCP server and a developer API, and the natural assumption is that they are two routes to the same data.
They are not.
The MCP server reads what Wispr Notetaker captured. The API converts audio into formatted text. Choosing between them means deciding which job your agent has, and then solving a different credential problem for each. Here's the framework.
Both surfaces are official and production-facing. They sit on opposite sides of your agent: one feeds it context, the other feeds it input.
Wispr Flow MCP is a remote server hosted by Wispr at the /connect/mcp path on Wispr's API host. It is the agent interface to Wispr Notetaker, which launched on Mac on August 5, 2026 and on Windows on September 15, 2026. Wispr includes MCP access in its existing plans, including Free.
The server is read-only. It searches and reads meetings, transcripts, summaries, attendees, calendar events, upcoming-meeting pre-reads, recurring series, Scratchpad notes, and shared-note links. It cannot read dictation history or Notetaker chat history. Authentication is browser-based OAuth against the user's Wispr account, with no API key option. Wispr documents the setup in its help center article on connecting an MCP client to Wispr Flow.
The Voice Interface API is Wispr's speech-to-text product for developers. It converts audio into formatted text across 100+ languages, with the same kind of cleanup the Flow apps apply: filler words removed, self-corrections resolved, and names fixed from context you pass in.
It offers a streaming WebSocket endpoint, which Wispr recommends, and a slower REST endpoint capped at 25MB or 6 minutes of 16kHz WAV audio per request. Access is by approval only. Approved orgs get an fl- prefixed API key that either authenticates backend calls directly or mints short-lived client JWTs. The published reference has no endpoints for meetings, transcripts, or notes. Wispr documents all of this in its Voice Interface API documentation.
For most tools in this series, the capability table shows a large API surface and a smaller MCP subset. Wispr Flow's table shows two disjoint surfaces.
Tool names are Wispr's; Scalekit's connector prefixes each with wisprflowmcp_.
The voice side of the table flips, and the event row stays empty on both.
Read-only is the headline constraint. An agent can find the action items from Tuesday's sync, but it cannot mark one done, write a note back, or rename a meeting.
The subtler limits show up in retrieval. search_meetings matches titles, summaries, and notes, not transcripts, so a phrase spoken only in conversation will not match. Notes and transcripts come back in character-bounded ranges with continuation offsets, and list tools page with cursors up to a 1,000-result cap. Attendee filters only match recordings linked to a calendar event. Pre-reads exist only when the desktop app generated one; otherwise the field is null.
The Voice Interface API has no memory of meetings. It does not read what Notetaker recorded; it is an input layer, audio in and text out. Usage is billed on tokens, and the developer platform has no configurable API limits, so spend control lives in your code.
This is where the two surfaces diverge hardest. MCP delegates a user's Wispr account to your agent. The API delegates your org's Wispr contract to your users.
Every user completes a browser OAuth flow against their own Wispr account. Scalekit's connector page classifies the flow as OAuth 2.1 with Dynamic Client Registration (DCR). Three details matter for agent builders:
After consent, the user's Flow app does not need to be open, because the server reads from the Wispr account itself.
The API credential belongs to your organization, not your user. Your backend holds the fl- key and calls /generate_access_token with a client_id, a duration_secs lifetime, and optional metadata. The returned JWT goes to the client as Authorization: Bearer <JWT> on /client_api, or as the client_key query parameter on /client_ws.
You can revoke a single token or every token for a client_id. Expired, revoked, and mismatched tokens all return 401 with a specific detail message. The API never asks your end users to sign in to Wispr; they are identified only by the client_id you choose.
Both paths require per-user credential isolation in a multi-tenant B2B agent. They put it in different places.
On the MCP path, 500 users means 500 OAuth grants, each governed by the policies of that user's own Wispr organization. An admin can switch MCP off, lock Cloud Sync, or sign a HIPAA BAA that blocks meeting and calendar reads. An enterprise IP allowlist can reject calls from outside approved networks.
On the API path, isolation is entirely your job. Each JWT's client_id must be your tenant-scoped user ID, and a leaked org key can run up usage with no dashboard cap to stop it. Neither path solves storage, rotation, or revocation. For the broader pattern, see how tool calling auth changes from single-tenant to multi-tenant.
Wispr hosts both surfaces. What lands on your side of the line differs.
Wispr maintains the server, the tool schemas, and the tool descriptions. You own the per-user token lifecycle and the error cases that come from Wispr's policy model: Cloud Sync disabled, HIPAA-covered accounts, org-disabled MCP or Scratchpad access, IP-allowlist denials, and too-many-requests responses, where Wispr advises waiting about a minute.
Two data-lifecycle gaps also reach your agent. If a user's connected calendar needs reauthorization, that provider's events drop out of MCP results until they reconnect. Retention policies delete transcripts and audio but keep summaries, so an agent can find a summary whose transcript no longer exists. There is no version pin for the tool contract; it changes when Wispr updates the server.
Wispr writes long, prescriptive tool descriptions. They tell the model to convert UTC timestamps to local time, prefer transcripts over summaries, and keep paging until has_more is false. That guidance improves answers. It also costs context.
Across the 14 tools on Scalekit's connector page, descriptions total roughly 2,200 words before any parameter schema. Three tools — search_meetings, get_meeting, and get_meeting_by_calendar_id — run past 300 words each. An agent that loads the full server pays that on every request. Scoping the surface, covered below, is the fix; see MCP token costs and why they matter for the math.
You own the audio pipeline. The API expects base64-encoded 16kHz, 16-bit mono PCM WAV; browsers record WebM by default, so conversion is on you. WebSocket sessions need equal-duration chunks, ideally 1 second each, a running position counter, and a final commit message with total_packets. Partial transcripts arrive roughly every 30 seconds.
You also own key custody, JWT minting and revocation, the warm-up call if latency matters, and spend control, since the platform has no usage caps.
Because the surfaces do not overlap, this is less a choice than a job description. Some agents need both.
A voice-driven meeting assistant is the natural combination. The user speaks a question into your app through the Voice API, and the agent answers it from Notetaker data over MCP. Keep one user identifier across both: the Scalekit identifier on the MCP side and the Wispr client_id on the API side. If your voice layer is a realtime agent rather than dictation, the sales call prep agent guide covers authenticated tool calling for that shape.
Neither Wispr surface gives you a vault, a refresh scheduler, or a revocation flow. The shape of the problem differs by path.
Fifty customers with twenty users each is 1,000 Wispr OAuth grants. Each needs encrypted, tenant-isolated storage, proactive refresh, and a clean path back to consent when a grant dies. The mechanics are in how to handle token refresh for AI agents in production.
Because Wispr offers no disconnect control, your product is also where users expect to revoke access. If your app does not delete the stored grant, nothing does.
The API path has one credential that matters: the org key. Keep it in a backend secret manager, never in a client bundle. Keep JWT lifetimes short with duration_secs, and wire /revoke_client_tokens into offboarding so a departed user's tokens die with their account.
Scalekit's Wispr Flow MCP connector handles the MCP side. Because the connector uses DCR, Scalekit handles OAuth client registration with Wispr; prefilled client fields in the connection form confirm it. Each user becomes a connected account: Scalekit runs the consent flow, stores and refreshes tokens, and delete_connected_account revokes stored credentials when a user disconnects in your app. Credentials never touch the agent runtime.
Scalekit does not ship a Voice Interface API connector. The org key stays in your backend, and the pattern is to reuse one user identifier on both sides.
Recommended Reading: Granola MCP vs Granola API for AI Agents (2026) works through the same decision for another meeting notetaker.
The example uses Python and the Anthropic Python SDK. Python matters here: Scalekit's Node SDK does not mint Virtual MCP session tokens yet, and the multi-tool section below needs them. The sequence is the one that keeps per-user agents correct: authorize, retrieve the scoped tools, then execute. The Scalekit Anthropic example shows the same pattern for other connectors.
The connection name in your code must match the connection name in the dashboard exactly. This example assumes wisprflowmcp. A mismatch is the most common integration error.
Each user connects once. The helper looks up the user's connected account and, if it is not ACTIVE, generates an authorization link. In production, render that link in your UI. Tell users to sign in to Wispr with Google, Apple, Microsoft, or SSO, since email-and-password accounts cannot finish Wispr's MCP consent.
The agent is not loading a connector catalog here. list_scoped_tools returns the tools bound to this user's connected account: the exact surface this user authorized, and nothing more. The filter is required; the API rejects a call without one.
execute_tool runs each call Claude requests as this user, through Scalekit, with the stored Wispr token. The system prompt carries today's date, because Wispr's search tools take ISO timestamps and the model cannot otherwise turn "the last 7 days" into a since value.
Tool failures go back to Claude as is_error results instead of crashing the loop, so a Cloud Sync or HIPAA block becomes an answer the user can act on.
The Voice API path runs outside Scalekit, but it should share the same user identity. Your backend mints a short-lived Wispr JWT using the Scalekit identifier as the Wispr client_id. Set WISPR_API_BASE to the /api/v1/dash base path from Wispr's API reference.
The browser opens /client_ws with that token in the client_key query parameter, streams audio, and hands the final text to the agent loop above as the user's message.
Most Wispr agents do something with meeting context. A follow-up agent reads Wispr and files issues in Linear. Connected naively, that is 14 Wispr tools plus 62 tools from the Linear connector: 76 tool definitions in context, including linear_issue_delete, before the agent does any work.
A Virtual MCP server fixes the bloat and the overreach at once. One server definition declares exactly which tools the agent sees, and each run gets a short-lived session token scoped to one user's connected accounts.
This runs once, not once per user. It exposes six tools: three from Wispr and three from Linear. The Wispr descriptions alone drop from about 2,200 words to about 700.
Before each run, confirm that the user's Wispr and Linear accounts are both active, then mint a token for that user only. Claude's MCP connector calls the Virtual MCP server from Anthropic's side, so there is no client-side tool loop. Full setup details are in the Virtual MCP server documentation.
Recommended Reading: What Is a Virtual MCP Server? Scoped Tools and Per-User Auth
A meeting agent fails quietly. The user asks what they owe a customer, the agent returns nothing, and nobody knows whether the transcript was deleted, the grant expired, or the customer's admin switched MCP off.
AgentKit logs tool calls across all connectors, including Wispr Flow MCP. The overview shows total calls, success rate, connector errors, and API errors over windows from 1 hour to 30 days, filterable to one connector. Each log record carries the timestamp, connection, tool, user identifier, and latency, plus the error code and full message for failures. The launch post on agent tool observability walks through the dashboard.
Scalekit separates two failure classes. An API error means the call never left Scalekit: invalid parameters, an expired token, a misconfigured tool. A connector error means the call reached Wispr and Wispr rejected it, which is where Cloud Sync, HIPAA, org-policy, and rate-limit responses fall.
Because every record carries the user identifier, "which customers lost meeting access this morning" becomes a filter, not an investigation.
If your agent needs to know what was said in a meeting, build against Wispr Flow MCP; there is no API alternative for Notetaker data. If your product needs voice input, build against the Voice Interface API; the MCP server cannot hear anything. If your agent does both, run both under one user identifier.
The harder decision is where credentials live. MCP leaves you holding an OAuth grant for every user, each subject to someone else's admin policy. The API leaves you holding one org key with no spending cap. Both are infrastructure problems. The Wispr surface you pick decides which one you inherit, not whether you have one. For a deep look at credential ownership across agent tool-calling patterns, see our dedicated guide.
Building a meeting-aware or voice-driven agent on Wispr Flow? Talk to us for immediate help with connector setup, Virtual MCP design, and multi-tenant auth.