Algolia Crawler

Live

BASIC AUTH

WEB CRAWLING

Search

Algolia Crawler gives your agent control of web crawling for your Algolia indexes: create and configure crawlers, start reindexes, crawl specific URLs, test extraction, and inspect crawl runs and URL stats.

  • Per-user credentials: each call uses the actual user's credentials, never a shared bot.
  • Encrypted per-tenant vault: AES-256, resolved at request time, never in LLM context.
  • Scoped before every call: pre-call scope check, 90-day SIEM-exportable audit chain.
Algolia Crawler
agent · Acme Q3
Run
Which URLs failed in the last crawl of our docs crawler?
S
algoliacrawler_get_url_stats
98ms
Search ops agent
The docs crawler has 12 FAILED URLs, mostly fetch errors on pages that return 404. 1,840 URLs finished as DONE and 37 were SKIPPED.
Sources: 1 crawler, URL stats
algoliacrawler
12 failed URLs
18:29
Message Claude...

Tools your search ops agent reaches for on Algolia Crawler, scoped per user.

CALL ANY TOOL
20 tools for web crawling into Algolia: list and inspect crawlers, run, pause, and reindex them, crawl or test single URLs, and version crawler configurations.
algoliacrawler_cancel_task
Cancel task
Cancel a blocking task on a crawler so its schedule can resume. Returns an empty successful response when the task is cancelled. Use this when a task ran into an error and is blocking the crawler's schedule. Use get_task_status to check whether a task is still pending before cancelling. Requires a crawler id from list_crawlers and the id of the blocking task.
Parameters
Name
Type
Required
Description
id
string
Required
Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
taskID
string
Required
Unique ID (UUID) of the blocking task to cancel. Tasks that ran into an error block the crawler schedule until cancelled. Example: 98458796-b7bb-4703-8b1b-785c1080b110.
algoliacrawler_crawl_urls
Crawl urls
algoliacrawler_create_crawler
Create crawler
algoliacrawler_delete_crawl_runs
Delete crawl runs
algoliacrawler_delete_crawler
Delete crawler
algoliacrawler_get_config_version
Get config version
algoliacrawler_get_crawl_run_file
Get crawl run file
algoliacrawler_get_crawler
Get crawler
algoliacrawler_get_task_status
Get task status
algoliacrawler_get_url_stats
Get url stats
algoliacrawler_list_config_versions
List config versions
algoliacrawler_list_crawl_runs
List crawl runs
algoliacrawler_list_crawlers
List crawlers
algoliacrawler_list_domains
List domains
algoliacrawler_pause_crawler
Pause crawler
algoliacrawler_run_crawler
Run crawler
algoliacrawler_start_reindex
Start reindex
algoliacrawler_test_url
Test url
algoliacrawler_update_crawler
Update crawler
algoliacrawler_update_crawler_config
Update crawler config
Build your Agent
Same auth pattern across LangChain, OpenAI, Anthropic, and Google ADK.
Python · LlamaIndex
import { ScalekitClient } from "@scalekit-sdk/node";
import { createReactAgent } from "@langchain/langgraph/prebuilt";

const sk = new ScalekitClient(env.SCALEKIT_ENV_URL, env.SCALEKIT_CLIENT_ID, env.SCALEKIT_CLIENT_SECRET);

// Algolia Crawler tools scoped to this user
const { tools } = await sk.tools.listScopedTools("user_123", {
  filter: { connectionNames: ["algoliacrawler"], toolNames: [
    "algoliacrawler_list_crawlers",
    "algoliacrawler_get_url_stats",
    "algoliacrawler_crawl_urls"] },
  pageSize: 100,
});

const agent = createReactAgent({ llm, tools });
await agent.invoke({ messages: [{ role: "user", content: "Which URLs failed in the last crawl of our docs crawler?" }] });
import OpenAI from "openai";
import { ScalekitClient } from "@scalekit-sdk/node";

const sk = new ScalekitClient(env.SCALEKIT_ENV_URL, env.SCALEKIT_CLIENT_ID, env.SCALEKIT_CLIENT_SECRET);
const openai = new OpenAI();

const { tools } = await sk.tools.listScopedTools("user_123", {
  filter: { connectionNames: ["algoliacrawler"] }, pageSize: 100,
});

const res = await openai.chat.completions.create({
  model: "gpt-5",
  messages: [{ role: "user", content: "Which URLs failed in the last crawl of our docs crawler?" }],
  tools,
});

// Execute the tool call with the user's vaulted Algolia Crawler username and password
await sk.tools.executeTool(res.choices[0].message.tool_calls[0], "user_123");
import Anthropic from "@anthropic-ai/sdk";
import { ScalekitClient } from "@scalekit-sdk/node";

const sk = new ScalekitClient(env.SCALEKIT_ENV_URL, env.SCALEKIT_CLIENT_ID, env.SCALEKIT_CLIENT_SECRET);
const anthropic = new Anthropic();

const { tools } = await sk.tools.listScopedTools("user_123", {
  filter: { connectionNames: ["algoliacrawler"] }, pageSize: 100,
});

const msg = await anthropic.messages.create({
  model: "claude-sonnet-5",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Which URLs failed in the last crawl of our docs crawler?" }],
  tools,
});

// Tool call runs with the user's vaulted Algolia Crawler username and password
await sk.tools.executeTool(msg.content, "user_123");
import { Agent } from "@google/adk/agents";
import { ScalekitClient } from "@scalekit-sdk/node";

const sk = new ScalekitClient(env.SCALEKIT_ENV_URL, env.SCALEKIT_CLIENT_ID, env.SCALEKIT_CLIENT_SECRET);

const { tools } = await sk.tools.listScopedTools("user_123", {
  filter: { connectionNames: ["algoliacrawler"] }, pageSize: 100,
});

const agent = new Agent({
  name: "search_ops_agent",
  model: "gemini-2.5-pro",
  instruction: "Algolia Crawler tools scoped to this user",
  tools,
});

await agent.run("Which URLs failed in the last crawl of our docs crawler?");
Try these prompts
Copy any prompt into your agent. Each maps directly to a Algolia Crawler tool. Click to copy, paste into your agent, done.
Crawlers
Copy the prompt
Copied
List the crawlers on our Algolia account.
Copy the prompt
Copied
Is the docs crawler running, reindexing, or blocked?
Copy the prompt
Copied
Pause the blog crawler while we fix its config.
Crawl runs and health
Copy the prompt
Copied
Show crawl runs for the docs crawler from the last 7 days.
Copy the prompt
Copied
Break down the last crawl by URL status and reason.
Copy the prompt
Copied
Check whether the reindex task is still pending.
Recrawls and config
Copy the prompt
Copied
Recrawl these 5 pricing page URLs.
Copy the prompt
Copied
Test crawl this URL and show the records it would extract.
Copy the prompt
Copied
List the config versions of the docs crawler and who changed them.
SEE HOW AUTH WORKS
Each user connects their own Algolia Crawler username and password once; Scalekit sends them with every call. Credentials stay vaulted, every call is scope checked, and every action is logged.
1
Authorize
Your user connects
Algolia Crawler
once. We tie it to their identity and the meetings they approved — no shared bot account, no org-wide access
Who:
user ‘A’
when:
Once per user
access:
Limited to user
2
Store
Their
Algolia Crawler
token lives in a vault scoped to them. User A's meetings are never reachable by an agent acting for user B, even on the same connection
vault:
encrypted
scope:
per-user
tokens:
auto-refreshed
3
Resolve
When your agent calls a
Algolia Crawler
tool, we fetch the right token server-side. It never touches your agent, never appears in the LLM context, never shows up in your logs
speed:
~40ms
check:
before every call
seen by:
nobody
4
Audit
Every
Algolia Crawler
tool call is logged — who triggered it, which meeting was fetched, what came back. 90 days of history, tied to the user who authorized it
history:
90 days
export:
SIEM-ready
logged:
every call
Test other agents
See the same per-user auth pattern across other search and crawling connectors.
Engineering Teams
DevOps assistant agent
Polls GitHub for failing checks and stale PRs, opens Linear issues for the ones that need work, and posts a daily digest to Slack. It acts as the engineer, not a shared service account.
Engineering Teams
Engineering standup agent
Pulls commits from GitHub and GitLab, tracks issue movement in Jira, and posts a per-engineer standup brief to Slack. Each engineer's activity is read on their own delegated OAuth.
Engineering Teams
Slack triage
Polls Slack for new messages, classifies bugs and support requests with a LangGraph router, files GitHub issues or Zendesk tickets, and confirms in the thread.
Engineering Teams
Auto release notes agent
Reads merged GitHub PRs, groups them into structured release notes, publishes the page to Notion, and announces the release in Slack. Every call runs on the engineer's own delegated OAuth.
Test other agents
See the same per-user auth pattern across other search and crawling connectors.
ENGINEERING
DevOps assistant agent
Poll GitHub for failing checks and stale pull requests, open Linear issues for the ones that need work, and digest to Slack.
ENGINEERING
Engineering standup agent
Pull commits from GitHub and GitLab, track Jira issue movement, and post a per-engineer standup brief to Slack.
ENGINEERING
Slack triage agent
Classify new Slack messages as bugs or support requests, file the GitHub issue or Zendesk ticket, and reply in the thread.
ENGINEERING
Auto-release notes agent
Group merged GitHub PRs into structured release notes, publish the page to Notion, and announce the release in Slack.
Why Scalekit
Secure your agent's access. Connectors ship in minutes
01.
Shared tokens break per-user analytics
One shared Algolia Crawler login looks fine in a demo. In production every crawler change looks like one service account, and you cannot tell which user triggered it. Scalekit resolves the credential of the actual user who triggered the agent, never a shared bot.
// shared login
audit → bot_service_account

// scalekit
audit → user_abc ✓
02.
Authentication is not authorization
03.
Multi-tenancy is architectural
04.
Algolia Crawler today. Ten connectors tomorrow.
“Our agents act across Salesforce, Gong, Google Drive, and more, on behalf of every customer. Scalekit behind the scenes meant we can keep adding tools without ever rebuilding how credentials or tool calling work.”
Venu Madhav Kattagoni
Head of Engineering / Von
FAQs
Frequently Asked Questions
Does the agent access Algolia Crawler as the user or as a shared key?
As the user. Each user connects their own Algolia Crawler username and password once, and Scalekit sends them with every call. Audit logs attribute every action to that user, not a shared service account.
Where are the Algolia Crawler credentials stored?
In Scalekit's managed AES-256 token vault, namespaced per tenant. To rotate them, reconnect with new credentials. Revocation is a single dashboard action. Credentials never appear in prompts, logs, or LLM context.
Can I limit what the agent does in Algolia Crawler?
Yes. Filter by tool name in listScopedTools to expose only what you want. Scalekit also enforces scope checks before every API call.
What happens when a user revokes Algolia Crawler access?
The connection is invalidated on the next tool call. Subsequent requests for that user fail closed with a clear error. Other users in the tenant remain unaffected. The event is logged for audit.
Can the agent delete crawlers or crawl history?
Only if you expose those tools. Of the 20 tools, 9 read, 9 write, and 2 are destructive: algoliacrawler_delete_crawler and algoliacrawler_delete_crawl_runs. Leave them out of listScopedTools and the agent can still inspect crawlers, run reindexes, and test URLs.
Start in your coding agent
Up and running in one command
Install the Scalekit skill in your editor of choice. Connector, auth, tools, prompt, all wired up
Claude Code REPL
/plugin marketplace add scalekit-inc/claude-code-authstack
/plugin install agentkit@scalekit-auth-stack
Cursor Code REPL
# ~/.cursor/mcp.json
{
""mcpServers"": {
""algoliacrawler"": {
""url"": ""https://mcp.scalekit.com/algoliacrawler"",
""headers"": { ""Authorization"": ""Bearer $SCALEKIT_TOKEN"" }
}
}
}
Codex Code REPL
# ~/.codex/config.toml
[mcp_servers.algoliacrawler]
url = ""https://mcp.scalekit.com/algoliacrawler""
auth_env = ""SCALEKIT_TOKEN""
Copilot Code REPL
# .vscode/mcp.json
{
""servers"": {
""algoliacrawler"": {
""url"": ""https://mcp.scalekit.com/algoliacrawler"",
""type"": ""http""
}
}
}