REST API Reference
Base URL
Section titled “Base URL”All endpoints are prefixed with the configured base path:
{origin}/apiAuthentication
Section titled “Authentication”All endpoints require a bearer token in the Authorization header:
Authorization: Bearer <token>Two kinds of token are accepted, and they are told apart by their signature and audience, never by their shape:
| Token | Who signs it | Lives | Get one |
|---|---|---|---|
| Session | Supabase Auth | ~1 hour, refreshed by the app | The access_token of a signed-in session, or POST {SUPABASE_URL}/auth/v1/token?grant_type=password |
| Access token | Chatty itself | 1–365 days (default 90), revocable | Settings → Access tokens, or POST /api/me/tokens from a session |
An access token acts as the person who minted it. It cannot mint further
tokens (POST /me/tokens needs a session), and revoking it takes effect
on the next request.
Access tokens
Section titled “Access tokens”GET /api/me/tokens # the person's tokens, revoked ones includedPOST /api/me/tokens # { "name": "laptop", "expires_in_days": 90 } → the row plus `secret`, ONCEDELETE /api/me/tokens/{id} # revokeSessions
Section titled “Sessions”List sessions
Section titled “List sessions”Returns all sessions for the authenticated user, including their chats and messages.
GET /api/sessionsResponse 200 OK
[ { "id": "uuid", "user_id": "uuid", "title": "My Session", "master_chat_index": 0, "created_at": "2026-03-23T10:00:00Z", "updated_at": "2026-03-23T10:30:00Z", "chats": [ { "id": "uuid", "session_id": "uuid", "position": 0, "model_name": "GPT-4", "context_text": "", "title": "GPT-4 Chat", "created_at": "2026-03-23T10:00:00Z", "messages": [ { "id": "uuid", "chat_id": "uuid", "position": 0, "role": "user", "content": "Hello", "created_at": "2026-03-23T10:01:00Z" } ] } ] }]Create session
Section titled “Create session”POST /api/sessionsContent-Type: application/jsonRequest body
{ "title": "Compare Models" }The title field is optional; defaults to "New Session".
Response 200 OK
{ "id": "uuid", "user_id": "uuid", "title": "Compare Models", "master_chat_index": 0, "created_at": "2026-03-23T10:00:00Z", "updated_at": "2026-03-23T10:00:00Z"}Delete session
Section titled “Delete session”Deletes the session and all its chats and messages (cascade).
DELETE /api/sessions/{id}Response 200 OK
{ "deleted": true }Update master chat index
Section titled “Update master chat index”PATCH /api/sessions/{id}/masterContent-Type: application/jsonRequest body
{ "master_chat_index": 1 }Response 200 OK — Returns the updated session object.
Add chat to session
Section titled “Add chat to session”POST /api/sessions/{session_id}/chatsContent-Type: application/jsonRequest body
{ "model_name": "GPT-4", "title": "GPT-4 Chat", "context_text": "You are a helpful assistant."}| Field | Type | Required | Description |
|---|---|---|---|
model_name | string | Yes | Must match a registered model name |
title | string | Yes | Display title for the chat tab |
context_text | string | No | System prompt injected as first message |
Response 200 OK — Returns ChatWithMessages (chat object flattened with messages array).
Remove chat from session
Section titled “Remove chat from session”DELETE /api/sessions/{session_id}/chats/{position}Response 200 OK
{ "deleted": true }Query chat (non-streaming)
Section titled “Query chat (non-streaming)”Sends a user message and returns the complete assistant response.
POST /api/chats/{chat_id}/queryContent-Type: application/jsonRequest body
{ "content": "What is quantum computing?" }Response 200 OK — Returns ChatWithMessages with all messages including the new user and assistant messages.
Sync chat
Section titled “Sync chat”Truncates messages after a given position and re-queries from that point.
POST /api/chats/{chat_id}/syncContent-Type: application/jsonRequest body
{ "msg_idx": 2, "content": "Actually, explain it differently"}| Field | Type | Description |
|---|---|---|
msg_idx | integer | Position after which to truncate messages |
content | string | New user message to insert after truncation |
Response 200 OK — Returns ChatWithMessages with the updated message history.
Import chat
Section titled “Import chat”Imports a chat with full message history into an existing session.
POST /api/sessions/{session_id}/chats/importContent-Type: application/jsonRequest body
{ "model_name": "GPT-4", "title": "Imported Chat", "context_text": "Optional system prompt", "messages": [ { "role": "user", "content": "Hello" }, { "role": "assistant", "content": "Hi there!" } ]}Response 200 OK — Returns ChatWithMessages with the imported chat and all messages.
Models
Section titled “Models”List available models
Section titled “List available models”Returns all models registered in the plugin registry.
GET /api/modelsResponse 200 OK
[ { "name": "GPT-4", "provider": "openai", "model_id": "gpt-4" }, { "name": "Mistral Large", "provider": "mistral", "model_id": "mistral-large-latest" }, { "name": "llama3", "provider": "ollama", "model_id": "llama3:latest" }]Context Items
Section titled “Context Items”List context items
Section titled “List context items”GET /api/context-items?offset=0&limit=50| Parameter | Type | Default | Description |
|---|---|---|---|
offset | integer | 0 | Pagination offset |
limit | integer | 50 | Maximum items to return |
Response 200 OK
[ { "id": "uuid", "title": "Medical Expert", "category": "medical", "tags": ["doctor", "expert"], "owner": "all", "content": "You are a medical expert...", "created_at": "2026-03-23T10:00:00Z" }]Create context item
Section titled “Create context item”POST /api/context-itemsContent-Type: application/jsonRequest body
{ "title": "Code Reviewer", "category": "development", "tags": ["code", "review"], "content": "You are an expert code reviewer..."}| Field | Type | Required | Description |
|---|---|---|---|
title | string | Yes | Display title |
category | string | Yes | Category label |
tags | string[] | No | Tag array |
content | string | Yes | Context content (used for embedding generation) |
Response 200 OK — Returns the created ContextItem.
Search context items (semantic)
Section titled “Search context items (semantic)”GET /api/context-items/search?q=heart+disease&limit=10| Parameter | Type | Default | Description |
|---|---|---|---|
q | string | (required) | Search query text |
limit | integer | 10 | Maximum results |
Response 200 OK
[ { "id": "uuid", "title": "Cardiology Expert", "category": "medical", "tags": ["cardiology"], "owner": "all", "content": "You are a cardiologist...", "created_at": "2026-03-23T10:00:00Z", "similarity": 0.89 }]Suggestions
Section titled “Suggestions”Search query suggestions
Section titled “Search query suggestions”GET /api/suggestions?q=transformers&limit=5| Parameter | Type | Default | Description |
|---|---|---|---|
q | string | (required) | Search query text |
limit | integer | 5 | Maximum results |
Response 200 OK
[ { "id": "uuid", "query": "How do transformer models work?", "similarity": 0.92 }, { "id": "uuid", "query": "Explain attention mechanisms", "similarity": 0.87 }]Schedules
Section titled “Schedules”A schedule fires a graph on a five-field cron. All routes are workspace-scoped and behind authentication.
List, create, edit, delete
Section titled “List, create, edit, delete”GET /api/schedulesPOST /api/schedulesGET /api/schedules/{id}PATCH /api/schedules/{id}DELETE /api/schedules/{id}Create body
| Field | Type | Default | Description |
|---|---|---|---|
graph_id | uuid | (required) | The graph to run |
cron | string | (required) | Five fields — minute, hour, day of month, month, day of week. No seconds |
tz | string | UTC | IANA zone the cron is written in |
name | string | "" | What it is for |
input | string | "" | What the run starts with |
overlap | string | skip | skip or queue when the previous run is still alive |
A cron that does not parse is a 400 with the reason, not a schedule saved disabled.
Response 200 OK
{ "id": "uuid", "graph_id": "uuid", "name": "the morning report", "cron": "0 9 * * *", "tz": "Europe/Madrid", "overlap": "skip", "enabled": true, "next_at": "2026-09-06T07:00:00Z", "last_fired_at": null, "last_run_id": null}There is no “fire now” route on purpose: running a graph already has one
(POST /api/graphs/{id}/execute), and a second path would report a person’s run as a
schedule’s.
Doorbells
Section titled “Doorbells”A doorbell lets a system outside start a graph with a signed POST.
Managing them
Section titled “Managing them”GET /api/hooksPOST /api/hooksGET /api/hooks/{id}PATCH /api/hooks/{id}POST /api/hooks/{id}/rotateDELETE /api/hooks/{id}POST /api/hooks takes { "graph_id": "uuid", "name": "orders" }. Create and rotate are
the only two responses that carry the secret, and they carry it once — it is encrypted at
rest and never serialised again.
{ "id": "uuid", "graph_id": "uuid", "name": "orders", "enabled": true, "last_rang_at": null, "secret": "chy_hook_…"}Ringing one
Section titled “Ringing one”POST /api/hooks/{id}/ringX-Chatty-Signature: sha256=<HMAC-SHA256 of the raw body, keyed with the secret>The only unauthenticated route in the API. What grants access is not who you are, it is that the signature matches. The signature covers the raw bytes and is verified before the body is parsed. The body is capped at 1 MiB: it must be read whole before the signature can be checked, so it is the one place where someone with no credential makes the process allocate.
Response 202 Accepted
{ "run_id": "uuid", "status": "queued" }202, not 200 with a result: what has happened when this returns is that the work is
queued. Read the outcome from /api/graphs/runs/{id}.
Unknown hook, disabled hook, missing header and bad signature all answer the same 401.
Distinguishing them would hand a prober an oracle; the difference goes to the log.
A doorbell that has spent its hour answers 429 instead — see
Automations.
The run acts with the authority of whoever created the doorbell. The caller supplies the body and nothing else — not the graph, not the scope, not who it acts as.
Research
Section titled “Research”A research is the folder the rest of this hangs from: it gathers experiments, labelling jobs, graphs and datasets, keeps a logbook, and records the reasoning that ties them together.
GET /api/researchPOST /api/researchGET /api/research/{id}PUT /api/research/{id}DELETE /api/research/{id}Creating one also creates its own graph repository — a research is where versions live, so it is born with somewhere to put them.
What hangs from it
Section titled “What hangs from it”Four kinds of thing, all attached and detached the same way: PUT (or
POST, for the two that are not created here) to attach, DELETE to
let go.
GET /api/research/{id}/experimentsPUT /api/research/{id}/experiments/{exp_id}DELETE /api/research/{id}/experiments/{exp_id}
GET /api/research/{id}/jobsPUT /api/research/{id}/jobs/{job_id}DELETE /api/research/{id}/jobs/{job_id}
GET /api/research/{id}/graphsPOST /api/research/{id}/graphs/{graph_id}DELETE /api/research/{id}/graphs/{graph_id}
GET /api/research/{id}/datasetsPOST /api/research/{id}/datasets/{dataset_id}DELETE /api/research/{id}/datasets/{dataset_id}Detaching is not deleting. A graph attached to a research keeps existing when you let it go: attaching is a link, not ownership — which is why deleting the research does not take the graphs with it.
GET /api/research/{id}/results is the read the results screen makes:
every experiment’s numbers, per version, in one response.
The logbook
Section titled “The logbook”GET /api/research/{id}/docPUT /api/research/{id}/docPOST /api/research/{id}/doc/publishGET /api/research/{id}/doc/versionsGET /api/research/{id}/doc/versions/{n}PUT saves the working copy; publish mints a numbered version of it.
The ordinal in the path is that number — versions are addressed by
where they are in the line, not by an id.
The reasoning
Section titled “The reasoning”Moves are the argument: what was claimed, under what, citing what.
GET /api/research/{id}/movesPOST /api/research/{id}/movesPUT /api/moves/{id}PUT /api/moves/{child}/under/{parent}POST /api/moves/{id}/saysPOST /api/moves/{id}/citesunder re-hangs a move, says rewords it, cites attaches evidence —
a trial, a version, a job. They are separate routes and not one PATCH
because they are separate acts: moving an argument, changing what it
says, and backing it up are three different edits to the record.
Experiments
Section titled “Experiments”An experiment runs a graph against material, over and over, with a parameter space. What it produces is trials.
GET /api/experimentsPOST /api/experimentsGET /api/experiments/{id}PUT /api/experiments/{id}DELETE /api/experiments/{id}GET /api/experiments/{id}/modelsPOST /api/experiments/{id}/pinPOST /api/experiments/{id}/clone{ "graph_id": "uuid", "name": "Two prompts", "param_space": { "type": "grid", "params": [] }, "dataset_id": "uuid", "dataset_version_id": "uuid", "branch": "main"}An experiment pins what it runs against — both halves. The graph
(a version of it, on branch, which defaults to main) and the
material: dataset_version_id names material that already exists, and
dataset_id commits the live table as it stands. Neither, and input
is used as a single case. That is what makes a number from last week
mean something today.
The two pins move for different reasons, so they move separately:
pin takes graph_version_id, dataset_version_id, or both, and is
for an experiment nothing has been run against yet. Once it has trials,
clone is what it gets instead — a second experiment starting where
this one is, inheriting whichever pin you leave out.
GET /{id}/models reports what it actually ran against: a model
alias moves and every number moves with it without the configuration
changing, so it is recorded per run rather than being part of a
version’s name.
Running it
Section titled “Running it”POST /api/experiments/{id}/run → { "experiment_id": "uuid", "expected_total": 120 }POST /api/experiments/{id}/preview → { "trial_id": "uuid", "item_count": 5 }GET /api/experiments/{id}/eventsrun returns as soon as the work is queued, with how many trials to
expect. preview runs a handful of rows first — { "sample_size": 5, "overrides": {…} } — which is how you find out the configuration is
wrong before spending the whole grid.
events is Server-Sent Events, not JSON: one data: line per
event as trials finish, closing itself when the experiment does.
Three other routes answer the same way, and for the same reason — work
that takes minutes and has something to say while it runs:
POST /api/settings/reembed-system,
POST /api/collections/{name}/reembed and
POST /api/models/ollama/pull. All four stream progress rather than
making you poll; none of them is a WebSocket, because nothing is being
said back.
Trials
Section titled “Trials”GET /api/experiments/{id}/trials # ?include_preview=true to see the previews tooGET /api/trials/{id}GET /api/trials/{id}/runsPOST /api/trials/{id}/rerun → { "trial_id": "uuid", "item_count": 12 }POST /api/trials/{id}/annotation-jobPOST /api/trials/annotation-jobrerun takes the rows to run again, by their position in the version
the experiment pinned — which is why it can name them at all.
The last two turn results into something a person judges: one trial’s outputs as a labelling job, or two trials side by side for a comparison. See Labelling jobs.
Datasets
Section titled “Datasets”GET /api/datasetsPOST /api/datasetsGET /api/datasets/{id}PUT /api/datasets/{id}DELETE /api/datasets/{id}POST /api/datasets/{id}/copy-to-workspaceGET /api/datasets/{id}/items # ?offset=&limit= — a page, not the tablePATCH /api/datasets/{id}/items/{iid}DELETE /api/datasets/{id}/items/{iid}GET /api/datasets/{id}/profileitems pages on purpose: a dataset of fifty thousand rows is not a
response. profile is the per-column summary — nulls, cardinality,
range — computed in SQL rather than pulled out to be counted.
Columns
Section titled “Columns”GET /api/datasets/{id}/columnsPOST /api/datasets/{id}/columnsPUT /api/datasets/{id}/columns/{cid}DELETE /api/datasets/{id}/columns/{cid}POST /api/datasets/{id}/columns/reorderOperations
Section titled “Operations”POST /api/datasets/{id}/opsAn operation does not edit the table in place: it mints a version and tells you which.
{ "type": "filter", "column": "score", "test": { … }, "message": "only the scored ones" }{ "type": "sample", "n": 200, "seed": "optional" }Omit the seed of a sample and the server mints one and writes it into the version’s parameters. A sample whose seed is not recorded is not reproducible, and then it is not evidence of anything.
Vectorising one
Section titled “Vectorising one”GET /api/datasets/{id}/embeddingsPOST /api/datasets/{id}/embeddings → 201GET /api/datasets/{id}/embeddings/{eid}DELETE /api/datasets/{id}/embeddings/{eid}GET /api/datasets/{id}/embeddings/{eid}/points{ "columns": ["question", "answer"], "version_id": "uuid, optional", "model": "provider:model, optional", "method": "pca"}Embedding a dataset pins a version of it first, exactly as an
experiment does — the same rows have to come back next time. Without
model it uses the workspace’s embedder; without version_id it
commits the live table. points returns the projected coordinates,
which is what the scatter plot draws.
POST answers 201 and the work happens behind it: poll the row until
it says it is done.
The Vault
Section titled “The Vault”What the workspace knows, as a corpus: chats, runs and documents embedded and clustered into topics.
GET /api/vault/recall?q=…&k=8GET /api/vault/reasoningGET /api/vault/graphGET /api/vault/pendingPOST /api/vault/backfillPUT /api/vault/objects/{kind}/{id}/topicsPOST /api/vault/topics/recomputerecall is the interesting one, and it is what the research_recall
tool calls. k is how many moves come in by similarity, before
expanding — each one brings its neighbourhood along, so what comes
back is a good deal larger than this number. Eight by default.
reasoning is the other half: the whole argument, to read. recall is
the piece of it that answers a question.
pending says how much of the corpus has no vector yet and backfill
embeds a slice of it — it is a slice and not the lot because embedding
costs money, so it is paid for in visible instalments. topics on an
object pins it to a theme by hand (and a hand-pinned object stops being
re-assigned automatically); recompute refits the clustering.
Deployments
Section titled “Deployments”A deployment is a graph reachable from outside, under its own
credentials and its own ceilings. It is the third way a graph leaves the
platform, and the only one that answers in the same request: a schedule
runs it on a clock, a doorbell starts it and hands you a run_id, a
deployment holds a conversation.
One per person, for now. A second POST answers 409.
Two keys, and what each one opens
Section titled “Two keys, and what each one opens”| Key | Shape | Where it lives | What it opens |
|---|---|---|---|
| Public | chy_pk_… | The HTML of your page, in plain sight | The browser door only — the WebSocket, and only from an origin you listed |
| Secret | chy_sk_… | Your server | The HTTP door only |
The split is the same one Stripe and Intercom make, and for the same
reason: the public key is readable by anyone who views source, so what
stops another site from using it is not secrecy — it is the Origin the
browser sends, which a page cannot forge. The secret never reaches a
browser, so it needs no origin: holding it is the credential.
They do not cross. The public key in the HTTP door is a 404, and the
secret in the browser door is a 404. The secret is stored as a
SHA-256 digest and is serialised once, when you create or rotate it.
Managing them
Section titled “Managing them”GET /api/deploymentsPOST /api/deploymentsGET /api/deployments/{id}PATCH /api/deployments/{id}POST /api/deployments/{id}/rotateDELETE /api/deployments/{id}GET /api/deployments/{id}/visitorsPOST /api/deployments takes the graph, a name, the origins allowed to
open the browser door, and any ceilings you want to change:
{ "graph_id": "uuid", "name": "Support", "origins": ["https://my-site.com"], "limits": { "runs_per_day": 200 }}Create and rotate are the only two responses that carry the secret.
{ "id": "uuid", "graph_id": "uuid", "name": "Support", "public_key": "chy_pk_…", "status": "asleep", "origins": ["https://my-site.com"], "limits": { "runs_per_day": 200 }, "turns_total": 0, "secret": "chy_sk_…"}PATCH takes name, status (live or paused — see below),
origins, limits and graph_version_id. Rotating mints both keys:
the old pair stops working at once.
Which version answers
Section titled “Which version answers”By default a deployment runs the tip of its graph: every save in the
Studio goes live to whoever is chatting. graph_version_id pins a
version instead, and then the door keeps serving that one however much
the graph moves.
The field has three states in PATCH, and a null is not the same as
leaving it out:
| Body | What happens |
|---|---|
{"graph_version_id": "uuid"} | Pins that version. It has to be a version of this deployment’s own graph, or the call answers 400 |
{"graph_version_id": null} | Unpins: back to the tip |
| field absent | Left as it was |
A pinned version that cannot be read any more — deleted, or an id that was never one — does not silence the door: the turn falls back to the live graph and the backend logs a warning. A deployment is something other people are talking to, and answering with today’s graph beats not answering.
Awake, asleep, paused
Section titled “Awake, asleep, paused”status is three words and only two of them are yours.
| Status | What it means |
|---|---|
asleep | Nobody has spoken to it lately. It is not off: the first message wakes it, if there is room |
live | Awake and answering |
paused | You turned it off. Both doors answer 403 until you set it back to live |
A deployment is born asleep and goes back to sleep after idle_minutes
without anybody. How many can be awake at once is a ceiling of the
installation (DEPLOYMENTS_AWAKE_MAX, 20 by default), which is what
keeps a dozen widgets from saturating one backend; when there is no room,
waking answers 429 and the next message tries again.
Ceilings
Section titled “Ceilings”Every value has a default, and the row stores only what you changed.
| Limit | Default | What it counts |
|---|---|---|
concurrency | 2 | Conversations answered at the same time |
runs_per_visitor_hour | 60 | Messages one visitor may send in an hour |
runs_per_day | 1000 | Messages the deployment answers in a day |
idle_minutes | 30 | Without anybody, before it sleeps |
max_input_chars | 4000 | Length of one message |
Going over any of them answers 429. A turn that takes more than 120
seconds is stopped and reported as an error.
Every conversation runs on your behalf, with your workspace’s models and keys: what a deployment spends is yours, which is what these ceilings are for.
The server door
Section titled “The server door”POST /api/deployments/{id}/chatAuthorization: Bearer chy_sk_…{ "input": "Hello", "visitor": "uuid, from a previous reply" }Response 200
{ "visitor": "uuid", "chat_id": "uuid", "run_id": "uuid", "output": "…" }This one blocks until the graph answers — unlike a doorbell, which
queues. Keep the visitor and send it back: it is what makes the next
message a reply rather than a new conversation. Omit it and the server
mints a new one.
Like the doorbell ring, this route takes no session, so it sits behind the 5 req/s ceiling of the table below rather than the 30.
Unknown deployment, wrong key and a key of the wrong kind all answer the
same 404: telling them apart would hand a prober an oracle. Paused
is the exception — 403, and only to a caller who already proved they
hold a key.
The browser door
Section titled “The browser door”GET {origin}/ws/deployments/{id}Origin: https://my-site.comA WebSocket, because a browser cannot set an Authorization header on
one and a key in the URL ends up in everybody’s logs. The first frame is
the credential:
{ "type": "auth", "key": "chy_pk_…", "visitor": "uuid, optional" }Anything else as a first frame closes the socket. On success the server answers with the visitor it minted — keep it, it is the thread of the conversation:
{ "type": "deployment_ready", "visitor": "uuid", "deployment": "Support" }From then on, one frame per message:
{ "type": "run", "input": "Hello" }and what comes back is what the product’s own chat draws: token
({ chat_id, content }) as the answer streams, done
({ chat_id, full_content }) when it finishes, and error
({ chat_id, message }) when it does not. Progress frames
(graph_node_*) also arrive, for whoever wants to show them.
An error whose chat_id is null is about the connection and not
about a message: a key that does not open, an origin that is not listed,
a deployment with no room to wake. Those arrive instead of
deployment_ready, and the socket closes.
Origin is checked against the deployment’s list on every connection,
trailing slash and case ignored. A request with no Origin — curl, a
script — does not open this door at all: that is the point of it.
Error responses
Section titled “Error responses”All endpoints return errors in a consistent format:
{ "error": "Description of what went wrong" }| Status | Meaning |
|---|---|
401 Unauthorized | Missing or invalid JWT token |
403 Forbidden | You are asking for something you are not allowed to do — a team resource you are not an admin of, for instance. Unlike 429, retrying will not help |
404 Not Found | Resource not found or not owned by the user |
400 Bad Request | Invalid request body or model not available |
429 Too Many Requests | You went over a ceiling. Carries Retry-After in seconds — see below |
500 Internal Server Error | Database or provider error |
Export a labelling job
Section titled “Export a labelling job”GET /api/jobs/{id}/export?format=csv|jsonlReturns the file itself (not JSON), with Content-Disposition: attachment. Accepts the same query
filters as /api/jobs/{id}/admin/records, so ?disagreement=true&format=jsonl gives you only the
units the annotators split on.
Behind the same guard as the aggregated results: annotators get 403 while labelling. See
Labelling jobs for the two shapes and what
the file does and does not carry.
Rate limits
Section titled “Rate limits”There are two ceilings in front of this API, and they are keyed by the calling address:
| What | Sustained | Burst |
|---|---|---|
| Endpoints reachable without a session (doorbell rings, the deployment HTTP door, invitation lookups, tunnel handshake, the landing demo) | 5 req/s | 100 |
| Everything else | 30 req/s | 300 |
Going over either answers 429 with a Retry-After header in seconds. Wait that long and
retry; the same request will then succeed. This is not a permission problem — that is 403.
The numbers are deliberately loose. They are not there to shape normal use: opening the Studio fires dozens of requests at once, and the burst allowance exists so that a page which loads twelve things does not throttle itself. They are there so that a caller with no credential at all cannot make the server work indefinitely.
Two things are never throttled: /api/health, because a throttled probe reads as an unhealthy
container, and calls arriving from the machine itself — which is how the
platform MCP server reaches this API, once per agent tool call.