Skip to content

REST API Reference

All endpoints are prefixed with the configured base path:

{origin}/api

All endpoints require a bearer token in the Authorization header:

Authorization: Bearer <token>

Two kinds of token are accepted, and they are told apart by their signature and audience, never by their shape:

TokenWho signs itLivesGet one
SessionSupabase Auth~1 hour, refreshed by the appThe access_token of a signed-in session, or POST {SUPABASE_URL}/auth/v1/token?grant_type=password
Access tokenChatty itself1–365 days (default 90), revocableSettings → Access tokens, or POST /api/me/tokens from a session

An access token acts as the person who minted it. It cannot mint further tokens (POST /me/tokens needs a session), and revoking it takes effect on the next request.

GET /api/me/tokens # the person's tokens, revoked ones included
POST /api/me/tokens # { "name": "laptop", "expires_in_days": 90 } → the row plus `secret`, ONCE
DELETE /api/me/tokens/{id} # revoke

Returns all sessions for the authenticated user, including their chats and messages.

GET /api/sessions

Response 200 OK

[
{
"id": "uuid",
"user_id": "uuid",
"title": "My Session",
"master_chat_index": 0,
"created_at": "2026-03-23T10:00:00Z",
"updated_at": "2026-03-23T10:30:00Z",
"chats": [
{
"id": "uuid",
"session_id": "uuid",
"position": 0,
"model_name": "GPT-4",
"context_text": "",
"title": "GPT-4 Chat",
"created_at": "2026-03-23T10:00:00Z",
"messages": [
{
"id": "uuid",
"chat_id": "uuid",
"position": 0,
"role": "user",
"content": "Hello",
"created_at": "2026-03-23T10:01:00Z"
}
]
}
]
}
]
POST /api/sessions
Content-Type: application/json

Request body

{ "title": "Compare Models" }

The title field is optional; defaults to "New Session".

Response 200 OK

{
"id": "uuid",
"user_id": "uuid",
"title": "Compare Models",
"master_chat_index": 0,
"created_at": "2026-03-23T10:00:00Z",
"updated_at": "2026-03-23T10:00:00Z"
}

Deletes the session and all its chats and messages (cascade).

DELETE /api/sessions/{id}

Response 200 OK

{ "deleted": true }
PATCH /api/sessions/{id}/master
Content-Type: application/json

Request body

{ "master_chat_index": 1 }

Response 200 OK — Returns the updated session object.


POST /api/sessions/{session_id}/chats
Content-Type: application/json

Request body

{
"model_name": "GPT-4",
"title": "GPT-4 Chat",
"context_text": "You are a helpful assistant."
}
FieldTypeRequiredDescription
model_namestringYesMust match a registered model name
titlestringYesDisplay title for the chat tab
context_textstringNoSystem prompt injected as first message

Response 200 OK — Returns ChatWithMessages (chat object flattened with messages array).

DELETE /api/sessions/{session_id}/chats/{position}

Response 200 OK

{ "deleted": true }

Sends a user message and returns the complete assistant response.

POST /api/chats/{chat_id}/query
Content-Type: application/json

Request body

{ "content": "What is quantum computing?" }

Response 200 OK — Returns ChatWithMessages with all messages including the new user and assistant messages.

Truncates messages after a given position and re-queries from that point.

POST /api/chats/{chat_id}/sync
Content-Type: application/json

Request body

{
"msg_idx": 2,
"content": "Actually, explain it differently"
}
FieldTypeDescription
msg_idxintegerPosition after which to truncate messages
contentstringNew user message to insert after truncation

Response 200 OK — Returns ChatWithMessages with the updated message history.

Imports a chat with full message history into an existing session.

POST /api/sessions/{session_id}/chats/import
Content-Type: application/json

Request body

{
"model_name": "GPT-4",
"title": "Imported Chat",
"context_text": "Optional system prompt",
"messages": [
{ "role": "user", "content": "Hello" },
{ "role": "assistant", "content": "Hi there!" }
]
}

Response 200 OK — Returns ChatWithMessages with the imported chat and all messages.


Returns all models registered in the plugin registry.

GET /api/models

Response 200 OK

[
{ "name": "GPT-4", "provider": "openai", "model_id": "gpt-4" },
{ "name": "Mistral Large", "provider": "mistral", "model_id": "mistral-large-latest" },
{ "name": "llama3", "provider": "ollama", "model_id": "llama3:latest" }
]

GET /api/context-items?offset=0&limit=50
ParameterTypeDefaultDescription
offsetinteger0Pagination offset
limitinteger50Maximum items to return

Response 200 OK

[
{
"id": "uuid",
"title": "Medical Expert",
"category": "medical",
"tags": ["doctor", "expert"],
"owner": "all",
"content": "You are a medical expert...",
"created_at": "2026-03-23T10:00:00Z"
}
]
POST /api/context-items
Content-Type: application/json

Request body

{
"title": "Code Reviewer",
"category": "development",
"tags": ["code", "review"],
"content": "You are an expert code reviewer..."
}
FieldTypeRequiredDescription
titlestringYesDisplay title
categorystringYesCategory label
tagsstring[]NoTag array
contentstringYesContext content (used for embedding generation)

Response 200 OK — Returns the created ContextItem.

GET /api/context-items/search?q=heart+disease&limit=10
ParameterTypeDefaultDescription
qstring(required)Search query text
limitinteger10Maximum results

Response 200 OK

[
{
"id": "uuid",
"title": "Cardiology Expert",
"category": "medical",
"tags": ["cardiology"],
"owner": "all",
"content": "You are a cardiologist...",
"created_at": "2026-03-23T10:00:00Z",
"similarity": 0.89
}
]

GET /api/suggestions?q=transformers&limit=5
ParameterTypeDefaultDescription
qstring(required)Search query text
limitinteger5Maximum results

Response 200 OK

[
{ "id": "uuid", "query": "How do transformer models work?", "similarity": 0.92 },
{ "id": "uuid", "query": "Explain attention mechanisms", "similarity": 0.87 }
]

A schedule fires a graph on a five-field cron. All routes are workspace-scoped and behind authentication.

GET /api/schedules
POST /api/schedules
GET /api/schedules/{id}
PATCH /api/schedules/{id}
DELETE /api/schedules/{id}

Create body

FieldTypeDefaultDescription
graph_iduuid(required)The graph to run
cronstring(required)Five fields — minute, hour, day of month, month, day of week. No seconds
tzstringUTCIANA zone the cron is written in
namestring""What it is for
inputstring""What the run starts with
overlapstringskipskip or queue when the previous run is still alive

A cron that does not parse is a 400 with the reason, not a schedule saved disabled.

Response 200 OK

{
"id": "uuid",
"graph_id": "uuid",
"name": "the morning report",
"cron": "0 9 * * *",
"tz": "Europe/Madrid",
"overlap": "skip",
"enabled": true,
"next_at": "2026-09-06T07:00:00Z",
"last_fired_at": null,
"last_run_id": null
}

There is no “fire now” route on purpose: running a graph already has one (POST /api/graphs/{id}/execute), and a second path would report a person’s run as a schedule’s.


A doorbell lets a system outside start a graph with a signed POST.

GET /api/hooks
POST /api/hooks
GET /api/hooks/{id}
PATCH /api/hooks/{id}
POST /api/hooks/{id}/rotate
DELETE /api/hooks/{id}

POST /api/hooks takes { "graph_id": "uuid", "name": "orders" }. Create and rotate are the only two responses that carry the secret, and they carry it once — it is encrypted at rest and never serialised again.

{
"id": "uuid",
"graph_id": "uuid",
"name": "orders",
"enabled": true,
"last_rang_at": null,
"secret": "chy_hook_…"
}
POST /api/hooks/{id}/ring
X-Chatty-Signature: sha256=<HMAC-SHA256 of the raw body, keyed with the secret>

The only unauthenticated route in the API. What grants access is not who you are, it is that the signature matches. The signature covers the raw bytes and is verified before the body is parsed. The body is capped at 1 MiB: it must be read whole before the signature can be checked, so it is the one place where someone with no credential makes the process allocate.

Response 202 Accepted

{ "run_id": "uuid", "status": "queued" }

202, not 200 with a result: what has happened when this returns is that the work is queued. Read the outcome from /api/graphs/runs/{id}.

Unknown hook, disabled hook, missing header and bad signature all answer the same 401. Distinguishing them would hand a prober an oracle; the difference goes to the log.

A doorbell that has spent its hour answers 429 instead — see Automations.

The run acts with the authority of whoever created the doorbell. The caller supplies the body and nothing else — not the graph, not the scope, not who it acts as.


A research is the folder the rest of this hangs from: it gathers experiments, labelling jobs, graphs and datasets, keeps a logbook, and records the reasoning that ties them together.

GET /api/research
POST /api/research
GET /api/research/{id}
PUT /api/research/{id}
DELETE /api/research/{id}

Creating one also creates its own graph repository — a research is where versions live, so it is born with somewhere to put them.

Four kinds of thing, all attached and detached the same way: PUT (or POST, for the two that are not created here) to attach, DELETE to let go.

GET /api/research/{id}/experiments
PUT /api/research/{id}/experiments/{exp_id}
DELETE /api/research/{id}/experiments/{exp_id}
GET /api/research/{id}/jobs
PUT /api/research/{id}/jobs/{job_id}
DELETE /api/research/{id}/jobs/{job_id}
GET /api/research/{id}/graphs
POST /api/research/{id}/graphs/{graph_id}
DELETE /api/research/{id}/graphs/{graph_id}
GET /api/research/{id}/datasets
POST /api/research/{id}/datasets/{dataset_id}
DELETE /api/research/{id}/datasets/{dataset_id}

Detaching is not deleting. A graph attached to a research keeps existing when you let it go: attaching is a link, not ownership — which is why deleting the research does not take the graphs with it.

GET /api/research/{id}/results is the read the results screen makes: every experiment’s numbers, per version, in one response.

GET /api/research/{id}/doc
PUT /api/research/{id}/doc
POST /api/research/{id}/doc/publish
GET /api/research/{id}/doc/versions
GET /api/research/{id}/doc/versions/{n}

PUT saves the working copy; publish mints a numbered version of it. The ordinal in the path is that number — versions are addressed by where they are in the line, not by an id.

Moves are the argument: what was claimed, under what, citing what.

GET /api/research/{id}/moves
POST /api/research/{id}/moves
PUT /api/moves/{id}
PUT /api/moves/{child}/under/{parent}
POST /api/moves/{id}/says
POST /api/moves/{id}/cites

under re-hangs a move, says rewords it, cites attaches evidence — a trial, a version, a job. They are separate routes and not one PATCH because they are separate acts: moving an argument, changing what it says, and backing it up are three different edits to the record.


An experiment runs a graph against material, over and over, with a parameter space. What it produces is trials.

GET /api/experiments
POST /api/experiments
GET /api/experiments/{id}
PUT /api/experiments/{id}
DELETE /api/experiments/{id}
GET /api/experiments/{id}/models
POST /api/experiments/{id}/pin
POST /api/experiments/{id}/clone
{
"graph_id": "uuid",
"name": "Two prompts",
"param_space": { "type": "grid", "params": [] },
"dataset_id": "uuid",
"dataset_version_id": "uuid",
"branch": "main"
}

An experiment pins what it runs against — both halves. The graph (a version of it, on branch, which defaults to main) and the material: dataset_version_id names material that already exists, and dataset_id commits the live table as it stands. Neither, and input is used as a single case. That is what makes a number from last week mean something today.

The two pins move for different reasons, so they move separately: pin takes graph_version_id, dataset_version_id, or both, and is for an experiment nothing has been run against yet. Once it has trials, clone is what it gets instead — a second experiment starting where this one is, inheriting whichever pin you leave out.

GET /{id}/models reports what it actually ran against: a model alias moves and every number moves with it without the configuration changing, so it is recorded per run rather than being part of a version’s name.

POST /api/experiments/{id}/run → { "experiment_id": "uuid", "expected_total": 120 }
POST /api/experiments/{id}/preview → { "trial_id": "uuid", "item_count": 5 }
GET /api/experiments/{id}/events

run returns as soon as the work is queued, with how many trials to expect. preview runs a handful of rows first — { "sample_size": 5, "overrides": {…} } — which is how you find out the configuration is wrong before spending the whole grid.

events is Server-Sent Events, not JSON: one data: line per event as trials finish, closing itself when the experiment does.

Three other routes answer the same way, and for the same reason — work that takes minutes and has something to say while it runs: POST /api/settings/reembed-system, POST /api/collections/{name}/reembed and POST /api/models/ollama/pull. All four stream progress rather than making you poll; none of them is a WebSocket, because nothing is being said back.

GET /api/experiments/{id}/trials # ?include_preview=true to see the previews too
GET /api/trials/{id}
GET /api/trials/{id}/runs
POST /api/trials/{id}/rerun → { "trial_id": "uuid", "item_count": 12 }
POST /api/trials/{id}/annotation-job
POST /api/trials/annotation-job

rerun takes the rows to run again, by their position in the version the experiment pinned — which is why it can name them at all.

The last two turn results into something a person judges: one trial’s outputs as a labelling job, or two trials side by side for a comparison. See Labelling jobs.


GET /api/datasets
POST /api/datasets
GET /api/datasets/{id}
PUT /api/datasets/{id}
DELETE /api/datasets/{id}
POST /api/datasets/{id}/copy-to-workspace
GET /api/datasets/{id}/items # ?offset=&limit= — a page, not the table
PATCH /api/datasets/{id}/items/{iid}
DELETE /api/datasets/{id}/items/{iid}
GET /api/datasets/{id}/profile

items pages on purpose: a dataset of fifty thousand rows is not a response. profile is the per-column summary — nulls, cardinality, range — computed in SQL rather than pulled out to be counted.

GET /api/datasets/{id}/columns
POST /api/datasets/{id}/columns
PUT /api/datasets/{id}/columns/{cid}
DELETE /api/datasets/{id}/columns/{cid}
POST /api/datasets/{id}/columns/reorder
POST /api/datasets/{id}/ops

An operation does not edit the table in place: it mints a version and tells you which.

{ "type": "filter", "column": "score", "test": { … }, "message": "only the scored ones" }
{ "type": "sample", "n": 200, "seed": "optional" }

Omit the seed of a sample and the server mints one and writes it into the version’s parameters. A sample whose seed is not recorded is not reproducible, and then it is not evidence of anything.

GET /api/datasets/{id}/embeddings
POST /api/datasets/{id}/embeddings → 201
GET /api/datasets/{id}/embeddings/{eid}
DELETE /api/datasets/{id}/embeddings/{eid}
GET /api/datasets/{id}/embeddings/{eid}/points
{
"columns": ["question", "answer"],
"version_id": "uuid, optional",
"model": "provider:model, optional",
"method": "pca"
}

Embedding a dataset pins a version of it first, exactly as an experiment does — the same rows have to come back next time. Without model it uses the workspace’s embedder; without version_id it commits the live table. points returns the projected coordinates, which is what the scatter plot draws.

POST answers 201 and the work happens behind it: poll the row until it says it is done.


What the workspace knows, as a corpus: chats, runs and documents embedded and clustered into topics.

GET /api/vault/recall?q=…&k=8
GET /api/vault/reasoning
GET /api/vault/graph
GET /api/vault/pending
POST /api/vault/backfill
PUT /api/vault/objects/{kind}/{id}/topics
POST /api/vault/topics/recompute

recall is the interesting one, and it is what the research_recall tool calls. k is how many moves come in by similarity, before expanding — each one brings its neighbourhood along, so what comes back is a good deal larger than this number. Eight by default.

reasoning is the other half: the whole argument, to read. recall is the piece of it that answers a question.

pending says how much of the corpus has no vector yet and backfill embeds a slice of it — it is a slice and not the lot because embedding costs money, so it is paid for in visible instalments. topics on an object pins it to a theme by hand (and a hand-pinned object stops being re-assigned automatically); recompute refits the clustering.


A deployment is a graph reachable from outside, under its own credentials and its own ceilings. It is the third way a graph leaves the platform, and the only one that answers in the same request: a schedule runs it on a clock, a doorbell starts it and hands you a run_id, a deployment holds a conversation.

One per person, for now. A second POST answers 409.

KeyShapeWhere it livesWhat it opens
Publicchy_pk_…The HTML of your page, in plain sightThe browser door only — the WebSocket, and only from an origin you listed
Secretchy_sk_…Your serverThe HTTP door only

The split is the same one Stripe and Intercom make, and for the same reason: the public key is readable by anyone who views source, so what stops another site from using it is not secrecy — it is the Origin the browser sends, which a page cannot forge. The secret never reaches a browser, so it needs no origin: holding it is the credential.

They do not cross. The public key in the HTTP door is a 404, and the secret in the browser door is a 404. The secret is stored as a SHA-256 digest and is serialised once, when you create or rotate it.

GET /api/deployments
POST /api/deployments
GET /api/deployments/{id}
PATCH /api/deployments/{id}
POST /api/deployments/{id}/rotate
DELETE /api/deployments/{id}
GET /api/deployments/{id}/visitors

POST /api/deployments takes the graph, a name, the origins allowed to open the browser door, and any ceilings you want to change:

{
"graph_id": "uuid",
"name": "Support",
"origins": ["https://my-site.com"],
"limits": { "runs_per_day": 200 }
}

Create and rotate are the only two responses that carry the secret.

{
"id": "uuid",
"graph_id": "uuid",
"name": "Support",
"public_key": "chy_pk_…",
"status": "asleep",
"origins": ["https://my-site.com"],
"limits": { "runs_per_day": 200 },
"turns_total": 0,
"secret": "chy_sk_…"
}

PATCH takes name, status (live or paused — see below), origins, limits and graph_version_id. Rotating mints both keys: the old pair stops working at once.

By default a deployment runs the tip of its graph: every save in the Studio goes live to whoever is chatting. graph_version_id pins a version instead, and then the door keeps serving that one however much the graph moves.

The field has three states in PATCH, and a null is not the same as leaving it out:

BodyWhat happens
{"graph_version_id": "uuid"}Pins that version. It has to be a version of this deployment’s own graph, or the call answers 400
{"graph_version_id": null}Unpins: back to the tip
field absentLeft as it was

A pinned version that cannot be read any more — deleted, or an id that was never one — does not silence the door: the turn falls back to the live graph and the backend logs a warning. A deployment is something other people are talking to, and answering with today’s graph beats not answering.

status is three words and only two of them are yours.

StatusWhat it means
asleepNobody has spoken to it lately. It is not off: the first message wakes it, if there is room
liveAwake and answering
pausedYou turned it off. Both doors answer 403 until you set it back to live

A deployment is born asleep and goes back to sleep after idle_minutes without anybody. How many can be awake at once is a ceiling of the installation (DEPLOYMENTS_AWAKE_MAX, 20 by default), which is what keeps a dozen widgets from saturating one backend; when there is no room, waking answers 429 and the next message tries again.

Every value has a default, and the row stores only what you changed.

LimitDefaultWhat it counts
concurrency2Conversations answered at the same time
runs_per_visitor_hour60Messages one visitor may send in an hour
runs_per_day1000Messages the deployment answers in a day
idle_minutes30Without anybody, before it sleeps
max_input_chars4000Length of one message

Going over any of them answers 429. A turn that takes more than 120 seconds is stopped and reported as an error.

Every conversation runs on your behalf, with your workspace’s models and keys: what a deployment spends is yours, which is what these ceilings are for.

POST /api/deployments/{id}/chat
Authorization: Bearer chy_sk_…
{ "input": "Hello", "visitor": "uuid, from a previous reply" }

Response 200

{ "visitor": "uuid", "chat_id": "uuid", "run_id": "uuid", "output": "…" }

This one blocks until the graph answers — unlike a doorbell, which queues. Keep the visitor and send it back: it is what makes the next message a reply rather than a new conversation. Omit it and the server mints a new one.

Like the doorbell ring, this route takes no session, so it sits behind the 5 req/s ceiling of the table below rather than the 30.

Unknown deployment, wrong key and a key of the wrong kind all answer the same 404: telling them apart would hand a prober an oracle. Paused is the exception — 403, and only to a caller who already proved they hold a key.

GET {origin}/ws/deployments/{id}
Origin: https://my-site.com

A WebSocket, because a browser cannot set an Authorization header on one and a key in the URL ends up in everybody’s logs. The first frame is the credential:

{ "type": "auth", "key": "chy_pk_…", "visitor": "uuid, optional" }

Anything else as a first frame closes the socket. On success the server answers with the visitor it minted — keep it, it is the thread of the conversation:

{ "type": "deployment_ready", "visitor": "uuid", "deployment": "Support" }

From then on, one frame per message:

{ "type": "run", "input": "Hello" }

and what comes back is what the product’s own chat draws: token ({ chat_id, content }) as the answer streams, done ({ chat_id, full_content }) when it finishes, and error ({ chat_id, message }) when it does not. Progress frames (graph_node_*) also arrive, for whoever wants to show them.

An error whose chat_id is null is about the connection and not about a message: a key that does not open, an origin that is not listed, a deployment with no room to wake. Those arrive instead of deployment_ready, and the socket closes.

Origin is checked against the deployment’s list on every connection, trailing slash and case ignored. A request with no Origin — curl, a script — does not open this door at all: that is the point of it.


All endpoints return errors in a consistent format:

{ "error": "Description of what went wrong" }
StatusMeaning
401 UnauthorizedMissing or invalid JWT token
403 ForbiddenYou are asking for something you are not allowed to do — a team resource you are not an admin of, for instance. Unlike 429, retrying will not help
404 Not FoundResource not found or not owned by the user
400 Bad RequestInvalid request body or model not available
429 Too Many RequestsYou went over a ceiling. Carries Retry-After in seconds — see below
500 Internal Server ErrorDatabase or provider error
GET /api/jobs/{id}/export?format=csv|jsonl

Returns the file itself (not JSON), with Content-Disposition: attachment. Accepts the same query filters as /api/jobs/{id}/admin/records, so ?disagreement=true&format=jsonl gives you only the units the annotators split on.

Behind the same guard as the aggregated results: annotators get 403 while labelling. See Labelling jobs for the two shapes and what the file does and does not carry.

There are two ceilings in front of this API, and they are keyed by the calling address:

WhatSustainedBurst
Endpoints reachable without a session (doorbell rings, the deployment HTTP door, invitation lookups, tunnel handshake, the landing demo)5 req/s100
Everything else30 req/s300

Going over either answers 429 with a Retry-After header in seconds. Wait that long and retry; the same request will then succeed. This is not a permission problem — that is 403.

The numbers are deliberately loose. They are not there to shape normal use: opening the Studio fires dozens of requests at once, and the burst allowance exists so that a page which loads twelve things does not throttle itself. They are there so that a caller with no credential at all cannot make the server work indefinitely.

Two things are never throttled: /api/health, because a throttled probe reads as an unhealthy container, and calls arriving from the machine itself — which is how the platform MCP server reaches this API, once per agent tool call.