MCP server
Chatty exposes itself as an MCP server. An agent connected to it can do what a person does in the app — open a research, write its reasoning, build a dataset, commit a version, run an experiment, set up a job that judges the trials — as the user whose token it carries. This is the other direction from MCP servers in the Library, where Chatty is the client.
The tools do not reach into the database. Each one calls the REST API with the caller’s credential, so scope, permissions and the rules of the domain are exactly the ones the interface and the Python package see. An agent cannot do through MCP anything its user could not do by hand.
Connecting
Section titled “Connecting”The server is served two ways. Both carry the same tools.
Every deployment serves it at {origin}/api/mcp over streamable HTTP,
behind the same authentication as the rest of the API. Pass the Supabase
access token as a bearer, and X-Workspace-Id when working in a team
workspace rather than your personal one.
claude mcp add --transport http chatty https://chatty-lab.com/api/mcp \ --header "Authorization: Bearer $CHATTY_TOKEN"The token is either the session token the REST API takes (it expires in an hour) or an access token minted in Settings → Access tokens, which lives for months and can be revoked. See Authentication.
chatty-mcp speaks MCP over stdio and forwards to a backend over HTTP.
MCP_TOKEN=$CHATTY_TOKEN MCP_BACKEND_URL=https://chatty-lab.com chatty-mcp| Variable | Meaning |
|---|---|
MCP_TOKEN | A user’s access token. |
MCP_BACKEND_URL | The backend, base path included (default http://localhost:5050). |
MCP_WORKSPACE | Team id, when the work is in a team workspace. |
MCP_SERVICE_USER_ID + SUPABASE_JWT_SECRET | Instead of MCP_TOKEN: mint a year-long HS256 token. Only a self-hosted Supabase verifies those; Supabase Cloud signs with ES256 and rejects them. |
The agent inside the app
Section titled “The agent inside the app”Every run of the engine gets the platform registered as an MCP server of
its own, under a fixed id, with a credential the backend mints for that
run (a run token: hours, not months) and the run’s workspace in the
header. An agent definition grants it like any other server — it shows up
as Chatty (this platform) in the MCP picker — and its tools arrive
namespaced as mcp_chatty_add_move, mcp_chatty_read_reasoning, and so
on. Nothing else about the run changes: the agent calls the same API,
as the run’s user, with the same rules.
A chatty-runner runs in another process without an HTTP server, so it
needs the API’s address in PLATFORM_MCP_URL (for example
https://api.chatty-lab.com/api/mcp); the backend itself uses loopback.
What the server says about itself
Section titled “What the server says about itself”The initialize handshake carries instructions the model reads before the
tool list. They describe the shape of the platform in a few lines: a research
is a repository that owns a graph, links datasets, and files experiments and
jobs; its reasoning is a DAG of moves — question, hypothesis, attempt (which
cites the trial it ran), finding (which says what it answers, validates or
refutes), decision — and the agent is asked to write moves as it works, so
the person reviewing the research sees the same reasoning the agent used.
The tools
Section titled “The tools”Researches and reasoning
Section titled “Researches and reasoning”| Tool | Does |
|---|---|
list_research, get_research | The researches in the workspace; one with its graphs, datasets, experiments and jobs. |
create_research, update_research | A research comes with its own graph. |
link_dataset_to_research, assign_experiment_to_research | File material under it. |
research_results | Every finished trial of its experiments, with metrics, by graph version. |
read_reasoning | Moves, under, says, citations, and the derived standing of every question and hypothesis. |
add_move | Write a question, hypothesis, attempt, finding or decision (with course), hanging it under other moves in the same gesture. |
say | A finding answers a question, validates or refutes a hypothesis; an attempt combines attempts. Saying the same verb to the same target again corrects the edge. |
cite | Evidence on an attempt or a finding: trial, experiment, graph_version, dataset_version, job, run… Only those two kinds may cite; the others are refused with the reason. |
hang_move, reword_move | Nothing is ever deleted. |
recall | What the workspace has learned about a query, across researches, with one hop of context. Needs an embedding model configured. |
workspace_reasoning | The reasoning of every research in the workspace, with standings. |
Datasets and versions
Section titled “Datasets and versions”| Tool | Does |
|---|---|
list_datasets, get_dataset | Columns and row count. |
create_dataset | With its columns in one call: {name, type, widget, required, settings}, defaults text / input. |
add_column, add_rows, read_rows | Rows are objects keyed by column name. |
commit_dataset, commit_graph | A version; committing with nothing changed returns the same one. Jobs and experiments pin versions. |
dataset_log, graph_log, diff_versions | The lineage, and what changed between two versions. |
Experiments and trials
Section titled “Experiments and trials”| Tool | Does |
|---|---|
list_experiments, get_experiment, create_experiment | Pinned to the graph’s current version; optionally over a dataset and filed under a research. |
run_experiment | Returns at once; one trial per point of the parameter space. |
list_trials, get_trial | Status, metrics, and the runs with their outputs and cost. |
| Tool | Does |
|---|---|
list_jobs, get_job, create_job | Over a dataset version; unit_spec says what a unit is (row, group, fields with blind). |
judge_trial, compare_trials | Evaluation jobs born from trials: one trial judged row by row, or several side by side. |
list_job_templates, apply_job_template, add_question | pairwise, ranking, rubric, or questions one by one. |
publish_job | Freezes the questionnaire and materialises the units. |
job_agreement, job_evaluation | Agreement per question; results by graph (wins, scores, a Bradley-Terry ranking when the wins connect). |
ascend_job | The verdict enters the reasoning as a finding citing the job and its trials, optionally saying what it validates or refutes. |
Graphs
Section titled “Graphs”list_graphs, get_graph, create_graph, update_graph, update_node,
delete_graph, validate_graph, describe_node_types, list_models,
execute_graph (waits for the run), get_run, list_runs.
A session, end to end
Section titled “A session, end to end”-
create_research→add_movea question. -
create_datasetwith columns,add_rows,commit_dataset→ a version id. -
add_movea hypothesis under the question;create_experimentover the research’s graph and the dataset;run_experiment;list_trialsuntil the trial iscompleted. -
add_movean attempt under the hypothesis andcitethe trial. -
add_movea finding under the attempt;saythat itvalidatesorrefutesthe hypothesis;read_reasoningnow shows the hypothesis standing. -
Two trials to compare?
compare_trials,apply_job_template pairwise,publish_job; when the annotators are done,job_evaluation, thenascend_jobto put the verdict into the reasoning.
Errors
Section titled “Errors”A refused request comes back as the tool’s error text, verbatim from the API:
HTTP 400 Bad Request: {"error":"un question habla de movimientos y no de versiones ni de ensayos: citar es de attempt o de finding"}. The domain
speaks in full sentences on purpose — the agent is meant to read them and
change course, not retry.