Skip to content

MCP server

Chatty exposes itself as an MCP server. An agent connected to it can do what a person does in the app — open a research, write its reasoning, build a dataset, commit a version, run an experiment, set up a job that judges the trials — as the user whose token it carries. This is the other direction from MCP servers in the Library, where Chatty is the client.

The tools do not reach into the database. Each one calls the REST API with the caller’s credential, so scope, permissions and the rules of the domain are exactly the ones the interface and the Python package see. An agent cannot do through MCP anything its user could not do by hand.

The server is served two ways. Both carry the same tools.

Every deployment serves it at {origin}/api/mcp over streamable HTTP, behind the same authentication as the rest of the API. Pass the Supabase access token as a bearer, and X-Workspace-Id when working in a team workspace rather than your personal one.

Terminal window
claude mcp add --transport http chatty https://chatty-lab.com/api/mcp \
--header "Authorization: Bearer $CHATTY_TOKEN"

The token is either the session token the REST API takes (it expires in an hour) or an access token minted in Settings → Access tokens, which lives for months and can be revoked. See Authentication.

Every run of the engine gets the platform registered as an MCP server of its own, under a fixed id, with a credential the backend mints for that run (a run token: hours, not months) and the run’s workspace in the header. An agent definition grants it like any other server — it shows up as Chatty (this platform) in the MCP picker — and its tools arrive namespaced as mcp_chatty_add_move, mcp_chatty_read_reasoning, and so on. Nothing else about the run changes: the agent calls the same API, as the run’s user, with the same rules.

A chatty-runner runs in another process without an HTTP server, so it needs the API’s address in PLATFORM_MCP_URL (for example https://api.chatty-lab.com/api/mcp); the backend itself uses loopback.

The initialize handshake carries instructions the model reads before the tool list. They describe the shape of the platform in a few lines: a research is a repository that owns a graph, links datasets, and files experiments and jobs; its reasoning is a DAG of moves — question, hypothesis, attempt (which cites the trial it ran), finding (which says what it answers, validates or refutes), decision — and the agent is asked to write moves as it works, so the person reviewing the research sees the same reasoning the agent used.

ToolDoes
list_research, get_researchThe researches in the workspace; one with its graphs, datasets, experiments and jobs.
create_research, update_researchA research comes with its own graph.
link_dataset_to_research, assign_experiment_to_researchFile material under it.
research_resultsEvery finished trial of its experiments, with metrics, by graph version.
read_reasoningMoves, under, says, citations, and the derived standing of every question and hypothesis.
add_moveWrite a question, hypothesis, attempt, finding or decision (with course), hanging it under other moves in the same gesture.
sayA finding answers a question, validates or refutes a hypothesis; an attempt combines attempts. Saying the same verb to the same target again corrects the edge.
citeEvidence on an attempt or a finding: trial, experiment, graph_version, dataset_version, job, run… Only those two kinds may cite; the others are refused with the reason.
hang_move, reword_moveNothing is ever deleted.
recallWhat the workspace has learned about a query, across researches, with one hop of context. Needs an embedding model configured.
workspace_reasoningThe reasoning of every research in the workspace, with standings.
ToolDoes
list_datasets, get_datasetColumns and row count.
create_datasetWith its columns in one call: {name, type, widget, required, settings}, defaults text / input.
add_column, add_rows, read_rowsRows are objects keyed by column name.
commit_dataset, commit_graphA version; committing with nothing changed returns the same one. Jobs and experiments pin versions.
dataset_log, graph_log, diff_versionsThe lineage, and what changed between two versions.
ToolDoes
list_experiments, get_experiment, create_experimentPinned to the graph’s current version; optionally over a dataset and filed under a research.
run_experimentReturns at once; one trial per point of the parameter space.
list_trials, get_trialStatus, metrics, and the runs with their outputs and cost.
ToolDoes
list_jobs, get_job, create_jobOver a dataset version; unit_spec says what a unit is (row, group, fields with blind).
judge_trial, compare_trialsEvaluation jobs born from trials: one trial judged row by row, or several side by side.
list_job_templates, apply_job_template, add_questionpairwise, ranking, rubric, or questions one by one.
publish_jobFreezes the questionnaire and materialises the units.
job_agreement, job_evaluationAgreement per question; results by graph (wins, scores, a Bradley-Terry ranking when the wins connect).
ascend_jobThe verdict enters the reasoning as a finding citing the job and its trials, optionally saying what it validates or refutes.

list_graphs, get_graph, create_graph, update_graph, update_node, delete_graph, validate_graph, describe_node_types, list_models, execute_graph (waits for the run), get_run, list_runs.

  1. create_research → add_move a question.

  2. create_dataset with columns, add_rows, commit_dataset → a version id.

  3. add_move a hypothesis under the question; create_experiment over the research’s graph and the dataset; run_experiment; list_trials until the trial is completed.

  4. add_move an attempt under the hypothesis and cite the trial.

  5. add_move a finding under the attempt; say that it validates or refutes the hypothesis; read_reasoning now shows the hypothesis standing.

  6. Two trials to compare? compare_trials, apply_job_template pairwise, publish_job; when the annotators are done, job_evaluation, then ascend_job to put the verdict into the reasoning.

A refused request comes back as the tool’s error text, verbatim from the API: HTTP 400 Bad Request: {"error":"un question habla de movimientos y no de versiones ni de ensayos: citar es de attempt o de finding"}. The domain speaks in full sentences on purpose — the agent is meant to read them and change course, not retry.