Research preview

Build agents.
Prove what they can do.

From first idea to evidence. Design agent systems, inspect every step, and discover which version works best. All in one lab.

01

Shape it.

Connect models, tools, and agents. On a shared canvas or in Python.

02

See inside.

Every step, every output, every cost. Follow the run as it unfolds.

03

Improve with evidence.

Change a variable. Compare versions. Keep the trail behind every result.

From idea to evidence

Every iteration answers more.

Design, run, evaluate. Change what matters and test again. Your graph, results, and history stay connected.

The system you test in a conversation is the one you evaluate across a dataset. Take a promising idea to a conclusion you can stand behind.

One continuous research loop
01

Design

connect agents and tools

02

Run

with a trace for every node

03

Evaluate

quality, cost, latency, failure

04

Branch

change one thing, keep the history

05

Conclude

with evidence you can revisit

Every conclusion points back to the exact system that produced it.reproducible by construction

Your workspace

Everything connected. You in control.

A canvas to think. Code to go further. Experiments to decide with data.

quality cost latency failures

Visual Studio

Design with your team in real time. Connect agents and run their branches in parallel.

Python SDK

Build and version the same graph from code. Canvas and SDK are two views of one object.

Experiments

Run datasets through graph versions with grid, random or TPE search and keep every trial trace.

Human evaluation

Turn outputs into typed labeling jobs, collect multiple judgments and measure agreement.

Agent runtime

Tools, skills, memory, panels and orchestrators run in a sandbox with explicit capabilities.

Architectures to explore

Open a system. See how it thinks.

From a research agent to a debating panel. Explore six architectures with their connections, decisions, and tools in view.

01 / ReAct · tool use

Reason, act, observe — and stop before the irreversible bit

An agent investigates, uses tools, and reviews what it finds. Follow each turn in the trace and add human approval before it acts on the world.

Loading diagram…
Explore step by step

Conversation

The last twenty messages, so the agent inherits the thread instead of restarting cold.

Researcher

The ReAct loop itself. Reason → tool → observe, capped at the agent's iteration budget, with its tools resolved from the workspace's MCP servers and sandboxes.

Needs sign-off?

An LLM judge reads the proposed action and routes anything that writes, sends or spends down the human arm. It answers in one word, because the arm is picked by matching that answer against the branch.

Approval

The run pauses here. It survives a page reload, and resumes with whatever the reviewer typed.

Answer

Both arms land on the same output, so the caller never has to know which one ran.

Reach for it when ·The task needs live information or side effects, and some of those side effects are ones you would rather approve by hand.

02 / Agentic RAG

Decide what to look for, then answer only from what came back

From question to sources: a model prepares the search, the retriever finds the passages, and the system decides whether there is enough evidence to answer.

Loading diagram…
Explore step by step

Query rewrite

Turns the question into a retrieval query: expands acronyms, strips pleasantries, keeps the entities.

Retrieve

Semantic search over the workspace’s own context items — five passages, nothing below 0.35 similarity.

Evidence enough?

A regex on the retriever’s output. Empty result set, no answer — the branch is explicit rather than buried in a prompt.

Grounded answer

Answers from the retrieved passages and cites them. Temperature 0.2: this is a reading task, not a writing one.

Abstain

Returns a plain “not in the corpus” instead of a confident guess. Cheaper to read than a hallucination is to catch.

Reach for it when ·You have a corpus that is the source of truth, and a wrong answer costs more than no answer.

03 / Reflection · evaluator–optimizer

One model writes it, another one tries to break it

Separate creation from critique. One model drafts, another finds flaws, and a final pass addresses the objections. Compare the draft, revision, and their costs.

Loading diagram…
Explore step by step

Draft

First attempt, warm. This one is allowed to be wrong.

Critique

A different model, temperature 0, told to list defects and nothing else — or to reply APPROVED if it finds none.

Approved?

A regex over the critique. Clean drafts skip the rewrite entirely instead of paying for a second generation.

Revision

Sees the task, its own draft and every objection, and has to address them point by point.

Final text

Whichever version survived. The trace keeps both, so you can see what the critique bought you.

Reach for it when ·Quality matters more than latency: long-form writing, code, anything where a second pass reliably beats a longer prompt.

04 / Supervisor · hierarchical delegation

A planner that picks the specialist for each step

A supervisor breaks down the goal, delegates to specialists, and adapts the plan with each result. Explore who does what and trace every delegation.

Loading diagram…
Explore step by step

Supervisor

Plans, delegates, and decides when the goal is met — up to six delegations, re-planning after each one.

Researcher

A member agent, enrolled by the dashed ◇ edge. Its own tools, memory and comms policy come from its definition in the library.

Analyst

Same enrolment, different specialisation. The supervisor chooses between them per step, not per run.

Writer

Members never call each other here; everything routes through the supervisor, which is what makes the delegation tree readable in the trace.

Report

The supervisor's synthesis, with the per-member token counts and cost attached to each branch of the span tree.

Reach for it when ·The work decomposes into genuinely different kinds of subtask, and you want one place deciding the order.

05 / Multi-agent debate

A round table that has to reach a verdict

Several agents, different perspectives, and a moderator. The panel debates over a bounded number of rounds and delivers a conclusion that preserves disagreements.

Loading diagram…
Explore step by step

Case conference

The moderator: opens each round, decides whether another one is worth it, and synthesises the conclusion. Four rounds is the ceiling, not the target.

Generalist

Argues from the whole picture. Members speak in turn and see what the others said in the previous round.

Specialist

Narrow and deep. Enrolled the same way — a ◇ membership edge from the panel.

Devil’s advocate

Instructed to attack the emerging consensus. Its dissent is what the schema below preserves.

Verdict

Structured output with a schema: conclusion, confidence, and the dissent that did not make it into the conclusion. Retried against the schema if the model wanders.

Reach for it when ·The question is genuinely contested, and you want the disagreement on the record — diagnosis, review boards, anything where a lone confident answer is a warning sign.

06 / Evaluation run

The same graph, over a dataset, with the numbers kept

Test the same system across a dataset. Score each answer, swap a model, and repeat. Compare quality and cost across both versions.

Loading diagram…
Explore step by step

Dataset

A CSV read in batch, one record per row. TREC, XML, JSON and whole folders read the same way.

Per row

For-each over the rows. Each iteration gets {{loop_item}} and its fields; a sample size lets you rehearse on twenty rows before paying for two thousand.

Answer

The system under test — here a local Ollama model, so a sweep costs GPU time instead of API credit.

Score

Exact match and cosine similarity between the answer and the row’s reference. ROUGE, BLEU, NDCG, MRR, F1 and the correlation metrics are all one checkbox away.

Scores

Accumulates every row’s metrics across iterations instead of overwriting them.

Results

A dashboard tab in the run view. The same numbers land in the experiment’s trials, next to token counts and cost.

Reach for it when ·You are choosing between two prompts, two models or two graph shapes, and you would like to choose on evidence.

Interactive graphs built with the components you’ll use in the Studio. Explore their nodes and configuration.

Born in research

A question worth building a lab for.

The original brief

A small demonstration for an accepted SEPLN paper.

Stopping at prompts and models

How do we know whether one agent system is better than another? That question turned a paper demo into Chatty Lab.

I built it at CiTIUS to study complete systems: how they are organised, where they fail, what they cost, and how they improve. A place to build an idea and put it to the test.

Manuel Couto Pintos

Researcher and builder · CiTIUS, Universidade de Santiago de Compostela

Follow the build
CiTIUS — Centro Singular de Investigación en Tecnoloxías Intelixentes

Give your next idea a system.

Connect your first agents. Ask a good question. See how far they go.