Visual Studio
Design with your team in real time. Connect agents and run their branches in parallel.
From first idea to evidence. Design agent systems, inspect every step, and discover which version works best. All in one lab.
Connect models, tools, and agents. On a shared canvas or in Python.
Every step, every output, every cost. Follow the run as it unfolds.
Change a variable. Compare versions. Keep the trail behind every result.
From idea to evidence
Design, run, evaluate. Change what matters and test again. Your graph, results, and history stay connected.
The system you test in a conversation is the one you evaluate across a dataset. Take a promising idea to a conclusion you can stand behind.
Design
connect agents and tools
Run
with a trace for every node
Evaluate
quality, cost, latency, failure
Branch
change one thing, keep the history
Conclude
with evidence you can revisit
Your workspace
A canvas to think. Code to go further. Experiments to decide with data.
Design with your team in real time. Connect agents and run their branches in parallel.
Build and version the same graph from code. Canvas and SDK are two views of one object.
Run datasets through graph versions with grid, random or TPE search and keep every trial trace.
Turn outputs into typed labeling jobs, collect multiple judgments and measure agreement.
Tools, skills, memory, panels and orchestrators run in a sandbox with explicit capabilities.
Architectures to explore
From a research agent to a debating panel. Explore six architectures with their connections, decisions, and tools in view.
01 / ReAct · tool use
An agent investigates, uses tools, and reviews what it finds. Follow each turn in the trace and add human approval before it acts on the world.
Conversation
The last twenty messages, so the agent inherits the thread instead of restarting cold.
Researcher
The ReAct loop itself. Reason → tool → observe, capped at the agent's iteration budget, with its tools resolved from the workspace's MCP servers and sandboxes.
Needs sign-off?
An LLM judge reads the proposed action and routes anything that writes, sends or spends down the human arm. It answers in one word, because the arm is picked by matching that answer against the branch.
Approval
The run pauses here. It survives a page reload, and resumes with whatever the reviewer typed.
Answer
Both arms land on the same output, so the caller never has to know which one ran.
02 / Agentic RAG
From question to sources: a model prepares the search, the retriever finds the passages, and the system decides whether there is enough evidence to answer.
Query rewrite
Turns the question into a retrieval query: expands acronyms, strips pleasantries, keeps the entities.
Retrieve
Semantic search over the workspace’s own context items — five passages, nothing below 0.35 similarity.
Evidence enough?
A regex on the retriever’s output. Empty result set, no answer — the branch is explicit rather than buried in a prompt.
Grounded answer
Answers from the retrieved passages and cites them. Temperature 0.2: this is a reading task, not a writing one.
Abstain
Returns a plain “not in the corpus” instead of a confident guess. Cheaper to read than a hallucination is to catch.
03 / Reflection · evaluator–optimizer
Separate creation from critique. One model drafts, another finds flaws, and a final pass addresses the objections. Compare the draft, revision, and their costs.
Draft
First attempt, warm. This one is allowed to be wrong.
Critique
A different model, temperature 0, told to list defects and nothing else — or to reply APPROVED if it finds none.
Approved?
A regex over the critique. Clean drafts skip the rewrite entirely instead of paying for a second generation.
Revision
Sees the task, its own draft and every objection, and has to address them point by point.
Final text
Whichever version survived. The trace keeps both, so you can see what the critique bought you.
04 / Supervisor · hierarchical delegation
A supervisor breaks down the goal, delegates to specialists, and adapts the plan with each result. Explore who does what and trace every delegation.
Supervisor
Plans, delegates, and decides when the goal is met — up to six delegations, re-planning after each one.
Researcher
A member agent, enrolled by the dashed ◇ edge. Its own tools, memory and comms policy come from its definition in the library.
Analyst
Same enrolment, different specialisation. The supervisor chooses between them per step, not per run.
Writer
Members never call each other here; everything routes through the supervisor, which is what makes the delegation tree readable in the trace.
Report
The supervisor's synthesis, with the per-member token counts and cost attached to each branch of the span tree.
05 / Multi-agent debate
Several agents, different perspectives, and a moderator. The panel debates over a bounded number of rounds and delivers a conclusion that preserves disagreements.
Case conference
The moderator: opens each round, decides whether another one is worth it, and synthesises the conclusion. Four rounds is the ceiling, not the target.
Generalist
Argues from the whole picture. Members speak in turn and see what the others said in the previous round.
Specialist
Narrow and deep. Enrolled the same way — a ◇ membership edge from the panel.
Devil’s advocate
Instructed to attack the emerging consensus. Its dissent is what the schema below preserves.
Verdict
Structured output with a schema: conclusion, confidence, and the dissent that did not make it into the conclusion. Retried against the schema if the model wanders.
06 / Evaluation run
Test the same system across a dataset. Score each answer, swap a model, and repeat. Compare quality and cost across both versions.
Dataset
A CSV read in batch, one record per row. TREC, XML, JSON and whole folders read the same way.
Per row
For-each over the rows. Each iteration gets {{loop_item}} and its fields; a sample size lets you rehearse on twenty rows before paying for two thousand.
Answer
The system under test — here a local Ollama model, so a sweep costs GPU time instead of API credit.
Score
Exact match and cosine similarity between the answer and the row’s reference. ROUGE, BLEU, NDCG, MRR, F1 and the correlation metrics are all one checkbox away.
Scores
Accumulates every row’s metrics across iterations instead of overwriting them.
Results
A dashboard tab in the run view. The same numbers land in the experiment’s trials, next to token counts and cost.
01/06
Interactive graphs built with the components you’ll use in the Studio. Explore their nodes and configuration.
Born in research
The original brief
A small demonstration for an accepted SEPLN paper.
How do we know whether one agent system is better than another? That question turned a paper demo into Chatty Lab.
I built it at CiTIUS to study complete systems: how they are organised, where they fail, what they cost, and how they improve. A place to build an idea and put it to the test.
Manuel Couto Pintos
Researcher and builder · CiTIUS, Universidade de Santiago de Compostela
Connect your first agents. Ask a good question. See how far they go.