Skip to content

WebSocket API Reference

Opening a socket takes two steps: ask for a ticket over HTTP, then use the ticket in the WebSocket URL.

POST /api/ws/ticket
Authorization: Bearer {jwt}
X-Workspace-Id: {team uuid} # optional; omit for your personal workspace
{ "ticket": "chy_ws_…", "expires_in": 60 }

A ticket is single-use and lives for expires_in seconds. It opens one socket and nothing else — it is not accepted anywhere in the REST API. The workspace is fixed here, from the authenticated request, not later by the socket.

ws[s]://{host}/ws/chat?session_id={uuid}&ticket={ticket}
ParameterTypeDescription
session_idUUIDThe session to operate on
ticketstringSingle-use ticket from POST /api/ws/ticket
const protocol = window.location.protocol === 'https:' ? 'wss:' : 'ws:';
const host = window.location.host;
const { ticket } = await fetch(`${apiBase}/ws/ticket`, {
method: 'POST',
headers: { Authorization: `Bearer ${accessToken}` }
}).then((r) => r.json());
const wsUrl = `${protocol}//${host}/ws/chat?session_id=${sessionId}&ticket=${ticket}`;

Ask for a new ticket on every attempt, reconnects included. A ticket is single-use, so a retry that reuses the URL is rejected.

CodeMeaning
4401The credential was rejected — expired, already used, or not valid. Do not retry with the same one; ask for a new ticket, or sign in again.

Messages sent from the frontend to the backend.

Send a message to a single chat and stream the response.

{
"type": "query",
"chat_id": "550e8400-e29b-41d4-a716-446655440000",
"content": "Explain quantum computing"
}
FieldTypeDescription
type"query"Message type discriminator
chat_idUUID stringTarget chat
contentstringUser message text

Backend behavior:

  1. Verifies the chat belongs to the authenticated user
  2. Inserts the user message into the messages table
  3. Loads the full conversation history
  4. Calls provider.send_stream() with the history
  5. Sends token messages as tokens arrive
  6. Inserts the assistant message and sends done

Send the same message to multiple chats simultaneously.

{
"type": "query_all",
"content": "Explain quantum computing",
"chat_ids": [
"550e8400-e29b-41d4-a716-446655440000",
"550e8400-e29b-41d4-a716-446655440001",
"550e8400-e29b-41d4-a716-446655440002"
]
}
FieldTypeDescription
type"query_all"Message type discriminator
contentstringUser message text (sent to all chats)
chat_idsUUID string[]Target chats

Backend behavior: Spawns a separate tokio::spawn task per chat_id. Each task executes the same flow as a single query. Responses stream concurrently and are interleaved on the same WebSocket connection, tagged with chat_id.

Messages sent from the backend to the frontend.

A single token (or token chunk) from the LLM provider’s streaming response.

{
"type": "token",
"chat_id": "550e8400-e29b-41d4-a716-446655440000",
"content": "Quantum"
}
FieldTypeDescription
type"token"Message type discriminator
chat_idUUID stringWhich chat this token belongs to
contentstringToken text (may be one or more characters)

Signals that streaming is complete for a chat. Includes the full assembled response.

{
"type": "done",
"chat_id": "550e8400-e29b-41d4-a716-446655440000",
"full_content": "Quantum computing is a type of computation..."
}
FieldTypeDescription
type"done"Message type discriminator
chat_idUUID stringWhich chat completed
full_contentstringComplete assistant response

An error occurred during processing.

{
"type": "error",
"chat_id": "550e8400-e29b-41d4-a716-446655440000",
"message": "Model 'GPT-4' not available"
}
FieldTypeDescription
type"error"Message type discriminator
chat_idUUID string or nullWhich chat the error relates to (null for connection-level errors)
messagestringHuman-readable error description

Common errors:

MessageCause
"Unauthorized"The ticket is unknown, expired or already used — or, on the deprecated path, the JWT is invalid or expired. Followed by close code 4401.
"Chat not found"Chat ID does not exist or is not owned by the user
"Model 'X' not available"Model is not in the plugin registry
"Stream error: ..."LLM provider returned an error during streaming
"Invalid message: ..."Client sent malformed JSON
sequenceDiagram
participant C as Client
participant S as Server
C->>S: Connect ws://host/ws/chat?session_id=...&token=...
Note over S: Validate JWT
S-->>C: Connection established
C->>S: {"type":"query","chat_id":"abc","content":"Hello"}
S-->>C: {"type":"token","chat_id":"abc","content":"Hi"}
S-->>C: {"type":"token","chat_id":"abc","content":" there"}
S-->>C: {"type":"token","chat_id":"abc","content":"!"}
S-->>C: {"type":"done","chat_id":"abc","full_content":"Hi there!"}
sequenceDiagram
participant C as Client
participant S as Server
C->>S: {"type":"query_all","content":"Hello","chat_ids":["a","b"]}
Note over S: Spawn task for chat "a" (GPT-4)
Note over S: Spawn task for chat "b" (Mistral)
S-->>C: {"type":"token","chat_id":"a","content":"Hi"}
S-->>C: {"type":"token","chat_id":"b","content":"Hello"}
S-->>C: {"type":"token","chat_id":"a","content":" there"}
S-->>C: {"type":"token","chat_id":"b","content":" there"}
S-->>C: {"type":"done","chat_id":"b","full_content":"Hello there"}
S-->>C: {"type":"token","chat_id":"a","content":"!"}
S-->>C: {"type":"done","chat_id":"a","full_content":"Hi there!"}

Tokens from different chats are interleaved. The client demultiplexes using the chat_id field.

sequenceDiagram
participant C as Client
participant S as Server
C->>S: {"type":"query","chat_id":"abc","content":"Hello"}
S-->>C: {"type":"token","chat_id":"abc","content":"Hi"}
S-->>C: {"type":"error","chat_id":"abc","message":"Stream error: connection reset"}

When an error occurs mid-stream, the partial response is not saved to the database.

The frontend WebSocket client automatically reconnects after 3 seconds when the connection drops:

  • The connected store is set to false
  • Pending streaming state is preserved in Svelte stores
  • A new connection is established with the same session and token

There is no built-in rate limiting on the WebSocket endpoint. In production, consider implementing rate limiting at the nginx reverse proxy level or in the application.

/ws/assistant is the platform assistant’s channel (RFC-015). It is opened with the same ticket as /ws/chat, without a session_id: the thread is per person and lives in assistant_turn, readable over REST at GET /api/assistant/turns.

wss://{host}/ws/assistant?ticket={ticket}

One client message:

{ "type": "turn", "turn_id": "<uuid minted by the client>", "text": "…", "state": { "route": "/research/…", "research_id": "…" } }

The client mints turn_id; it comes back as the chat_id of the same token, done and error frames a chat streams, so no new frame is needed to introduce it. state is whatever the host application knows about where the person is — the graph reads it as {{state}} (JSON) and {{state.<field>}} for each scalar field, and the last answered turns as {{history}}. A socket runs one turn at a time; the next waits.

Ceilings: 60 turns per person per hour, 16 in flight across everyone, 300 seconds per turn. A refused turn arrives as an error frame with the turn’s chat_id and the reason in words.