Skip to content

Memory strategies

Each agent has a memory strategy that decides what conversation history it sees on each turn. Memory is scoped per (agent, session), so an agent resumes with the right history when it runs again in the same chat or graph run.

StrategyWhat it does
NoneNo memory — every turn starts fresh.
WindowKeep only the last N messages.
Full transcriptStore and replay the entire conversation.
SummarizedKeep a rolling summary, refreshed every N messages by a chosen model.
EpisodicRetrieve relevant past messages from a vector collection (top_k, threshold).
SemanticExtract structured facts (via an extractor prompt) before storing them.
StateCarry a structured state (Σ) with a fixed set of keys instead of the thread. Each turn the model sends a patch; the runtime validates it and throws the reasoning away.
CompositeStack several strategies as ordered layers, consulted and written in order.

State is the one strategy that does not carry a transcript at all. The agent declares a fixed set of keys; what it sees each turn is the spec, the current Σ as JSON, and the last exchange verbatim. Between turns a model folds what happened into a patch — only the keys that changed — which the runtime validates before applying:

  • a key that is not in the schema refuses the whole patch;
  • a key set to null is deleted (that is how the agent forgets);
  • a patch that would push Σ past max_chars is refused.

A refused patch is retried once with the reason, and if it fails again Σ is left as it was and the pending messages are dropped. Losing a fold loses memory; keeping an impossible fold would retry it forever, growing each time. Not being able to reach the model is different — a network that went down comes back, so the queue is kept and folded on the next turn.

{
"kind": "state",
"spec": "Keep what this person is working on and what is still open.",
"fields": [
{ "name": "goal", "description": "what they are trying to achieve" },
{ "name": "working_on", "description": "the ids in play right now" },
{ "name": "open", "description": "what is still unresolved" }
],
"max_chars": 4000
}

model is optional: without it the fold uses memory_extractor_model if registered, else the first provider available — the same fallback Semantic uses.

  • Agent default: every agent definition carries a memory_strategy (default None).
  • Per-slot override: an Agent (SubAgent) node can set a memory_override for that one slot in the graph.

The effective strategy is memory_override.unwrap_or(agent.memory_strategy) — the node-level override wins for that slot, otherwise the agent’s own default applies.

  • Episodic relies on embeddings — configure your embedding model first. It uses the system embedder, whichever provider that is; it is not tied to Ollama. Without one configured, episodes are still stored but nothing is recalled, and the run log says so.
  • Semantic and State need a chat model, not an embedder: Semantic extracts facts with one and stores no vectors at all, and State folds turns into Σ — without a model it keeps only the recent tail and degrades to a window.
  • Folding happens once per run, not once per model round-trip: an agent that takes thirty steps costs one extra call, and the growth of its own loop inside that run is unchanged.
  • Memory writes happen automatically as the agent records each message; you don’t manage storage.
  • The memory inspector shows Σ above whatever has not been folded into it yet.
  • The canonical definition is MemoryStrategy in crates/chatty-domain/src/memory.rs; the behaviour lives in crates/chatty-engine/src/agents/memory/.