Ollama & embeddings
The Ollama & embeddings section of Settings does two separate things that share a screen: it connects a local Ollama backend, and it picks the embedder your workspace uses. Embeddings turn text into vectors, which is what every semantic feature here runs on.
The two are independent. An embedder does not have to be an Ollama — it is a
provider:model pair against any provider you have configured, and Ollama is one of the
options rather than the assumption.
Local Ollama backend
Section titled “Local Ollama backend”Point Chatty at a running Ollama instance to use local models. Once configured, Ollama serves both chat/completion models and the embedding model.
The embedder
Section titled “The embedder”Each workspace has two embedder settings, resolved separately:
| Setting | What it embeds |
|---|---|
| The system embedder | Context items, prompt search, the documentation index |
| The Vault embedder | The research corpus |
Each is a provider:model pair — novita:baai/bge-m3, ollama:nomic-embed-text,
nvidia:nvidia/nv-embedqa-e5-v5 — so the provider is part of the choice. Setting one and
leaving the other empty is a real state and a common surprise: the docs index works and
the Vault does not embed, in the same workspace on the same day.
-
Open Settings and go to Ollama & embeddings.
-
Pick an embedder for each of the two. If you are using an Ollama one, the model has to be pulled on that instance; for a hosted provider, the workspace needs its key.
-
If you changed a model, use Re-vectorize now so existing content is re-embedded.
What embeddings power
Section titled “What embeddings power”The system embedder drives:
- Semantic context — finding relevant context items by meaning for RAG.
- VectorDB Query nodes — similarity search in graphs; see the VectorDB Query node.
- Episodic agent memory — how an agent recalls relevant past turns; see Memory strategies. Each episode is stored with the model that embedded it, so changing the embedder here restarts recall for conversations already under way. (Semantic memory does not use embeddings — it extracts facts with a chat model.)