Quickstart
Five minutes from pip install to your first working memory layer. No API keys required if you use the local defaults.
1. Install
The local extra pulls FAISS, the Ollama client, and sentence-transformers (the quotes keep zsh from expanding the brackets). You need Ollama installed and running; llama3.1:8b is the default model and weighs 4.9 GB. The first run also downloads the all-MiniLM-L6-v2 embedding model (about 90 MB). After that, set HF_HUB_OFFLINE=1 and widemem runs fully offline. Python 3.10+.
What local costs: sentence-transformers brings PyTorch (about 550 MB on macOS; more on Linux with CUDA wheels), and each add() makes up to two LLM calls, extract then resolve. On an Apple M4 with 32 GB, an add() took 5 to 17 seconds and a search 0.04 seconds.
For the Claude Code MCP server and skill integration:
The MCP server defaults to Ollama too. Since v2.0.0 it needs mcp 2.x. If another tool pins you to mcp 1.x, pin widemem-ai<2 instead.
For LangChain, the langchain extra adds WidememRetriever, a standard BaseRetriever that drops into any chain:
2. First add and search
One import, no YAML, no setup wizard, no cloud account, no API key.
WideMemory() boots the default local pipeline: Ollama llama3.1:8b for fact extraction and conflict resolution, sentence-transformers all-MiniLM-L6-v2 (384 dimensions) for embeddings, FAISS for the vector index, SQLite for history. Override any of these via MemoryConfig.
Without a path, FAISS keeps vectors in RAM and they are gone when the process exits; the SQLite history persists either way. To keep them, give FAISS a directory:
Cloud providers, if you want them
Nothing leaves your machine unless you configure a cloud provider. OpenAI and Anthropic are optional extras, and each picks its own default model when you leave model unset.
Set OPENAI_API_KEY or ANTHROPIC_API_KEY. Mixing is fine: a cloud LLM with local embeddings keeps every vector on your machine.
The defaults moved from OpenAI to local in 2.0. A 1536-dimension FAISS index or Qdrant collection built with OpenAI embeddings will not open under the 384-dimension default embedder. Set provider="openai" on both LLMConfig and EmbeddingConfig to keep the old behavior.
3. Conflict resolution is automatic
Add a contradicting fact and the resolver handles it in a single LLM call. No manual dedup logic, no stale entries.
4. Pin what must not be forgotten
For facts that should never decay (allergies, API keys, critical safety constraints), use pin():
Pinned memories start at high importance and resist the decay function. See the YMYL section for the automatic classification that also pins health, financial, and legal facts.
5. Confidence-aware retrieval
Every search() call returns a confidence level so your agent knows whether it has an answer or is guessing.
Four levels: high, moderate, low, none. Pass the confidence and a mode to build_uncertainty_guidance() to decide what to say when confidence is low: strict (refuse), helpful (hedge with context), or creative (offer to guess).
6. Context manager for cleanup
Use a with block to guarantee SQLite connections close cleanly. Matters in long-running services and tests.
Next
Configure providers, retrieval modes, and the scoring function: /docs/configuration.
Deploy to production: /docs/self-hosting.
See how widemem scores against Mem0, Zep, and LangMem on the LoCoMo benchmark: /benchmarks.