widemem
ProductDocsSecurityPricingBlog
Open sourceenterprise-readyTalk to us
← Back to home
Docs

Quickstart

OverviewQuickstartConfigurationSelf-hosting

Five minutes from pip install to your first working memory layer. No API keys required if you use the local defaults.

1. Install

$bash
pip install "widemem-ai[local]"
ollama pull llama3.1:8b

The local extra pulls FAISS, the Ollama client, and sentence-transformers (the quotes keep zsh from expanding the brackets). You need Ollama installed and running; llama3.1:8b is the default model and weighs 4.9 GB. The first run also downloads the all-MiniLM-L6-v2 embedding model (about 90 MB). After that, set HF_HUB_OFFLINE=1 and widemem runs fully offline. Python 3.10+.

What local costs: sentence-transformers brings PyTorch (about 550 MB on macOS; more on Linux with CUDA wheels), and each add() makes up to two LLM calls, extract then resolve. On an Apple M4 with 32 GB, an add() took 5 to 17 seconds and a search 0.04 seconds.

For the Claude Code MCP server and skill integration:

$bash
pip install "widemem-ai[mcp,local]"
python -m widemem.mcp_server

The MCP server defaults to Ollama too. Since v2.0.0 it needs mcp 2.x. If another tool pins you to mcp 1.x, pin widemem-ai<2 instead.

For LangChain, the langchain extra adds WidememRetriever, a standard BaseRetriever that drops into any chain:

$python
# pip install "widemem-ai[langchain,local]"
from widemem import WideMemory
from widemem.integrations.langchain import WidememRetriever

retriever = WidememRetriever(memory=WideMemory(), user_id="alice", top_k=5)
docs = retriever.invoke("where does alice live")

2. First add and search

One import, no YAML, no setup wizard, no cloud account, no API key.

$python
from widemem import WideMemory

memory = WideMemory()

# Store a fact
memory.add("I live in San Francisco and work as a software engineer", user_id="alice")

# Search for it
results = memory.search("where does alice live", user_id="alice")
for r in results:
    print(f"{r.memory.content} (score: {r.final_score:.2f})")

WideMemory() boots the default local pipeline: Ollama llama3.1:8b for fact extraction and conflict resolution, sentence-transformers all-MiniLM-L6-v2 (384 dimensions) for embeddings, FAISS for the vector index, SQLite for history. Override any of these via MemoryConfig.

Keep the vectors

Without a path, FAISS keeps vectors in RAM and they are gone when the process exits; the SQLite history persists either way. To keep them, give FAISS a directory:

$python
from widemem import WideMemory, MemoryConfig
from widemem.core.types import VectorStoreConfig

memory = WideMemory(MemoryConfig(
    vector_store=VectorStoreConfig(provider="faiss", path="./widemem_data"),
))

Cloud providers, if you want them

Nothing leaves your machine unless you configure a cloud provider. OpenAI and Anthropic are optional extras, and each picks its own default model when you leave model unset.

$bash
pip install "widemem-ai[openai,faiss]"   # or [anthropic,local] for Claude + local embeddings
$python
from widemem import WideMemory, MemoryConfig
from widemem.core.types import EmbeddingConfig, LLMConfig

# OpenAI for both: gpt-4o-mini and text-embedding-3-small
memory = WideMemory(MemoryConfig(
    llm=LLMConfig(provider="openai"),
    embedding=EmbeddingConfig(provider="openai"),
))

# Claude for extraction; embeddings and vectors stay local
memory = WideMemory(MemoryConfig(llm=LLMConfig(provider="anthropic")))

Set OPENAI_API_KEY or ANTHROPIC_API_KEY. Mixing is fine: a cloud LLM with local embeddings keeps every vector on your machine.

Upgrading from 1.x

The defaults moved from OpenAI to local in 2.0. A 1536-dimension FAISS index or Qdrant collection built with OpenAI embeddings will not open under the 384-dimension default embedder. Set provider="openai" on both LLMConfig and EmbeddingConfig to keep the old behavior.

3. Conflict resolution is automatic

Add a contradicting fact and the resolver handles it in a single LLM call. No manual dedup logic, no stale entries.

$python
memory.add("I just moved to Boston", user_id="alice")

# The San Francisco memory gets updated or replaced,
# not silently appended alongside the new one.
results = memory.search("where does alice live", user_id="alice")

4. Pin what must not be forgotten

For facts that should never decay (allergies, API keys, critical safety constraints), use pin():

$python
memory.pin(
    "Allergic to penicillin",
    user_id="alice",
    importance=9.0,
)

Pinned memories start at high importance and resist the decay function. See the YMYL section for the automatic classification that also pins health, financial, and legal facts.

5. Confidence-aware retrieval

Every search() call returns a confidence level so your agent knows whether it has an answer or is guessing.

$python
results = memory.search("what is alice's blood type", user_id="alice")

if results.confidence == "none":
    reply = "I don't have that stored."
elif results.confidence == "high":
    reply = f"Alice's blood type is {results[0].memory.content}"
else:
    reply = "I think so, but I'm not 100% sure."

Four levels: high, moderate, low, none. Pass the confidence and a mode to build_uncertainty_guidance() to decide what to say when confidence is low: strict (refuse), helpful (hedge with context), or creative (offer to guess).

6. Context manager for cleanup

Use a with block to guarantee SQLite connections close cleanly. Matters in long-running services and tests.

$python
with WideMemory() as memory:
    memory.add("I live in San Francisco", user_id="alice")
    results = memory.search("where does alice live", user_id="alice")
# Connection closed automatically.

Next

Configure providers, retrieval modes, and the scoring function: /docs/configuration.

Deploy to production: /docs/self-hosting.

See how widemem scores against Mem0, Zep, and LangMem on the LoCoMo benchmark: /benchmarks.

← PreviousOverviewNext →Configuration
widemem

The open-source memory layer for LLM agents. Local-first, importance-scored, auditable.

Apache-2.0 · v2.0.0 · Python 3.10+

Product

FeaturesHow it worksDeployInstall

Company

AboutTalk to usPricingCareers

Resources

DocsBenchmarksSecurityGitHub
No cookies. Privacy is the default, not a setting.© 2026 widemem · Terms