Your AI Should Know When It Doesn't Know
A user asks your agent: "What's my doctor's name?" The agent has never been told this. But the vector store returns something with a 0.38 similarity score, maybe a mention of a hospital visit, maybe a prescription refill. The agent confidently says "Dr. Martinez." It made that up. The model did not invent the name from nothing. The memory layer handed it a weak match with no warning attached.
The memory system returned low-quality results and the agent treated them as facts. The system had no way to say "I don't have good information about this."
I hit this repeatedly while building widemem. The retrieval was working. The scoring was working. But the system was answering questions it had no business answering, because every query got results and every result looked the same. v1.4.0 added confidence scoring, uncertainty modes, and retrieval depth controls. Here is what changed and why.
Update, Oct 4, 2026: an earlier version of this post described an API that was never wired into search: a WideMemory(uncertainty_mode=...) argument that filtered results per mode, a pin(memory_id=...) call that removed decay, and frustration handling inside add(). None of that exists. The examples below are rewritten against widemem 2.0, and every code output below comes from a real run on widemem-ai 2.0.1 with the default local stack.
THE PROBLEM WITH ALWAYS ANSWERING
Most vector stores return the top K results for any query, regardless of quality. Ask for the user's blood type and you get back their favorite color, because that was the closest vector in the space. The similarity score might be 0.1, which is noise, but the system returns it anyway.
The downstream LLM sees retrieved context and assumes it is relevant. It weaves that context into its response. The user gets an answer that looks authoritative but is built on a foundation of "closest thing I had, which was not close at all."
This is worse than returning nothing. A blank response tells the user the system does not know. A confident wrong response tells them the system knows and is lying. The second failure mode erodes trust much faster.
I tracked this in testing with a synthetic user who mentioned living in San Francisco, working as a data engineer, and being allergic to peanuts. When I asked about their spouse's name, something never mentioned, the system still pulled back the peanut allergy fact, and the agent worked it into a response. The store returned something, and the agent had no way to tell it was not an answer.
CONFIDENCE LEVELS: HIGH, MODERATE, LOW, NONE
By default, search() returns a SearchResult. It behaves like a list of results, and it also carries one confidence level for the whole retrieval. Not the raw number. A human-readable verdict:
| Level | Top similarity (MiniLM) | Meaning |
|---|---|---|
| HIGH | ≥ 0.60 | Strong match. Use with confidence. |
| MODERATE | 0.30 to 0.59 | Relevant but may be incomplete or tangential. |
| LOW | 0.20 to 0.29 | Weak match. Might be useful, might be noise. |
| NONE | < 0.20 | No meaningful match found. |
Those bands are the 2.0.1 defaults for the default embedder, all-MiniLM-L6-v2. Other embedders, including OpenAI's text-embedding-3-small, keep 0.60 / 0.50 / 0.30. The WIDEMEM_CONFIDENCE_HIGH, WIDEMEM_CONFIDENCE_MODERATE, and WIDEMEM_CONFIDENCE_LOW environment variables override each level.
The level comes from the top result's raw cosine similarity, not from the composite score. That matters. The composite score blends similarity with importance and recency, so a fresh, important, irrelevant fact can still score well. Similarity is the only signal that says whether the memory is about the question at all.
The search still hands back three facts, and a composite score of 0.47 looks respectable. The confidence says what the number hides: none of this is about a doctor. The agent can respond with "I don't have information about your doctor" instead of hallucinating a name from weak context.
THREE UNCERTAINTY MODES
Confidence levels tell the agent how good the results are. Uncertainty modes tell the application what to do about it. search() never filters by mode; it always returns the ranked results plus the confidence. You pick the behavior by passing the confidence and a mode to build_uncertainty_guidance. It returns None on HIGH (answer normally) and otherwise a dict with an action (answer, hedge, refuse, or offer_guess), a message, and sometimes a short related list.
MemoryConfig has an uncertainty_mode field, but nothing reads it. Pass the mode explicitly where you build the response.
Strict mode
Refuse on NONE and LOW. Hedge on MODERATE. Only HIGH gets a plain answer. If there is no strong match, the agent says so and stops.
Strict mode is the right choice for YMYL domains. If you are building a health assistant, returning nothing is better than returning something you are not sure about. The user can rephrase or provide more context. The system does not guess.
Helpful mode
Refuse on NONE. On LOW, hedge and hand back up to three related memories. On MODERATE, answer with a caveat. This is the reasonable default for general-purpose agents and productivity tools.
The stored fact is "live in San Francisco", with the subject stripped, and it reads MODERATE rather than HIGH. Helpful mode answers with a caveat. Only LOW guidance carries a related list, hence .get().
Creative mode
Never refuse outright. On NONE and LOW, offer to guess from what is there. On MODERATE, answer. Useful when you want the system to make connections that might not be obvious, like a recommendation engine or a creative writing tool.
The mode name is "creative" but the real use case is exploration. Sometimes you want the system to surface weak signals because a human or a capable LLM can reason about them. The key is that the confidence is always there. The agent is never tricked into thinking a 0.3 match is a 0.8 match.
MEM.PIN(): MEMORIES THAT STAY ON TOP
Temporal decay is useful for most facts. Your lunch order from three months ago should fade. But some things should stay at the top of every relevant search. pin() takes the text itself, runs it through the normal extraction pipeline, and raises each resulting memory to an importance of at least 9 (out of 10).
Pinning is an importance boost, not a decay exemption. Recency still counts in the score. What actually switches decay off is YMYL. Enable it with YMYLConfig(enabled=True); decay immunity is on by default after that. A fact strongly classified as health, financial, legal, or another YMYL category then gets full recency no matter its age. For an allergy, you want both: the YMYL flag keeps it from fading, and the pin keeps it ranked above routine facts. There is no unpin call and no pinned flag in search results; a pinned memory is an ordinary memory with high importance.
Use pinning sparingly. The point of importance scoring is that most facts are not critical. If you pin everything, you have rebuilt the flat vector store that created the problems in the first place. Pin the things that would cause real harm if forgotten: allergies, critical preferences, identity facts the user has explicitly confirmed.
FRUSTRATION DETECTION
There is a specific interaction pattern that signals a memory failure: the user says something like "I already told you this" or "we talked about this last week." This means the system had the information and failed to retrieve it, or had it and let it decay when it should not have.
widemem does not watch for this on its own. add() stores the message like any other. The detection lives in build_frustration_response, which your agent calls on the incoming message along with the confidence of a search for it:
The full phrase list lives in widemem/retrieval/uncertainty.py. The function returns None when there is no frustration signal. When there is one, it picks between three actions. If the search found something at HIGH, it says reassure: the memory is there, so look again. MODERATE is not enough. Here "moved to Boston" had not been stored yet, and the San Francisco fact matched at MODERATE: related to the user, but not the fact they are restating. Reassuring on that would drop it. If not HIGH, and it can pull the fact out of the sentence, it says recover_and_pin and hands you the fact to pin at importance 9. Otherwise it says apologize_and_ask.
Treat the complaint as a bug report: the user is telling you, in real time, that curation dropped a fact it should have kept.
RETRIEVAL MODES: FAST, BALANCED, DEEP
Not every query needs the same level of effort. Asking "what's their name?" should be cheap. Asking "summarize everything we know about this user" justifies more computation. search() takes a mode argument, and each mode is a preset:
Fast returns up to 10 results and skips hierarchy routing. Balanced, the default, returns up to 25. Deep returns up to 50. For short factual questions, deeper modes also pull a wider candidate pool from the vector store (3x, 4x, or 5x the result count) and give raw similarity a bigger boost in ranking. Balanced and deep route the query across the fact, summary, and theme tiers when you have built them. Every mode gets the same confidence verdict.
The two pinned facts lead. Note what else is in there: "live in San Francisco" survived next to "moved to Boston." In this run, with the default local model (llama3.1:8b), the conflict resolver did not connect the two. Confidence scoring does not fix contradictions; it only tells you how well the store matched the question.
A practical pattern: use fast mode for inline autocomplete and real-time suggestions. Use balanced for normal conversation turns. Use deep for explicit "tell me everything" queries or when the agent detects it needs more context to answer well.
PUTTING IT TOGETHER
Here is the intended flow when the agent wires these pieces together, with YMYL enabled (illustrative, not a recorded run):
Week 1
User mentions they live in San Francisco and work at a startup. Agent stores both facts.
Week 3
User asks "do you remember my wife's name?" If the search comes back at NONE, helpful-mode guidance says refuse, and the agent replies: "I don't have that information. Could you tell me?" Under MiniLM, other facts about the user can still lift it to MODERATE; see the last section.
Without the check: the agent would have used the closest match and hallucinated a name.
Week 6
User mentions they moved to Boston. The agent stores it, and checks whether the conflict resolver retired the San Francisco fact. In the run above it had not, so the agent flags the conflict to the user. User also says "I'm severely allergic to shellfish." YMYL flagged. Agent pins it.
Week 10
User says "I told you I moved to Boston!" The agent runs the message through build_frustration_response. The Boston fact was stored in week 6, so the search matches it strongly and the guidance is reassure: the miss was in the agent's cached context, not in the store. The agent re-reads and answers from Boston. Below HIGH, recover_and_pin is for a fact that was never stored, as in the run above.
Week 14
Application runs a deep retrieval: "summarize this user." The shellfish allergy leads (pinned, and decay-immune because YMYL is on). Boston is next. The startup fact has slid down as its recency decayed, still there but no longer near the top.
WHAT THIS DOES NOT SOLVE
Confidence calibration. The original thresholds (0.60 for HIGH, 0.50 for MODERATE, 0.30 for LOW, defined in widemem/retrieval/uncertainty.py) were set for OpenAI's text-embedding-3-small. Under all-MiniLM-L6-v2, the default since 2.0, short extracted facts landed in LOW even when they answered the question. 2.0.1 recalibrates for MiniLM (HIGH 0.60, MODERATE 0.30, LOW 0.20), and most short correct facts now read MODERATE: recall at MODERATE or above on subject-stripped answer facts was 77% on the training queries and 58% on the holdout. 2.0.1 ships measured thresholds for MiniLM only; other models fall back to the OpenAI-calibrated values. Per-model auto-calibration does not exist yet.
"The answer is stored" versus "something about this person is stored." Under MiniLM, confidence cannot tell these apart, because the name or subject in the query dominates similarity. With only "I live in San Francisco" and "I work at a startup" stored, "where does my sister live?" matched the San Francisco fact at 0.33 on 2.0.1. That reads MODERATE, and helpful mode answers with a caveat instead of refusing. Separating the two is follow-up research. Until then, strict mode is the safer choice when a wrong answer costs something.
Confidence is retrieval strength. It measures how close the best memory is to the question. It is not a probability that the final answer is right. A HIGH match can still feed a wrong answer.
Frustration detection is pattern-based. It catches explicit signals like "I told you this" but misses subtle frustration. A user who quietly rephrases the same question three times is frustrated too. Detecting that requires tracking query patterns across a session, which is a different problem.
Mode selection is manual. The application chooses fast/balanced/deep, and it applies the uncertainty mode itself. Ideally the system would auto-detect retrieval depth based on query complexity. A simple name lookup should automatically use fast mode. A broad "tell me about this user" should automatically use deep. This is on the roadmap but not shipped yet.
THE HONEST MEMORY THESIS
The core argument is simple: a memory system that says "I don't know" is more useful than one that always answers. False confidence destroys the trust that makes memory useful in the first place.
If your agent confidently tells a user they live in San Francisco when they moved to Boston three months ago, the user will stop trusting the agent's memory entirely. They will start re-stating facts every conversation. At that point, the memory system costs more than it gives: it still retrieves and injects stale context that the LLM has to override.
The deep run above did exactly this. Confidence scoring would not have caught it, which is why contradictions need their own fix.
Confidence scoring, uncertainty modes, and retrieval depth are all ways of saying the same thing: be honest about what you know and how well you know it. The user can handle uncertainty. They cannot handle confidently wrong.
widemem is open source (Apache 2.0) on GitHub and PyPI. Confidence scoring and the guidance helpers shipped in v1.4.0; the examples above run on 2.0.1. If you have thoughts on confidence calibration or automatic mode selection, I would like to hear about it.
READ RELATED
I BUILT A MEMORY LAYER THAT FORGETS ONLY WHAT DOESN'T MATTER
Why forgetting is harder than remembering, and how batch conflict resolution, importance decay, and YMYL safety work under the hood.
YOUR AI FORGOT SOMEONE'S MEDICATION. NOW WHAT?
YMYL safety for AI memory: why some facts should never decay, and the edge cases that still need solving.
THE CONTRADICTION PROBLEM IN AI MEMORY
What happens when AI agents accumulate conflicting facts, and why vector similarity alone cannot detect it.