When we claim Dakera scores 88.2% Recall@20 on LoCoMo, that number deserves an explanation. What is LoCoMo? What does that score actually measure for an AI agent remembering real conversations?
This post covers all of it. No hand-waving — just the methodology and the data behind the number.
What is LoCoMo?
LoCoMo (Long-Context Memory) is a benchmark for evaluating how well an AI memory system retains and recalls information from long conversational histories. It was introduced by Maharana et al. (2024) — researchers at UNC Chapel Hill, USC and Snap Inc. — and has become the de-facto standard evaluation for dedicated AI memory engines. The dataset and evaluation code are published in the snap-research/locomo repository.
The benchmark works by providing a memory system with a long conversational history — months of simulated interactions between two people — then asking questions that require recalling specific facts from that history. The questions are designed to reflect the kinds of recall failures that break real agents in production:
Across 1,536 questions spanning all four categories, Dakera v0.11.107 scores 88.2% Recall@20. The metric is LLM-judged retrieval recall: an LLM judge checks whether the ground-truth answer is present in the top-20 memories Dakera retrieves for each question.
Why 1,536 questions? The public LoCoMo release consists of 10 conversations (up to ~35 sessions each). Its QA set is 1,536 evaluated questions across four categories — single-hop, multi-hop, temporal and open-domain — with the adversarial category excluded. Our harness evaluates the 1,536 questions that parse to a scorable ground-truth answer. We run the full set rather than a sampled subset: two percentage points on a 100-question eval is statistically meaningless, while on 1,536 questions it represents a real signal.
Dakera's score: the full picture
We publish our evaluation methodology in full so you can audit these numbers. Scores change with releases — track them on our benchmark page.
What drives the score
Temporal inference at 73.9% is our hardest category and the most active area of development. Questions like "how long did they discuss the apartment renovation?" or "what did they decide after the trip to Japan?" require the memory system to reason about time sequences — not just retrieve stored facts.
Three architectural choices drive Dakera's overall score:
1. Hybrid retrieval
Vector similarity alone fails for queries with specific terms ("that API endpoint with the dry_run flag"). BM25 alone fails for semantic queries ("what was that issue with the server last month?"). Dakera runs both in parallel and combines them with a configurable scoring function:
2. Temporal re-ranking
After hybrid retrieval returns candidates, Dakera applies a non-LLM temporal post-reranker that scales scores multiplicatively based on memory age and recency signals in the query. A question about "the recent camping trip" down-weights memories from years ago, even if they're semantically similar to the query. The temporal re-ranker added +2.2pp on Cat3; subsequent improvements including session-date consensus year injection added another +1.0pp, and stronger temporal proximity scoring added another +2.2pp — together bringing Cat3 from ~67% to the current 73.9%.
3. Importance-weighted recall
Not all memories are equal. A memory stored with importance=0.95 (a critical configuration, a key decision) outranks a 0.4-importance observation even when both match the query semantically. The importance score compounds with decay — memories that prove useful over time stay sharp; memories that are never recalled fade away.
Together these three systems explain why Dakera outperforms pure vector-database approaches. Vector DBs handle single-hop recall well but struggle with temporal inference and importance weighting. Dakera's retrieval pipeline is purpose-built for agent memory workloads, not generic similarity search.
Our evaluation setup
Every Dakera release that touches retrieval runs the full LoCoMo evaluation against a live server instance. The pipeline is:
- Spin up a fresh Dakera instance with default configuration
- Ingest all 10 LoCoMo conversations into the memory store
- Run the 1,536 evaluated questions against the memory API
- For each question, an LLM judge (GPT-4o) checks whether the ground-truth answer is present in the top-20 retrieved memories
- Report overall Recall@20 and the per-category breakdown
We benchmark the same default configuration we ship — there is no separate benchmark-only profile. The scores you see are what you'd get if you dropped the Dakera binary on a server and ran the benchmark against it.
Temporal inference: the hard problem
A 73.9% score on temporal inference deserves more explanation, because it's the category most relevant to production agent workloads.
Temporal questions ask things like: "Who was mentioned more recently — Alex or Sam?" or "How many months passed between the camping trip and the house purchase?" These require the memory system to reason about time sequences across dozens of stored memories — not just retrieve one fact, but synthesize a temporal narrative from multiple fragments.
This is genuinely hard. Temporal questions often depend on memories that are not the closest semantic matches — the relevant fact is anchored to a date rather than to the wording of the query. The judge checks whether the memories that pin down the timeline (for example "camping trip: July 2023" and "house purchase: November 2023") are actually surfaced in the top-20, not just whether something topically similar was retrieved.
Our temporal re-ranker and session-date consensus year injection and stronger temporal proximity scoring (`INFERENCE_TEMPORAL_MULT_BETA` 0.5→0.65) have brought this from ~67% to 73.9% — multiplicatively boosting temporally-relevant memories and anchoring year inference to session metadata. Temporal inference remains our most actively developed category.
Why publish the hard category? Because developers evaluating memory engines deserve real numbers. A vendor that only reports their best categories isn't giving you useful information. We report all four categories in every release. Temporal inference at 73.9% is where we are today — and it's better than where we were last quarter.
What 88.2% means in practice
Benchmark scores are proxies. What matters for production is whether your agents recall the right context at the right time. A few things the score tells you — and doesn't tell you:
| Question | What the score tells you |
|---|---|
| Will my agent remember facts accurately? | Yes — 88.2% on a standardized long-context eval is strong coverage of factual recall |
| Will it handle temporal reasoning? | Partially — 73.9% on temporal questions; complex timeline reasoning is harder |
| Will it work for your specific domain? | Unknown — LoCoMo uses general conversational data. Run domain-specific evals if precision matters |
| Is 88.2% better than rolling your own? | Yes — pure vector DB approaches typically score 60–70% on LoCoMo; hybrid retrieval matters |
The benchmark is a starting point, not an end state. We publish it with every release precisely so you can watch the trajectory — not just trust a static number.
For the full category breakdown and evaluation details, see the dedicated benchmark page.
Run Dakera against your own workload
Dakera is live in public alpha. Drop a single binary, point your agents at it, and measure recall on your own data.
