NewDakera v0.11.107 — temporal precision re-ranking, session-cohesion scoring, and reliability fixesSee what's new →

How We Benchmark Memory: Dakera on LoCoMo

Our methodology for measuring long-context memory with LoCoMo, how we reached 88.2% Recall@20, and exactly what that number measures.

How We Benchmark Memory: Dakera on LoCoMo

When we claim Dakera scores 88.2% Recall@20 on LoCoMo, that number deserves an explanation. What is LoCoMo? What does that score actually measure for an AI agent remembering real conversations?

LoCoMo benchmark methodology showing 1,536 evaluated questions across 4 categories from 10 conversations

This post covers all of it. No hand-waving — just the methodology and the data behind the number.

What is LoCoMo?

LoCoMo (Long-Context Memory) is a benchmark for evaluating how well an AI memory system retains and recalls information from long conversational histories. It was introduced by Maharana et al. (2024) — researchers at UNC Chapel Hill, USC and Snap Inc. — and has become the de-facto standard evaluation for dedicated AI memory engines. The dataset and evaluation code are published in the snap-research/locomo repository.

The benchmark works by providing a memory system with a long conversational history — months of simulated interactions between two people — then asking questions that require recalling specific facts from that history. The questions are designed to reflect the kinds of recall failures that break real agents in production:

Cat1 — Single-hop recall
Direct facts from recent or distant sessions. Tests whether the memory system stores and retrieves faithfully.
86.9%
282 questions — Dakera v0.11.107
Cat2 — Multi-hop reasoning
Questions requiring inference across two or more stored memories. Tests whether recall compounds correctly across sessions.
85.4%
321 questions — improved by entity graph traversal
Cat3 — Temporal inference
Questions about sequences, durations, and "what happened before/after". Tests temporal reasoning across stored memories.
73.9%
92 questions — hardest category, actively improving
Cat4 — Open-domain recall
Free-form questions spanning multiple topics and entities. Tests breadth and versatility of the retrieval system.
91.0%
841 questions — best category

Across 1,536 questions spanning all four categories, Dakera v0.11.107 scores 88.2% Recall@20. The metric is LLM-judged retrieval recall: an LLM judge checks whether the ground-truth answer is present in the top-20 memories Dakera retrieves for each question.

Why 1,536 questions? The public LoCoMo release consists of 10 conversations (up to ~35 sessions each). Its QA set is 1,536 evaluated questions across four categories — single-hop, multi-hop, temporal and open-domain — with the adversarial category excluded. Our harness evaluates the 1,536 questions that parse to a scorable ground-truth answer. We run the full set rather than a sampled subset: two percentage points on a 100-question eval is statistically meaningless, while on 1,536 questions it represents a real signal.

Dakera's score: the full picture

Dakera v0.11.107
88.2%
Recall@20 — no LLM in the retrieval path
Questions evaluated
1,536
10 conversations, adversarial category excluded
Evaluated
Sep 2026
Version v0.11.107 — scores tracked per release

We publish our evaluation methodology in full so you can audit these numbers. Scores change with releases — track them on our benchmark page.

What drives the score

Temporal inference at 73.9% is our hardest category and the most active area of development. Questions like "how long did they discuss the apartment renovation?" or "what did they decide after the trip to Japan?" require the memory system to reason about time sequences — not just retrieve stored facts.

Three architectural choices drive Dakera's overall score:

1. Hybrid retrieval

Vector similarity alone fails for queries with specific terms ("that API endpoint with the dry_run flag"). BM25 alone fails for semantic queries ("what was that issue with the server last month?"). Dakera runs both in parallel and combines them with a configurable scoring function:

score = α × vector_similarity + β × bm25_score + γ × importance Defaults α=0.5, β=0.3, γ=0.2 — the configuration shipped in every binary. Configurable per-query.

2. Temporal re-ranking

After hybrid retrieval returns candidates, Dakera applies a non-LLM temporal post-reranker that scales scores multiplicatively based on memory age and recency signals in the query. A question about "the recent camping trip" down-weights memories from years ago, even if they're semantically similar to the query. The temporal re-ranker added +2.2pp on Cat3; subsequent improvements including session-date consensus year injection added another +1.0pp, and stronger temporal proximity scoring added another +2.2pp — together bringing Cat3 from ~67% to the current 73.9%.

3. Importance-weighted recall

Not all memories are equal. A memory stored with importance=0.95 (a critical configuration, a key decision) outranks a 0.4-importance observation even when both match the query semantically. The importance score compounds with decay — memories that prove useful over time stay sharp; memories that are never recalled fade away.

Together these three systems explain why Dakera outperforms pure vector-database approaches. Vector DBs handle single-hop recall well but struggle with temporal inference and importance weighting. Dakera's retrieval pipeline is purpose-built for agent memory workloads, not generic similarity search.

Our evaluation setup

Every Dakera release that touches retrieval runs the full LoCoMo evaluation against a live server instance. The pipeline is:

  1. Spin up a fresh Dakera instance with default configuration
  2. Ingest all 10 LoCoMo conversations into the memory store
  3. Run the 1,536 evaluated questions against the memory API
  4. For each question, an LLM judge (GPT-4o) checks whether the ground-truth answer is present in the top-20 retrieved memories
  5. Report overall Recall@20 and the per-category breakdown

We benchmark the same default configuration we ship — there is no separate benchmark-only profile. The scores you see are what you'd get if you dropped the Dakera binary on a server and ran the benchmark against it.

Temporal inference: the hard problem

A 73.9% score on temporal inference deserves more explanation, because it's the category most relevant to production agent workloads.

Temporal questions ask things like: "Who was mentioned more recently — Alex or Sam?" or "How many months passed between the camping trip and the house purchase?" These require the memory system to reason about time sequences across dozens of stored memories — not just retrieve one fact, but synthesize a temporal narrative from multiple fragments.

This is genuinely hard. Temporal questions often depend on memories that are not the closest semantic matches — the relevant fact is anchored to a date rather than to the wording of the query. The judge checks whether the memories that pin down the timeline (for example "camping trip: July 2023" and "house purchase: November 2023") are actually surfaced in the top-20, not just whether something topically similar was retrieved.

Our temporal re-ranker and session-date consensus year injection and stronger temporal proximity scoring (`INFERENCE_TEMPORAL_MULT_BETA` 0.5→0.65) have brought this from ~67% to 73.9% — multiplicatively boosting temporally-relevant memories and anchoring year inference to session metadata. Temporal inference remains our most actively developed category.

Why publish the hard category? Because developers evaluating memory engines deserve real numbers. A vendor that only reports their best categories isn't giving you useful information. We report all four categories in every release. Temporal inference at 73.9% is where we are today — and it's better than where we were last quarter.

What 88.2% means in practice

Benchmark scores are proxies. What matters for production is whether your agents recall the right context at the right time. A few things the score tells you — and doesn't tell you:

Question What the score tells you
Will my agent remember facts accurately? Yes — 88.2% on a standardized long-context eval is strong coverage of factual recall
Will it handle temporal reasoning? Partially — 73.9% on temporal questions; complex timeline reasoning is harder
Will it work for your specific domain? Unknown — LoCoMo uses general conversational data. Run domain-specific evals if precision matters
Is 88.2% better than rolling your own? Yes — pure vector DB approaches typically score 60–70% on LoCoMo; hybrid retrieval matters

The benchmark is a starting point, not an end state. We publish it with every release precisely so you can watch the trajectory — not just trust a static number.

For the full category breakdown and evaluation details, see the dedicated benchmark page.

Run Dakera against your own workload

Dakera is live in public alpha. Drop a single binary, point your agents at it, and measure recall on your own data.

Stay sharp on agent memory
Benchmark releases, engineering deep-dives, and product updates. Once a week max, no fluff.
✓ Subscribed. Thanks!