NewRelease 0.12.2 is out: Engine v0.12.2, Dashboard 0.5.1, Helm 0.12.2, the four SDKs, CLI 0.8.1 and MCP 0.12.2Release 0.12.2 is outSee what's new →
Comparisons

Mem0 Alternatives in 2026: Self-Hosted Agent Memory Compared (Benchmarks)

Open-source benchmarks, real recall accuracy numbers, and an honest assessment of when each tool wins — for teams who need memory that runs on their own infrastructure.

Dakera AI TeamPublished 9 min read
On this page
  1. The quick summary
  2. Why teams move off Mem0
  3. Dakera
  4. Mem0
  5. Zep / Graphiti
  6. Letta
  7. LangMem
  8. Decision framework: which tool to use
  9. Migrating from Mem0 to Dakera
  10. The benchmark landscape is fragmented — and that's intentional

Mem0 dominates the agent memory conversation, but it's not the only option — and depending on your requirements for self-hosting, data privacy, recall accuracy, and budget, it may not be the best one. This post compares four serious alternatives: Zep, Letta, Dakera, and LangMem. We include real benchmark numbers where they exist and flag where "benchmark" is just a vendor claim.

Methodology note: Different tools use different benchmarks (LoCoMo, LongMemEval, internal evals). We use the benchmarks each tool has published. Direct cross-benchmark comparison isn't meaningful — but the data does show what each team chose to optimize for, which is itself revealing.

The quick summary

ToolSelf-hostedBest forPublished benchmarkFree tier
Dakera ✓ Fully Accuracy-first, privacy-sensitive teams 88.2% Recall@20 on LoCoMo ✓ Unlimited (self-host)
Mem0 Partial (graph tier = $249/mo cloud) Fast personalization, cloud-first 49.0% LongMemEval Limited (100 memories)
Zep Graphiti open source; hosted product retired self-hosted CE in 2025 Temporal KG, conversation-centric 63.8% LongMemEval (Graphiti) ✓ Graphiti (Apache 2.0)
Letta ✓ Fully (Apache 2.0) Long-running autonomous agents ~83.2% LongMemEval ✓ Unlimited (self-host)
LangMem Tied to LangGraph/LangChain LangGraph agents, zero infra No public benchmark ✓ Open source

Why teams move off Mem0

Mem0 has strong mindshare and excellent documentation. Teams typically outgrow it for three reasons:

  1. The graph memory paywall. Mem0's vector recall is available in the free tier, but the knowledge graph tier — which significantly improves accuracy for complex, multi-fact queries — requires the $249/month plan. Self-hosted graph memory requires paying for the managed tier.
  2. Accuracy on temporal and relational queries. Mem0 scores 49.0% on LongMemEval overall, dropping sharply on questions that require temporal reasoning ("what changed since last month?") or entity relationships ("who owns the data pipeline?"). For applications where wrong recall is worse than no recall, this is a hard limit.
  3. Data residency. Professional teams working on proprietary codebases or regulated data (healthcare, finance, legal) often can't route agent memory through a third-party US cloud. Mem0's cloud is US-based with no data residency options on lower tiers.
See recall accuracy in action
Run live store/recall/knowledge graph scenarios against Dakera — no account needed
Try the Playground →

Dakera

Dakera Best overall for self-hosted teams

Dakera is an open-core memory server written in Rust. It combines vector search, BM25 full-text indexing, importance-weighted scoring, and a knowledge graph layer into a single self-hosted binary. The core is open source (Apache 2.0); the cloud offering is optional.

Benchmark: 88.2% overall accuracy on the LoCoMo benchmark — a 30-turn conversation recall test that evaluates semantic recall (Cat1), entity tracking (Cat2), temporal inference (Cat3), and multi-fact reasoning (Cat4). Dakera is the only open-core memory server to publish a score on LoCoMo, which tests recall across longer conversation spans than LongMemEval. The hardest category — temporal inference (Cat3) — scores 68.5%.

Dakera
88.2%
Letta
~83.2%
Zep/Graphiti
63.8%
Mem0
49.0%

Note: Dakera runs LoCoMo; Letta, Zep, and Mem0 run LongMemEval. Benchmarks are not directly comparable — see methodology note above.

Architecture: Rust binary. ONNX embedding models (768-dim, run locally). HNSW vector index. BM25 for full-text. Knowledge graph with entity extraction. Memory decay by importance. MCP server with 14 core tools (86+ via profiles). The entire stack runs in a single Docker container with no external dependencies.

Setup:

# Start server
docker run -d -p 3000:3000 \
  -e DAKERA_ROOT_API_KEY=my-key \
  -e DAKERA_STORAGE=filesystem \
  -e DAKERA_STORAGE_PATH=/data \
  -v dakera-data:/data \
  ghcr.io/dakera-ai/dakera:latest

# Install MCP bridge
npm install -g @dakera-ai/dakera-mcp

# Test connection
dakera-mcp --test-connection

Python SDK:

pip install dakera

from dakera import AsyncDakeraClient
import asyncio

async def main():
    client = AsyncDakeraClient(base_url="http://localhost:3000", api_key="my-key")
    await client.memories.store(
        agent_id="my-agent",
        content="User always uses async/await, never callbacks",
        importance=0.85
    )
    results = await client.memories.recall(
        agent_id="my-agent",
        query="user's coding style preferences",
        top_k=5
    )

asyncio.run(main())

Best for: Teams that need the highest available recall accuracy, full data residency, and don't want to pay a per-memory or per-seat cloud fee. The free self-hosted tier is unlimited in memory count, agent count, and request volume.

Watch out for: Requires running a server process. Not a library you drop into a Python project — it's infrastructure. If your team has no ops capacity, the Dakera cloud offering provides a managed path.

Mem0

Mem0 Best for cloud-first teams who need fast setup

Mem0 is the best-documented, most widely integrated agent memory tool. Its strength is speed: the managed API is production-ready in under a minute, with SDKs for Python, TypeScript, and LangChain integration out of the box.

Benchmark: 49.0% on LongMemEval (GPT-4o). LongMemEval is a multi-session memory benchmark from Microsoft Research that tests recall accuracy across longer conversation histories. Mem0's score reflects its optimization focus: fast personalization recall for shorter histories, at the cost of complex temporal and relational queries.

The graph paywall: Mem0's basic vector recall is available free (up to 100 memories). The knowledge graph tier — which substantially improves multi-fact and relational recall — requires the paid plan starting at $249/month. Self-hosted graph memory requires the managed tier. Teams migrating off Mem0 typically do so when they hit this ceiling.

Best for: Proof-of-concepts, startups that need zero ops overhead, and applications where quick personalization (preferences, session summaries) matters more than deep temporal or relational recall.

Watch out for: The 49% LongMemEval score means roughly half of complex multi-session queries return incorrect or irrelevant results. For production applications where wrong recall is worse than no recall — medical, legal, financial — this accuracy floor is a real constraint.

Zep / Graphiti

Zep / Graphiti Best for conversation-centric temporal reasoning

Zep originally built a memory layer for conversation AI, and its approach — a temporal knowledge graph that tracks fact changes over time — is genuinely differentiated. When a user says "I moved from Seattle to Austin," Zep/Graphiti records both the old and new state with timestamps, enabling accurate answers to "where did they used to live?"

In 2025, Zep retired the self-hosted Community Edition of the Zep product. However, the underlying temporal knowledge graph engine, Graphiti, remains open source (Apache 2.0) and actively maintained on GitHub. Teams building on Zep should be clear about which product they're adopting.

Benchmark: Graphiti scores 63.8% on LongMemEval with GPT-4o. This is notably stronger than Mem0's 49.0% on the same benchmark — primarily because temporal fact changes are Graphiti's core design intent.

Best for: Applications where the same entities change state over time (user location changes, preference changes, relationship changes). Customer support, long-running personal assistants, any agent that needs to answer "what changed?"

Watch out for: Graphiti requires an LLM at write time for entity extraction (additional cost per memory write). The Zep hosted product's pricing starts at $99/month. If you need pure self-hosted: use Graphiti directly (Python library), not the Zep product.

Letta

Letta Best for long-running, stateful autonomous agents

Letta (formerly MemGPT) takes a fundamentally different architectural approach: instead of external memory storage, Letta builds memory management into the agent's runtime itself. The agent manages its own memory blocks (core memory, archival memory, recall memory) and decides what to remember, summarize, or forget.

Benchmark: Approximately 83.2% on LongMemEval — significantly outperforming Mem0 and Zep on the same benchmark. This reflects Letta's architecture: the agent itself is reasoning about what to remember, not just storing vectors.

Self-hosted: Letta is Apache 2.0 and fully self-hostable. The Docker deployment is well-documented and production-ready. This is the clearest pure open-source competitor to Dakera for teams that need to run everything themselves.

Best for: Long-running autonomous agents (days to weeks), workflows where the agent needs to self-edit its context, research agents that accumulate domain knowledge over many sessions.

Watch out for: Letta's architecture means the agent (and the LLM behind it) is actively involved in memory management decisions. This increases per-request LLM calls and latency. It also makes the memory system harder to query programmatically — you can't do a simple semantic search across all memories without going through the agent. For applications where you want deterministic, low-latency retrieval without agent involvement, Dakera or Zep are a better fit.

LangMem

LangMem Best for teams already deep in the LangGraph ecosystem

LangMem is LangChain's memory solution, designed to integrate natively with LangGraph agent workflows. It's a library, not a server — memory is stored in a backend you configure (Postgres, Redis, vector DB). This makes it the lowest-infrastructure-overhead option for teams already using LangGraph.

Benchmark: No public benchmark published. LangMem's accuracy is entirely dependent on the underlying storage backend and embedding model you configure.

Best for: LangGraph agents where you want native memory integration without deploying a separate server. The zero-new-infrastructure property is its main selling point.

Watch out for: The abstraction is broad — you're assembling accuracy by choosing the right storage backend, embedding model, and retrieval strategy. If you need production-grade accuracy without tuning all of those, a purpose-built memory server (Dakera, Letta) will outperform a general-purpose library.

Decision framework: which tool to use

Your situationRecommendedReason
Need highest recall accuracy, self-hosted, no per-memory fees Dakera 88.2% Recall@20 on LoCoMo, unlimited self-hosted, Rust binary with no external deps
Need temporal fact tracking ("what changed?") in conversations Zep / Graphiti Purpose-built for fact mutation tracking over time
Building long-running autonomous agents that self-manage context Letta Agent-managed memory blocks, strong benchmark, Apache 2.0
Need zero infra overhead and already use LangGraph LangMem Native integration, library not server, zero deployment burden
POC/prototype, cloud is fine, need fastest time-to-working Mem0 Best documentation, fastest SDK integration, works immediately
Regulated data (healthcare, legal, finance), must be fully on-prem Dakera or Letta Both fully self-hosted with no data leaving your infrastructure. Dakera sends anonymous version telemetry (opt out: DAKERA_TELEMETRY=off). Letta is Apache 2.0 open-source.

Migrating from Mem0 to Dakera

If you're currently using Mem0 and want to migrate, the transition is straightforward. Dakera's Python SDK mirrors Mem0's core API surface:

# Mem0 (before)
from mem0 import MemoryClient
client = MemoryClient(api_key="m0-xxx")
client.add("User prefers dark mode", user_id="alice")
results = client.search("UI preferences", user_id="alice")

# Dakera (after — same pattern, self-hosted)
from dakera import AsyncDakeraClient
import asyncio

async def main():
    client = AsyncDakeraClient(base_url="http://localhost:3000", api_key="my-key")
    await client.memories.store(
        agent_id="alice",
        content="User prefers dark mode",
        importance=0.8
    )
    results = await client.memories.recall(
        agent_id="alice",
        query="UI preferences",
        top_k=5
    )
asyncio.run(main())

The key differences: user_id → agent_id (can be any identifier), and the importance parameter controls how long the memory survives Dakera's importance-based decay. Start all memories at 0.8 and tune from there.

For the knowledge graph features (entity extraction, relationship traversal) that Mem0 gates behind its $249/month plan, use Dakera's power or all MCP profile — available at no additional cost in the self-hosted version.

MCP migration: If you're using Mem0's MCP server, replacing it with Dakera's is a config file change. Same MCP protocol, different server, 14 core tools in the default profile. See the MCP server guide for the full tool reference.

The benchmark landscape is fragmented — and that's intentional

One thing that stands out when comparing these tools: everyone uses a different benchmark. Mem0 and Zep/Graphiti both cite LongMemEval. Dakera publishes LoCoMo results. Others publish no benchmark at all.

This fragmentation isn't accidental. Each benchmark tests different capabilities. LongMemEval is a multi-session conversation recall test. LoCoMo includes a harder temporal inference category (Cat3) that tests whether a system can reason about when events happened, not just recall that they happened. Letta's ~83.2% on LongMemEval reflects its strength in conversation history recall; Dakera's 88.2% Recall@20 on LoCoMo reflects its strength in temporal reasoning and importance-weighted retrieval.

The honest recommendation: if accuracy is your primary criterion, run your own evaluation on a sample of production queries before committing. Ask each vendor for their benchmark methodology, test set, and what model they used. Benchmark scores on standardized datasets are a useful starting signal, not a substitute for testing on your data.

Try Dakera before you decide

Seven live scenarios against the real memory engine — temporal reasoning, knowledge graphs, semantic recall. No account, no setup. Or deploy your own instance in 5 minutes.

Stay sharp on agent memory
Benchmark releases, engineering deep-dives, and product updates. Once a week max, no fluff.
✓ Subscribed. Thanks!
Prefer a reader? RSS · Atom