NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →
COMPARE · INTERPRETATION VS EVIDENCE

Honcho asks an LLM what your user is like. Dakera remembers what happened.

Honcho, from Plastic Labs, stores conclusions that reasoning models draw about each person your agent talks to, and answers questions about them with another LLM call. Dakera stores the source memories with provenance, importance, decay and a knowledge graph, ranks them deterministically in milliseconds, and leaves the reasoning to your own model at answer time. One is interpretation, the other is evidence.

Reviewed October 2026

How Honcho works

Honcho calls itself persistent, reasoning-based memory for users and other peers. Its data model has four primitives: workspaces contain peers (entities that persist over time, such as a user or an agent) and sessions (interaction threads), and messages are the data units that trigger reasoning. Work then happens at three different moments.

1 · On writesynchronous · no LLM
session.add_messages()A message is stored and a reasoning task is enqueued in the same request; the API returns immediately
Workspace, peers, sessions, messagesThe Storage service
2 · In the backgroundasynchronous · LLM calls
DeriverLLMExtracts explicit and deductive conclusions about each peer
SummarizerLLMShort and long session summaries (every 20 and 60 messages)
DreamerLLMConsolidates conclusions, adds inductive ones, writes peer cards
3 · On queryinline · LLM call
peer.chat()LLMThe Dialectic agent searches conclusions, pulls supporting messages and writes an answer
session.context()Formats stored summaries and messages; no LLM call
Stored in Postgres with pgvector. Messages and core entities live in Postgres; conclusions are vector-embedded documents keyed by observer and observed peer.
No LLM callLLM call
Honcho's pipeline as described in its architecture, reasoning and dreaming docs, and the repository README.
  • On write. Per the architecture docs, a message is stored and a reasoning task is enqueued in the same request, and the API returns immediately. Reasoning runs afterwards in a queue consumed by the deriver worker, with session-level ordering.
  • In the background. The deriver extracts explicit and deductive conclusions from messages. A summarizer rolls up session summaries. The dreamer runs later, once at least 50 new conclusions exist and 8 hours have passed, and uses deduction and induction agents to consolidate conclusions and write the peer card, a profile of stable facts capped at 40 per card.
  • Representations. A peer representation is the evolving, reasoned model of a peer: conclusions (deductive, inductive, abductive), summaries and the peer card. With observe_others, one peer can hold its own view of another based only on interactions it shared.
  • On query. peer.chat(), the dialectic API, runs the agent's full reasoning loop with an LLM and returns a synthesised answer. session.context() is a read without an LLM call.

The result is a rich, evolving profile of a person, with an LLM in three places: background derivation, dreaming and the query itself. The README lists LLM provider API keys as a requirement.

What Dakera does instead

Dakera keeps the evidence. A store embeds the text and extracts entities on-device, and a recall runs HNSW and BM25, fuses them with RRF and reranks with a cross-encoder, all in one container with the models included. Nothing is rewritten, so each result is a memory you stored, with its importance, timestamps and session. Your own model reasons over that evidence when it answers.

Dakerano LLM on store or recall
Storeimportance, tags, session, agent
Embed and extract entitiesmultilingual, on-device
Hybrid recall and rerankHNSW, BM25, RRF, cross-encoder
Ranked source memoriesscores and a rerank_report
Your model answersreasons over the evidence
Your infrastructureThird-party API

Cost and latency of one question

An interpretation layer pays at query time. A retrieval layer pays at the index.

HONCHO · ONE DIALECTIC QUESTION

An LLM call

Managed-service price per query, by the reasoning_level you choose:

$0.001minimal
$0.01low
$0.05medium
$0.10high
$0.50max

Deeper levels think longer and use more tools. Ingestion is priced separately at $2.00 per million tokens. Bar heights are steps, not to scale. Pricing

DAKERA · ONE RECALL

0.87 ms p50

Vector search, 50k vectors, CPU only, recall@10 0.998 (p95 1.17 ms). One stage of recall, which also runs BM25, RRF fusion and a cross-encoder rerank.

  • No LLM and no external API in the path
  • No per-call or per-token fee: your server cost, flat
  • Works with the network cut

Two different things: Honcho's figures are dollars for a model-written answer, Dakera's is retrieval time for ranked memories. They are shown side by side, not compared as one metric.

The same question in code

Both snippets store a preference and ask what the user prefers. The Honcho side is from its quickstart and the Dakera side uses the Python SDK.

Honcho
from honcho import Honcho

honcho = Honcho(
    workspace_id="my-app",
    api_key=os.environ["HONCHO_API_KEY"],
)
alice = honcho.peer("alice")
session = honcho.session("session-1")
session.add_messages([
    alice.message(
        "I prefer email to phone calls"),
])

# an LLM reasons over stored conclusions
answer = alice.chat(
    "what does the user prefer?")

# returns a model-written answer (str):
# "Based on conclusions, the user
#  prefers ..."
Dakera
from dakera import DakeraClient

client = DakeraClient(
    "http://localhost:3000", api_key=KEY)

client.store_memory(
    agent_id="alice",
    content="I prefer email to phone calls",
    importance=0.8,
)

# hybrid retrieval + rerank, on-device
response = client.recall(
    agent_id="alice",
    query="what does the user prefer?",
    top_k=5,
)
for m in response.memories:
    print(round(m.score, 2), m.content)

# returns ranked source memories with
# scores, plus a rerank_report

Honcho returns a written answer you trust or re-verify. Dakera returns the memories themselves, so the answer can be checked against them.

What you can audit

Honcho keeps the original messages and can return the conclusions and messages it consulted, so both systems retain source text. They differ in what sits between the source and the answer.

What you can audit
CapabilityHonchoDakera v0.12.0
Original messages are storedHoncho keeps messages in Postgres; Dakera keeps each memory as writtenYesYes
Evidence behind an answer is retrievableinclude_evidence on peer.chat(); in Dakera the answer is the memoriesYesYes
No LLM call on the query pathpeer.chat() runs an LLM reasoning loop; Dakera recall does notNoYes
No LLM provider needed to runHoncho's README lists provider keys as a requirement; Dakera ships its modelsNoYes
Conclusions about each person, written by an LLMHoncho's core strengthYesNo
Open-source serverHoncho is AGPL-3.0. Dakera's SDKs, CLI and MCP server are MIT; the server binary is proprietary with no usage feesYesPartly
What only Dakera adds on the audit side. Every recall returns a score per memory and a rerank_report (whether the cross-encoder ran, over how many candidates), and the Recall lab below shows why each memory ranked where it did. The same query over the same memories ranks the same way, because no model is sampled in the retrieval path. Start the container with the network cut and it still runs, and memory text never leaves your server, which simplifies data-residency and compliance reviews.
The Dakera Dashboard Recall lab listing ranked memories with scores and the rerank report
The Dakera Dashboard Recall lab: run a recall for an agent and inspect each result's score and the rerank report.

Where Honcho is strong, and using both

If you are building something that has to understand people, such as a tutor, companion or assistant that notices patterns over months, Honcho is designed for exactly that: formal-logic conclusions, peer cards and per-peer perspectives are things a store of memories does not produce on its own. It is AGPL-3.0 open source with a managed service.

The two also sit at different layers. Dakera can be the evidence store, holding exactly what happened with ranked, explainable recall, while your LLM, or Honcho, interprets it. This is an architecture pattern, not a tested integration.

Choose Honcho if

  • Your product is about understanding people: how a user, a team or another agent changes over time.
  • You want an LLM to reason over conversations and keep conclusions and peer cards.
  • You want an AGPL-3.0 server you can read, with a managed service for convenience.

Choose Dakera if

  • You want the exact memories back, with scores, and answers you can check against them.
  • You want millisecond recall with no per-question fee and no LLM provider to run.
  • The host is air-gapped, or memory text must stay on your servers.
  • You need multilingual, multimodal or knowledge-graph memory shared by Python, TypeScript, Go, Rust and MCP clients.

Prefer managed? Dakera Cloud is coming. Join the waitlist.

On benchmarks, Dakera reports LoCoMo Recall@20 of 88.2% (does retrieval surface the answer) while Honcho reports QA accuracy (does an LLM answer correctly). Different measurements. Method

Frequently asked questions

How does Honcho work?

Messages in a session trigger background reasoning. A deriver worker extracts conclusions about each peer, a dreamer consolidates them and writes peer cards, and the dialectic API (peer.chat()) answers questions about a peer with an LLM. Data is stored in Postgres with pgvector.

Does Honcho need an LLM provider?

Per its README, yes: you need API keys for at least one LLM provider. It names Google Gemini and Anthropic for background and dialectic work, and OpenAI for embeddings.

What does a Honcho query cost?

On the managed service, reasoning queries are priced from $0.001 (minimal) to $0.50 (max) each, and memory ingestion is $2.00 per million tokens, per honcho.dev.

Can Dakera build user models like Honcho?

Dakera stores and recalls source memories without an LLM. If you want a model's interpretation, your agent can call its own LLM with Dakera supplying the ranked memories.

What does Dakera need to run?

One container and one secret protecting your own server. No database, LLM or embedding provider is required.

Self-hosted is free

Run Dakera on your own infrastructure.

One binary with its models baked in and no per-call fees. Start it with one Docker command and see it for yourself.

Prefer managed hosting? Join the Dakera Cloud waitlist →