Honcho asks an LLM what your user is like. Dakera remembers what happened.
Honcho, from Plastic Labs, stores conclusions that reasoning models draw about each person your agent talks to, and answers questions about them with another LLM call. Dakera stores the source memories with provenance, importance, decay and a knowledge graph, ranks them deterministically in milliseconds, and leaves the reasoning to your own model at answer time. One is interpretation, the other is evidence.
Reviewed October 2026
How Honcho works
Honcho calls itself persistent, reasoning-based memory for users and other peers. Its data model has four primitives: workspaces contain peers (entities that persist over time, such as a user or an agent) and sessions (interaction threads), and messages are the data units that trigger reasoning. Work then happens at three different moments.
- On write. Per the architecture docs, a message is stored and a reasoning task is enqueued in the same request, and the API returns immediately. Reasoning runs afterwards in a queue consumed by the deriver worker, with session-level ordering.
- In the background. The deriver extracts explicit and deductive conclusions from messages. A summarizer rolls up session summaries. The dreamer runs later, once at least 50 new conclusions exist and 8 hours have passed, and uses deduction and induction agents to consolidate conclusions and write the peer card, a profile of stable facts capped at 40 per card.
- Representations. A peer representation is the evolving, reasoned model of a peer: conclusions (deductive, inductive, abductive), summaries and the peer card. With
observe_others, one peer can hold its own view of another based only on interactions it shared. - On query.
peer.chat(), the dialectic API, runs the agent's full reasoning loop with an LLM and returns a synthesised answer.session.context()is a read without an LLM call.
The result is a rich, evolving profile of a person, with an LLM in three places: background derivation, dreaming and the query itself. The README lists LLM provider API keys as a requirement.
What Dakera does instead
Dakera keeps the evidence. A store embeds the text and extracts entities on-device, and a recall runs HNSW and BM25, fuses them with RRF and reranks with a cross-encoder, all in one container with the models included. Nothing is rewritten, so each result is a memory you stored, with its importance, timestamps and session. Your own model reasons over that evidence when it answers.
Cost and latency of one question
An interpretation layer pays at query time. A retrieval layer pays at the index.
HONCHO · ONE DIALECTIC QUESTION
Managed-service price per query, by the reasoning_level you choose:
Deeper levels think longer and use more tools. Ingestion is priced separately at $2.00 per million tokens. Bar heights are steps, not to scale. Pricing
DAKERA · ONE RECALL
Vector search, 50k vectors, CPU only, recall@10 0.998 (p95 1.17 ms). One stage of recall, which also runs BM25, RRF fusion and a cross-encoder rerank.
- No LLM and no external API in the path
- No per-call or per-token fee: your server cost, flat
- Works with the network cut
Two different things: Honcho's figures are dollars for a model-written answer, Dakera's is retrieval time for ranked memories. They are shown side by side, not compared as one metric.
The same question in code
Both snippets store a preference and ask what the user prefers. The Honcho side is from its quickstart and the Dakera side uses the Python SDK.
from honcho import Honcho
honcho = Honcho(
workspace_id="my-app",
api_key=os.environ["HONCHO_API_KEY"],
)
alice = honcho.peer("alice")
session = honcho.session("session-1")
session.add_messages([
alice.message(
"I prefer email to phone calls"),
])
# an LLM reasons over stored conclusions
answer = alice.chat(
"what does the user prefer?")
# returns a model-written answer (str):
# "Based on conclusions, the user
# prefers ..."from dakera import DakeraClient
client = DakeraClient(
"http://localhost:3000", api_key=KEY)
client.store_memory(
agent_id="alice",
content="I prefer email to phone calls",
importance=0.8,
)
# hybrid retrieval + rerank, on-device
response = client.recall(
agent_id="alice",
query="what does the user prefer?",
top_k=5,
)
for m in response.memories:
print(round(m.score, 2), m.content)
# returns ranked source memories with
# scores, plus a rerank_reportHoncho returns a written answer you trust or re-verify. Dakera returns the memories themselves, so the answer can be checked against them.
What you can audit
Honcho keeps the original messages and can return the conclusions and messages it consulted, so both systems retain source text. They differ in what sits between the source and the answer.
| Capability | Honcho | Dakera v0.12.0 |
|---|---|---|
| Original messages are storedHoncho keeps messages in Postgres; Dakera keeps each memory as written | Yes | Yes |
Evidence behind an answer is retrievableinclude_evidence on peer.chat(); in Dakera the answer is the memories | Yes | Yes |
No LLM call on the query pathpeer.chat() runs an LLM reasoning loop; Dakera recall does not | No | Yes |
| No LLM provider needed to runHoncho's README lists provider keys as a requirement; Dakera ships its models | No | Yes |
| Conclusions about each person, written by an LLMHoncho's core strength | Yes | No |
| Open-source serverHoncho is AGPL-3.0. Dakera's SDKs, CLI and MCP server are MIT; the server binary is proprietary with no usage fees | Yes | Partly |
rerank_report (whether the cross-encoder ran, over how many candidates), and the Recall lab below shows why each memory ranked where it did. The same query over the same memories ranks the same way, because no model is sampled in the retrieval path. Start the container with the network cut and it still runs, and memory text never leaves your server, which simplifies data-residency and compliance reviews.
Where Honcho is strong, and using both
If you are building something that has to understand people, such as a tutor, companion or assistant that notices patterns over months, Honcho is designed for exactly that: formal-logic conclusions, peer cards and per-peer perspectives are things a store of memories does not produce on its own. It is AGPL-3.0 open source with a managed service.
The two also sit at different layers. Dakera can be the evidence store, holding exactly what happened with ranked, explainable recall, while your LLM, or Honcho, interprets it. This is an architecture pattern, not a tested integration.
Choose Honcho if
- Your product is about understanding people: how a user, a team or another agent changes over time.
- You want an LLM to reason over conversations and keep conclusions and peer cards.
- You want an AGPL-3.0 server you can read, with a managed service for convenience.
Choose Dakera if
- You want the exact memories back, with scores, and answers you can check against them.
- You want millisecond recall with no per-question fee and no LLM provider to run.
- The host is air-gapped, or memory text must stay on your servers.
- You need multilingual, multimodal or knowledge-graph memory shared by Python, TypeScript, Go, Rust and MCP clients.
Prefer managed? Dakera Cloud is coming. Join the waitlist.
On benchmarks, Dakera reports LoCoMo Recall@20 of 88.2% (does retrieval surface the answer) while Honcho reports QA accuracy (does an LLM answer correctly). Different measurements. Method
Frequently asked questions
How does Honcho work?
Messages in a session trigger background reasoning. A deriver worker extracts conclusions about each peer, a dreamer consolidates them and writes peer cards, and the dialectic API (peer.chat()) answers questions about a peer with an LLM. Data is stored in Postgres with pgvector.
Does Honcho need an LLM provider?
Per its README, yes: you need API keys for at least one LLM provider. It names Google Gemini and Anthropic for background and dialectic work, and OpenAI for embeddings.
What does a Honcho query cost?
On the managed service, reasoning queries are priced from $0.001 (minimal) to $0.50 (max) each, and memory ingestion is $2.00 per million tokens, per honcho.dev.
Can Dakera build user models like Honcho?
Dakera stores and recalls source memories without an LLM. If you want a model's interpretation, your agent can call its own LLM with Dakera supplying the ranked memories.
What does Dakera need to run?
One container and one secret protecting your own server. No database, LLM or embedding provider is required.
Run Dakera on your own infrastructure.
One binary with its models baked in and no per-call fees. Start it with one Docker command and see it for yourself.
Prefer managed hosting? Join the Dakera Cloud waitlist →