Dakera is a self-hosted memory engine for AI agents: one Rust server that stores agent memories and documents, embeds them on the machine it runs on, and recalls them with vector, full-text and hybrid search, ranking and a knowledge graph. It is served over REST (port 3000) and gRPC (port 50051), with SDKs for Python, TypeScript, Go and Rust, the dk CLI and an MCP server.
These pages document v0.12.0. Start with the Quick Start if you are new to Dakera, or with the upgrade guide if you run v0.11.108; the table below points to the page for each task.
Public alpha — the self-hosted server, SDKs, CLI and MCP server are available now. Dakera Cloud (managed hosting) is coming next — join the waitlist →
Dakera v0.12.0 — You are reading the latest docs. New: opt-in multimodal memory (attachments, speech to text, image indexing, records), late interaction, multilingual search and RaBitQ; GET /v1/capabilities; gRPC authentication; enforced namespace quotas; and dakera downgrade to go back. What's new in v0.12.0 → · Upgrading from v0.11.108 → · Rolling back to v0.11 → · Looking for the previous release? v0.11 docs →
HNSW vector index with SIMD-accelerated distances, plus an opt-in RaBitQ search mode
Elasticsearch · OpenSearch
BM25 full-text search engine with per-namespace indexes
OpenAI / Cohere embeddings API
On-device ONNX inference (MiniLM, BGE, E5; opt-in bge-m3 multilingual and colbert-small late interaction) — zero API calls
Redis / Postgres memory layer
Decay-weighted agent memory with sessions, importance scoring, and 6 decay strategies
Neo4j knowledge graph
Memory graph with 4 edge types (related, shared entity, precedes, explicit link) and a cross-agent network
Mem0 / Zep memory services
Import and export in Mem0 and Zep formats (and JSONL, CSV)
Separate NER service
GLiNER zero-shot named entity extraction with multi-provider support
88.2% Recall@20 on LoCoMo — Dakera v0.11.107 scored 88.2% Recall@20 on the LoCoMo benchmark (10 conversations, 1,536 evaluated questions); this figure was measured on v0.11.107, not on v0.12.0 (v0.12.0 release measurements are in What's new). Read the methodology →
Key capabilities
Capability
Description
Hybrid Retrieval
Vector ANN and BM25 full-text combined (min-max or Reciprocal Rank Fusion), then cross-encoder reranking. HNSW search p50 0.87 ms on BEIR Quora 50k; see Performance.
On-device Inference
Embeddings (MiniLM, BGE, E5), reranking, and entity extraction via ONNX Runtime — zero external API calls.
Memory Decay
Six decay strategies. Per-type TTLs. Spaced repetition. Configurable half-life.
Knowledge Graphs
Memory graph with 4 edge types, BFS traversal, shortest paths and a cross-agent network.
HA Clustering
Gossip membership and leader election; every node holds a full copy and accepts writes, and versioned replication (hybrid logical clock, tombstones) converges all nodes. Details →
Multilingual Search
Opt-in bge-m3 embeddings, per-language full-text stemming and stop words, CJK bigram indexing, dates understood in seven query languages, and a per-request lang. Details →
Multimodal Memory
Opt-in attachments, speech to text, image-page indexing with visual recall, and multi-vector records. Every media job reserves its memory first. Details →
Late Interaction & RaBitQ
Opt-in per-token MaxSim reranking (colbert-small) and a RaBitQ search mode for latency. Details →
One-command Rollback
dakera downgrade converts a stopped deployment's data back to what v0.11.108 reads. Details →
Enterprise Security
AES-256-GCM encryption at rest with key rotation, scoped API keys, rate limiting, input validation against injection patterns, audit logging.
AutoPilot
Background deduplication and consolidation of low-importance memories.
MCP Native
14 core tools (99 in total across profiles) for Claude Desktop, Claude Code, Cursor and Windsurf, with no code to write. Details →
Dual Protocol
REST API on port 3000 and gRPC API on port 50051. Both share the same engine, storage, and auth layers.
Query Classification
Each query is classified (keyword, semantic, hybrid, temporal, multi-hop) and routed to a matching retrieval strategy. Rule-based by default; an embedding-based classifier is opt-in (DAKERA_ML_CLASSIFIER).
Session Management
Group memories by agent session, recall within one session, and end a session with a summary (written by you or generated by the server).
Entity Extraction
GLiNER zero-shot NER with a rule-based pre-pass (dates, URLs, UUIDs, emails, IPs). Other extraction providers (OpenAI, Anthropic, OpenRouter, Ollama) are configurable per namespace.
Backup & Recovery
Scheduled backups, compressed (zstd) or encrypted. Write-ahead log and snapshots. Import/export (Mem0, Zep, JSONL, CSV).
Cross-encoder Reranking
bge-reranker-v2-m3 ONNX model (shipped in the image) for relevance scoring. Applied after initial retrieval for higher-precision results.
Dakera sends product telemetry by default. Since v0.12 it includes the machine's hostname, and PostHog records the connecting IP (GeoIP). Opt out: DAKERA_TELEMETRY=0 or DO_NOT_TRACK=1. Details →
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.
Lock in founder pricing
Dakera Cloud — managed hosting, SLA, and team monitoring — is launching soon. Join now to secure early access pricing before public launch.
No spam. Unsubscribe anytime.
✓ You're on the list!We'll reach out when beta slots open.
Lock in founder pricing
Dakera Cloud — managed hosting, SLA, and team monitoring — is launching soon. Join the waitlist now to secure early access pricing before public launch.