Dakera v0.12.0: multilingual, multimodal, multi-vector memory
v0.12.0 lets an agent remember in more languages and more media: text in other languages, audio, document pages, and several vectors per record. Recall costs less time and memory, security is tighter, and going back to v0.11.108 is one command.
This page is the release overview for teams evaluating or running Dakera. v0.12.0 upgrades in place from v0.11.108, and every new feature is opt-in: a deployment that sets none of the new variables keeps its behavior, apart from the changes listed under What changes when you upgrade.
CPU only. Search and index memory: BEIR Quora, 50k vectors. Reranking memory: 4 vCPU / 8 GiB container. Details in Measured performance.
What's in the release
Multilingual search
The bge-m3 embedding model (1024 dimensions), full-text stemming and stop words per language, character-bigram indexing for Chinese, Japanese and Korean, dates understood in seven query languages (English, German, French, Spanish, Italian, Portuguese, Dutch), and a per-request lang.
Multimodal memory
Attachments referenced from memories, speech to text that turns a WAV recording into a memory (a choice of five Whisper models: the default, whisper-base, is multilingual and detects the spoken language; lighter English-only models are one setting away), and image indexing with visual recall over document pages (colmodernvbert). Each media job reserves its memory first, and waits or answers 503 with Retry-After instead of risking an out-of-memory kill.
Multi-vector records
One indexed vector plus named extra representations (dense, token multivector, patch multivector), stored as f32, f16 or i8 and kept beside it by every storage backend, snapshots, export/import and replication.
Late interaction and search controls
Late interaction with colbert-small reranks with per-token MaxSim. RaBitQ search mode cuts search latency. Rerank controls cap how many candidates the cross-encoder scores, and GET /v1/capabilities tells clients what the server supports.
Each is opt-in: a deployment that does not turn them on runs as before. Visual recall is measured on two ViDoRe sets (nDCG@5 0.643 on TabFQuAD, 0.7705 on Shift Project).
Measured performance
Each figure is a recorded measurement. “Before” is v0.11.108 unless noted.
| What | Before | v0.12.0 | Conditions |
|---|---|---|---|
Reranked recall, top_k 16 | baseline | about 2× faster, no GPU | same CPU, 4 vCPU / 8 GiB container |
| Ingest | baseline | 2.7× faster | same workload, same machine |
| Multimodal stack (text, speech, vision models loaded) | — | ~530 MiB working memory | CPU only |
| Container memory after reranking | 4.98 GB | 0.65 GB, flat across top_k 1–32 | 4 CPU / 8 GiB container |
| Reranker load | 4.6 s | 0.6 s | 4 CPU / 8 GiB container |
| HNSW memory per 1024-d vector | 4.81 KB | 2.76 KB, recall@10 0.998 unchanged | BEIR Quora, 50k vectors |
| HNSW search p50 / p95 | 1.07 / 1.44 ms | 0.87 / 1.17 ms | BEIR Quora, 50k vectors |
ANN indexes are now saved and loaded at start instead of rebuilt, an expired memory no longer forces a full index rebuild, and concurrent writes are flushed to disk together. The CPU Docker image is 783 MB (amd64) and is ready 4.4 s after start, with no download in air-gapped or read-only starts.
More figures and methods: Performance. Retrieval quality is on the benchmark page.
Security hardening
- gRPC requires an API key and is rate limited, time-bounded and audited like REST.
- Encrypted values are bound to their record, with keys in a replicated keyring, rotation one namespace at a time, and older backups that still restore.
- Namespace-bound decryption: memory content is opened only under the namespace it was read from, and clients cannot store sealed values.
- Cluster traffic is authenticated with
DAKERA_CLUSTER_SECRET, and gossip packets are HMAC-authenticated. - Tighter permissions: keys pinned to namespaces no longer reach node-wide admin routes, and backup download, upload and restore need global
super_admin. - Smaller attack surface: with authentication off, only local web pages may call the server; knowledge-graph edges are private to their agent; imports, extractor responses and graph searches are bounded; secrets are kept out of logs and
/health.
Reliability and operations
- Cluster replication is versioned and merges instead of overwriting, with a durable outbox and tombstones.
- Tiered storage keeps a durable local hot tier and keeps serving through object-store outages.
- Startup: the server binds its port while models download, with
/health/liveand/health/ready. Configuration mistakes are caught at startup, anddakera --check-configruns that check alone. - Model management:
dakera models list | pull | prunemanages model files, and the images ship their default models for air-gapped starts. - Data lifecycle: backups, TTL, forget, deduplication and consolidation are hardened end to end.
Operations → · Models and Docker →
Going back is one command
dakera downgrade, shipped in the v0.12.0 binary, converts a stopped deployment's data back to what v0.11.108 reads: encrypted values, full-text indexes, the write-ahead log, tiered writes and caches. No v0.11 patch release is needed. It refuses, changing nothing, when a server still runs on the data or when the data uses something v0.11.108 lacks (bge-m3, colbert-small, the visual lane).
docker run --rm <env and volumes> <v0.12 image> downgrade
What changes when you upgrade
An unchanged v0.11.108 deployment upgrades in place. Check these before you switch:
- gRPC clients need an API key in the call metadata when authentication is on.
- Keys pinned to namespaces get
403on node-wide routes; backup tooling needs globalsuper_admin. - Namespace quotas are enforced: writes over a
hardquota get413. - Ranking follows the query more closely and
smart_scorechanges scale: re-check any absolute cut on it. - Probes: point readiness at
/health/readyand liveness at/health/live. - Telemetry, when on, sends the machine's hostname, and PostHog keeps the connecting IP (GeoIP). Opt out with
DAKERA_TELEMETRY=0orDO_NOT_TRACK=1.
Stored memories are re-embedded once in the background after the upgrade (about 35 minutes per 10,000 memories on bge-large); recall is served throughout. The full list, with what to do for each item, is in the upgrade guide and the changelog.
Get started
Reference: Configuration · REST API · Telemetry · High availability