NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

What's new in v0.12.0

Dakera v0.12.0, released on 2026-10-01, adds multilingual search, multimodal memory, multi-vector records and late interaction; makes recall faster and leaner; hardens security and cluster replication; and goes back to v0.11.108 with one command. This page summarizes the release for developers and operators deciding what to use and how to upgrade.

Every new feature is opt-in. An unchanged v0.11.108 deployment upgrades in place: stop v0.11.108, start v0.12.0 on the same data and environment. The operator guide is Upgrading from v0.11.108, and every change is listed in the Changelog.

Highlights

Multilingual search

Four independent, opt-in layers. A server that sets none of them behaves as v0.11.108 did (English).

Details: Multilingual search.

Multimodal memory

Off by default; each feature turns on with one variable.

Every media job reserves the memory it needs before it starts, and waits or answers 503 with Retry-After instead of risking an out-of-memory kill. With every model loaded, the process heap measured 530 MiB (memory-mapped model files are not counted). Details: Multimodal memory.

Multi-vector records

DAKERA_RECORDS turns on POST /v1/namespaces/{ns}/records and GET /v1/namespaces/{ns}/records/{id}. A record is one indexed vector plus named extra representations (dense, token_multivector, patch_multivector, stored as f32, f16 or i8), kept atomically beside it by every storage backend, snapshots, export/import and replication. Use records when you compute multi-vector embeddings yourself and want Dakera to keep them with the document. Details: Records and multi-vector.

Late interaction and search controls

Details: Search and ranking.

Performance

Every figure is a measurement; "Before" is the same measurement without the change (v0.11.108's behavior unless noted). More in Performance.

What Before v0.12.0 Measured on
Vector search p50 / p95 at recall@10 0.998 1.07 / 1.44 ms 0.87 / 1.17 ms BEIR Quora, 50k vectors, CPU only
Reranked recall, top_k 16 baseline about 2× faster on the same CPU, no GPU 4 vCPU / 8 GiB container, bge-reranker-v2-m3
Container memory after reranking 4.98 GB 0.65 GB, flat across top_k 1 to 32 same
Reranker load 4.6 s 0.6 s same
HNSW memory per 1024-d vector 4.81 KB 2.76 KB (−43 %), same recall@10 (0.998) BEIR Quora, 50k vectors
Ingest speed / memory baseline / 2.4 GB 2.7× faster / 0.97 GB Same workload on v0.11.108 and v0.12.0, same machine
CPU Docker image (amd64) n/a 783 MB; ready 4.4 s after start, with no download in air-gapped or read-only starts Docker run of the release image
Release binary (x64) n/a 106,228,848 bytes (about 101 MB) Release build

Also: every CPU model runs on one memory-mapped copy with no idle arena; stored vectors no longer depend on what else was in the batch; ANN indexes are saved and reloaded instead of rebuilt; an expired memory no longer forces an index rebuild; concurrent writes into one namespace share the write-ahead log's fsync.

Security hardening

Details: Security and Upgrading: security and access.

One-command rollback

dakera downgrade ships in the v0.12.0 binary. Run it once on the stopped deployment, with the server's environment and volumes, and it converts the data back to what v0.11.108 reads: encrypted values, full-text indexes, the write-ahead log, tiered writes, caches and re-embedded models included. No v0.11 patch release is needed.

stop v0.12.0  ->  dakera downgrade  ->  start v0.11.108

Details: Rolling back to v0.11.

Reliability and operations

Telemetry disclosure

When telemetry is on, v0.12.0 sends lifecycle events and a 6-hourly heartbeat. The machine's hostname is sent, and PostHog keeps the connecting IP address and derives a coarse location from it (GeoIP). Opt out with DAKERA_TELEMETRY=0 or DO_NOT_TRACK=1. See Telemetry.

What changes for existing deployments

Read these before upgrading. Each is explained, with what to do, in Upgrading from v0.11.108.

  1. gRPC clients need an API key in the call metadata (x-api-key or authorization: Bearer) when authentication is on.
  2. Permissions: keys pinned to namespaces get 403 on node-wide routes; backup download, upload and restore need global super_admin.
  3. Namespace quotas are enforced. Quotas set on v0.11 were never checked; writes over a hard quota now get 413.
  4. Ranking follows the query more closely. A recall no longer raises the importance of what it returns, the access count is no longer a score term, and recency counts from when a memory was created. smart_score values change scale.
  5. A stored, enabled backup schedule starts running at its next slot (v0.11 stored it and never ran it).
  6. Health checks: the port answers while models load; point readiness at /health/ready and liveness at /health/live.
  7. Configuration: an upgraded deployment starts with warnings, never refusals (/health config_warnings); a fresh install with the same mistakes is refused. Run dakera --check-config first.
  8. Stored memories are re-embedded once in the background (they were embedded with the query instruction): about 35 minutes per 10,000 memories on bge-large; recall is served throughout.
  9. Clusters: leave DAKERA_CLUSTER_SECRET unset until every node runs v0.12.0, keep the mixed period short, and give each node its own S3 bucket.
  10. Speech to text defaults to whisper-base (multilingual, language auto-detected), not English-only whisper-tiny.en. Speech is opt-in; a deployment that turns it on without DAKERA_WHISPER_MODEL transcribes with whisper-base. Set DAKERA_WHISPER_MODEL=whisper-tiny.en to keep the lightest English model.
  11. Errors: every error body is JSON, every 503 carries Retry-After, configuration errors are 501, over-size bodies are 413.
  12. Namespace config: PATCH merges and refuses unknown fields; PUT replaces.
  13. Telemetry, when on: lifecycle events and a 6-hourly heartbeat; the hostname is sent and the connecting IP is kept (see above).
  14. POST /ops/shutdown really shuts the server down (in v0.11 it only set a flag).

Operating notes

Where to go next

Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.