NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →
Released 2026-10-01

Dakera v0.12.0: multilingual, multimodal, multi-vector memory

v0.12.0 lets an agent remember in more languages and more media: text in other languages, audio, document pages, and several vectors per record. Recall costs less time and memory, security is tighter, and going back to v0.11.108 is one command.

This page is the release overview for teams evaluating or running Dakera. v0.12.0 upgrades in place from v0.11.108, and every new feature is opt-in: a deployment that sets none of the new variables keeps its behavior, apart from the changes listed under What changes when you upgrade.

1.17 msvector search p95 at recall@10 0.998 (0.87 ms p50)
7.6×less memory after reranking: 4.98 GB to 0.65 GB
−43 %HNSW memory per 1024-d vector, same recall@10
1 commandto roll back to v0.11.108

CPU only. Search and index memory: BEIR Quora, 50k vectors. Reranking memory: 4 vCPU / 8 GiB container. Details in Measured performance.

What's in the release

Multilingual search

The bge-m3 embedding model (1024 dimensions), full-text stemming and stop words per language, character-bigram indexing for Chinese, Japanese and Korean, dates understood in seven query languages (English, German, French, Spanish, Italian, Portuguese, Dutch), and a per-request lang.

Multilingual search →

Multimodal memory

Attachments referenced from memories, speech to text that turns a WAV recording into a memory (a choice of five Whisper models: the default, whisper-base, is multilingual and detects the spoken language; lighter English-only models are one setting away), and image indexing with visual recall over document pages (colmodernvbert). Each media job reserves its memory first, and waits or answers 503 with Retry-After instead of risking an out-of-memory kill.

Multimodal memory →

Multi-vector records

One indexed vector plus named extra representations (dense, token multivector, patch multivector), stored as f32, f16 or i8 and kept beside it by every storage backend, snapshots, export/import and replication.

Records and multi-vector →

Late interaction and search controls

Late interaction with colbert-small reranks with per-token MaxSim. RaBitQ search mode cuts search latency. Rerank controls cap how many candidates the cross-encoder scores, and GET /v1/capabilities tells clients what the server supports.

Search and ranking →

Each is opt-in: a deployment that does not turn them on runs as before. Visual recall is measured on two ViDoRe sets (nDCG@5 0.643 on TabFQuAD, 0.7705 on Shift Project).

Measured performance

Each figure is a recorded measurement. “Before” is v0.11.108 unless noted.

WhatBeforev0.12.0Conditions
Reranked recall, top_k 16baselineabout 2× faster, no GPUsame CPU, 4 vCPU / 8 GiB container
Ingestbaseline2.7× fastersame workload, same machine
Multimodal stack (text, speech, vision models loaded)—~530 MiB working memoryCPU only
Container memory after reranking4.98 GB0.65 GB, flat across top_k 1–324 CPU / 8 GiB container
Reranker load4.6 s0.6 s4 CPU / 8 GiB container
HNSW memory per 1024-d vector4.81 KB2.76 KB, recall@10 0.998 unchangedBEIR Quora, 50k vectors
HNSW search p50 / p951.07 / 1.44 ms0.87 / 1.17 msBEIR Quora, 50k vectors

ANN indexes are now saved and loaded at start instead of rebuilt, an expired memory no longer forces a full index rebuild, and concurrent writes are flushed to disk together. The CPU Docker image is 783 MB (amd64) and is ready 4.4 s after start, with no download in air-gapped or read-only starts.

More figures and methods: Performance. Retrieval quality is on the benchmark page.

Security hardening

Security →

Reliability and operations

Operations → · Models and Docker →

Going back is one command

dakera downgrade, shipped in the v0.12.0 binary, converts a stopped deployment's data back to what v0.11.108 reads: encrypted values, full-text indexes, the write-ahead log, tiered writes and caches. No v0.11 patch release is needed. It refuses, changing nothing, when a server still runs on the data or when the data uses something v0.11.108 lacks (bge-m3, colbert-small, the visual lane).

docker run --rm <env and volumes> <v0.12 image> downgrade

Rolling back to v0.11 →

What changes when you upgrade

An unchanged v0.11.108 deployment upgrades in place. Check these before you switch:

Stored memories are re-embedded once in the background after the upgrade (about 35 minutes per 10,000 memories on bge-large); recall is served throughout. The full list, with what to do for each item, is in the upgrade guide and the changelog.

Get started

Reference: Configuration · REST API · Telemetry · High availability