What's new in v0.12.0
Dakera v0.12.0, released on 2026-10-01, adds multilingual search, multimodal memory, multi-vector records and late interaction; makes recall faster and leaner; hardens security and cluster replication; and goes back to v0.11.108 with one command. This page summarizes the release for developers and operators deciding what to use and how to upgrade.
Every new feature is opt-in. An unchanged v0.11.108 deployment upgrades in place: stop v0.11.108, start v0.12.0 on the same data and environment. The operator guide is Upgrading from v0.11.108, and every change is listed in the Changelog.
Highlights
bge-m3 embeddings, full-text in 18 stemmed languages plus CJK bigrams, dates in seven languages.
Multimodal memoryAttachments, speech to text and visual recall over document pages.
Multi-vector recordsOne indexed vector plus named token, patch or dense representations.
Late interactioncolbert-small with a MUVERA first stage and MaxSim rerank; RaBitQ search mode.
Faster, leaner engineVector search 1.17 ms p95 at recall@10 0.998; 7.6× less memory after reranking; reranking about 2× faster on the same CPU.
Security hardeningAuthenticated gRPC and cluster traffic, record-bound encryption, tighter permissions.
One-command rollbackdakera downgrade converts the data back to what v0.11.108 reads.
Multilingual search
Four independent, opt-in layers. A server that sets none of them behaves as v0.11.108 did (English).
- Embeddings:
DAKERA_MODEL=bge-m3puts text from 100+ languages in one 1024-dimensional space (XLM-RoBERTa, CPU ONNX backend only). - Full-text:
DAKERA_FULLTEXT_LANGUAGEselects per-language stemming and stop words;DAKERA_FULLTEXT_CJK_BIGRAMSindexes Chinese, Japanese, Korean, Thai and similar scripts as character bigrams (v0.11.108 did not index such text at all).POST /admin/fulltext/reindexswitches an existing namespace. - Dates and query routing in seven languages (
DAKERA_QUERY_LANG:en,de,fr,es,it,pt,nl, orauto). - Per-request
langon recall, search, store, batch store, update and extraction.
Details: Multilingual search.
Multimodal memory
Off by default; each feature turns on with one variable.
- Attachments (
DAKERA_ATTACHMENTS): files stored per namespace, content-addressed, referenced from memories, counted by quotas, backed up and replicated. - Speech to text (
DAKERA_WHISPER_MODEL): a WAV attachment becomes a memory through a background job. Five models:whisper-base(the default, multilingual, recommended),whisper-tiny.en(lightest English),whisper-base.en(better English),whisper-tiny(lightest multilingual) andwhisper-small(quality). The multilingual models detect the spoken language and record it as the memory'slang. See the model table. - Image indexing and visual recall (
DAKERA_VISION,colmodernvbert): recall over document pages by the text of a question. ViDoRe nDCG@5: TabFQuAD 0.643, Shift Project 0.7705 on CPU.
Every media job reserves the memory it needs before it starts, and waits or answers 503 with Retry-After instead of risking an out-of-memory kill. With every model loaded, the process heap measured 530 MiB (memory-mapped model files are not counted). Details: Multimodal memory.
Multi-vector records
DAKERA_RECORDS turns on POST /v1/namespaces/{ns}/records and GET /v1/namespaces/{ns}/records/{id}. A record is one indexed vector plus named extra representations (dense, token_multivector, patch_multivector, stored as f32, f16 or i8), kept atomically beside it by every storage backend, snapshots, export/import and replication. Use records when you compute multi-vector embeddings yourself and want Dakera to keep them with the document. Details: Records and multi-vector.
Late interaction and search controls
- Late interaction:
DAKERA_MODEL=colbert-smallwithDAKERA_SCORING_STRATEGY=late-interactionshortlists by a fixed-dimensional encoding (MUVERA) and reranks with per-token MaxSim. Steady recall p50 is 0.10 s at 1k memories and 0.23 s at 10k. Switch it on where token-level precision matters. - RaBitQ search mode (
DAKERA_SEARCH_MODE=rabitq,DAKERA_RABITQ_BITS1 to 8) for lower search latency. - Rerank controls:
rerank_candidatesper request andDAKERA_RERANK_MAX_CANDIDATESserver-wide cap how many candidates the cross-encoder scores. Recall and search answer arerank_report(applied,candidates,skipped). GET /v1/capabilitiestells clients what the server supports before they depend on it: models, index kinds, on-disk format version, query languages, the scoring strategy, the multimodal and record switches, andreembed_pending.
Details: Search and ranking.
Performance
Every figure is a measurement; "Before" is the same measurement without the change (v0.11.108's behavior unless noted). More in Performance.
| What | Before | v0.12.0 | Measured on |
|---|---|---|---|
| Vector search p50 / p95 at recall@10 0.998 | 1.07 / 1.44 ms | 0.87 / 1.17 ms | BEIR Quora, 50k vectors, CPU only |
Reranked recall, top_k 16 |
baseline | about 2× faster on the same CPU, no GPU | 4 vCPU / 8 GiB container, bge- |
| Container memory after reranking | 4.98 GB | 0.65 GB, flat across top_k 1 to 32 |
same |
| Reranker load | 4.6 s | 0.6 s | same |
| HNSW memory per 1024-d vector | 4.81 KB | 2.76 KB (−43 %), same recall@10 (0.998) | BEIR Quora, 50k vectors |
| Ingest speed / memory | baseline / 2.4 GB | 2.7× faster / 0.97 GB | Same workload on v0.11.108 and v0.12.0, same machine |
| CPU Docker image (amd64) | n/a | 783 MB; ready 4.4 s after start, with no download in air-gapped or read-only starts | Docker run of the release image |
| Release binary (x64) | n/a | 106,228,848 bytes (about 101 MB) | Release build |
Also: every CPU model runs on one memory-mapped copy with no idle arena; stored vectors no longer depend on what else was in the batch; ANN indexes are saved and reloaded instead of rebuilt; an expired memory no longer forces an index rebuild; concurrent writes into one namespace share the write-ahead log's fsync.
Security hardening
- gRPC requires an API key and is rate limited, time-bounded and audited like REST.
- Encrypted data is bound to its record and keys live in a replicated keyring: rotate one namespace at a time, and old backups still restore. Existing values are re-sealed in the background.
- Namespace-bound decryption: memory content is opened only under the namespace it was read from, and clients cannot store sealed values.
- Tighter permissions: keys pinned to namespaces no longer reach node-wide admin routes; backup download, upload and restore need global
super_admin. - Cluster traffic is authenticated with
DAKERA_CLUSTER_SECRET. - Safer defaults: with authentication off, only local web pages may call the server; knowledge graph edges are private to their agent; imports, extractor responses and graph searches are bounded; secrets are kept out of logs and
/health. - Backups can be encrypted (AES-256-GCM) and compressed (zstd).
Details: Security and Upgrading: security and access.
One-command rollback
dakera downgrade ships in the v0.12.0 binary. Run it once on the stopped deployment, with the server's environment and volumes, and it converts the data back to what v0.11.108 reads: encrypted values, full-text indexes, the write-ahead log, tiered writes, caches and re-embedded models included. No v0.11 patch release is needed.
stop v0.12.0 -> dakera downgrade -> start v0.11.108
- It prints a JSON report and exits
0(done),1(something still needs fixing; run it again) or78(refused, nothing changed). - It refuses while a server still runs on the data, and for deployments on
bge-m3,colbert-smallor the visual lane until they are moved to a v0.11 model.
Details: Rolling back to v0.11.
Reliability and operations
- Cluster replication is versioned and merges instead of overwriting, with a durable outbox and tombstones. See High availability.
- Tiered storage keeps a durable local hot tier and keeps serving through object-store outages.
- Startup: the server binds its port while models download, with
/health/liveand/health/ready;/healthaddsdegraded,config_warningsandembed_migration. dakera --check-configruns the startup configuration check alone and exits0(would start) or78(would refuse).dakera models list | pull | pruneshows, pre-fetches, bakes into an image or cleans models. The images ship their default models, so a default start needs no network. See Models and Docker.- Observability: Prometheus alert rules and a rebuilt Grafana dashboard, published in dakera-deploy
monitoring/with the dakera-deploy v0.12.0 release. See Operations. - Data lifecycle: backups, TTL, forget, deduplication and consolidation are hardened end to end.
Telemetry disclosure
When telemetry is on, v0.12.0 sends lifecycle events and a 6-hourly heartbeat. The machine's hostname is sent, and PostHog keeps the connecting IP address and derives a coarse location from it (GeoIP). Opt out with DAKERA_TELEMETRY=0 or DO_NOT_TRACK=1. See Telemetry.
What changes for existing deployments
Read these before upgrading. Each is explained, with what to do, in Upgrading from v0.11.108.
- gRPC clients need an API key in the call metadata (
x-api-keyorauthorization: Bearer) when authentication is on. - Permissions: keys pinned to namespaces get
403on node-wide routes; backup download, upload and restore need globalsuper_admin. - Namespace quotas are enforced. Quotas set on v0.11 were never checked; writes over a
hardquota now get413. - Ranking follows the query more closely. A recall no longer raises the importance of what it returns, the access count is no longer a score term, and recency counts from when a memory was created.
smart_scorevalues change scale. - A stored, enabled backup schedule starts running at its next slot (v0.11 stored it and never ran it).
- Health checks: the port answers while models load; point readiness at
/health/readyand liveness at/health/live. - Configuration: an upgraded deployment starts with warnings, never refusals (
/healthconfig_warnings); a fresh install with the same mistakes is refused. Rundakera --check-configfirst. - Stored memories are re-embedded once in the background (they were embedded with the query instruction): about 35 minutes per 10,000 memories on
bge-large; recall is served throughout. - Clusters: leave
DAKERA_CLUSTER_SECRETunset until every node runs v0.12.0, keep the mixed period short, and give each node its own S3 bucket. - Speech to text defaults to
whisper-base(multilingual, language auto-detected), not English-onlywhisper-tiny.en. Speech is opt-in; a deployment that turns it on withoutDAKERA_WHISPER_MODELtranscribes withwhisper-base. SetDAKERA_WHISPER_MODEL=whisper-tiny.ento keep the lightest English model. - Errors: every error body is JSON, every
503carriesRetry-After, configuration errors are501, over-size bodies are413. - Namespace config:
PATCHmerges and refuses unknown fields;PUTreplaces. - Telemetry, when on: lifecycle events and a 6-hourly heartbeat; the hostname is sent and the connecting IP is kept (see above).
POST /ops/shutdownreally shuts the server down (in v0.11 it only set a flag).
Operating notes
- Persistent storage: keep the data directory on a persistent volume, also when vectors are stored on S3.
- Visual store: use a dedicated data directory for a deployment that indexes document pages.
- Score thresholds:
smart_scorehas a new scale (w_vec * relevance + w_imp * importance + w_rec * recency); re-check any absolute cut on it. - Clients on v0.11 SDKs: use
GET /health/readyto wait for a starting server, andPUT /v1/namespaces/{ns}/configto clear a namespace'sentity_types. - After upgrading from v0.11.108: rebuild the full-text indexes once with
POST /admin/fulltext/reindex; see the upgrade step. - Restores: run them in a maintenance window; the server keeps serving while a restore runs.
Where to go next
- Upgrading from v0.11.108: the operator guide.
- Rolling back to v0.11:
dakera downgrade. - Changelog: every change in v0.12.0.
- Features: Multilingual search, Multimodal memory, Records and multi-vector, Search and ranking.
- Operations: Operations, Models and Docker, Performance, Telemetry.
- Reference: Configuration, Deployment, Security, High availability, REST API.