Performance in v0.12.0
This page lists the performance figures Dakera publishes for v0.12.0, the conditions each was measured under, and what the "before" column compares against. It also lists the changes that have no published figure, the measurements still to come, and how to measure your own deployment. It is for operators sizing a deployment or deciding whether to upgrade from v0.11.108.
How to read the numbers:
- Each figure is a measurement taken for this release. Nothing is extrapolated to other hardware, data sizes or models.
- "Before" is the same measurement without the change, which is v0.11.108's behavior unless the row says otherwise.
- A figure holds for the conditions stated with it. Your hardware, model choice and data will give different absolute numbers; see How to measure it on your hardware.
Reranked recall
Measured CPU-only, in a container limited to 4 vCPUs and 8 GiB of memory, with the default reranker (bge-reranker-v2-m3, a 568M-parameter cross-encoder) scoring the candidates of a top_k 16 recall, compared with v0.11.108 on the same machine. These are CPU figures: a GPU or a hosted reranking service is a different deployment class, and its latencies are not comparable with these. The CPU reranker is now a pool of single-threaded workers over one memory-mapped copy of the model, scoring one pair per run, so scores are deterministic.
| What | Before | v0.12.0 |
|---|---|---|
Recall with reranking, top_k 16 |
11.5 s, 45 CPU-s | 6.2 s, 24 CPU-s |
| Two concurrent reranked recalls | 19.6 s | 11.2 s |
| Container memory after reranking | 4.98 GB | 0.65 GB, flat across top_k 1 to 32 |
| Reranker load time | 4.6 s | 0.6 s |
To bound the work a rerank does per request, see rerank_candidates and rerank_report in Search and ranking.
HNSW memory and latency
The in-memory HNSW graph now keeps node vectors in f16 (with a SIMD kernel) instead of f32, and re-scores the returned candidates against the stored f32 vectors, so the scores in a response are exact. There is no setting and no on-disk format change.
Measured on BEIR Quora, 50,000 vectors of 1024 dimensions:
| What | Before (f32) |
v0.12.0 (f16 nodes) |
|---|---|---|
| Memory per vector | 4.81 KB | 2.76 KB (−43 %) |
| recall@10 | 0.998 | 0.998 (unchanged) |
| Search p50 | 1.07 ms | 0.87 ms |
| Search p95 | 1.44 ms | 1.17 ms |
The recall@10 here is approximate-nearest-neighbor recall against exact search on a public dataset. It shows that the lower-precision graph did not cost search recall; it is not a memory-quality score.
Recall, ingest and memory in a paired run
The same conversations were run through v0.11.108 and v0.12.0 on the same harness: LoCoMo, three conversations, 382 questions.
| What | v0.11.108 | v0.12.0 |
|---|---|---|
| Recall latency, p50 | 3.8 s | 1.8 s |
| Ingest time | 124 s | 46 s |
| Server memory | 2.4 GB | 0.97 GB |
This comparison covers latency, ingest time and memory. Answer quality is checked separately: the release gate compares retrieval on the same paired run and blocks publishing when it fails.
Embedding
No figure is published for these changes; they explain where the gains above come from.
- The CPU embedder embeds one text per run. Workers are single-threaded over one memory-mapped model. On real ingest this was faster than batching.
- A stored vector depends on its text only. With batched INT8 inference, a vector also depended on the other texts in its batch. Embedding one text per run removes that.
- One session configuration for every CPU model. The embedder, reranker, entity extractor, speech-to-text and visual models run without ONNX Runtime's arena, memory pattern, spin-waiting or prepacking, each over one memory-mapped copy of the model, so an idle model holds no per-session arena.
- Queries do not wait behind ingest. Queries take their own permits and a separate query session, so a batch ingest no longer stalls a recall into a
503. DAKERA_ONNX_POOL_SIZEandDAKERA_ONNX_BATCH_SIZEapply to the GPU path only.
Index persistence and rebuilds
- ANN indexes are saved and loaded at start instead of rebuilt. Writes, TTL sweeps, decay and the re-embed cycle update the cached index in place.
- An expired memory no longer forces a full index rebuild.
Late interaction
Opt-in (DAKERA_SCORING_STRATEGY=late-interaction with DAKERA_MODEL=colbert-small; see Search and ranking):
| What | Value |
|---|---|
| Steady recall p50 at 1,000 memories | 0.10 s |
| Steady recall p50 at 10,000 memories | 0.23 s |
A recall right after a bulk ingest is served by the dense first stage while the late-interaction stage rebuilds in the background; the response reports which stage it used. Late interaction is opt-in: switch it on where token-level precision matters.
Visual indexing
With DAKERA_VISION on, indexing a document page takes about 10.7 s on CPU. Pages are indexed one at a time. Measured retrieval quality on the visual lane is listed in Multimodal memory.
Memory footprint with every model loaded
With attachments, speech to text, image indexing and the other optional models enabled, the measured footprint is 530 MiB of process heap (anonymous memory, excluding the memory-mapped model files) with every model loaded. Model weights are memory-mapped (page cache: shared and reclaimable); anonymous memory is what the process holds privately. The optional vision, speech-to-text and entity-extraction models are unloaded after DAKERA_HNSW_CACHE_TTI_SECS (default 3600 s) without use.
Container image
The CPU image (ghcr.io/dakera-ai/dakera, amd64) is about 0.8 GB (783 MB measured). Measured on a Docker run: the server is ready 4.4 s after start, and an air-gapped or read-only start downloads nothing because the default models ship in the image. The x64 release binary is 106,228,848 bytes (about 101 MB).
Re-embedding after the upgrade
On the first start after an upgrade, stored memories that may carry the old query-side embedding are re-embedded in the background. Recall is served throughout. The time it takes grows with the number of stored memories; Upgrading from v0.11.108 gives the expected rate. Progress: /health embed_migration, GET /admin/reembed/migration and the metric dakera_reembed_migration_remaining.
Changes without a published figure
These changes are in the release; no number is claimed for them.
- Write-ahead log group commit. One shared
fdatasyncoutside the writer lock, and concurrent writes into one namespace share it. Durability is unchanged: a write is acknowledged only once it is durable. - Incremental BM25 persistence. A base blob plus an append-only, sealed delta log per namespace: a write stores only its changed documents instead of rewriting the whole index. Postings no longer carry unused term positions, and idle indexes are evicted.
- No whole-namespace scans on request paths. Forget, attachment lookups, session memories and the recency listing read a maintained in-memory index. Periodic jobs page with a cursor on every backend, RocksDB iterators are bounded, and deep agent-listing pages cost in proportion to
offset + limitrather than growing with its square. - Knowledge graph. Per-memory edge lookups are index probes, and recall's graph reads (personalized PageRank, supersede, traversals) are batched, namespace-scoped and run off the async workers.
- Memory backpressure counts only memory the kernel cannot reclaim and no longer drops ANN indexes. The ANN cache and the L1 vector cache are bounded by bytes.
- Cluster. Reconciliation reads nothing from an unchanged cluster (cached digests), outbox I/O runs on its own thread, and repeated recalls of one memory cost one storage write and one replication instead of one each.
- RaBitQ (
DAKERA_SEARCH_MODE=rabitq) is a latency mode for faster search.
How to measure it on your hardware
Dakera exposes Prometheus metrics at GET /metrics, which is unauthenticated by default for internal scraping (keep it off the public network), and at GET /v1/ops/metrics, which needs a global admin key. Useful series:
| Metric | What it tells you |
|---|---|
dakera_ |
Request latency by route; refusals (429, 401/403, 408, 413, 503) are counted too |
dakera_ |
Request counts by method, route and status |
dakera_ |
Requests in flight |
dakera_ |
Where a recall spends its time, per stage (embed, search, storage_get, rerank, scoring and others) |
dakera_ |
The same for the store and batch-store path |
dakera_ |
Model inference time |
dakera_, dakera_ |
Model load times and counts |
dakera_ |
Working memory reserved by media jobs and entity extraction |
dakera_ |
Which first stage late-interaction recalls used |
dakera_ |
Memories left to re-embed after the upgrade |
dakera_ |
ANN index saves and loads |
A short recipe:
- Ingest a representative sample, then run your real recall queries with
curl -w '%{time_total}'or your client, at the concurrency you expect. - Read the per-stage histograms (p50 and p95 per stage) to see whether embedding, search, storage or rerank dominates.
- Check
GET /healthfordegradedcomponents (for example a reranker that failed to load) before you read a latency result. - Compare the container's memory from
docker statsor the cgroup files with the figures above. Allow for page cache: model weights are memory-mapped.
Measure on the same machine before and after you change a setting, with the same data, and report both runs, as the tables above do.