NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

Performance in v0.12.0

This page lists the performance figures Dakera publishes for v0.12.0, the conditions each was measured under, and what the "before" column compares against. It also lists the changes that have no published figure, the measurements still to come, and how to measure your own deployment. It is for operators sizing a deployment or deciding whether to upgrade from v0.11.108.

How to read the numbers:

Reranked recall

Measured CPU-only, in a container limited to 4 vCPUs and 8 GiB of memory, with the default reranker (bge-reranker-v2-m3, a 568M-parameter cross-encoder) scoring the candidates of a top_k 16 recall, compared with v0.11.108 on the same machine. These are CPU figures: a GPU or a hosted reranking service is a different deployment class, and its latencies are not comparable with these. The CPU reranker is now a pool of single-threaded workers over one memory-mapped copy of the model, scoring one pair per run, so scores are deterministic.

What Before v0.12.0
Recall with reranking, top_k 16 11.5 s, 45 CPU-s 6.2 s, 24 CPU-s
Two concurrent reranked recalls 19.6 s 11.2 s
Container memory after reranking 4.98 GB 0.65 GB, flat across top_k 1 to 32
Reranker load time 4.6 s 0.6 s

To bound the work a rerank does per request, see rerank_candidates and rerank_report in Search and ranking.

HNSW memory and latency

The in-memory HNSW graph now keeps node vectors in f16 (with a SIMD kernel) instead of f32, and re-scores the returned candidates against the stored f32 vectors, so the scores in a response are exact. There is no setting and no on-disk format change.

Measured on BEIR Quora, 50,000 vectors of 1024 dimensions:

What Before (f32) v0.12.0 (f16 nodes)
Memory per vector 4.81 KB 2.76 KB (−43 %)
recall@10 0.998 0.998 (unchanged)
Search p50 1.07 ms 0.87 ms
Search p95 1.44 ms 1.17 ms

The recall@10 here is approximate-nearest-neighbor recall against exact search on a public dataset. It shows that the lower-precision graph did not cost search recall; it is not a memory-quality score.

Recall, ingest and memory in a paired run

The same conversations were run through v0.11.108 and v0.12.0 on the same harness: LoCoMo, three conversations, 382 questions.

What v0.11.108 v0.12.0
Recall latency, p50 3.8 s 1.8 s
Ingest time 124 s 46 s
Server memory 2.4 GB 0.97 GB

This comparison covers latency, ingest time and memory. Answer quality is checked separately: the release gate compares retrieval on the same paired run and blocks publishing when it fails.

Embedding

No figure is published for these changes; they explain where the gains above come from.

Index persistence and rebuilds

Late interaction

Opt-in (DAKERA_SCORING_STRATEGY=late-interaction with DAKERA_MODEL=colbert-small; see Search and ranking):

What Value
Steady recall p50 at 1,000 memories 0.10 s
Steady recall p50 at 10,000 memories 0.23 s

A recall right after a bulk ingest is served by the dense first stage while the late-interaction stage rebuilds in the background; the response reports which stage it used. Late interaction is opt-in: switch it on where token-level precision matters.

Visual indexing

With DAKERA_VISION on, indexing a document page takes about 10.7 s on CPU. Pages are indexed one at a time. Measured retrieval quality on the visual lane is listed in Multimodal memory.

Memory footprint with every model loaded

With attachments, speech to text, image indexing and the other optional models enabled, the measured footprint is 530 MiB of process heap (anonymous memory, excluding the memory-mapped model files) with every model loaded. Model weights are memory-mapped (page cache: shared and reclaimable); anonymous memory is what the process holds privately. The optional vision, speech-to-text and entity-extraction models are unloaded after DAKERA_HNSW_CACHE_TTI_SECS (default 3600 s) without use.

Container image

The CPU image (ghcr.io/dakera-ai/dakera, amd64) is about 0.8 GB (783 MB measured). Measured on a Docker run: the server is ready 4.4 s after start, and an air-gapped or read-only start downloads nothing because the default models ship in the image. The x64 release binary is 106,228,848 bytes (about 101 MB).

Re-embedding after the upgrade

On the first start after an upgrade, stored memories that may carry the old query-side embedding are re-embedded in the background. Recall is served throughout. The time it takes grows with the number of stored memories; Upgrading from v0.11.108 gives the expected rate. Progress: /health embed_migration, GET /admin/reembed/migration and the metric dakera_reembed_migration_remaining.

Changes without a published figure

These changes are in the release; no number is claimed for them.

How to measure it on your hardware

Dakera exposes Prometheus metrics at GET /metrics, which is unauthenticated by default for internal scraping (keep it off the public network), and at GET /v1/ops/metrics, which needs a global admin key. Useful series:

Metric What it tells you
dakera_http_request_duration_seconds Request latency by route; refusals (429, 401/403, 408, 413, 503) are counted too
dakera_http_requests_total Request counts by method, route and status
dakera_active_requests Requests in flight
dakera_recall_pipeline_stage_duration_seconds{stage} Where a recall spends its time, per stage (embed, search, storage_get, rerank, scoring and others)
dakera_ingest_stage_duration_seconds{stage} The same for the store and batch-store path
dakera_inference_duration_seconds Model inference time
dakera_model_load_duration_seconds{model}, dakera_model_loads_total Model load times and counts
dakera_memory_budget_reserved_bytes Working memory reserved by media jobs and entity extraction
dakera_late_interaction_first_stage_total{stage} Which first stage late-interaction recalls used
dakera_reembed_migration_remaining Memories left to re-embed after the upgrade
dakera_ann_index_persist_total{event} ANN index saves and loads

A short recipe:

  1. Ingest a representative sample, then run your real recall queries with curl -w '%{time_total}' or your client, at the concurrency you expect.
  2. Read the per-stage histograms (p50 and p95 per stage) to see whether embedding, search, storage or rerank dominates.
  3. Check GET /health for degraded components (for example a reranker that failed to load) before you read a latency result.
  4. Compare the container's memory from docker stats or the cgroup files with the figures above. Allow for page cache: model weights are memory-mapped.

Measure on the same machine before and after you change a setting, with the same data, and report both runs, as the tables above do.

Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.