NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

Search and ranking

This page is for developers and operators who tune retrieval or depend on scores. It covers the retrieval and ranking controls of v0.12.0: late interaction (MUVERA first stage and MaxSim rerank), the RaBitQ search mode, the rerank controls and rerank_report, GET /v1/capabilities, and the ranking changes that existing clients can notice (the smart score, recall admission, distance metric resolution). Late interaction and RaBitQ are opt-in; the rest applies to every v0.12.0 server.

Related pages: Multilingual search, Multimodal memory (the visual lane uses late interaction), Records and multi-vector, Configuration and Upgrading from v0.11.108.

At a glance

Control Variable or field Default Effect
Late interaction DAKERA_SCORING_STRATEGY=late-interaction with DAKERA_MODEL=colbert-small single-vector Recall shortlists by a fixed-dimensional encoding (FDE) and reranks with MaxSim over per-token vectors
RaBitQ DAKERA_SEARCH_MODE=rabitq, DAKERA_RABITQ_BITS hybrid, 1 bit Quantized graph walk, then an exact float re-rank; a latency option
Rerank candidate cap rerank_candidates (recall), DAKERA_RERANK_MAX_CANDIDATES whole pool Bounds how many candidates the cross-encoder scores
Rerank report rerank_report in recall and search responses always present What the cross-encoder did, and why not when it did not run
Reranker DAKERA_RERANKER_MODEL, DAKERA_RERANK_WARMUP bge-reranker-v2-m3, warm-up on Which cross-encoder runs, and whether it loads at start
Capabilities GET /v1/capabilities always on What this server supports and its late-interaction counters

Late interaction

What it is

With the default single-vector strategy a memory is one dense vector. With late interaction a memory also keeps one vector per token. Recall:

  1. embeds the query into token vectors;
  2. shortlists candidates with a fixed-dimensional encoding (FDE, the MUVERA construction) of each memory, which turns the token set into one vector that can be searched with an ordinary index;
  3. reranks the shortlist with MaxSim: for each query token the best matching memory token, summed over the query tokens.

The FDE parameters are fixed in the build (k_sim 6, d_proj 8, reps 20, so an FDE has 10 240 components), and the first stage shortlists 100 candidates (the MaxSim rerank scores max(top_k, 100)).

When to use it

When token-level matching matters more than a single pooled vector, and you accept an opt-in model with its own constraints (below). It is opt-in; measure it on your own data before switching.

Configuration

Variable Value Notes
DAKERA_SCORING_STRATEGY single-vector (default), late-interaction An unrecognized value keeps single-vector and is reported once; on a fresh install a typo refuses the start
DAKERA_MODEL colbert-small 96-dimensional token vectors; the one embedding model with a late-interaction recipe

The store records the model that embedded it. Starting an existing bge-large store with colbert-small is refused until you migrate: start once with DAKERA_ALLOW_MODEL_CHANGE=1, then POST /admin/namespaces/migrate-dimensions with reembed_same_dimension: true (GET /v1/capabilities reports reembed_pending until it completes). Or start on a fresh store. See Configuration.

With the visual lane (DAKERA_VISION), the same strategy reads the patch and patch.fde slots of indexed pages instead; see Multimodal memory.

How it behaves

Constraints

Verify it works

curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" \
  | jq '{strategy: .scoring.strategy, li: .scoring.late_interaction, stats: .late_interaction_stats | {searches, reranked, dense_fallbacks, fde_background_builds}}'

After a few recalls, searches grows and reranked grows with it. reranked at 0 while searches is above 0 means MaxSim never ran; dense_fallbacks equal to searches means no candidate had a usable token block. The counters are described under Capabilities.

RaBitQ search mode

What it is

An alternative to the default hybrid quantized search. The HNSW graph is walked on 1-bit (or extended, up to 8-bit) RaBitQ codes with an unbiased distance estimator (random Hadamard rotation plus a per-code correction), the shortlist is over-selected, and the final order comes from an exact float re-rank. Answered scores are exact. It is metric-aware: cosine, euclidean and dot product each use their own estimator.

When to use it

When search latency matters and the memory for the codes is acceptable. It is not a memory saving: the codes live in memory next to the float vectors, and they are rebuilt after a restart.

Configuration

Variable Accepted Default Notes
DAKERA_SEARCH_MODE hybrid, binary, float, scalar (alias sq), rabitq hybrid Process-wide; not selectable per request. An unrecognized value keeps hybrid and is reported
DAKERA_RABITQ_BITS 1 to 8 1 Above 1 is Extended RaBitQ. An out-of-range number is clamped to 1..8 with a startup warning; a non-number is an invalid value

The shortlist over-selection depends on the code width: the graph beam is ef_search times 10, 6, 4, 3 and 2 for 1, 2, 3, 4 and 5 to 8 bits (more bits give a tighter estimate, so less surplus is needed). ef_search is DAKERA_HNSW_EF_SEARCH (default 100).

Verify it works

curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" | jq '{search_mode, search_modes_accepted}'
# {"search_mode":"rabitq","search_modes_accepted":"hybrid, binary, float, scalar (alias sq), rabitq"}

Rerank controls and rerank_report

Recall (POST /v1/memory/recall) reranks its candidates with a cross-encoder by default (rerank defaults to true). POST /v1/memory/search reranks only when the request sets rerank: true.

Capping the work

The cross-encoder's cost grows with the number of candidates it scores, not with top_k.

curl -s -X POST localhost:3000/v1/memory/recall -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"agent_id": "my-agent", "query": "what did we decide about the launch date?", "top_k": 5, "rerank_candidates": 24}'

The report

Recall and /v1/memory/search answer rerank_report; recall also answers the same object as rerank (the name used in the v0.12 pre-releases).

"rerank_report": { "applied": true, "candidates": 24 }
Field Meaning
applied The cross-encoder rescored candidates and its order was used
candidates How many candidates it scored
skipped Present when it did not apply: the reason

skipped values:

Value Reason
disabled The request set rerank: false (or a search did not set rerank: true)
no_query A search without a query
no_candidates Retrieval returned nothing to rerank
bm25_routing A keyword-routed (BM25) query that is not temporal: its retrieval order is already reliable, so it is not reranked
low_confidence The cross-encoder ran but its confidence gate reverted to the retrieval order (temporal inference queries); applied is then false and candidates is what it scored
visual_lane The visual late-interaction lane is on; the text cross-encoder does not apply
unavailable The reranker could not be loaded (it also shows in /health degraded)
failed The reranker failed on this request

Choosing and warming the reranker

Variable Accepted Default
DAKERA_RERANKER_MODEL bge-reranker-v2-m3 (alias v2-m3), bge-reranker-base (alias base) bge-reranker-v2-m3
DAKERA_RERANK_WARMUP false or 0 turns the start-up warm-up off on

The default reranker ships in the image. It loads in the background after the server is ready: until then reranked recalls are served without the rerank (skipped: "unavailable") and /health lists reranker under degraded. Turn the warm-up off only for deployments that never rerank. bge-reranker-base is a smaller English-centric alternative (about 279 MB download). See Models and Docker.

Resource implications

Measured CPU-only in a 4 vCPU / 8 GiB container with the default reranker (bge-reranker-v2-m3, a 568M-parameter cross-encoder), on the same machine for both versions: a top_k 16 recall took 11.5 s before v0.12.0 and 6.2 s now (GPU or hosted rerankers are a different deployment class), and container memory after reranking went from 4.98 GB to 0.65 GB, flat across top_k 1 to 32 (Performance). Lowering the candidate count lowers the CPU a recall takes in proportion.

Capabilities

GET /v1/capabilities (any key with Read scope) tells a client what the running server supports, before the client depends on it. Lists are generated from the server's registries; clients must ignore unknown fields and unknown strings. capabilities_version changes only when the document is reshaped in a breaking way (it is 1).

curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY"
{
  "capabilities_version": 1,
  "server_version": "0.12.0",
  "api_versions": ["v1"],
  "default_model": "bge-large",
  "models": [
    { "name": "bge-large", "aliases": ["bge-large-en", "bge-large-en-v1.5"], "dimension": 1024,
      "max_seq_length": 512, "effective_max_seq_length": 512, "active": true, "modality": "text" },
    { "name": "modernbert-embed-base", "aliases": ["modernbert", "modern-bert"], "dimension": 768,
      "max_seq_length": 8192, "effective_max_seq_length": 8192, "active": false,
      "mrl_dimensions": [256, 768], "modality": "text" }
  ],
  "index_kinds": ["..."],
  "vector_index_kinds": ["..."],
  "live_vector_index_kinds": ["hnsw"],
  "distance_metrics": ["cosine", "euclidean", "dot_product"],
  "search_mode": "hybrid",
  "search_modes_accepted": "hybrid, binary, float, scalar (alias sq), rabitq",
  "fulltext_language": "en",
  "on_disk_format_version": 1,
  "records": { "enabled": false, "representation_kinds": ["dense", "token_multivector", "patch_multivector"],
               "dtypes": ["f32", "f16", "i8"], "max_representations": 8, "max_vectors": 4096, "max_bytes": 8388608 },
  "scoring": { "strategy": "single-vector", "strategies_accepted": "single-vector, late-interaction",
    "late_interaction": { "enabled": false, "model_supported": false, "lane": "text",
      "token_slot": "colbert", "fde_slot": "colbert.fde", "fde_k_sim": 6, "fde_d_proj": 8,
      "fde_reps": 20, "fde_dim": 10240, "candidates": 100 } },
  "attachments": { "enabled": false, "max_bytes": 26214400,
    "transcription": { "model": "whisper-tiny.en", "models": ["whisper-tiny.en"],
      "media_types": ["audio/wav", "audio/x-wav"], "languages": ["en"], "sample_rate_hz": 16000 } },
  "vision": { "enabled": false, "model": "colmodernvbert", "models": ["colmodernvbert"],
    "media_types": ["image/png"], "dimension": 128, "patch_slot": "patch", "patch_fde_slot": "patch.fde",
    "tile_size": 512, "longest_edge": 2048, "image_seq_len": 64, "max_tiles": 17 },
  "query_languages": ["en", "de", "fr", "es", "it", "pt", "nl"],
  "reembed_pending": false,
  "unreadable_records": 0,
  "late_interaction_stats": { "searches": 0, "fde_first_stage": 0, "dense_first_stage": 0, "reranked": 0, "...": 0 }
}

(The example is shortened: models lists every embedding model the build can load, and late_interaction_stats has more counters.)

Field Meaning
default_model The model that embeds new memories (the default model under DAKERA_TIERED=1)
models[] Every embedding model this build can load. active marks the one in use; a request naming another model is a 400. effective_max_seq_length is what the server embeds before truncating (DAKERA_MAX_SEQ_LENGTH applied), mrl_dimensions is present only for Matryoshka models
live_vector_index_kinds The index kinds the engine actually builds (HNSW)
search_mode, search_modes_accepted The configured DAKERA_SEARCH_MODE and the accepted values
fulltext_language The analyzer language of new full-text indexes (DAKERA_FULLTEXT_LANGUAGE)
records Whether the record routes answer (DAKERA_RECORDS) and what they accept
scoring The strategy and the late-interaction lane: lane is text or visual (chosen by DAKERA_VISION), the slot names, the FDE parameters and the first-stage shortlist size
attachments, vision Whether each lane is on and its limits; with the lane off the routes answer 501 FEATURE_DISABLED
query_languages The languages the per-request lang field accepts; any other value is a 400
reembed_pending A model change was acknowledged but the store is not fully re-embedded yet: recall mixes two embedding spaces until POST /admin/namespaces/migrate-dimensions with reembed_same_dimension completes
unreadable_records Records this node skipped because a newer Dakera wrote them. Non-zero means searches can return fewer results than the store holds until the node is upgraded
late_interaction_stats Counters since start: searches, fde_first_stage, dense_first_stage, dense_union_first_stage, reranked, dense_fallbacks, foreign_fde_slots, the FDE build counters (fde_index_builds, fde_exact_builds, fde_background_builds, fde_loads, fde_saves, dense_while_building, ...) and build timings in milliseconds. All zero on a single-vector deployment

Ranking changes that affect clients

These apply to every deployment that upgrades from v0.11.108, with no setting.

Recall admission

At most two recalls per CPU core (and at least four) run at once. The rest queue in arrival order. A request that has waited 20 s in total (at most half of DAKERA_REQUEST_TIMEOUT) is answered 503 with Retry-After, below the official SDKs' 30 s client timeout, so the 503 arrives first. The limit is derived from the core count; there is no setting. POST /v1/memories/recall/batch takes an admission slot too, and cuts a limit above 10 000 (answering truncated: true).

Distance metric resolution

An omitted distance_metric resolves to the namespace's metric (request value, then the namespace's distance, then cosine) on query, batch query, multi-vector, unified, hybrid and explain requests and on gRPC Query (an empty gRPC distance_metric means the namespace's). In v0.11 an omitted value was always cosine. Memory recall is unaffected. Accepted values are cosine, euclidean and dot_product.

Other response fields

Field Where Meaning
effective_top_k POST /v1/memory/search The top_k the search could honor. A search with a query ranks at most 500 candidates, so a larger top_k returns at most 500 memories; without a query it equals top_k
x-dakera-limit Response header of listings that cap their page The page size actually used (for example the agent memory listing at 1 000 and the session listings)
rerank_report, rerank Recall and search See above
late_interaction_first_stage Recall under late interaction See above
resource Every 404 body What was not found
Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.