Search and ranking
This page is for developers and operators who tune retrieval or depend on scores. It covers the retrieval and ranking controls of v0.12.0: late interaction (MUVERA first stage and MaxSim rerank), the RaBitQ search mode, the rerank controls and rerank_report, GET /v1/capabilities, and the ranking changes that existing clients can notice (the smart score, recall admission, distance metric resolution). Late interaction and RaBitQ are opt-in; the rest applies to every v0.12.0 server.
Related pages: Multilingual search, Multimodal memory (the visual lane uses late interaction), Records and multi-vector, Configuration and Upgrading from v0.11.108.
At a glance
| Control | Variable or field | Default | Effect |
|---|---|---|---|
| Late interaction | DAKERA_ with DAKERA_ |
single-vector |
Recall shortlists by a fixed-dimensional encoding (FDE) and reranks with MaxSim over per-token vectors |
| RaBitQ | DAKERA_, DAKERA_ |
hybrid, 1 bit |
Quantized graph walk, then an exact float re-rank; a latency option |
| Rerank candidate cap | rerank_ (recall), DAKERA_ |
whole pool | Bounds how many candidates the cross-encoder scores |
| Rerank report | rerank_report in recall and search responses |
always present | What the cross-encoder did, and why not when it did not run |
| Reranker | DAKERA_, DAKERA_ |
bge-, warm-up on |
Which cross-encoder runs, and whether it loads at start |
| Capabilities | GET / |
always on | What this server supports and its late-interaction counters |
Late interaction
What it is
With the default single-vector strategy a memory is one dense vector. With late interaction a memory also keeps one vector per token. Recall:
- embeds the query into token vectors;
- shortlists candidates with a fixed-dimensional encoding (FDE, the MUVERA construction) of each memory, which turns the token set into one vector that can be searched with an ordinary index;
- reranks the shortlist with MaxSim: for each query token the best matching memory token, summed over the query tokens.
The FDE parameters are fixed in the build (k_sim 6, d_proj 8, reps 20, so an FDE has 10 240 components), and the first stage shortlists 100 candidates (the MaxSim rerank scores max(top_k, 100)).
When to use it
When token-level matching matters more than a single pooled vector, and you accept an opt-in model with its own constraints (below). It is opt-in; measure it on your own data before switching.
Configuration
| Variable | Value | Notes |
|---|---|---|
DAKERA_ |
single-vector (default), late- |
An unrecognized value keeps single-vector and is reported once; on a fresh install a typo refuses the start |
DAKERA_MODEL |
colbert-small |
96-dimensional token vectors; the one embedding model with a late-interaction recipe |
The store records the model that embedded it. Starting an existing bge-large store with colbert-small is refused until you migrate: start once with DAKERA_ALLOW_MODEL_CHANGE=1, then POST /admin/namespaces/migrate-dimensions with reembed_same_dimension: true (GET /v1/capabilities reports reembed_pending until it completes). Or start on a fresh store. See Configuration.
With the visual lane (DAKERA_VISION), the same strategy reads the patch and patch.fde slots of indexed pages instead; see Multimodal memory.
How it behaves
- Slots. A memory written under the strategy carries two record extras next to its primary vector:
colbert(the token vectors, stored as f16) andcolbert.fde(the document FDE). The visual lane usespatchandpatch.fde. The primary vector is the same onesingle-vectorwould store for the same text. - The first stage is never built on the query path. The FDE index of a namespace is built in the background (one shard per core) and persisted in the ANN index store, so a restart loads it instead of rebuilding. Until it is ready, recall serves the dense first stage and still reranks with MaxSim.
- First-stage reporting. Recall answers
late_interaction_first_stage(absent undersingle-vector) with the stage that served its first search:fde_graph(cached FDE graph, above the ANN threshold),fde_exact(cached exact scan, at or below it),dense_while_building(the FDE stage is still being built),dense_no_fde(the namespace holds no usable FDE of the query's width) orcold_fallback(the FDE stage was built from a partial scan because the cold tier was unreachable). The Prometheus counterdakera_late_interaction_first_stage_total{stage}counts the same stages. - Latency. Steady recall p50 was 0.10 s at 1 000 memories and 0.23 s at 10 000. A recall right after a bulk ingest is served by the dense first stage while the late-interaction stage rebuilds.
Constraints
- It cannot be combined with
DAKERA_TIERED=1(the tiered engine pins a static write tier with no token vectors). - A configuration that cannot produce token vectors (a model without a late-interaction recipe, or
DAKERA_TIERED=1) is answered501 NOT_IMPLEMENTED, with a message naming the settings to change. It is a501, not a503, because a retry cannot fix it. Memory stores and recalls both return it. GET /v1/capabilitiesshows the state before you depend on it:scoring.late_interaction.enabledis true when the strategy is selected, andmodel_supportedis false when the active model has no recipe.- Deployments on
colbert-smallor the visual lane cannot be rolled back to v0.11.108; move to a v0.11 model first (Rolling back to v0.11).
Verify it works
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" \
| jq '{strategy: .scoring.strategy, li: .scoring.late_interaction, stats: .late_interaction_stats | {searches, reranked, dense_fallbacks, fde_background_builds}}'
After a few recalls, searches grows and reranked grows with it. reranked at 0 while searches is above 0 means MaxSim never ran; dense_fallbacks equal to searches means no candidate had a usable token block. The counters are described under Capabilities.
RaBitQ search mode
What it is
An alternative to the default hybrid quantized search. The HNSW graph is walked on 1-bit (or extended, up to 8-bit) RaBitQ codes with an unbiased distance estimator (random Hadamard rotation plus a per-code correction), the shortlist is over-selected, and the final order comes from an exact float re-rank. Answered scores are exact. It is metric-aware: cosine, euclidean and dot product each use their own estimator.
When to use it
When search latency matters and the memory for the codes is acceptable. It is not a memory saving: the codes live in memory next to the float vectors, and they are rebuilt after a restart.
Configuration
| Variable | Accepted | Default | Notes |
|---|---|---|---|
DAKERA_ |
hybrid, binary, float, scalar (alias sq), rabitq |
hybrid |
Process-wide; not selectable per request. An unrecognized value keeps hybrid and is reported |
DAKERA_ |
1 to 8 |
1 |
Above 1 is Extended RaBitQ. An out-of-range number is clamped to 1..8 with a startup warning; a non-number is an invalid value |
The shortlist over-selection depends on the code width: the graph beam is ef_search times 10, 6, 4, 3 and 2 for 1, 2, 3, 4 and 5 to 8 bits (more bits give a tighter estimate, so less surplus is needed). ef_search is DAKERA_HNSW_EF_SEARCH (default 100).
Verify it works
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" | jq '{search_mode, search_modes_accepted}'
# {"search_mode":"rabitq","search_modes_accepted":"hybrid, binary, float, scalar (alias sq), rabitq"}
Rerank controls and rerank_report
Recall (POST /v1/memory/recall) reranks its candidates with a cross-encoder by default (rerank defaults to true). POST /v1/memory/search reranks only when the request sets rerank: true.
Capping the work
The cross-encoder's cost grows with the number of candidates it scores, not with top_k.
rerank_candidateson a recall request: score at most this many candidates, the best by retrieval score. The rest keep their retrieval order just below the scored ones.DAKERA_RERANK_MAX_CANDIDATES: a server-wide cap on the same number, for recall and/v1/memory/search. It must be a positive integer. A request value above it is lowered to it.- Unset, both mean the whole candidate pool, as in v0.11.
curl -s -X POST localhost:3000/v1/memory/recall -H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"agent_id": "my-agent", "query": "what did we decide about the launch date?", "top_k": 5, "rerank_candidates": 24}'
The report
Recall and /v1/memory/search answer rerank_report; recall also answers the same object as rerank (the name used in the v0.12 pre-releases).
"rerank_report": { "applied": true, "candidates": 24 }
| Field | Meaning |
|---|---|
applied |
The cross-encoder rescored candidates and its order was used |
candidates |
How many candidates it scored |
skipped |
Present when it did not apply: the reason |
skipped values:
| Value | Reason |
|---|---|
disabled |
The request set rerank: false (or a search did not set rerank: true) |
no_query |
A search without a query |
no_candidates |
Retrieval returned nothing to rerank |
bm25_routing |
A keyword-routed (BM25) query that is not temporal: its retrieval order is already reliable, so it is not reranked |
low_confidence |
The cross-encoder ran but its confidence gate reverted to the retrieval order (temporal inference queries); applied is then false and candidates is what it scored |
visual_lane |
The visual late-interaction lane is on; the text cross-encoder does not apply |
unavailable |
The reranker could not be loaded (it also shows in /health degraded) |
failed |
The reranker failed on this request |
Choosing and warming the reranker
| Variable | Accepted | Default |
|---|---|---|
DAKERA_ |
bge- (alias v2-m3), bge- (alias base) |
bge- |
DAKERA_ |
false or 0 turns the start-up warm-up off |
on |
The default reranker ships in the image. It loads in the background after the server is ready: until then reranked recalls are served without the rerank (skipped: "unavailable") and /health lists reranker under degraded. Turn the warm-up off only for deployments that never rerank. bge-reranker-base is a smaller English-centric alternative (about 279 MB download). See Models and Docker.
Resource implications
Measured CPU-only in a 4 vCPU / 8 GiB container with the default reranker (bge-reranker-v2-m3, a 568M-parameter cross-encoder), on the same machine for both versions: a top_k 16 recall took 11.5 s before v0.12.0 and 6.2 s now (GPU or hosted rerankers are a different deployment class), and container memory after reranking went from 4.98 GB to 0.65 GB, flat across top_k 1 to 32 (Performance). Lowering the candidate count lowers the CPU a recall takes in proportion.
Capabilities
GET /v1/capabilities (any key with Read scope) tells a client what the running server supports, before the client depends on it. Lists are generated from the server's registries; clients must ignore unknown fields and unknown strings. capabilities_version changes only when the document is reshaped in a breaking way (it is 1).
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY"
{
"capabilities_version": 1,
"server_version": "0.12.0",
"api_versions": ["v1"],
"default_model": "bge-large",
"models": [
{ "name": "bge-large", "aliases": ["bge-large-en", "bge-large-en-v1.5"], "dimension": 1024,
"max_seq_length": 512, "effective_max_seq_length": 512, "active": true, "modality": "text" },
{ "name": "modernbert-embed-base", "aliases": ["modernbert", "modern-bert"], "dimension": 768,
"max_seq_length": 8192, "effective_max_seq_length": 8192, "active": false,
"mrl_dimensions": [256, 768], "modality": "text" }
],
"index_kinds": ["..."],
"vector_index_kinds": ["..."],
"live_vector_index_kinds": ["hnsw"],
"distance_metrics": ["cosine", "euclidean", "dot_product"],
"search_mode": "hybrid",
"search_modes_accepted": "hybrid, binary, float, scalar (alias sq), rabitq",
"fulltext_language": "en",
"on_disk_format_version": 1,
"records": { "enabled": false, "representation_kinds": ["dense", "token_multivector", "patch_multivector"],
"dtypes": ["f32", "f16", "i8"], "max_representations": 8, "max_vectors": 4096, "max_bytes": 8388608 },
"scoring": { "strategy": "single-vector", "strategies_accepted": "single-vector, late-interaction",
"late_interaction": { "enabled": false, "model_supported": false, "lane": "text",
"token_slot": "colbert", "fde_slot": "colbert.fde", "fde_k_sim": 6, "fde_d_proj": 8,
"fde_reps": 20, "fde_dim": 10240, "candidates": 100 } },
"attachments": { "enabled": false, "max_bytes": 26214400,
"transcription": { "model": "whisper-tiny.en", "models": ["whisper-tiny.en"],
"media_types": ["audio/wav", "audio/x-wav"], "languages": ["en"], "sample_rate_hz": 16000 } },
"vision": { "enabled": false, "model": "colmodernvbert", "models": ["colmodernvbert"],
"media_types": ["image/png"], "dimension": 128, "patch_slot": "patch", "patch_fde_slot": "patch.fde",
"tile_size": 512, "longest_edge": 2048, "image_seq_len": 64, "max_tiles": 17 },
"query_languages": ["en", "de", "fr", "es", "it", "pt", "nl"],
"reembed_pending": false,
"unreadable_records": 0,
"late_interaction_stats": { "searches": 0, "fde_first_stage": 0, "dense_first_stage": 0, "reranked": 0, "...": 0 }
}
(The example is shortened: models lists every embedding model the build can load, and late_interaction_stats has more counters.)
| Field | Meaning |
|---|---|
default_model |
The model that embeds new memories (the default model under DAKERA_TIERED=1) |
models[] |
Every embedding model this build can load. active marks the one in use; a request naming another model is a 400. effective_ is what the server embeds before truncating (DAKERA_ applied), mrl_dimensions is present only for Matryoshka models |
live_ |
The index kinds the engine actually builds (HNSW) |
search_mode, search_ |
The configured DAKERA_ and the accepted values |
fulltext_ |
The analyzer language of new full-text indexes (DAKERA_) |
records |
Whether the record routes answer (DAKERA_RECORDS) and what they accept |
scoring |
The strategy and the late-interaction lane: lane is text or visual (chosen by DAKERA_VISION), the slot names, the FDE parameters and the first-stage shortlist size |
attachments, vision |
Whether each lane is on and its limits; with the lane off the routes answer 501 FEATURE_ |
query_languages |
The languages the per-request lang field accepts; any other value is a 400 |
reembed_pending |
A model change was acknowledged but the store is not fully re-embedded yet: recall mixes two embedding spaces until POST / with reembed_ completes |
unreadable_ |
Records this node skipped because a newer Dakera wrote them. Non-zero means searches can return fewer results than the store holds until the node is upgraded |
late_ |
Counters since start: searches, fde_first_stage, dense_, dense_, reranked, dense_fallbacks, foreign_, the FDE build counters (fde_, fde_, fde_, fde_loads, fde_saves, dense_, ...) and build timings in milliseconds. All zero on a single-vector deployment |
Ranking changes that affect clients
These apply to every deployment that upgrades from v0.11.108, with no setting.
- The smart score.
smart_score = w_vec·relevance + w_imp·importance + w_rec·recency. The default-route weights are 0.55, 0.20 and 0.13 (sum 0.88; the importance-weighted route also sums to 0.88). Multi-hop queries sum to 0.73 and temporal-inference queries to 0.98. v0.11.108 gave the remainder to an access-count term, which is gone. - Recency counts from
created_at, not fromlast_accessed_at. - A recall no longer raises the importance of what it returns. v0.11 added
0.05 + 0.05 × importanceto every memory a recall returned, so whatever one recall returned ranked higher in every later one. A recall still recordsaccess_countandlast_accessed_at(decay and wake-up use them);POST /v1/memory/feedbackstill raises importance. smart_scorevalues change scale. A memory that no recall, feedback or session promotion had touched scores exactly as before; one that v0.11 recalls had returned can score up to 0.27 lower, and one last used long after it was stored up to 0.13 lower. Re-check any absolute cut calibrated on v0.11:DAKERA_ABSTAIN_MIN_SMART_SCORE(unset by default; it drops results below the value) and any client-side filter on the score. Importance boosted by past v0.11 recalls is kept and decays from there.- Importance decay no longer compounds. A decay pass follows the configured curve from where the importance started, so decayed importances are higher than before. An importance set by a client or by feedback decays from the moment it was set.
Recall admission
At most two recalls per CPU core (and at least four) run at once. The rest queue in arrival order. A request that has waited 20 s in total (at most half of DAKERA_REQUEST_TIMEOUT) is answered 503 with Retry-After, below the official SDKs' 30 s client timeout, so the 503 arrives first. The limit is derived from the core count; there is no setting. POST /v1/memories/recall/batch takes an admission slot too, and cuts a limit above 10 000 (answering truncated: true).
Distance metric resolution
An omitted distance_metric resolves to the namespace's metric (request value, then the namespace's distance, then cosine) on query, batch query, multi-vector, unified, hybrid and explain requests and on gRPC Query (an empty gRPC distance_metric means the namespace's). In v0.11 an omitted value was always cosine. Memory recall is unaffected. Accepted values are cosine, euclidean and dot_product.
Other response fields
| Field | Where | Meaning |
|---|---|---|
effective_top_k |
POST / |
The top_k the search could honor. A search with a query ranks at most 500 candidates, so a larger top_k returns at most 500 memories; without a query it equals top_k |
x-dakera-limit |
Response header of listings that cap their page | The page size actually used (for example the agent memory listing at 1 000 and the session listings) |
rerank_report, rerank |
Recall and search | See above |
late_ |
Recall under late interaction | See above |
resource |
Every 404 body |
What was not found |