Records and multi-vector
A record is one vector plus named extra representations. The one vector (the primary, a plain dense embedding) is what is indexed and searched, exactly like any vector written with POST /v1/namespaces/{ns}/vectors. The extras sit beside it and are stored, replicated, backed up and deleted with it: per-token vectors for late interaction (a token multivector), per-patch vectors for a page image (a patch multivector), or another dense vector. Use records when you compute multi-vector embeddings yourself and want Dakera to keep them with the document. Memories written by the server under late interaction, and image pages indexed through the multimodal lane, are records of the same shape.
The routes are off by default (DAKERA_RECORDS): with the switch off both answer 501 FEATURE_DISABLED before touching any state, and no other route changes.
Configuration
| Variable | Default | Meaning |
|---|---|---|
DAKERA_RECORDS |
off | Turns the two record routes on (boolean grammar 1/true/yes/on). Read on every request |
DAKERA_ |
4096 |
Most vectors in one extra representation. Must be greater than 0. Over it: 413 |
DAKERA_ |
8388608 (8 MiB) |
Packed payload bytes of all extra representations of one record together. Must be greater than 0. Over it: 413 |
One structural limit is compiled in and has no variable: at most 8 extra representations per record (the primary is not counted). GET /v1/capabilities reports records.enabled, records.representation_kinds, records.dtypes, records.max_representations, records.max_vectors and records.max_bytes; with the routes off it still reports what they would accept once enabled.
Write records
POST /v1/namespaces/{ns}/records (write scope). The namespace is created if it does not exist.
curl -s -X POST localhost:3000/v1/namespaces/docs/records \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"records": [{
"id": "r1",
"values": [0.5, 0.25, 0.25, 0.5],
"representations": [{
"name": "tokens",
"kind": "token_multivector",
"vectors": [[0.5, -0.25], [0.125, 1.0]],
"store_as": "f16"
}],
"metadata": {"source": "demo"}
}]}'
# 200 {"upserted_count": 1}
Request body: {"records": [ … ]}, between 1 and 10,000 records.
| Field of a record | Type | Notes |
|---|---|---|
id |
string | Required. The same id rules as vectors: at most 512 characters, starting with a letter or digit and containing only letters, digits, _, -, : and .. Unique within the request |
values |
float array | Required. The primary vector: non-empty, at most 65,536 values, all finite, the same length for every record of the request and the namespace's dimension |
representations |
array | Optional extras (below). A record with none is stored byte-for-byte as the same POST …/vectors write |
metadata |
object | Optional, validated like vector metadata |
ttl_seconds |
integer | Optional expiry, in seconds from now |
| Field of a representation | Type | Notes |
|---|---|---|
name |
string | Required. The slot name: unique within the record, not empty, not dense (reserved for the primary), and not a server-derived slot (colbert.fde, patch.fde) |
kind |
string | dense (default), token_ or patch_. Any other string is 400 |
model |
string | Optional registry name of the model that produced the vectors; empty means the namespace's default model. Returned in the manifest |
vectors |
array of float arrays | Required. Row-major: every row has the same non-zero length and every value is finite |
store_as |
string | How the server packs the block on disk: f32 (default), f16 or i8 |
Response: {"upserted_count": N}. A record is one vector with extras, so the count is of records and of vectors alike. An upsert replaces the whole record: rewriting a record without representations removes the extras it had.
Representation kinds
kind |
What it holds | Typical slot name |
|---|---|---|
dense |
One or more ordinary vectors (a second embedding of the document, a title vector) | any |
token_ |
One vector per text token of the document, scored with MaxSim by late interaction | colbert |
patch_ |
One vector per image patch of a page | patch |
Storage types (store_as)
The wire format is always plain nested float arrays; store_as only chooses how the server packs the block.
store_as |
Bytes per value | Precision |
|---|---|---|
f32 |
4 | Lossless |
f16 |
2 | Half the bytes; the rounding error is far below what a MaxSim ranking can see on L2-normalized vectors. A value with magnitude above 65504 is 400 |
i8 |
1 | A quarter of the bytes: signed 8-bit codes with one scale per block (value = code x scale, scale = largest magnitude / 127), so the error per value is at most half a step |
The payload of a representation is dim x count x bytes per value; it is what bytes reports and what counts against DAKERA_RECORD_MAX_BYTES. For example 128-wide token vectors stored as f16 take 256 bytes each, so 4,096 of them (the vector limit) are 1 MiB; 1024-wide vectors as f32 take 4 KiB each, so the 8 MiB byte limit is reached at 2,048 of them.
Late-interaction slots
Two slot names have a meaning to the engine. They are the slots a memory written by the server carries, so a record written through this route is searched exactly like one:
| Slot | Kind | Written by | Used for |
|---|---|---|---|
colbert |
token_ |
you, or the server for text memories under late interaction | MaxSim over per-token vectors (text lane) |
colbert.fde |
dense |
the server | The fixed-dimensional encoding (MUVERA) that shortlists candidates before MaxSim |
patch |
patch_ |
you, or the image indexing job | MaxSim over per-patch vectors (visual lane) |
patch.fde |
dense |
the server | The page's fixed-dimensional encoding |
Under DAKERA_SCORING_STRATEGY=late-interaction the server derives the .fde slot from your colbert (kind token_multivector) or patch (kind patch_multivector) slot when it stores the record, from rows normalized to unit length for the encoding only (your stored rows are not changed). Sending an .fde slot yourself is 400 in every configuration, because only the server's encoder produces one the first stage can search. A record written while the strategy was single-vector has no .fde slot and is reached through its primary vector until it is written again. See Search and ranking for the scoring strategy.
Read a record
GET /v1/namespaces/{ns}/records/{id} (read scope).
curl -s localhost:3000/v1/namespaces/docs/records/r1 -H "Authorization: Bearer $KEY"
# {"id": "r1",
# "dimension": 4,
# "representations": [{"name": "tokens", "kind": "token_multivector",
# "dim": 2, "count": 2, "dtype": "f16", "bytes": 8}],
# "metadata": {"source": "demo"}}
curl -s "localhost:3000/v1/namespaces/docs/records/r1?include_vectors=true" -H "Authorization: Bearer $KEY"
# {"id": "r1", "values": [0.5, 0.25, 0.25, 0.5], "dimension": 4,
# "representations": [{"name": "tokens", "kind": "token_multivector", "dim": 2, "count": 2,
# "dtype": "f16", "bytes": 8, "vectors": [[0.5, -0.25], [0.125, 1.0]]}],
# "metadata": {"source": "demo"}}
By default a read returns a manifest: the primary's dimension and, per extra representation, its name, kind, model (omitted when empty), dim, count, dtype and packed bytes. A token block is 20 to 100 times the size of the dense vector beside it and most reads only want to know what a record has, so the payloads (values and each representation's vectors, decoded back to floats) come only with include_vectors=true. Fields that are empty are omitted: a plain dense record has no representations, and ttl_seconds and expires_at appear only on a record that expires. unsupported_representations (a count) appears when the record holds slots written by a newer Dakera with a kind or dtype this binary cannot read; they are skipped, never served and never deleted.
A sidecar that no longer matches its primary (the memory's content was updated, the vector was re-embedded, or the id was rewritten through the plain vector route) is stale and is not served: the record reads back as the dense record it also is. The next TTL sweep removes it.
An unknown record is 404 (VECTOR_NOT_FOUND, the v0.11 code kept).
Delete a record
There is deliberately no record delete route. Delete the id through the vector routes, which are document level: the extras go with the primary.
curl -s -X POST localhost:3000/v1/namespaces/docs/vectors/delete \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" -d '{"ids": ["r1"]}'
# {"deleted_count": 1}
Atomicity and storage
The extras are kept as a sidecar value next to the primary on every backend. RocksDB writes both in one batch (atomic); the object store writes the sidecar first and the primary second, with the primary as the commit point, so a reader never sees extras without their primary. The sidecar is bound to the exact primary vector it was written with, which is why a rewrite of the primary orphans it rather than serving mismatched vectors. Snapshots, export/import (JSONL carries the representations; CSV is dense only), backups and cluster replication carry records whole, so a peer never keeps only the dense half of a record. Records go through the same write path as vectors, so namespace quotas (max_storage_bytes counts the extras' bytes; an overwrite releases the old extras), ANN invalidation and the cluster broadcast apply unchanged. WAL, caches and tiered storage forward records, and the tiered flush carries the sidecar.
v0.11.108 does not know the extras: after dakera downgrade it serves a record as its dense vector and the sidecars are left in place for a later upgrade.
Errors
| Status | Code | When |
|---|---|---|
501 |
FEATURE_ |
DAKERA_RECORDS is off (details: set DAKERA_) |
400 |
INVALID_REQUEST |
Empty records array; more than 10,000 records; empty or duplicate record id; empty, non-finite or mixed-length values; a slot with an empty or duplicate name, the name dense or a derived .fde name; kind unknown or the container kind; ragged or empty vectors; a non-finite value; a value out of range for f16; an unknown store_as |
400 |
DIMENSION_ |
values does not match the namespace's dimension |
413 |
PAYLOAD_ |
More than 8 extra representations, more than DAKERA_ vectors in one, or more than DAKERA_ in all |
413 |
QUOTA_EXCEEDED |
The write would pass the namespace's max_vectors or max_ quota |
404 |
VECTOR_ |
Reading a record that does not exist |
The message names the record and slot: records[0] (id 'r1'): representation 'tokens': rows differ in length: expected 2, found 1. A well-formed record that is too large is 413; anything malformed is 400.
Verify it works
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" | jq .recordsshows"enabled": trueand the limits in force.- Write the example record above and read it back: the manifest shows
dtype: "f16"andbytes: 8. - Read it with
include_vectors=trueand compare the rows (exactly representable values come back unchanged; others come back rounded to the chosen type). - Write it again without
representationsand read it: the extras are gone. - Delete it with
POST …/vectors/deleteand read it:404.