NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

Records and multi-vector

A record is one vector plus named extra representations. The one vector (the primary, a plain dense embedding) is what is indexed and searched, exactly like any vector written with POST /v1/namespaces/{ns}/vectors. The extras sit beside it and are stored, replicated, backed up and deleted with it: per-token vectors for late interaction (a token multivector), per-patch vectors for a page image (a patch multivector), or another dense vector. Use records when you compute multi-vector embeddings yourself and want Dakera to keep them with the document. Memories written by the server under late interaction, and image pages indexed through the multimodal lane, are records of the same shape.

The routes are off by default (DAKERA_RECORDS): with the switch off both answer 501 FEATURE_DISABLED before touching any state, and no other route changes.

Configuration

Variable Default Meaning
DAKERA_RECORDS off Turns the two record routes on (boolean grammar 1/true/yes/on). Read on every request
DAKERA_RECORD_MAX_VECTORS 4096 Most vectors in one extra representation. Must be greater than 0. Over it: 413
DAKERA_RECORD_MAX_BYTES 8388608 (8 MiB) Packed payload bytes of all extra representations of one record together. Must be greater than 0. Over it: 413

One structural limit is compiled in and has no variable: at most 8 extra representations per record (the primary is not counted). GET /v1/capabilities reports records.enabled, records.representation_kinds, records.dtypes, records.max_representations, records.max_vectors and records.max_bytes; with the routes off it still reports what they would accept once enabled.

Write records

POST /v1/namespaces/{ns}/records (write scope). The namespace is created if it does not exist.

curl -s -X POST localhost:3000/v1/namespaces/docs/records \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"records": [{
        "id": "r1",
        "values": [0.5, 0.25, 0.25, 0.5],
        "representations": [{
          "name": "tokens",
          "kind": "token_multivector",
          "vectors": [[0.5, -0.25], [0.125, 1.0]],
          "store_as": "f16"
        }],
        "metadata": {"source": "demo"}
      }]}'
# 200 {"upserted_count": 1}

Request body: {"records": [ … ]}, between 1 and 10,000 records.

Field of a record Type Notes
id string Required. The same id rules as vectors: at most 512 characters, starting with a letter or digit and containing only letters, digits, _, -, : and .. Unique within the request
values float array Required. The primary vector: non-empty, at most 65,536 values, all finite, the same length for every record of the request and the namespace's dimension
representations array Optional extras (below). A record with none is stored byte-for-byte as the same POST …/vectors write
metadata object Optional, validated like vector metadata
ttl_seconds integer Optional expiry, in seconds from now
Field of a representation Type Notes
name string Required. The slot name: unique within the record, not empty, not dense (reserved for the primary), and not a server-derived slot (colbert.fde, patch.fde)
kind string dense (default), token_multivector or patch_multivector. Any other string is 400
model string Optional registry name of the model that produced the vectors; empty means the namespace's default model. Returned in the manifest
vectors array of float arrays Required. Row-major: every row has the same non-zero length and every value is finite
store_as string How the server packs the block on disk: f32 (default), f16 or i8

Response: {"upserted_count": N}. A record is one vector with extras, so the count is of records and of vectors alike. An upsert replaces the whole record: rewriting a record without representations removes the extras it had.

Representation kinds

kind What it holds Typical slot name
dense One or more ordinary vectors (a second embedding of the document, a title vector) any
token_multivector One vector per text token of the document, scored with MaxSim by late interaction colbert
patch_multivector One vector per image patch of a page patch

Storage types (store_as)

The wire format is always plain nested float arrays; store_as only chooses how the server packs the block.

store_as Bytes per value Precision
f32 4 Lossless
f16 2 Half the bytes; the rounding error is far below what a MaxSim ranking can see on L2-normalized vectors. A value with magnitude above 65504 is 400
i8 1 A quarter of the bytes: signed 8-bit codes with one scale per block (value = code x scale, scale = largest magnitude / 127), so the error per value is at most half a step

The payload of a representation is dim x count x bytes per value; it is what bytes reports and what counts against DAKERA_RECORD_MAX_BYTES. For example 128-wide token vectors stored as f16 take 256 bytes each, so 4,096 of them (the vector limit) are 1 MiB; 1024-wide vectors as f32 take 4 KiB each, so the 8 MiB byte limit is reached at 2,048 of them.

Late-interaction slots

Two slot names have a meaning to the engine. They are the slots a memory written by the server carries, so a record written through this route is searched exactly like one:

Slot Kind Written by Used for
colbert token_multivector you, or the server for text memories under late interaction MaxSim over per-token vectors (text lane)
colbert.fde dense the server The fixed-dimensional encoding (MUVERA) that shortlists candidates before MaxSim
patch patch_multivector you, or the image indexing job MaxSim over per-patch vectors (visual lane)
patch.fde dense the server The page's fixed-dimensional encoding

Under DAKERA_SCORING_STRATEGY=late-interaction the server derives the .fde slot from your colbert (kind token_multivector) or patch (kind patch_multivector) slot when it stores the record, from rows normalized to unit length for the encoding only (your stored rows are not changed). Sending an .fde slot yourself is 400 in every configuration, because only the server's encoder produces one the first stage can search. A record written while the strategy was single-vector has no .fde slot and is reached through its primary vector until it is written again. See Search and ranking for the scoring strategy.

Read a record

GET /v1/namespaces/{ns}/records/{id} (read scope).

curl -s localhost:3000/v1/namespaces/docs/records/r1 -H "Authorization: Bearer $KEY"
# {"id": "r1",
#  "dimension": 4,
#  "representations": [{"name": "tokens", "kind": "token_multivector",
#                       "dim": 2, "count": 2, "dtype": "f16", "bytes": 8}],
#  "metadata": {"source": "demo"}}

curl -s "localhost:3000/v1/namespaces/docs/records/r1?include_vectors=true" -H "Authorization: Bearer $KEY"
# {"id": "r1", "values": [0.5, 0.25, 0.25, 0.5], "dimension": 4,
#  "representations": [{"name": "tokens", "kind": "token_multivector", "dim": 2, "count": 2,
#                       "dtype": "f16", "bytes": 8, "vectors": [[0.5, -0.25], [0.125, 1.0]]}],
#  "metadata": {"source": "demo"}}

By default a read returns a manifest: the primary's dimension and, per extra representation, its name, kind, model (omitted when empty), dim, count, dtype and packed bytes. A token block is 20 to 100 times the size of the dense vector beside it and most reads only want to know what a record has, so the payloads (values and each representation's vectors, decoded back to floats) come only with include_vectors=true. Fields that are empty are omitted: a plain dense record has no representations, and ttl_seconds and expires_at appear only on a record that expires. unsupported_representations (a count) appears when the record holds slots written by a newer Dakera with a kind or dtype this binary cannot read; they are skipped, never served and never deleted.

A sidecar that no longer matches its primary (the memory's content was updated, the vector was re-embedded, or the id was rewritten through the plain vector route) is stale and is not served: the record reads back as the dense record it also is. The next TTL sweep removes it.

An unknown record is 404 (VECTOR_NOT_FOUND, the v0.11 code kept).

Delete a record

There is deliberately no record delete route. Delete the id through the vector routes, which are document level: the extras go with the primary.

curl -s -X POST localhost:3000/v1/namespaces/docs/vectors/delete \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" -d '{"ids": ["r1"]}'
# {"deleted_count": 1}

Atomicity and storage

The extras are kept as a sidecar value next to the primary on every backend. RocksDB writes both in one batch (atomic); the object store writes the sidecar first and the primary second, with the primary as the commit point, so a reader never sees extras without their primary. The sidecar is bound to the exact primary vector it was written with, which is why a rewrite of the primary orphans it rather than serving mismatched vectors. Snapshots, export/import (JSONL carries the representations; CSV is dense only), backups and cluster replication carry records whole, so a peer never keeps only the dense half of a record. Records go through the same write path as vectors, so namespace quotas (max_storage_bytes counts the extras' bytes; an overwrite releases the old extras), ANN invalidation and the cluster broadcast apply unchanged. WAL, caches and tiered storage forward records, and the tiered flush carries the sidecar.

v0.11.108 does not know the extras: after dakera downgrade it serves a record as its dense vector and the sidecars are left in place for a later upgrade.

Errors

Status Code When
501 FEATURE_DISABLED DAKERA_RECORDS is off (details: set DAKERA_RECORDS to enable it)
400 INVALID_REQUEST Empty records array; more than 10,000 records; empty or duplicate record id; empty, non-finite or mixed-length values; a slot with an empty or duplicate name, the name dense or a derived .fde name; kind unknown or the container kind; ragged or empty vectors; a non-finite value; a value out of range for f16; an unknown store_as
400 DIMENSION_MISMATCH values does not match the namespace's dimension
413 PAYLOAD_TOO_LARGE More than 8 extra representations, more than DAKERA_RECORD_MAX_VECTORS vectors in one, or more than DAKERA_RECORD_MAX_BYTES in all
413 QUOTA_EXCEEDED The write would pass the namespace's max_vectors or max_storage_bytes quota
404 VECTOR_NOT_FOUND Reading a record that does not exist

The message names the record and slot: records[0] (id 'r1'): representation 'tokens': rows differ in length: expected 2, found 1. A well-formed record that is too large is 413; anything malformed is 400.

Verify it works

  1. curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" | jq .records shows "enabled": true and the limits in force.
  2. Write the example record above and read it back: the manifest shows dtype: "f16" and bytes: 8.
  3. Read it with include_vectors=true and compare the rows (exactly representable values come back unchanged; others come back rounded to the chosen type).
  4. Write it again without representations and read it: the extras are gone.
  5. Delete it with POST …/vectors/delete and read it: 404.
Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.