NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

Multimodal memory

This page is for developers and operators who want an agent's memories to carry files and be built from them. v0.12.0 adds three opt-in features: attachments (files stored per namespace and referenced from memories), speech to text (an audio attachment becomes a memory), and image indexing with visual recall (document pages become searchable by the text of a question). Every media job reserves the memory it needs before it starts, so a small machine answers 503 with Retry-After instead of being killed for running out of memory.

Everything on this page is off by default. A server that sets none of the variables below answers the routes with 501 FEATURE_DISABLED (the error's details name the variable that turns them on), loads none of the models and behaves exactly as v0.11.108 did. Each switch reads the common boolean grammar (1/true/yes/on, 0/false/no/off).

Related pages: Records and multi-vector (one vector plus named extra representations), Search and ranking (late interaction, RaBitQ, rerank controls, GET /v1/capabilities), Multilingual search (bge-m3, full-text languages, per-request lang) and Models and Docker (where the models come from, dakera models, air-gapped installs).

At a glance

Feature Turn it on with Route Model Download
Attachments DAKERA_ATTACHMENTS …/attachments (upload, list, download, delete) none none
Speech to text DAKERA_ATTACHMENTS POST …/attachments/{ref}/transcribe whisper-base by default; five choices via DAKERA_WHISPER_MODEL (see Speech to text) see the model table (whisper-tiny.en: about 151 MB)
Image indexing and visual recall DAKERA_ATTACHMENTS and DAKERA_VISION (plus DAKERA_SCORING_STRATEGY=late-interaction to search the pages) POST …/attachments/{ref}/index colmodernvbert (DAKERA_VISION_MODEL) about 966 MB
Variable Default Meaning
DAKERA_ATTACHMENTS off Attachment routes, attachment_ref on memories, and speech to text
DAKERA_ATTACHMENT_MAX_BYTES 26214400 (25 MiB) Largest upload; over it the request is 413. Must be greater than 0
DAKERA_WHISPER_MODEL whisper-base The speech-to-text model: whisper-base (default), whisper-tiny.en, whisper-base.en, whisper-tiny or whisper-small (see Speech to text)
DAKERA_VISION off Image indexing (needs DAKERA_ATTACHMENTS too) and the visual recall lane
DAKERA_VISION_MODEL colmodernvbert The visual model (the registry's only row today)
DAKERA_MEM_HIGH_WATER_FRACTION 0.85 Share of the memory limit media jobs may fill (valid above 0 and up to 1; anything else falls back to the default)
DAKERA_MEM_BACKPRESSURE on 0 or false stops memory refusals (reservations are still counted)
DAKERA_HNSW_CACHE_TTI_SECS 3600 Idle time after which whisper, the visual model and GLiNER are unloaded

GET /v1/capabilities reports what the running server has on: attachments (enabled, max_bytes, transcription) and vision (enabled, model, media_types, dimension, tile_size, longest_edge, max_tiles). See Search and ranking.

Attachments

An attachment is a file an agent's memories can point at. It is stored content-addressed per namespace: its reference is attachment_ref = sha256: followed by the 64 lowercase hex characters of the SHA-256 of its bytes. Uploading the same bytes into the same namespace twice is the same object. Attachments are replicated like vectors in a cluster, counted by namespace quotas, carried by backups, and released with the last memory that references them.

Route Scope Answer
POST /v1/namespaces/{ns}/attachments write Upload. A multipart form with a file field, or a raw body with its Content-Type. 201 for new bytes, 200 when the namespace already held them
GET /v1/namespaces/{ns}/attachments read The namespace's attachments, without bytes
GET /v1/namespaces/{ns}/attachments/{ref} read The bytes, with the media type they were uploaded with and ETag: "<hex hash>"
DELETE /v1/namespaces/{ns}/attachments/{ref} write 204; 409 while a memory references it

Upload

curl -s -X POST localhost:3000/v1/namespaces/uploads/attachments \
  -H "Authorization: Bearer $KEY" -H "Content-Type: audio/wav" \
  --data-binary @note.wav
# 201 {"attachment_ref":"sha256:9f2c…","content_type":"audio/wav","size_bytes":480044,"created":true}

curl -s -X POST localhost:3000/v1/namespaces/uploads/attachments \
  -H "Authorization: Bearer $KEY" -F "[email protected];type=image/png"

List the namespace's attachments:

curl -s localhost:3000/v1/namespaces/uploads/attachments -H "Authorization: Bearer $KEY"
# {"attachments":[{"attachment_ref":"sha256:9f2c…","content_type":"audio/wav","size_bytes":480044}]}

Referencing an attachment from a memory

Send "attachment_ref": "sha256:…" on POST /v1/memory/store (and on the items of the batch store). The reference must name an attachment in the memory's own namespace, _dakera_agent_{agent_id}; otherwise the store is 404 (VECTOR_NOT_FOUND with resource: "attachment"). A malformed reference is 400. With DAKERA_ATTACHMENTS off, a memory carrying attachment_ref is refused 501.

The transcription and image jobs below take the attachment from any namespace the key can read and copy it into the agent's namespace for you.

Lifecycle

Speech to text

POST /v1/namespaces/{ns}/attachments/{ref}/transcribe turns an uploaded audio file into a memory. It needs DAKERA_ATTACHMENTS; the model is chosen with DAKERA_WHISPER_MODEL (default whisper-base; set whisper-tiny.en for the lightest English model):

Model Languages Role
whisper-base Multilingual (about 99 languages), language auto-detected The default; the recommended choice (74M parameters)
whisper-tiny.en English The lightest English model (39M)
whisper-base.en English Better English (74M)
whisper-tiny Multilingual, language auto-detected The lightest multilingual model (39M)
whisper-small Multilingual, language auto-detected The quality option (244M)

Models are pinned to a SHA-256 like the embedding models, dakera models list|pull|prune|--bake covers every variant, memory admission reserves each model's footprint, and idle models are unloaded.

curl -s -X POST "localhost:3000/v1/namespaces/uploads/attachments/$REF/transcribe" \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"agent_id": "my-agent", "tags": ["voice"], "importance": 0.7}'
# 202
# {"job_id": "job_1a2b3c4d_0",
#  "attachment_ref": "sha256:9f2c…",
#  "agent_id": "my-agent",
#  "memory_id": "mem_17f3…",
#  "model": "whisper-base",
#  "status_url": "/v1/namespaces/uploads/attachments/sha256:9f2c…/transcribe/job_1a2b3c4d_0"}

Request: agent_id is required; the other fields take the defaults of POST /v1/memory/store: id, memory_type, session_id, importance (0.0 to 1.0), tags, metadata, ttl_seconds, expires_at, lang. The key needs read on the upload's namespace and write on the agent's namespace. An unknown attachment is 404; an audio file that cannot be transcribed is 400 before any job exists.

What the job stores: one memory through the normal write path (embedded, full-text indexed, quotas applied) with the transcript as content, attachment_ref set to the audio, _dakera_source: "attachment" and _dakera_transcript in its metadata (model, detected language, duration, timed segments). lang selects how dates in the transcript are parsed; with a multilingual model and no lang, the detected language is used.

Image indexing and visual recall

POST /v1/namespaces/{ns}/attachments/{ref}/index needs both DAKERA_VISION and DAKERA_ATTACHMENTS (a server with neither names DAKERA_VISION). It has the same request and answer shape as …/transcribe, with one extra optional field, content: the caption stored as the memory's text (default [image sha256:…], at most 100,000 characters).

curl -s -X POST "localhost:3000/v1/namespaces/uploads/attachments/$PAGE_REF/index" \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"agent_id": "docs-agent", "content": "Annual report, page 12: revenue by region"}'
# 202 {"job_id":"job_1a2b3c4d_1","attachment_ref":"sha256:…","agent_id":"docs-agent",
#      "memory_id":"mem_…","model":"colmodernvbert",
#      "status_url":"/v1/namespaces/uploads/attachments/sha256:…/index/job_1a2b3c4d_1"}

Visual recall

Searching the pages needs the late-interaction strategy on top of DAKERA_VISION (DAKERA_SCORING_STRATEGY=late-interaction, see Search and ranking). Recall then embeds the query with the visual model's text side, shortlists pages by their fixed-dimensional encoding (patch.fde; the pooled 128-d page vector serves while that stage is built) and reranks with MaxSim over the page patches. The text cross-encoder is not applied on this lane (rerank_report.skipped is visual_lane).

The visual lane is a store of its own. One store is one embedding space: the agent namespaces then hold 128-d page vectors, not text embeddings, and the text late-interaction lane (colbert slots) is replaced by the visual one (patch slots). Use a deployment (a data root) dedicated to it, and do not turn DAKERA_VISION on over an existing text store. dakera downgrade refuses a store that uses the visual lane.

Measured quality (ViDoRe, nDCG@5, visual lane): TabFQuAD 0.643 and Shift Project 0.7705. The other ViDoRe sets have not been measured.

Jobs

Transcription and image jobs run in the background and are polled until they finish.

curl -s "localhost:3000$STATUS_URL" -H "Authorization: Bearer $KEY"
# {"id":"job_1a2b3c4d_0","job_type":"transcription","status":"Running","created_at":1790000000,
#  "started_at":1790000000,"completed_at":null,"progress":42,"message":"transcribing",
#  "metadata":{"namespace":"uploads","attachment_ref":"sha256:9f2c…","agent_id":"my-agent",
#              "memory_id":"mem_17f3…","model":"whisper-base"}}

Memory, CPU and disk

None of the multimodal models is in the image (the image carries bge-large and the reranker). Each downloads from Hugging Face into the model cache (HF_HOME, /app/models in the image; mount it as a volume to download once), is converted once to an ORT-format copy next to it (disk is about the model's size again), and is then memory-mapped. With DAKERA_ATTACHMENTS or DAKERA_VISION on, whisper and the visual model are downloaded and converted in the background right after the server is ready, so the first job does not wait for them. dakera models pull whisper vision fetches them ahead of time and RUN dakera models pull --bake whisper vision bakes them into your own image; see Models and Docker.

Model Download Loaded when
whisper-base (speech to text, default) see the model table the first transcription job (fetched at start)
whisper-tiny.en about 151 MB the first transcription job, when chosen with DAKERA_WHISPER_MODEL
whisper-base.en, whisper-tiny, whisper-small (chosen with DAKERA_WHISPER_MODEL) see the model table the first transcription job
colmodernvbert (image indexing, visual recall) about 966 MB the first image job or visual recall (fetched at start)
colbert-small (late interaction, DAKERA_MODEL) about 34 MB startup (it is the embedding model)
GLiNER (entity extraction, per-namespace extract_entities) about 782 MB the first store that extracts entities (it waits up to 10 s, then 503 with Retry-After while the model downloads)

A download that is cut short continues where it stopped on the next attempt or the next start; while a job waits for its model, its message shows the download.

Memory is admitted, not assumed. Every media job (and every GLiNER pass on the store path) reserves its estimated peak working memory before it decodes anything: for audio, the samples plus one 30-second window's activations (128 MiB); for an image, the raster and decode buffers read from its header plus 128 MiB per tile (a page is at most a 4 x 4 grid plus a thumbnail, 17 tiles). The reservation is granted while resident memory plus reservations plus the request stay within the memory limit (the cgroup limit, else physical memory) times DAKERA_MEM_HIGH_WATER_FRACTION, the same reading the store routes' backpressure uses. A job that does not fit first triggers a reclaim (idle models unloaded, derived caches dropped), then waits up to 10 s for running jobs to give memory back, and is then answered 503 with Retry-After: 10. There is no fixed job cap. DAKERA_MEM_BACKPRESSURE=0 turns the refusals off. On a platform where no memory limit can be read, reservations are granted.

A model's first load converts its graph to the ORT format once per model and build, and that conversion reserves its peak from the same budget before it starts (the download before it reserves nothing). When it does not fit, the job fails 503 with a message naming what the conversion needs, what the budget has and the remedy, and /health lists speech_to_text, vision_model or entity_extraction under degraded until a load is admitted. dakera models pull whisper or dakera models pull vision converts the model outside the server once; a converted model loads without that reservation.

Idle models leave memory. Whisper, the visual model and GLiNER are unloaded after DAKERA_HNSW_CACHE_TTI_SECS (default 3600 s) without use, and earlier under memory pressure; the next use loads them again from the local cache. A job that uses a loaded model counts as use. The weights are memory-mapped (shared, reclaimable page cache); what each model holds privately is its working memory while it runs.

CPU is shared. Every model's work draws from one pool of compute tokens (one per core). Query-side work (query embeddings, rerank, visual queries) is served first; bulk work (document embeddings, transcription, image pages, entity extraction) uses at most cores - 1 at once, so a transcription never takes the whole machine from recall.

Measured footprint: with every model loaded, the process heap was 530 MiB (memory-mapped model files are not counted). A small box can run everything, one heavy job at a time: jobs that do not fit wait or get a 503 to retry. The official SDKs time requests out at 30 s; the server's queueing answers within 20 s so the 503 and its Retry-After arrive first.

Errors

Every error body is JSON: {"error": "...", "code": "...", "status": 400, "details": "..."} (a 404 also carries resource).

Status Code When
501 FEATURE_DISABLED The switch is off; details say set DAKERA_ATTACHMENTS to enable it (or DAKERA_VISION)
400 INVALID_REQUEST Malformed attachment_ref; empty upload; invalid media type; multipart body without a file part; audio or image the model cannot take; invalid agent_id, tags, importance, id, lang
404 VECTOR_NOT_FOUND, resource: "attachment" No such attachment in that namespace (the v0.11 code is kept; resource tells the cases apart)
404 JOB_NOT_FOUND, resource: "job" Unknown job, or one lost to a restart (the message says which)
409 CONFLICT Deleting an attachment a memory still references
413 PAYLOAD_TOO_LARGE Upload over DAKERA_ATTACHMENT_MAX_BYTES; transcript or caption over 100,000 characters
413 QUOTA_EXCEEDED The namespace's max_storage_bytes quota
503 SERVICE_UNAVAILABLE with Retry-After Memory admission, model not loadable, page queue full (30 s) or timed out, writes still referencing a deleted attachment (1 s), shutdown in progress

Verify it works

  1. curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" | jq '.attachments, .vision' shows enabled: true for the switches you set, with max_bytes and the model names. Without the switches the routes answer 501.
  2. Upload a small WAV: the answer is 201 with created: true; the same upload again is 200 with created: false.
  3. Start a transcription and poll status_url until status is Completed (the first job also downloads and converts the model; its message shows the download). Recall by the spoken words and check the memory's attachment_ref.
  4. For the visual lane, upload a PNG page, index it, then call POST /v1/memory/recall with a question about it (late-interaction strategy on); the rerank_report of the response says visual_lane.
  5. curl -s localhost:3000/health | jq .degraded lists speech_to_text, vision_model or entity_extraction only while a model cannot be loaded within the memory budget. dakera_memory_budget_reserved_bytes and dakera_model_loads_total are exported on /metrics.

Constraints

Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.