Multimodal memory
This page is for developers and operators who want an agent's memories to carry files and be built from them. v0.12.0 adds three opt-in features: attachments (files stored per namespace and referenced from memories), speech to text (an audio attachment becomes a memory), and image indexing with visual recall (document pages become searchable by the text of a question). Every media job reserves the memory it needs before it starts, so a small machine answers 503 with Retry-After instead of being killed for running out of memory.
Everything on this page is off by default. A server that sets none of the variables below answers the routes with 501 FEATURE_DISABLED (the error's details name the variable that turns them on), loads none of the models and behaves exactly as v0.11.108 did. Each switch reads the common boolean grammar (1/true/yes/on, 0/false/no/off).
Related pages: Records and multi-vector (one vector plus named extra representations), Search and ranking (late interaction, RaBitQ, rerank controls, GET /v1/capabilities), Multilingual search (bge-m3, full-text languages, per-request lang) and Models and Docker (where the models come from, dakera models, air-gapped installs).
At a glance
| Feature | Turn it on with | Route | Model | Download |
|---|---|---|---|---|
| Attachments | DAKERA_ |
…/attachments (upload, list, download, delete) |
none | none |
| Speech to text | DAKERA_ |
POST …/ |
whisper-base by default; five choices via DAKERA_ (see Speech to text) |
see the model table (whisper-tiny.en: about 151 MB) |
| Image indexing and visual recall | DAKERA_ and DAKERA_VISION (plus DAKERA_ to search the pages) |
POST …/ |
colmodernvbert (DAKERA_) |
about 966 MB |
| Variable | Default | Meaning |
|---|---|---|
DAKERA_ |
off | Attachment routes, attachment_ref on memories, and speech to text |
DAKERA_ |
26214400 (25 MiB) |
Largest upload; over it the request is 413. Must be greater than 0 |
DAKERA_ |
whisper-base |
The speech-to-text model: whisper-base (default), whisper-tiny.en, whisper-base.en, whisper-tiny or whisper-small (see Speech to text) |
DAKERA_VISION |
off | Image indexing (needs DAKERA_ too) and the visual recall lane |
DAKERA_ |
colmodernvbert |
The visual model (the registry's only row today) |
DAKERA_ |
0.85 |
Share of the memory limit media jobs may fill (valid above 0 and up to 1; anything else falls back to the default) |
DAKERA_ |
on | 0 or false stops memory refusals (reservations are still counted) |
DAKERA_ |
3600 |
Idle time after which whisper, the visual model and GLiNER are unloaded |
GET /v1/capabilities reports what the running server has on: attachments (enabled, max_bytes, transcription) and vision (enabled, model, media_types, dimension, tile_size, longest_edge, max_tiles). See Search and ranking.
Attachments
An attachment is a file an agent's memories can point at. It is stored content-addressed per namespace: its reference is attachment_ref = sha256: followed by the 64 lowercase hex characters of the SHA-256 of its bytes. Uploading the same bytes into the same namespace twice is the same object. Attachments are replicated like vectors in a cluster, counted by namespace quotas, carried by backups, and released with the last memory that references them.
| Route | Scope | Answer |
|---|---|---|
POST / |
write | Upload. A multipart form with a file field, or a raw body with its Content-Type. 201 for new bytes, 200 when the namespace already held them |
GET / |
read | The namespace's attachments, without bytes |
GET / |
read | The bytes, with the media type they were uploaded with and ETag: "<hex hash>" |
DELETE / |
write | 204; 409 while a memory references it |
Upload
curl -s -X POST localhost:3000/v1/namespaces/uploads/attachments \
-H "Authorization: Bearer $KEY" -H "Content-Type: audio/wav" \
--data-binary @note.wav
# 201 {"attachment_ref":"sha256:9f2c…","content_type":"audio/wav","size_bytes":480044,"created":true}
curl -s -X POST localhost:3000/v1/namespaces/uploads/attachments \
-H "Authorization: Bearer $KEY" -F "[email protected];type=image/png"
createdisfalse(HTTP200) when the namespace already held the same bytes. The upload is then a no-op and the media type stored the first time is kept.- Media type: a
type/subtypetoken without parameters; parameters are dropped (Audio/WAV; codecs=1is stored asaudio/wav). With noContent-Typethe bytes are stored asapplication/octet-stream. A value that is not a media type is400. - In a multipart body the first part named
file(or carrying a file name) is the attachment, and the part's ownContent-Typewins. A multipart body with no such part is400. - An empty body is
400. A body overDAKERA_ATTACHMENT_MAX_BYTESis413 PAYLOAD_TOO_LARGE(the default 25 MiB is about 13 minutes of 16 kHz, 16-bit mono WAV). - The namespace is created if it does not exist. An upload that would take a namespace over its
max_storage_bytesquota is413 QUOTA_EXCEEDED.
List the namespace's attachments:
curl -s localhost:3000/v1/namespaces/uploads/attachments -H "Authorization: Bearer $KEY"
# {"attachments":[{"attachment_ref":"sha256:9f2c…","content_type":"audio/wav","size_bytes":480044}]}
Referencing an attachment from a memory
Send "attachment_ref": "sha256:…" on POST /v1/memory/store (and on the items of the batch store). The reference must name an attachment in the memory's own namespace, _dakera_agent_{agent_id}; otherwise the store is 404 (VECTOR_NOT_FOUND with resource: "attachment"). A malformed reference is 400. With DAKERA_ATTACHMENTS off, a memory carrying attachment_ref is refused 501.
The transcription and image jobs below take the attachment from any namespace the key can read and copy it into the agent's namespace for you.
Lifecycle
- Release: when the last memory that references an attachment is forgotten, the attachment is released with it, in the namespace and on every cluster node. If writes that reference it are still in flight, the release waits up to 30 s and then keeps the bytes (an orphaned file costs storage; a released file under a live memory would lose data).
- Explicit delete:
DELETE …/attachments/{ref}is404for an unknown reference,409while any memory references it (the body names up to five of them: forget the memories instead), and503withRetry-After: 1while writes that reference it are in progress. - Consolidation refuses (
400) explicitmemory_idsthat include a memory with an attachment, so a merge never drops one. - Quotas: attachment bytes count toward the namespace's
max_storage_bytes; uploading bytes the namespace already holds costs nothing, and a delete frees the stored size. - Backups and restore: a backup carries each namespace's attachments with their bytes. A restore writes them back; they are counted toward the quota but a restore is never refused by it, as for vectors.
- Cluster: uploads and deletes are replicated to peers with versions, so a delete is not undone by a node that missed it.
- Rollback:
dakera downgradeleaves attachments in place; v0.11.108 ignores them and serves the memory without its file.
Speech to text
POST /v1/namespaces/{ns}/attachments/{ref}/transcribe turns an uploaded audio file into a memory. It needs DAKERA_ATTACHMENTS; the model is chosen with DAKERA_WHISPER_MODEL (default whisper-base; set whisper-tiny.en for the lightest English model):
| Model | Languages | Role |
|---|---|---|
whisper-base |
Multilingual (about 99 languages), language auto-detected | The default; the recommended choice (74M parameters) |
whisper-tiny.en |
English | The lightest English model (39M) |
whisper-base.en |
English | Better English (74M) |
whisper-tiny |
Multilingual, language auto-detected | The lightest multilingual model (39M) |
whisper-small |
Multilingual, language auto-detected | The quality option (244M) |
Models are pinned to a SHA-256 like the embedding models, dakera models list|pull|prune|--bake covers every variant, memory admission reserves each model's footprint, and idle models are unloaded.
curl -s -X POST "localhost:3000/v1/namespaces/uploads/attachments/$REF/transcribe" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"agent_id": "my-agent", "tags": ["voice"], "importance": 0.7}'
# 202
# {"job_id": "job_1a2b3c4d_0",
# "attachment_ref": "sha256:9f2c…",
# "agent_id": "my-agent",
# "memory_id": "mem_17f3…",
# "model": "whisper-base",
# "status_url": "/v1/namespaces/uploads/attachments/sha256:9f2c…/transcribe/job_1a2b3c4d_0"}
Request: agent_id is required; the other fields take the defaults of POST /v1/memory/store: id, memory_type, session_id, importance (0.0 to 1.0), tags, metadata, ttl_seconds, expires_at, lang. The key needs read on the upload's namespace and write on the agent's namespace. An unknown attachment is 404; an audio file that cannot be transcribed is 400 before any job exists.
What the job stores: one memory through the normal write path (embedded, full-text indexed, quotas applied) with the transcript as content, attachment_ref set to the audio, _dakera_source: "attachment" and _dakera_transcript in its metadata (model, detected language, duration, timed segments). lang selects how dates in the transcript are parsed; with a multilingual model and no lang, the detected language is used.
- Media: WAV audio, PCM 8/16/24/32-bit or float, any channel count and sample rate (it is mixed to mono and resampled to 16 kHz). Anything else is
400. The attachment must have been uploaded asaudio/wav,audio/x-wav,audio/wave,audio/vnd.waveor with no content type (application/octet-stream). - Language: the English models (
whisper-tiny.en,whisper-base.en) transcribe English. A multilingual model (whisper-base, the default,whisper-tiny,whisper-small) detects the spoken language itself and records it on the stored memory as itslang, so full-text stemming, date parsing and bge-m3 multilingual embeddings work downstream. An explicit per-requestlangoverrides detection.GET /v1/capabilitieslists the configured model'slanguagesandsample_rate_hz: 16000. - Limits of the result: audio with no speech fails
400("transcription produced no speech"); a transcript over 100,000 characters (the memory content limit) fails413. - Idempotent: with no
id, the memory id is derived from the agent, the attachment, its namespace and the rest of the request, so sending the same request again stores the same memory instead of a second one. Thememory_idin the202is how to find the result.
Image indexing and visual recall
POST /v1/namespaces/{ns}/attachments/{ref}/index needs both DAKERA_VISION and DAKERA_ATTACHMENTS (a server with neither names DAKERA_VISION). It has the same request and answer shape as …/transcribe, with one extra optional field, content: the caption stored as the memory's text (default [image sha256:…], at most 100,000 characters).
curl -s -X POST "localhost:3000/v1/namespaces/uploads/attachments/$PAGE_REF/index" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"agent_id": "docs-agent", "content": "Annual report, page 12: revenue by region"}'
# 202 {"job_id":"job_1a2b3c4d_1","attachment_ref":"sha256:…","agent_id":"docs-agent",
# "memory_id":"mem_…","model":"colmodernvbert",
# "status_url":"/v1/namespaces/uploads/attachments/sha256:…/index/job_1a2b3c4d_1"}
- Media: PNG only (
image/png), at most 64 megapixels (for example 8000 x 8000). Another media type is400before a job exists.GET /v1/capabilitiesreportsvision.media_types,tile_size(512),longest_edge(2048) andmax_tiles(17). - What is stored: the job embeds the pixels with
colmodernvbertand stores one memory whose primary vector is the pooled page vector and whose extra representations are the patch multivector (patch) and its fixed-dimensional encoding (patch.fde); see Records and multi-vector. Vectors are 128 wide. The caption only feeds the full-text index, so a caption never quietly becomes what is searched. The memory's metadata records_dakera_image(model, tile grid, patch counts). - Pages are indexed one at a time, in the order they were submitted, because the visual model has one page session. Submit a whole document's pages at once: up to 2,500 jobs wait for the session (a queued page holds no memory reservation and no decoded pixels, only its job record and a parked task, a few KB), and only past that is a new job answered
503withRetry-After: 30. A page waits up to 30 s for each page ahead of it and the one running; a page whose turn does not come in that time fails503and is started again later. - Speed: about 10.7 s per page on CPU.
- Progress: 2 % while the model loads, 5 % waiting for the page session, 10 % embedding the page, 90 % storing, 100 % done.
- Status:
GET …/attachments/{ref}/index/{job_id}.
Visual recall
Searching the pages needs the late-interaction strategy on top of DAKERA_VISION (DAKERA_SCORING_STRATEGY=late-interaction, see Search and ranking). Recall then embeds the query with the visual model's text side, shortlists pages by their fixed-dimensional encoding (patch.fde; the pooled 128-d page vector serves while that stage is built) and reranks with MaxSim over the page patches. The text cross-encoder is not applied on this lane (rerank_report.skipped is visual_lane).
The visual lane is a store of its own. One store is one embedding space: the agent namespaces then hold 128-d page vectors, not text embeddings, and the text late-interaction lane (colbert slots) is replaced by the visual one (patch slots). Use a deployment (a data root) dedicated to it, and do not turn DAKERA_VISION on over an existing text store. dakera downgrade refuses a store that uses the visual lane.
Measured quality (ViDoRe, nDCG@5, visual lane): TabFQuAD 0.643 and Shift Project 0.7705. The other ViDoRe sets have not been measured.
Jobs
Transcription and image jobs run in the background and are polled until they finish.
curl -s "localhost:3000$STATUS_URL" -H "Authorization: Bearer $KEY"
# {"id":"job_1a2b3c4d_0","job_type":"transcription","status":"Running","created_at":1790000000,
# "started_at":1790000000,"completed_at":null,"progress":42,"message":"transcribing",
# "metadata":{"namespace":"uploads","attachment_ref":"sha256:9f2c…","agent_id":"my-agent",
# "memory_id":"mem_17f3…","model":"whisper-base"}}
- Where to poll:
GET …/transcribe/{job_id}orGET …/index/{job_id}with any key that can read the namespace;GET /ops/jobs/{id}andGET /ops/jobsneed a key with global admin scope. Any node of a cluster answers a poll. - Status:
Pending,Running,Completed,Failed.job_typeistranscriptionorimage-index. - Progress is what the job has done. Transcription: 2 % while the model loads (the
messagesaysloading the speech-to-text model, with the download progress when a file is being fetched), 5 % to 80 % across the audio's 30-second windows, 85 % embedding the transcript, 95 % storing, 100 % done (message:memory <id> stored). - A failed job keeps the progress it reached and carries
error: {"status": 400, "code": "INVALID_REQUEST"}with the reason inmessage. The status and code are what the same work would have answered as a request:400 INVALID_REQUEST(the audio or image cannot be processed),503 SERVICE_UNAVAILABLE(the model cannot be loaded, memory does not admit the job, the page queue timed out, or the server is shutting down),413 PAYLOAD_TOO_LARGE(the transcript or caption is over the limit),500 INTERNAL_ERROR(the model failed on valid input; the cause is in the server log). - Jobs live in memory. Their ids are
job_<run>_<n>. After a restart a poll answers404 JOB_NOT_FOUND, saying the job is unknown because the server restarted (in a cluster: because no live node holds it). Whether it finished is not known; look up the memory by thememory_idthe202returned, or start the same request again, which stores that same memory. The job table keeps at most 10,000 jobs. - Shutdown (SIGTERM): jobs start no store step any more (a job reaching it fails
503) and the store steps already running are waited for, up to 15 s, before the server's final flushes. A job still transcribing or embedding stores nothing and must be started again.
Memory, CPU and disk
None of the multimodal models is in the image (the image carries bge-large and the reranker). Each downloads from Hugging Face into the model cache (HF_HOME, /app/models in the image; mount it as a volume to download once), is converted once to an ORT-format copy next to it (disk is about the model's size again), and is then memory-mapped. With DAKERA_ATTACHMENTS or DAKERA_VISION on, whisper and the visual model are downloaded and converted in the background right after the server is ready, so the first job does not wait for them. dakera models pull whisper vision fetches them ahead of time and RUN dakera models pull --bake whisper vision bakes them into your own image; see Models and Docker.
| Model | Download | Loaded when |
|---|---|---|
whisper-base (speech to text, default) |
see the model table | the first transcription job (fetched at start) |
whisper-tiny.en |
about 151 MB | the first transcription job, when chosen with DAKERA_ |
whisper-base.en, whisper-tiny, whisper-small (chosen with DAKERA_) |
see the model table | the first transcription job |
colmodernvbert (image indexing, visual recall) |
about 966 MB | the first image job or visual recall (fetched at start) |
colbert-small (late interaction, DAKERA_MODEL) |
about 34 MB | startup (it is the embedding model) |
GLiNER (entity extraction, per-namespace extract_) |
about 782 MB | the first store that extracts entities (it waits up to 10 s, then 503 with Retry-After while the model downloads) |
A download that is cut short continues where it stopped on the next attempt or the next start; while a job waits for its model, its message shows the download.
Memory is admitted, not assumed. Every media job (and every GLiNER pass on the store path) reserves its estimated peak working memory before it decodes anything: for audio, the samples plus one 30-second window's activations (128 MiB); for an image, the raster and decode buffers read from its header plus 128 MiB per tile (a page is at most a 4 x 4 grid plus a thumbnail, 17 tiles). The reservation is granted while resident memory plus reservations plus the request stay within the memory limit (the cgroup limit, else physical memory) times DAKERA_MEM_HIGH_WATER_FRACTION, the same reading the store routes' backpressure uses. A job that does not fit first triggers a reclaim (idle models unloaded, derived caches dropped), then waits up to 10 s for running jobs to give memory back, and is then answered 503 with Retry-After: 10. There is no fixed job cap. DAKERA_MEM_BACKPRESSURE=0 turns the refusals off. On a platform where no memory limit can be read, reservations are granted.
A model's first load converts its graph to the ORT format once per model and build, and that conversion reserves its peak from the same budget before it starts (the download before it reserves nothing). When it does not fit, the job fails 503 with a message naming what the conversion needs, what the budget has and the remedy, and /health lists speech_to_text, vision_model or entity_extraction under degraded until a load is admitted. dakera models pull whisper or dakera models pull vision converts the model outside the server once; a converted model loads without that reservation.
Idle models leave memory. Whisper, the visual model and GLiNER are unloaded after DAKERA_HNSW_CACHE_TTI_SECS (default 3600 s) without use, and earlier under memory pressure; the next use loads them again from the local cache. A job that uses a loaded model counts as use. The weights are memory-mapped (shared, reclaimable page cache); what each model holds privately is its working memory while it runs.
CPU is shared. Every model's work draws from one pool of compute tokens (one per core). Query-side work (query embeddings, rerank, visual queries) is served first; bulk work (document embeddings, transcription, image pages, entity extraction) uses at most cores - 1 at once, so a transcription never takes the whole machine from recall.
Measured footprint: with every model loaded, the process heap was 530 MiB (memory-mapped model files are not counted). A small box can run everything, one heavy job at a time: jobs that do not fit wait or get a 503 to retry. The official SDKs time requests out at 30 s; the server's queueing answers within 20 s so the 503 and its Retry-After arrive first.
Errors
Every error body is JSON: {"error": "...", "code": "...", "status": 400, "details": "..."} (a 404 also carries resource).
| Status | Code | When |
|---|---|---|
501 |
FEATURE_ |
The switch is off; details say set DAKERA_ (or DAKERA_VISION) |
400 |
INVALID_REQUEST |
Malformed attachment_ref; empty upload; invalid media type; multipart body without a file part; audio or image the model cannot take; invalid agent_id, tags, importance, id, lang |
404 |
VECTOR_, resource: "attachment" |
No such attachment in that namespace (the v0.11 code is kept; resource tells the cases apart) |
404 |
JOB_NOT_FOUND, resource: "job" |
Unknown job, or one lost to a restart (the message says which) |
409 |
CONFLICT |
Deleting an attachment a memory still references |
413 |
PAYLOAD_ |
Upload over DAKERA_; transcript or caption over 100,000 characters |
413 |
QUOTA_EXCEEDED |
The namespace's max_ quota |
503 |
SERVICE_ with Retry-After |
Memory admission, model not loadable, page queue full (30 s) or timed out, writes still referencing a deleted attachment (1 s), shutdown in progress |
Verify it works
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" | jq '.attachments, .vision'showsenabled: truefor the switches you set, withmax_bytesand the model names. Without the switches the routes answer501.- Upload a small WAV: the answer is
201withcreated: true; the same upload again is200withcreated: false. - Start a transcription and poll
status_urluntilstatusisCompleted(the first job also downloads and converts the model; itsmessageshows the download). Recall by the spoken words and check the memory'sattachment_ref. - For the visual lane, upload a PNG page, index it, then call
POST /v1/memory/recallwith a question about it (late-interaction strategy on); thererank_reportof the response saysvisual_lane. curl -s localhost:3000/health | jq .degradedlistsspeech_to_text,vision_modelorentity_extractiononly while a model cannot be loaded within the memory budget.dakera_memory_budget_reserved_bytesanddakera_model_loads_totalare exported on/metrics.
Constraints
- Speech to text takes WAV only; images are PNG only, at most 64 megapixels; uploads are at most
DAKERA_ATTACHMENT_MAX_BYTES. - Pages index one at a time (about 10.7 s each on CPU).
- Jobs are kept in memory and lost on restart; the memory they stored is kept.
- The visual lane needs a dedicated data root. A deployment that uses it, or
colbert-smallorbge-m3, cannot be downgraded to v0.11.108 until it is moved back to a v0.11 model. See Rolling back to v0.11.