Models and Docker
Dakera runs its models itself: an embedding model (always), a cross-encoder reranker (on by default), and — when you turn them on — speech-to-text (whisper), image indexing (the vision model) and entity extraction (GLiNER). This page says where they come from, how to have them ready before the server starts, and how to see and clean what is on disk.
It is written for operators of the Docker images, the Helm chart and the binary, including hosts without network access. The variables that select each model are listed in the Configuration reference.
Quick reference
# Ready in seconds, no network needed: the image ships its default models.
docker run -p 3000:3000 -e DAKERA_ROOT_API_KEY=dk-mykey ghcr.io/dakera-ai/dakera:0.12.0
# What is on disk, where, how big — and what can go.
docker exec <container> dakera models list
docker exec <container> dakera models prune --dry-run
# Another model, fetched ahead of the first start (into the model volume):
docker run --rm -v dakera-models:/app/models ghcr.io/dakera-ai/dakera:0.12.0 models pull bge-m3
# Your own image with more models baked in:
# FROM ghcr.io/dakera-ai/dakera:latest
# RUN dakera models pull --bake whisper vision
What the images ship
| Image | Models inside | Size |
|---|---|---|
ghcr. (:cpu, :<version>; amd64 + arm64) |
bge-large (embedding, INT8) and bge- (reranker, INT8), ready to map |
~0.8 GB (783 MB measured, amd64) |
ghcr. (:cuda; amd64, NVIDIA GPU required) |
bge-large FP32 and the reranker, for CUDA |
~4–5 GB |
The shipped models live in the image itself (/usr/local/share/dakera/models),
already converted to the form the server memory-maps, and are read before
anything else. A default start therefore:
- needs no network (
--network noneworks); - downloads, converts and writes nothing;
- works with a read-only root filesystem;
- is not affected by anything mounted at
/app/models— a volume kept from an older image, a bind mount, a Kubernetes volume.
Everything else is downloaded on first use into the model cache:
HF_HOME, which is /app/models in the images (a VOLUME). Mount a volume
there to download each model once instead of at every new container.
| Model | Name for dakera models |
Turned on by | Download | Loaded |
|---|---|---|---|---|
| bge-large (default embedding model) | bge-large |
default | shipped | at start |
| another embedding model | bge-m3, colbert-small, gte-, … (DAKERA_MODEL names) |
DAKERA_MODEL |
23 MB – 570 MB | at start |
| bge-reranker-v2-m3 | bge- (or reranker) |
default (DAKERA_) |
shipped | in the background after start |
| bge-reranker-base | bge- |
DAKERA_ |
279 MB | in the background after start |
| whisper-base (speech to text, default) | whisper-base (or whisper) |
DAKERA_ |
see the model table in Multimodal memory | first transcription (fetched in the background at start) |
other Whisper models (whisper-tiny.en ~151 MB, whisper-base.en, whisper-tiny, whisper-small) |
the model name | DAKERA_ and DAKERA_ |
see the model table in Multimodal memory | first transcription |
| colmodernvbert (image indexing) | colmodernvbert (or vision) |
DAKERA_VISION |
~1 GB | first image job (fetched in the background at start) |
| GLiNER (entity extraction) | gliner |
per namespace (extract_ + entity_types) |
~782 MB | first store that extracts entities |
Every model file is checked against a pinned SHA-256 before it is used; a corrupt or truncated file is deleted and fetched again once.
When is the server ready?
GET /health/liveanswers200as soon as the port is bound (use it for liveness).GET /health/readyanswers200once the embedding model is loaded and storage answers (use it for readiness). While starting it answers503with what it is doing and the progress of every download.- Nothing else gates readiness. The reranker loads in the background (reranked
recalls are served without the rerank until it is ready, and
/healthlists it underdegradedmeanwhile). Whisper and the vision model, when turned on, are downloaded and converted in the background right after the server is ready, and loaded by the first job; a job that arrives earlier shows the download in itsmessage. A store into a GLiNER namespace waits up to 10 seconds for the model, then gets503withRetry-Afterand the download progress. - A model's first load converts it to the ORT format once, and that
conversion reserves its peak from the memory budget (GLiNER ~780 MB, the
visual model ~1 GB). On a small memory limit it may never fit: the store or
job is then answered
503with what the conversion needs, what the budget has, and the fix, and/healthlistsentity_extraction,speech_to_textorvision_modelunderdegraded. Convert it outside the server —dakera models pull gliner(orwhisper,vision) into the model cache, an init container (dakera.models.pull), or a derived image — after which the load needs only its session; or raise the memory limit.
Journeys
docker run
docker run -d --name dakera -p 3000:3000 \
-v dakera-data:/data \
-v dakera-models:/app/models \
-e DAKERA_ROOT_API_KEY=dk-mykey -e DAKERA_STORAGE=filesystem \
ghcr.io/dakera-ai/dakera:0.12.0
/data holds your data; /app/models only what the image does not ship. With
the default models you can leave the model volume out.
Docker Compose
Mount one dakera-models volume at /app/models on every Dakera service. Replicas can share it: a model not in the image is
downloaded once (the other replicas wait for that download instead of
starting their own) and converted once. With a service named dakera:
docker compose run --rm dakera models pull bge-m3 # ahead of switching DAKERA_MODEL
docker compose run --rm dakera models list
Kubernetes (Helm)
dakera:
models:
pull: [configured] # an init container pulls what the env configures
persistence:
enabled: true # keep the model cache across pods
accessMode: ReadWriteMany # one cache for every replica
readOnlyRootFilesystem: true # the chart mounts the model cache and /tmp
extraEnv:
- name: DAKERA_MODEL
value: bge-m3
pull takes model names or configured (the embedding model, the reranker,
whisper with DAKERA_ATTACHMENTS, the vision model with DAKERA_VISION, as the
container's environment sets them). The pod starts with them on disk.
The pod runs as the image's user (uid/gid 1000) with fsGroup: 1000
(dakera.podSecurityContext, fsGroupChangePolicy: OnRootMismatch), so the
data and model-cache volumes are writable by it even on a provisioner whose
volume root is root:root 0755. A cluster that assigns its own uids (OpenShift's
restricted SCC) sets dakera.podSecurityContext: {} and lets the platform
choose.
Air-gapped
The default image needs nothing. For other models, on a machine with access:
docker run --rm -v "$PWD/models:/seed" ghcr.io/dakera-ai/dakera:<version> \
models pull --dir /seed bge-m3 whisper vision
Copy models/ to the air-gapped host and mount it at /app/models (or point
HF_HOME at it). Set HF_HUB_OFFLINE=1 so a file that is missing fails at
once with its name instead of waiting on the network. A cached model that does
not match its pin (another upstream revision, e.g. a v0.11 cache's reranker or
GLiNER) is left in place offline — never deleted with nothing to replace it —
and its load fails naming the file. Alternatively build a
derived image (below) where the network is available and ship the image.
The embedding model, whisper and the vision model can also be read from an
operator directory as before (DAKERA_MODEL_PATH, DAKERA_WHISPER_MODEL_PATH,
DAKERA_VISION_MODEL_PATH); a read-only directory there no longer costs a
conversion at every start (the converted copy goes to the model cache once).
Only when the model cache cannot be written either does the copy go to the
temp directory, and then to a private directory of the server's user
(<tmp>/dakera-ort-<uid>, mode 0700); a copy lying in the shared temp
directory itself is never used.
Behind a proxy, a mirror, gated models
The downloader reads the usual variables:
| Variable | Effect |
|---|---|
https_proxy / HTTPS_PROXY, then all_proxy / ALL_PROXY |
the proxy for the Hub's https:// downloads, chosen as curl does: the scheme's own variable first, lower case before upper case, an empty value counts as unset |
http_proxy / HTTP_PROXY, then all_proxy / ALL_PROXY |
the same for a plain-http:// mirror (HF_ENDPOINT) |
no_proxy / NO_PROXY |
hosts (and their subdomains), IP addresses and CIDR blocks (10.0.0.0/8) reached directly; * for all. Decided for every redirect hop, so a mirror listed here can redirect to a CDN that is reached through the proxy |
HF_ENDPOINT |
a Hugging Face mirror or internal proxy instead of https:// |
HF_TOKEN (or HUGGING_) |
sent to the Hub's own origin only, never to a redirect target |
HF_ |
never download; a missing file is an immediate error naming it |
Proxy URLs: http://host:port (no scheme means http://; port 1080 when none
is given, as with curl), and socks4://, socks4a://, socks5://,
socks5h:// (with SOCKS 5 the proxy resolves the Hub's name). Credentials go in
the URL (http://user:pass@host:port, percent-encoded). An https:// proxy —
a TLS connection to the proxy itself — is not supported, and neither is any
other scheme or an IPv6 proxy address: the download fails with an error that
names the variable and the scheme instead of silently connecting directly. Use
the proxy's http:// address; the download itself stays TLS end to end
through the proxy's CONNECT tunnel.
GPU
ghcr.io/dakera-ai/dakera:gpu ships the FP32 embedding model and the reranker
for CUDA and refuses to start without an NVIDIA GPU. Nothing is converted on
the GPU path. Speech, vision and GLiNER run on the CPU there too and download as
on the CPU image. With DAKERA_USE_GPU=1 (the GPU image's default) dakera models pull fetches what the CUDA sessions load — the FP32 embedding export
and the reranker's ONNX, nothing converted — and list reports those files. The
Helm chart's models init container runs the binary directly (command: ["dakera"]), not the GPU image's entrypoint, which refuses to start without a
GPU.
Switching the embedding model
Pull the new model first so the switch does not wait on a download. The store
records the model that embedded it, so a different one needs
DAKERA_ALLOW_MODEL_CHANGE=1 for one start and a re-embed with
POST /admin/namespaces/migrate-dimensions and reembed_same_dimension: true
(see Multilingual search):
docker run --rm -v dakera-models:/app/models ghcr.io/dakera-ai/dakera:<version> models pull bge-m3
Turning on attachments / vision later
Set DAKERA_ATTACHMENTS=1 (and DAKERA_VISION=1) and restart: right after the
server is ready it downloads and converts whisper (and the vision model) in the
background. To have them before the restart, models pull whisper vision into
the model volume, or bake them into your image.
Baking your own image
FROM ghcr.io/dakera-ai/dakera:0.12.0
RUN dakera models pull --bake whisper vision
--bake writes into the image's own model store, in the final form (the
converted copy only), with the image's own binary — so the copies match it
exactly and no volume can hide them. Pull a specific set (bge-m3,
gliner, …) the same way. Rebuild the derived image when you move to a new
Dakera version (a baked model is tied to the build that baked it).
Upgrading the image with a persisted model volume
Nothing to do: the shipped models are read from the new image, whatever the volume holds. Models downloaded into the volume are converted once for the new version at their first use. Afterwards, free what only the old version used:
docker run --rm -v dakera-models:/app/models ghcr.io/dakera-ai/dakera:<version> models prune --dry-run
docker run --rm -v dakera-models:/app/models ghcr.io/dakera-ai/dakera:<version> models prune
A volume kept from v0.11 holds a copy of the models the image now ships
(~0.9 GB, plus v0.11's converted copies): prune removes them. Those copies
are what v0.11.108 reads, so prune once you are staying on v0.12: a rollback
on a pruned volume downloads them again (prune's output says how much of what
it removed was such a copy; air-gapped, re-seed with models pull --dir /app/models default before starting v0.11.108).
Several replicas on one model volume
Supported: a file one replica is downloading is waited for by the others (and
continued if that replica dies; if it stops receiving bytes for two minutes —
a suspended dakera models pull, a hung process — the others download the file
themselves instead of waiting on it for good), a model is converted once (the
others wait for the conversion), and every file is published atomically. During a rolling
upgrade each version converts its own copy of the non-shipped models.
Disk-constrained hosts
- The default image needs no model volume at all.
dakera models listshows every model's size and the space taken by copies of other versions, abandoned partial downloads and unknown files.dakera models pruneremoves what no start of this version reads;--unusedalso removes the models this configuration does not use (GLiNER is kept: it is enabled per namespace).
dakera models reference
dakera models list [--json]
dakera models pull [--bake | --dir <root>] <model>... | default | configured
dakera models prune [--dry-run] [--unused] [--dir <root>]
- list — one line per model: where it is (
image,cache,cache (not converted),partial,-), its size on disk,*when this configuration uses it; then the model cache's other contents and whatprunewould free.--jsonfor scripts. WithDAKERA_USE_GPU=1(the GPU image) a model's status is that of the files a CUDA start loads (the FP32 embedding export, the reranker's ONNX). The converted copies this version made for an operator directory (dakera/local/…) and the GPU exports count as in use, not as files nothing uses. - pull — download, verify (pinned SHA-256; tokenizers must parse) and
convert, with retries that continue a cut download. Default target: the model
cache the server reads (
HF_HOME); a model the image already ships is skipped.--dir <root>fills another directory with the cache's layout (pointHF_HOMEat it, or mount it at/app/models).--bakefills the image's model store (forRUNin a Dockerfile). Words: model names and aliases (DAKERA_MODELnames,bge-reranker-v2-m3,bge-reranker-base,gliner,whisper-base,whisper-tiny.en,whisper-base.en,whisper-tiny,whisper-small,colmodernvbert),reranker,whisper(the default,whisper-base),vision,default(what the image ships),configured(what this environment loads). Exit code 1 when a model could not be pulled. - prune — removes from the model cache: converted copies made by other
versions, partial downloads and temp files nobody is writing, and the cache's
copies of models the image ships.
--unused: also the models this configuration does not use. Never touches the image, a file it does not recognize, a file a kept model shares, or anything under an operator model directory (DAKERA_MODEL_PATH,DAKERA_WHISPER_MODEL_PATH,DAKERA_VISION_MODEL_PATH— even one inside the cache), and never deletes through a symbolic link: a model directory (or file) in the cache that links elsewhere is left alone and listed as kept.--dry-runlists without removing. Replicas on the same volume running another version reconvert what it removed of theirs.
The command runs with the same environment as the server (HF_HOME,
HF_ENDPOINT, proxies, DAKERA_MODEL, …) and exits; the server is not
started.