Upgrading from v0.11.108
This guide is for operators moving a v0.11.108 deployment to v0.12.0. It lists the steps to take before the upgrade, what happens on the first start, and every behavior change an existing deployment can notice, with what to do about each.
An unchanged v0.11.108 deployment upgrades in place: stop v0.11.108, start v0.12.0 on the same data and the same environment. It starts, keeps its data and answers as before, except for the defects v0.12.0 fixes on purpose, which are listed below. Going back is supported for every deployment on a v0.11 embedding model, encrypted or not: stop v0.12.0, run dakera downgrade, start v0.11.108 (see Rolling back to v0.11). For a summary of the release, see What's new in v0.12.0.
Upgrade checklist
- Take a backup:
POST /admin/backups. - Check the configuration with the new image: run
dakera --check-configagainst the deployment's environment and data volume. It lists every warning the first start will give (see Checking a configuration first). - Give gRPC clients an API key when authentication is on (gRPC authentication).
- Check keys pinned to namespaces and backup tooling against the new permissions model.
- Check namespace quotas. They are now enforced: a quota set on v0.11 (which ignored it) refuses writes over it (Quotas).
- Decide about the backup schedule. A stored, enabled schedule starts running at its next slot; disable it first if you do not want that.
- Point probes at
/health/ready(readiness) and/health/live(liveness). The port now binds before the models load; the Helm chart and the compose file already do this. - Re-check any threshold on
smart_score(orDAKERA_ABSTAIN_MIN_SMART_SCORE) (Recall and ranking). - Clusters: plan a short mixed period and read Cluster upgrade. Leave
DAKERA_CLUSTER_SECRETunset until every node runs v0.12.0; an encrypted cluster holds sealed changes for v0.11 nodes until they are upgraded. - Model volumes: do not run
dakera models pruneuntil you are staying on v0.12 (it removes the files v0.11.108 reads). - Stop v0.11.108 and start v0.12.0 on the same data and environment.
- Rebuild the full-text indexes once with
POST /admin/fulltext/reindexand{"rebuild": true}, as soon as the server is ready. This is required after upgrading data from v0.11.108 (details and commands). - Watch
/health:config_warnings,degraded,embed_migration(and, with encryption,GET /admin/encryption/status).
After the upgrade: rebuild the full-text indexes
Full-text indexes built by v0.11 are re-analysed under v0.12's text analysis. The rebuild does that once, so keyword search and keyword-style recall return results again. It only rebuilds derived search indexes; your memories are not touched. On a production deployment with about 18,000 memories it took about 18 seconds. Run it once the server is ready (GET /health/ready answers 200), with a key of global admin scope:
curl -X POST https://<your-dakera>/admin/fulltext/reindex \
-H "x-api-key: <global admin key>" -H "content-type: application/json" \
-d '{"rebuild": true}'
Omit namespace to cover every agent memory namespace; add "namespace": "<ns>" to rebuild one. The response lists each namespace it processed (see Multilingual search for the fields).
Docker Compose
docker compose exec dakera curl -s -X POST http://localhost:3000/admin/fulltext/reindex \
-H "x-api-key: $DAKERA_ROOT_API_KEY" -H 'content-type: application/json' -d '{"rebuild":true}'
This needs curl inside the container; the v0.12 image's healthcheck uses it. If your image does not have it, run the host-side curl above against the published port (for example http://localhost:3000).
Kubernetes
Forward the service port and run the same curl from your workstation:
kubectl port-forward svc/dakera 3000:3000 &
curl -X POST http://localhost:3000/admin/fulltext/reindex \
-H "x-api-key: $ADMIN_KEY" -H "content-type: application/json" \
-d '{"rebuild": true}'
Or run it as a one-off Job from an image that has curl, pointing at the in-cluster service URL. Adjust the service name and namespace to your release.
Check that it worked
- A keyword search returns hits:
POST /v1/namespaces/<ns>/fulltext/searchwith a common word from your data. - A short keyword recall returns memories.
If either still returns nothing, see Troubleshooting.
What happens on the first start
Nothing below needs an operator step. Recall and ingest are served throughout.
| What | How it behaves | How to watch it |
|---|---|---|
| Configuration check | Names and values v0.12 cannot use are warnings, not refusals, because the data root already holds data. They are logged as CONFIGURATION WARNING and listed in GET /health under config_warnings. The same configuration on an empty data directory (a fresh install) is refused with exit code 78. |
startup log; /health config_warnings |
| Data locations | v0.12 derives every local path from one data root (DAKERA_). Data v0.11 kept elsewhere (the write-ahead log at ./data/wal, the image's old /app/data, the tiered defaults /data/hot and /) is used where it is while the new place is empty. |
/health config_warnings (data_location) |
| Knowledge graph | Lives under the data root and survives restarts. v0.11 kept it in memory unless DAKERA_DATA_DIR was set, so a deployment without it starts with an empty graph, exactly what a v0.11 restart gave. With DAKERA_DATA_DIR set, the graph stays where v0.11 kept it, and the first start migrates it once (edges are keyed by namespace), about 0.7 s per 500,000 edges. |
Data root resolved log line |
Text endpoints without model |
v0.11.108 embedded upsert-text, query-text and batch- with bge-large whatever DAKERA_MODEL said; v0.12 embeds with the configured model. On an upgraded deployment a text namespace that v0.11 filled keeps its model, so nothing breaks. |
nothing to do |
| Embed-side migration | Memories stored with the query instruction (single store, update, consolidation and session summaries in v0.11) are re-embedded on the document side in the background. Recall ranks move a little while it runs; that is the fix taking effect. About 35 minutes per 10,000 memories on bge-large. | /health embed_migration (state: pending, surveying, running, complete, not_needed); GET / |
| Encryption re-seal | With DAKERA_, every $enc$v1$ value is re-sealed to $enc$v2$ (bound to its namespace, record and field) by a background pass; v1 values keep decrypting until then. |
GET / |
| Embedding-model record | The store records its model and embedding recipe. The default model's vectors are unchanged. A v0.11 store has no record: the running model is adopted. gte- (now CLS-pooled) and modernbert- (now with its required prefixes) changed recipe: their stored memories and text vectors are re-embedded by the embed-side migration, with no refusal. |
startup log; /health embed_migration |
| Models | The image's models (bge-large, the reranker) are read from the image itself, already converted: nothing is downloaded or converted at the first start, whatever is mounted at /app/models. A model volume kept from v0.11 still holds that image's copies, now unused (about 0.9 GB): dakera models prune frees them, but they are what v0.11.108 reads, so prune only once you are staying on v0.12. Models you downloaded there (another DAKERA_MODEL, whisper, ...) are used as before and converted once for v0.12. |
dakera models list |
| Tiered hot tier | The default is now a persistent RocksDB hot tier (v0.11: in memory). It starts empty; reads fall through to warm and cold. DAKERA_ keeps the old tier. |
/ |
| Telemetry | Only when telemetry is on; DAKERA_ and DO_NOT_TRACK are read as v0.11 read them, and an opted-out deployment still sends nothing. v0.12.0 adds lifecycle events (engine_started, engine_stopped, engine_crashed, engine_) and sends the heartbeat every 6 hours instead of 24. engine_started and the heartbeat carry the machine's hostname, and PostHog keeps the connecting IP address and derives a coarse location from it (GeoIP). No content, ids, secrets or paths are sent. See Telemetry. |
startup log line; DAKERA_ |
Stored records are written byte-for-byte as in v0.11 (the binary record format is read-only in v0.12). Index snapshots and the BM25 base format are unchanged; HNSW keeps its in-RAM vectors in f16 with no on-disk change.
Startup and configuration changes
- Explicitly configured features refuse a fresh install instead of running without themselves.
DAKERA_REDIS_URLset but Redis unreachable, aDAKERA_DISK_CACHE_DIRthat cannot be opened, an explicitly enabled ML classifier that fails, a graph file that cannot be opened, an S3 backup store that cannot be opened, gossip that cannot start, a cluster outbox directory (DAKERA_CLUSTER_OUTBOX_DIR) that cannot be opened. An upgraded deployment starts as v0.11 did, without the component, and lists it as degraded in/health(redis_cache,disk_cache,ml_classifier,knowledge_graph,backups,cluster_gossip,cluster_outbox). Redis is checked every 15 s and leaves the list as soon as it answers. - Fresh install or upgrade is decided by the data root. A data root that already holds data (a write-ahead log or snapshot, filesystem or warm-tier namespaces, a RocksDB hot tier with data) is an upgrade. With S3 storage and nothing on local disk, the bucket is listed at startup (a namespace there means an upgrade); if it cannot be listed after three attempts the node starts as an upgraded deployment, with a
deploymententry inconfig_warnings, never as a fresh install. - One server per data root. A server locks its data root (
.dakera.lock) while it runs; a second server started on it refuses to start ("REFUSING TO START ... locked by a running Dakera process"). Not locked: plain S3 (nothing on disk) and an in-memory store withDAKERA_WAL=false. The lock isflock: on NFS it holds across hosts only where the mount supports locks. - The Helm chart sets a pod security context:
dakera.podSecurityContextdefaults to uid/gid 1000 (the image's user) withfsGroup: 1000andfsGroupChangePolicy: OnRootMismatch, so a volume whose root isroot:root 0755is writable. On the first rollout Kubernetes chowns a volume whose root does not match, once. Setdakera.podSecurityContext: {}where the platform assigns uids. POST /ops/shutdownshuts the server down with the same graceful shutdown as SIGTERM: both ports stop accepting, in-flight requests finish, then the shutdown flushes run and the process exits. It needs a key of globaladminscope. A container or service manager restarts it according to its restart policy; with none, the server stays down. Any tooling that called it as a health or smoke test must stop doing so. The graceful shutdown now drains gRPC as well as REST.
Health endpoints
The REST port now answers while the server starts. v0.11 bound the port only after the embedding model had loaded, so a model that was not in the image (bge-m3, colbert-small) downloaded behind a closed port, and a Docker HEALTHCHECK or Kubernetes liveness probe restarted the process.
| Endpoint | Answer |
|---|---|
GET / |
200 as soon as the port is bound. Use it for liveness. |
GET / |
200 once the embedding model is loaded and storage answers; 503 while starting, with "starting": true, a reason and downloads (each model file with received_bytes and total_bytes). Use it for readiness. |
GET /health |
The same startup answer while starting; afterwards status, plus degraded (components running degraded, each with component and reason) and config_warnings. Unauthenticated, so reasons show paths as <path> and credential-bearing URLs as <url>. |
Every other request answers 503 with Retry-After: 5 until the models are loaded. A client that waited for the connection to be accepted should wait for /health/ready (or retry the 503). A download that was cut short continues where it stopped when the server serves byte ranges. The images' HEALTHCHECK start period is now 10 minutes, and the Helm chart's liveness probe reads /health/live behind a 10-minute startup probe.
Checking a configuration first
dakera --check-config runs the startup configuration check alone (same environment, same data root, same decision), prints what it found and exits 0 (the server would start, warnings included) or 78 (it would refuse). It starts nothing and writes nothing.
docker run --rm -e DAKERA_ROOT_API_KEY=... -v dakera-data:/data \
ghcr.io/dakera-ai/dakera:0.12.0 --check-config
Once running, the gauge dakera_config_warnings counts the config_warnings entries; alert on > 0.
Security and access
gRPC authentication and validation
v0.11.108's gRPC port (DAKERA_GRPC_PORT, default 50051) had no authentication. With DAKERA_AUTH_ENABLED on (the default), every RPC except Health now needs a key the REST API accepts, sent in the call metadata:
x-api-key: <key>
authorization: Bearer <key>
- Without a key the call fails
UNAUTHENTICATED; a key without the scope or namespace the RPC needs failsPERMISSION_DENIED, as the REST twin answers401/403. - Every authenticated call takes a token from the server-wide rate limit (
DAKERA_RATE_LIMIT_RPS/DAKERA_RATE_LIMIT_BURST); when spent the call failsRESOURCE_EXHAUSTED. - Every RPC is bounded by
DAKERA_REQUEST_TIMEOUT, andQuerybyquery_timeout_msofPUT /admin/config(DEADLINE_EXCEEDED). - While maintenance mode rejects requests, every RPC but
HealthfailsUNAVAILABLE. Calls are written to the audit log. - gRPC validates like REST:
Querywithtop_kunset (proto3 sends0) isINVALID_ARGUMENT(settop_kbetween 1 and 10,000); an unknowndistance_metricisINVALID_ARGUMENT; anUpsertwhosemetadata_jsonis not valid JSON isINVALID_ARGUMENT.
What to do: give every gRPC client a key in its metadata before upgrading, or keep DAKERA_AUTH_ENABLED=false if the deployment ran without authentication.
Permissions
A key pinned to namespaces (for example an admin key minted by POST /v1/namespaces/{ns}/keys) used to pass every check that asked for a scope only. These routes now answer 403 where v0.11.108 did not:
| Route | v0.11.108 | v0.12 |
|---|---|---|
POST /, POST /, GET / |
admin |
super_admin with no namespace restriction (a bundle carries every API key record) |
Every other /admin/* and /v1/admin/* route |
admin |
admin with no namespace restriction, except DELETE /, POST /, POST / and POST /, which a pinned admin key may still call for its own namespaces |
/admin/keys/* |
super_admin |
super_admin with no namespace restriction |
/, /ops/jobs, /ops/compact, /ops/shutdown, /ops/events, /v1/ops/metrics, /debug/config |
admin |
admin with no namespace restriction |
/v1/analytics/*, /v1/audit*, /v1/kpis, /, POST / |
admin |
admin with no namespace restriction |
GET / |
read on _ |
also the agent's namespace |
GET /v1/export, POST /v1/import |
any key | read (export) / write (import) on the agent's namespace _ |
POST /, PUT /, GET / |
the scope only | the scope on that namespace |
POST / with extra_ |
any namespace | only namespaces the calling key can reach |
Deployments with authentication off (every request is super_admin) see none of these. An unrestricted admin key can no longer download, upload or restore a backup bundle: use a global super_admin key. POST /ops/shutdown needs a key of global admin scope.
Other security changes
- Sealed values from clients are refused. A vector write whose
_textorcontentmetadata starts with$enc$v1$or$enc$v2$is refused (400; gRPCINVALID_ARGUMENT), because storing a client-supplied ciphertext made the server a decryption oracle. - Importing a v0.11 export of an encrypted deployment needs a key of global
adminorsuper_adminscope (or authentication off) and the sameDAKERA_ENCRYPTION_KEYthat sealed it. Exports taken from v0.12 carry plaintext and import with any key allowed to write the agent. - With authentication off, CORS allows loopback origins only by default, and a state-changing request from a web page of another origin is refused with
403 CROSS_ORIGIN_REQUEST_REFUSED. SDKs and curl are unaffected. - Knowledge graph edges are per agent; the existing graph is migrated on the first start.
POST /v1/memories/{id}/linksanswers404unless both ids are live memories of the agent, and path searches (GET /v1/memories/{id}/path,GET /v1/knowledge/path) stop at 10 hops or 4096 memories and answer400saying so. - Audit body capture (
DAKERA_AUDIT_REQUEST_BODY,DAKERA_AUDIT_RESPONSE_BODY) andDAKERA_AUDIT_FORMAT=textnow work; with the defaults the audit output is unchanged. - Raw vector APIs return memory content as stored (unchanged from v0.11):
/v1/namespaces/{ns}/query, batch and multi-vector query,/v1/namespaces/{ns}/export,/v1/namespaces/{ns}/records/{id}and gRPCQueryreturn an agent memory'scontentmetadata in its at-rest form (z64:for content over 256 bytes,$enc$v2$with encryption). Read memories through the memory API, which returns plaintext. - A caller-chosen extractor
base_urlis not fetched through the system proxy, and provider responses are capped at 4 MiB. An uploaded backup bundle may inflate to at most 16 timesDAKERA_MAX_BODY_SIZE(413beyond). - Retired operator key masters can be revoked with
POST /admin/encryption/keys/{key_id}/revoke-master(global admin). Deleting a namespace removes its own encryption key assignment.
Storage, data and backups
Quotas are enforced
v0.11 stored, returned and replicated namespace quotas (/admin/quotas..., and max_vectors_per_namespace of PUT /admin/config) but no write consulted them. Every write path (REST, gRPC, memory, import, re-embed, background and admin writes) now goes through one quota check. With the default hard enforcement, a write that would take a namespace over max_vectors, max_storage_bytes, max_dimensions or max_metadata_bytes is refused with 413 QUOTA_EXCEEDED before anything is written. soft logs a warning and writes; none only tracks. A deployment that set quotas on v0.11 and relies on them not biting should raise them, or set enforcement to soft or none, before upgrading.
Backups
- A stored backup schedule now runs. A backup is taken at every slot of its
cron(UTC; on the leader in a cluster) and scheduled backups beyondretention_days/max_backupsare deleted (the newest completed one is always kept; manual and uploaded backups are never deleted by it). An upgraded deployment whose stored schedule is enabled gets abackup_scheduleentry inconfig_warnings; the first backup is taken at the first slot after the start. To keep v0.11's behavior, sendPOST /admin/backups/schedulewith{"enabled": false}before that slot. - Restore without
overwritebehaves as in v0.11; with"overwrite": trueit gives the exact point in time of the backup. A restore no longer restores the backup schedule: the live schedule is kept. - A backup interrupted by a stop is marked
failedat the next start instead of stayinginprogress. - Restore does not pause serving: restore in a maintenance window.
Other storage changes
- RocksDB hot-tier writes are fsynced (
DAKERA_ROCKSDB_SYNC, defaulttrue). - Concurrent writes into one namespace share the write-ahead log's fsync.
- Memories report the TTL they were stored with on every backend; spaced-repetition TTL extensions now reach the TTL sweep.
- Forget removes a memory's sentence sub-memories, and TTL expiry removes expired memories from full-text search, the graph and the counters.
- A write whose client gave up still applies, and replicates. In a cluster it is now replicated to the peers like any write, with its tombstone for a delete, instead of waiting for the next reconciliation (up to 5 minutes). A timed-out write may have taken effect: read it back before retrying a non-idempotent change (memory stores with an
idand deletes are safe to retry). A storage backend that stops answering mid-write is no longer cut loose by the request timeout: the write waits for it, is logged at error every 30 s and counted indakera_storage_write_stalls_total; alert on any increase. - Backup schedule details. The schedule runs with its namespaces, type,
encryptandcompression. A schedule with an invalid expression, or anencrypt/compressiona backup cannot honor (lz4,encryptwithoutDAKERA_ENCRYPTION_KEY), is refused with400when set. A stored one with an invalid expression logs an error and takes no backup.DELETE /admin/backups/{id}accepts aninprogressbackup that no node is taking. - A restore keeps the live schedule. When the backup's schedule differs, the restore logs it, and the operator adopts it with
POST /admin/backups/schedule. Revoked API keys stay revoked, and keyring records are restored only where missing. - Session ids belong to their agent.
POST /v1/sessions/startwith theidof another agent's session answers409(v0.11 replaced that session, owner included). - Deep agent listings are cheaper.
GET /v1/agents/{id}/memoriespages cost in proportion tooffset + limit(they grew with the square ofoffset), and a page read while the agent is written lists each memory once. - Import and export are complete.
POST /v1/importobeys the 100,000-byte content limit: an item over it is refused and named in the job'serrors.GET /v1/exportcarries plaintext content and each memory's TTL, and no longer exports sentence sub-memories.
Recall and ranking
Recall no longer ranks what earlier recalls returned higher. v0.11 raised the importance of every memory a recall returned and scored the access count; both are query-independent, so the results of one query crowded out more relevant memories in every later one. In v0.12 a recall records access_count and last_accessed_at but changes neither importance nor the score. The score's recency term counts from the memory's created_at. POST /v1/memory/feedback (upvote) still raises importance.
smart_score values move with this. The score is w_vec * relevance + w_imp * importance + w_rec * recency; the weights sum to:
| Route | Weight sum |
|---|---|
| Default and importance-weighted | 0.88 |
| Multi-hop queries | 0.73 |
| Temporal-inference queries | 0.98 |
v0.11 gave the rest to the frequency term. A memory that no recall, feedback or session promotion had touched scores exactly as in v0.11; one that v0.11 recalls had returned scores up to 0.27 lower, and one last used long after it was stored up to 0.13 lower. An absolute cut on smart_score calibrated on v0.11 (DAKERA_ABSTAIN_MIN_SMART_SCORE, or a client filtering on the score) can drop such memories: re-check it against v0.12's scores. Relative order within a query is what changed on purpose.
Other behavior changes:
- Importance decay no longer compounds. Decayed importances are higher than v0.11 would have left them.
- Auto-pilot dedup merges only true duplicates, and
POST /v1/knowledge/deduplicateuses the same safeguards (soduplicates_mergedcan be lower thanduplicates_found). - Recall admission: at most two recalls per core (at least four) run at once; the rest queue and answer
503withRetry-Afteronly after 20 s without progress (at most half ofDAKERA_REQUEST_TIMEOUTwhen that is shorter). - An omitted
distance_metricresolves to the namespace's metric instead of always cosine. - The rerank report is
rerank_report(applied,candidates,skipped);rerankin a request stays the on/off switch. - Page caps are stated in
x-dakera-limit;POST /v1/memories/recall/batchreturns at most 10,000 memories.
More recall and write behavior
- Feedback, importance and update writes apply to the memory as stored.
POST /v1/memory/feedback,POST /v1/memories/{id}/feedback,POST /v1/memory/importance,PATCH /v1/memories/{id}/importanceandPUT /v1/memory/update/{id}used to write back their read copy, losing an access count or content update that landed meanwhile. They now write only over an unchanged memory and retry; a memory that keeps changing through five attempts answers503withRetry-After: 1. On an expired memory they answer404. One that would take the namespace over its storage quota is refused. - Importance set by a client decays from the moment it was set. An importance a client sets, or feedback changes, decays from that moment. Importances boosted by past v0.11 recalls are kept and decay from there.
GET /v1/agents/{id}/wake-upstill orders by importance and recency of use. - Consolidation refuses memories it would lose data from.
POST /v1/memory/consolidatewith explicitmemory_idsanswers400for a memory with an attachment, a TTL, user metadata or that has expired. - Admission queueing ends by the 20 s mark. The 20 s counts the whole request: a recall that waited for its slot, or a store that waited for memory or a model load, waits only what is left of it. 20 s, not 30 s, because every official SDK times a request out at 30 s.
POST /v1/memories/recall/batchnow waits for an admission slot like every recall. - Stored memories are embedded on the document side, and
gte-modernbert-baseandmodernbert-embed-baseare corrected against upstream, so their vectors are re-embedded by the migration.
Errors and API responses
Every error body is JSON (
{"error", "code", "status"}), including framework refusals: invalid body (400), wrong shape (422), wrong content type (415), unknown path (404ROUTE_NOT_FOUND), wrong method (405), timeout (408) and the server-wide rate limit (429).Every
503carriesRetry-After. Two refusals that no retry can change are now501 NOT_IMPLEMENTEDwith the settings to change indetails:DAKERA_SCORING_STRATEGY=late-interactionon a model without a late-interaction recipe or underDAKERA_TIERED=1, and an invalidDAKERA_EXTRACTOR_*configuration.A body over the size limit is
413(attachments overDAKERA_ATTACHMENT_MAX_BYTES, imports overDAKERA_MAX_BODY_SIZE).Namespace config:
PATCH /v1/namespaces/{ns}/configmerges what it is sent and answers400naming an unknown field;PUTis a full replacement (omittedentity_typesclears the list). Persisted memory policies are re-applied at startup (v0.11 lost them at every restart).PUT /admin/configsettings take effect (max_vectors_per_namespace, cache settings, rate limit,query_timeout_ms), are re-applied at startup over their environment variable, are listed inruntime_overridesand replicate to every cluster node.query_timeout_msis a deadline on query, recall and search routes and gRPCQuery(504 QUERY_TIMEOUT; default0= none).default_index_typeaccepts onlyhnsw;GETreports the effective values.404bodies name theirresource, and error text is redacted everywhere.resourceis one ofnamespace,vector,memory,session,job,attachment,path,encryption_key,api_key. Memories, sessions, jobs, attachments and graph paths keep the v0.11 codeVECTOR_NOT_FOUND, and theirerrorloses the "Vector not found:" prefix. An unknown API key isAPI_KEY_NOT_FOUND; an unknown/ops/jobs/{id}isJOB_NOT_FOUND(v0.11 answered a barenull). Every403'sdetailsnames scopes as keys are created with them (required: admin, actual: read).New error codes:
ROUTE_NOT_FOUND,METHOD_NOT_ALLOWED,UNSUPPORTED_MEDIA_TYPE,REQUEST_TIMEOUT. Their statuses are unchanged (the timeout stays408), and the headers (Allow,Retry-After,X-RateLimit-*) are kept. The authentication middleware's401keeps its v0.11 fields and gainscodeandstatus.The re-embed drain answers inside the request timeout.
POST /admin/reembed/drainanswers after at most 4/5 ofDAKERA_REQUEST_TIMEOUTwithtimed_out: true; a second drain is409 CONFLICT(was503).Validation names its rule;
../is accepted in text. Metadata values, tags and filter values containing../or..\are no longer refused as path traversal (ids and keys still are).Maintenance mode is held by the node that receives
POST /admin/cluster/maintenance/enable: not replicated, not persisted, and it refuses requests only on that node, only whennode_idsis empty or lists it. To take several nodes into maintenance, send the request to each. Its503carriesRetry-After;reject_requestsis applied;POST /admin/cluster/maintenance/disablewithoutforceis409while jobs run.Smaller API corrections: the importance routes answer
409for a memory whose content does not open on this node; namespace event streams are capped at 128 per key (1,024 per process);POST /v1/namespaces/{ns}/explainwithexecute: truereturns realactual_stats; an admin import into another agent's namespace rewrites the memories'agent_idto that agent.
Logs, metrics and alerts
- Quieter logs. Per-request INFO lines of the data path are DEBUG; the audit log still records every request. A request without a valid key is a DEBUG line with at most one WARN per minute (
dakera_errors_total{type="auth"}counts every one). Log messages no longer carry internal ticket names: alerts matching those words need updating (for exampleSEC-3 keyringis nowencryption keyring). The compose dev profile logs atinfo. - Refusals are counted.
dakera_http_requests_totalanddakera_http_request_duration_secondsnow include429,401/403,408,413and unmatched paths (path="unmatched"), so expect those statuses in dashboards; the traffic itself is unchanged.dakera_active_requestsno longer drifts on disconnects. - New metrics:
dakera_model_loads_total{model,outcome},dakera_model_load_duration_seconds{model},dakera_backups_total{outcome},dakera_wal_write_failures_total,dakera_memory_read_failures_total,dakera_cache_redis_failed_writes_total,dakera_cache_redis_failed_publishes_total,dakera_cache_redis_failed_evictions_total,dakera_cluster_sync_held_entries,dakera_memory_budget_reserved_bytes,dakera_component_degraded{component}anddakera_config_warnings.wal_replayin/healthdegradedmeans the write-ahead-log replay dropped entries the store refused. - Alert rules and dashboard:
monitoring/dakera.rules.ymlin dakera-deploymonitoring/(v0.12.0 versions arrive with the dakera-deploy v0.12.0 release) covers these signals plus availability rules (target down, 5xx and 503 ratios, memory-budget refusals, cold-tier circuit and backlog, RocksDB write stop, dropped cluster changes, clock skew, corrupt records). The Grafana dashboard reads only metrics the server emits; panels that read names nothing emitted are gone.
Cluster upgrade
- Cluster mode requires
DAKERA_CLUSTER_SECRET(at least 16 characters, the same on every node) andDAKERA_CLUSTER_NODE_IDon fresh installs. v0.11 had no secret: node-to-node routes and gossip were open to anyone who reached the ports. An upgraded node without them starts as v0.11 ran it and isdegradedin/health(cluster_auth) until both are set. Without a node id it generates one once and keeps it in{data root}/cluster-node-id. - Rolling upgrade: leave the secret unset while v0.11 and v0.12 nodes are mixed (a v0.12 node with a secret refuses v0.11 peers, which send none). Once every node runs v0.12, set the same secret and a stable node id on all of them and restart them. Keep the mixed period short.
- An encrypted cluster (
DAKERA_ENCRYPTION_KEY) upgrades node by node, but v0.12 nodes hold changes carrying$enc$v2$text for peers that still run v0.11, in a durable outbox, and deliver them once the peer runs v0.12 (/healthlistscluster_sync_held; the metric isdakera_cluster_sync_held_entries). Clients reading from a v0.11 node do not see what was written on v0.12 nodes until it is upgraded. Deletes and key revocations are not held. - Cluster nodes must not share one S3 bucket as their store. Every cluster node on S3 keeps a heartbeat marker in its bucket. A fresh bucket that finds another node's live marker refuses to boot (exit 78); an existing one starts with a warning and
shared_storedegraded. The repository's composehaprofile now gives each node its own bucket. - Backups with one bucket per node: a backup's files live in the backup store of the node that took it (
node_idinGET /admin/backups). Restore it through that node, or give the nodes one shared backup bucket. - During a mixed period, a v0.12 node applies what a v0.11 node sends as made when it arrives, as v0.11 does. A change a v0.11 node has no route for (namespace and runtime configuration, attachments, record writes) is kept in the sending node's durable outbox and delivered once that node runs v0.12; it holds none of the node's other changes back.
- Encrypted clusters: v0.12 nodes tell v0.11 peers apart by the
X-Dakera-Peer-Protocolheader. Changes without encrypted text are not held: API key deletes and deactivations, quotas, memory, namespace and attachment deletes reach the v0.11 nodes, so a key revoked on a v0.12 node stops authenticating there too. With a stableDAKERA_CLUSTER_NODE_IDthe held changes are delivered when the upgraded node answers; without one, the upgraded node gets the data from its startup reconciliation instead. The outbox holds at most 64 MiB; what it evicts is repaired by reconciliation. - The compose
haprofile now gives each node its own bucket:dakera-1keepsdakera,dakera-2anddakera-3use the new bucketsdakera-2anddakera-3. To move an existing compose deployment, stop it (docker compose --profile ha down, volumes kept), pull the new file and start it again;dakera-2anddakera-3start on empty buckets and are filled from their peers by the startup reconciliation (watchGET /admin/cluster/status,replication.startup_reconciled, and comparevectors/countacross the nodes). Set your ownDAKERA_CLUSTER_SECRET: the compose files carry a development default. - Shared-bucket check: each node's marker is
_dakera_cluster/nodes/<node id>.json, refreshed every 30 s; another node's marker younger than 120 s means the bucket is shared. Standalone replicas (cluster mode off) sharing a bucket are not checked. - Deleting a backup that one node took also deletes its files: the node that took it removes them from its bucket when the delete reaches it. Backups taken before the upgrade have no
node_id. - Cluster timings now all follow
DAKERA_ELECTION_LEASE_MS(default 15000, accepted 1000 to 600000).
Environment variables
Names that now only warn
Set by v0.11 deployments but never read, or no longer read. On an upgraded deployment each is a warning; on a fresh install it refuses the boot. Remove them, or set what the right column says.
| Name | Set instead / why |
|---|---|
DAKERA_ |
DAKERA_ (the L2 read cache) |
DAKERA_ |
DAKERA_ |
DAKERA_, DAKERA_ |
AWS_, AWS_ |
DAKERA_ |
RUST_LOG |
DAKERA_API_KEY |
the client-side key of the SDKs and the MCP server (Dashboard 0.4.0 has no key: operators sign in with their own); the server's is DAKERA_ |
DAKERA_ |
DAKERA_STORAGE |
DAKERA_ |
DAKERA_ |
DAKERA_ |
DAKERA_ |
DAKERA_DEV_MODE, DAKERA_, DAKERA_ |
nothing: set by v0.11's compose and README, never read |
Removed variables
Read by a v0.11 release, not by v0.12 (a warning on an upgraded deployment):
| Variable | What happened |
|---|---|
DAKERA_ |
Removed: the selector only ever built HNSW, which every namespace uses |
DAKERA_RANKER, DAKERA_, DAKERA_ |
The opt-in GBDT ranker is removed; the linear ranker used by default is the only one |
DAKERA_ |
The CPU embedder runs one text per worker; DAKERA_ bounds a text |
DAKERA_ |
Build-time only |
DAKERA_, DAKERA_, DAKERA_, DAKERA_ |
Already removed during v0.11 |
Variables added in v0.12
All opt-in or deploy inputs; a deployment that sets none behaves as before apart from the changes above. DAKERA_HOT_TIER is not new, but its default changed to rocksdb (memory keeps the v0.11 tier).
| Variable | Purpose |
|---|---|
DAKERA_, DAKERA_, DAKERA_, DAKERA_ |
Attachments and speech to text (Multimodal memory) |
DAKERA_VISION, DAKERA_, DAKERA_ |
Image indexing and the visual recall lane |
DAKERA_RECORDS, DAKERA_, DAKERA_ |
Records: one vector plus extra representations |
DAKERA_ |
late- reranks with MaxSim |
DAKERA_ |
Truncation length of the text models (bge-m3 default 2048) |
DAKERA_, DAKERA_, DAKERA_ |
Multilingual full-text analysis and query routing |
DAKERA_ |
Bits per dimension (1-8) of the RaBitQ search mode, selected with DAKERA_ (an existing variable that gained the rabitq value) |
DAKERA_ |
Acknowledge that the store was embedded by another model |
DAKERA_ |
Skip pin verification for a deliberately different local model export |
DAKERA_ |
Cap how many candidates the cross-encoder scores |
DAKERA_, _MODEL, _BASE_URL, _TIMEOUT_SECS, _ |
Server-default entity extraction provider |
DAKERA_, DAKERA_, DAKERA_ |
Canonical names of the server-wide rate limit (RATE_LIMIT_RPS and RATE_ stay as deprecated aliases) |
DAKERA_, DAKERA_, DAKERA_ (default 604800), DAKERA_ (default 60000) |
Cluster authentication and replication bounds |
DAKERA_ (default 15000), DAKERA_ |
Cluster failure-detection scale and election |
DAKERA_ |
Knowledge graph file (:memory: for in-memory) |
DAKERA_ (default true) |
Fsync RocksDB hot-tier writes |
DAKERA_, DAKERA_ |
S3 deadlines |
DAKERA_ |
Start although the environment has invalid values or unknown names |
The cluster timings all follow DAKERA_ELECTION_LEASE_MS, accepted between 1000 and 600000 (a value outside is invalid like any other); scheduled reconciliation runs every 5 minutes. RATE_LIMIT_RPS and RATE_LIMIT_BURST still work as deprecated aliases, with a warning. Several tuning variables that appeared during v0.12 development (cluster timings other than the lease, DAKERA_RECALL_TRACK_ACCESS, DAKERA_EMBED_* threading knobs and others) were removed before release. A deployment that ran a v0.12 pre-release and still sets one starts with a warning if it holds v0.11 data; a deployment that a pre-release installed refuses to start naming the variable (remove it, or start once with DAKERA_CONFIG_LENIENT=1).
Names that do nothing where they are set
A variable set where it has no effect is logged and listed in config_warnings as env_ignored, never refused: DAKERA_MODEL other than bge-large together with DAKERA_TIERED=1 (the tiered embedding engine always embeds with bge-large; tiered storage is DAKERA_TIERED_STORAGE), tiered-storage settings without DAKERA_TIERED_STORAGE=true, DAKERA_NODE_ID different from DAKERA_CLUSTER_NODE_ID, and DAKERA_GRPC_PORT with gRPC off.
Booleans
A fresh v0.12 install reads 1/true/yes/on and 0/false/no/off (trimmed, any case) as written on every boolean variable. v0.11.108 read some spellings the other way (DAKERA_GRPC_ENABLED=yes disabled gRPC, DAKERA_WAL=off left the log on, DAKERA_AUTH_ENABLED=no left authentication on). An upgraded deployment keeps v0.11.108's reading for this release; each such value is listed in config_warnings (env_reading) with the true or false spelling that both versions read alike.
One value refuses on every deployment: DAKERA_WAL_SYNC=periodic:0, a write-ahead log that never fsyncs. Set everywrite, periodic:N with N greater than 0, or manual.
Docker model store and air-gapped installs
The default image ships bge-large and the reranker, already converted, and needs no network (--network none works), writes nothing at start, and works with a read-only root filesystem. Anything else is downloaded on first use into the model cache (HF_HOME, /app/models in the images).
- A model volume kept from v0.11 holds copies of the models the image now ships (about 0.9 GB):
dakera models prunefrees them, but they are what v0.11.108 reads, so prune only once you are staying on v0.12. - The reranker and GLiNER are now pinned by SHA-256. A v0.11 cache may hold another upstream revision: v0.12 replaces it at first load. With
HF_HUB_OFFLINE=1nothing can be fetched, so fetch the pinned files where the Hub is reachable (dakera models pull --dir <model cache> reranker gliner) and copy them in. - Downloads honor
HTTPS_PROXY,HTTP_PROXY,ALL_PROXY,NO_PROXY,HF_ENDPOINT(a mirror) andHF_HUB_OFFLINE. v0.11 connected tohuggingface.codirectly whatever the environment said. - To seed an air-gapped host, run
docker run --rm -v "$PWD/models:/seed" ghcr.io/dakera-ai/dakera:0.12.0 models pull --dir /seed bge-m3 whisper visionon a connected machine, copymodels/over, mount it at/app/modelsand setHF_HUB_OFFLINE=1.
The full reference is in Model store and air-gapped installs.
Models, media and inference
A store into a GLiNER namespace no longer waits for a first download: the model loads in the background, and a store waits up to 10 seconds, then answers
503withRetry-After. Rundakera models pull glinerahead of time.Speech to text defaults to
whisper-base. Speech is new in v0.12.0 and opt-in. The default model iswhisper-base(multilingual, language auto-detected), not English-onlywhisper-tiny.en: a deployment that turns on attachments and sets noDAKERA_WHISPER_MODELtranscribes withwhisper-base. SetDAKERA_WHISPER_MODEL=whisper-tiny.ento keep the lightest English model.Idle models (the visual model, whisper, GLiNER) are unloaded after
DAKERA_HNSW_CACHE_TTI_SECS(default 1 hour).Media jobs and entity extraction reserve their estimated peak memory against the memory limit before they start, and answer
503withRetry-Afterwhen it does not fit.DAKERA_MEM_BACKPRESSURE=0turns these refusals off.The reranker and GLiNER are pinned by SHA-256; the first load replaces another upstream revision with the pinned file (a download of about 570 MB and 780 MB). The one-time ORT-format conversion of a model's first load reserves its peak from the memory budget: where it does not fit, the request or job is answered
503naming the conversion's size, the budget and the remedy, and/healthshows the model as degraded. A server with about 2 GiB and entity types should rundakera models pull glineronce.Proxy variables: the proxy is chosen as curl chooses it (
https_proxyfirst for the Hub's URLs,all_proxyonly when neither is set); HTTP and SOCKS (4, 4a, 5, 5h) proxies work. A proxy value the downloader cannot use (anhttps://proxy URL, an unknown scheme, an IPv6 proxy address) fails the download naming the variable; use the proxy'shttp://address. Addhuggingface.cotoNO_PROXYif the proxy cannot reach it.Media jobs (new in v0.12): the first job of a model loads it in the job, so the
202no longer waits for a download;progressreflects the work; a failed job carrieserror: {status, code}; ids arejob_<run>_<n>, and a poll of a job lost to a restart answers404 JOB_NOT_FOUND. In a cluster any node answers a job poll. A job with noidstores under a derived id, so a resubmission replaces instead of duplicating. Media that cannot be decoded (a non-WAV transcription, an image over the pixel limit) is refused with400before a job exists.Vision lane: the visual model's prompt cap is its own (2,048 tokens for
colmodernvbert);DAKERA_MAX_SEQ_LENGTHapplies to the text models only.
Notes for upgrading
- v0.11 SDKs and the startup window: wait on
GET /health/readybefore sending traffic to a starting server. - v0.11 Go SDK and namespace config: clear
entity_typeswithPUT /v1/namespaces/{ns}/config. - Persistent storage: keep the data root on a persistent volume, also with S3 storage.
- Metrics: about 25 Prometheus series per active agent namespace; size the scrape for your agent count.
- Release gate: every release is published only after the upgrade, topology, cluster and rerank checks and the benchmark gates pass on the tagged commit.
See also What's new in v0.12.0, Configuration, Deployment and Security.