NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

Operations

This runbook is for operators who run Dakera v0.12.0 in production. It covers health and readiness probes, checking a configuration before it runs, backups and restores, the background re-embed migration, maintenance mode, shutdown, going back to v0.11.108, and what to scrape and alert on.

Health, readiness and liveness

Three unauthenticated endpoints, all GET:

Endpoint Use it for Answers
/health/live Liveness probe 200 {"alive": true, "version", "uptime_seconds"} as soon as the port is bound, even while models download
/health/ready Readiness probe 200 {"ready": true, "version", "checks": {...}} once the embedding model is loaded and storage answers; 503 otherwise
/health Dashboards, humans, alerts 200 with the status, degraded components and config_warnings below

The port binds before the models load. Until the embedding model (and, with DAKERA_TIERED=1, the tiered engine) has loaded, /health/live answers 200; /health/ready and /health answer 503 with Retry-After: 5; every other route answers 503 with Retry-After: 5. A server that is still downloading a model is therefore not restarted by a liveness probe. While starting, the body names what the server is doing and the progress of every download:

{
  "ready": false,
  "version": "0.12.0",
  "starting": true,
  "reason": "starting",
  "downloads": []
}

(/health answers {"service": "dakera", "status": "starting", "version", "reason", "downloads"}.) Point readiness at /health/ready and liveness at /health/live, and give the liveness probe a start period long enough for a model download (the image's HEALTHCHECK uses 600 s; the first success ends it). A probe that still points at /health sees 503 while the server starts.

/health fields

Field Meaning
service, version, build_sha Identity (build_sha only in image builds)
status healthy, or degraded when any component is listed in degraded, the node holds records it cannot read, its clock is too far ahead of its cluster peers, or its tiered cold tier is not taking writes
degraded Array of {component, reason}; always present, empty when nothing is degraded. Reasons never contain file paths or credentials (the server log has them in full)
config_warnings Array of {component, reason}; always present. Informational, never makes status degraded
embed_migration Progress of the background re-embed migration (see below); absent until it has started
unreadable_records, corrupt_records Records this node skipped since start because a newer Dakera wrote them (non-zero after a rollback or in a mixed-version fleet) or because they do not decode
clock_ahead_ms, advice Present while the node's clock is too far ahead of its cluster peers (writes are refused) or when a count above is non-zero

Degraded components

Group Components Meaning
Enabled by the server itself and running without reranker, ml_classifier, embedding_model_guard, embedding_engine, tiered_engine, cold_tier A part the server turns on by default failed to load (cold_tier: the S3 circuit is open)
Configured explicitly, on a deployment that already holds data redis_cache, disk_cache, backups, ml_classifier, cluster_gossip, knowledge_graph, cluster_outbox, cluster_auth The same mistake refuses the boot on a fresh install; an existing deployment starts without the component and lists it here
Found while running wal_replay, cluster_sync_held, shared_store, speech_to_text, vision_model, entity_extraction A lossy write-ahead-log replay at startup; changes held for peers still on v0.11; another cluster node uses this node's S3 bucket; a media model that the memory budget refused to load

Each is exported as dakera_component_degraded{component} except embedding_engine, tiered_engine and cold_tier.

config_warnings kinds

component Meaning
env_name A DAKERA_* name no code reads (with the closest real name)
env_value A value that cannot be honored; the setting runs on its default
env_reading A boolean this version reads differently from v0.11.108 (the deployment keeps v0.11.108's reading)
env_ignored A variable set where it has no effect (for example DAKERA_MODEL under DAKERA_TIERED)
data_location Data kept at its pre-0.12 location while the new one is empty
deployment The object store could not be listed at boot, so the node started as an existing deployment
topology Another cluster node uses this node's S3 bucket
backup_schedule A stored, enabled backup schedule that now runs

An upgraded deployment starts with these as warnings; a fresh install with the same mistakes is refused (exit 78). dakera_config_warnings (gauge) counts the entries.

Verify

curl -s localhost:3000/health/ready      # 200 when ready
curl -s localhost:3000/health | jq '{status, degraded, config_warnings, embed_migration}'

Check a configuration before it runs

docker run --rm <the deployment's env and volumes> ghcr.io/dakera-ai/dakera:0.12.0 --check-config

dakera --check-config runs the startup configuration check alone (same variable registry, same fresh-install-versus-existing-deployment decision, same policy) against the environment and data root it is given, prints its findings and starts nothing. Output lines: deployment: ... (fresh install, existing, or installed by v0.12), the unknown names and invalid values it found, configuration error: ..., warning: ..., and a final result: line. Exit code 0 means the server would start (warnings included; they will be in /health config_warnings), 78 means it would refuse. It does not check what only a running server can see: Redis reachability, the S3 bucket's contents, the stored backup schedule. Run it with the new image before an upgrade (Upgrading from v0.11.108).

DAKERA_CONFIG_LENIENT=1 turns refusals into warnings (and starts with DAKERA_WAL_SYNC=periodic:0 as everywrite).

Models

dakera models list | pull | prune shows, pre-fetches, bakes into an image and cleans the models with the server's own code; the images ship bge-large and the reranker already converted, so a default start downloads nothing. Flags, proxies, mirrors and air-gapped installs are in Models and Docker.

Backups

All backup routes need a global admin key, except download, upload and restore, which need a global super_admin key (a bundle carries every API key record).

Route Purpose
POST /admin/backups Create a backup: {"name", "backup_type"?, "namespaces"?, "encrypt"?, "compression"?}
GET /admin/backups, GET /admin/backups/{id} List and inspect (status inprogress, completed, failed, deleting; node_id of the node that took it)
DELETE /admin/backups/{id} Delete (an inprogress backup that no node is taking is accepted)
GET /admin/backups/{id}/download, POST /admin/backups/upload Move a bundle (an upload may inflate to at most 16 times DAKERA_MAX_BODY_SIZE)
POST /admin/backups/restore {"backup_id", "target_namespaces"?, "overwrite"?, "point_in_time"?}
GET /admin/backups/restore/{restore_id} Restore progress
GET/POST /admin/backups/schedule {"enabled", "cron", "backup_type", "retention_days", "max_backups", "namespaces", "encrypt", "compression"}

Metric: dakera_backups_total{outcome}; a shipped alert fires on failed backups.

curl -s -X POST localhost:3000/admin/backups -H "Authorization: Bearer $ADMIN_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "nightly", "backup_type": "full", "compression": "zstd", "encrypt": true}'

Re-embed migration

Memories stored before v0.12 were embedded with the query instruction; on the first start the server re-embeds every stored memory that may carry it in the background (resumable, at most 8 per second, yielding to ingest, on the leader only in a cluster, never over a newer user write). Expect roughly 35 minutes per 10,000 memory records with bge-large on conversational turns; recall is served throughout and mixes migrated and unmigrated vectors until it finishes. There is nothing to re-embed (state not_needed) for a model that embeds queries and documents identically, such as minilm, on the visual lane, or under late-interaction scoring. gte-modernbert-base and modernbert-embed-base stores are re-embedded too, because their embedding recipe was corrected.

Monitor it with /health embed_migration (state: pending, surveying, running, standby, waiting, complete, not_needed; remaining, reembedded, skipped, rate_per_sec, eta_secs), GET /admin/reembed/migration (adds reason, target, rate_limit_per_sec, cursor and per-namespace progress) and the dakera_reembed_migration_* metrics.

Maintenance mode

POST /admin/cluster/maintenance/enable with {"reason": "...", "reject_requests": true, "duration_minutes": 30} puts the node that receives the request in maintenance mode; GET /admin/cluster/maintenance shows it; POST /admin/cluster/maintenance/disable ends it (409 while maintenance jobs run, unless force). It is held in memory by that node alone: not replicated, and ended by a restart. With reject_requests, every request outside /admin, /v1/admin, /ops, /health, /metrics and /internal is answered 503 with Retry-After (the time to the scheduled end, or a minute); gRPC calls other than Health fail UNAVAILABLE.

Write-ahead log, TTL and checkpoints

Graceful shutdown

SIGTERM, Ctrl+C or POST /ops/shutdown (global admin; it really shuts down in v0.12, v0.11 answered and kept running) starts one shutdown for both REST and gRPC: both servers drain, then the final flushes run (access statistics, the sync outbox, full-text indexes, the cold tier). In a cluster the node drains its outbox once and announces a graceful leave. Media jobs start no new store step during shutdown, and stores already running are waited for up to 15 s; a job still transcribing or embedding stores nothing and must be started again.

Going back to v0.11.108

dakera downgrade converts a stopped deployment's data back to what v0.11.108 reads, with no v0.11 patch release.

  1. Stop v0.12 (a crash is covered). In a cluster, wait until every node's outbox is empty (GET /admin/cluster/status, replication.outbox_entries), then stop every node.
  2. Run it once with the server's environment and volumes, after the server has stopped: docker run --rm <env and volumes> <v0.12 image> downgrade, docker compose run --rm <service> downgrade, or a Kubernetes Job with args: ["downgrade"]. Do not run it as the server's own command in a rolling update.
  3. Read the JSON report (it is alone on stdout; its log goes to stderr). Exit 0: the data is v0.11.108's. Exit 1: something is not converted yet (values it could not open, or a step failed with completed: false); do not start v0.11.108, fix the cause and run it again. Exit 78: it refused and changed nothing.
  4. Apply what the report lists and start v0.11.108 on the same data and environment.

Report fields: namespaces_scanned, records_scanned, records_rewritten, contents_resealed_v1, texts_opened, values_unreadable, memories_reembedded, texts_reembedded, fulltext_indices_rewritten, fulltext_deltas_replayed, fulltext_records_removed, fulltext_unreadable, keyring_pass_records_removed, cold_writes_flushed, write_ahead_log, caches_cleared, environment_for_v011 (boolean variables to respell for v0.11.108), memory_policies_to_reapply (namespace to policy body; v0.11.108 never loads persisted policies, so PUT /v1/namespaces/{ns}/memory_policy each again), notes, errors and completed.

It converts $enc$v2$ values (key rotations since the upgrade included) to $enc$v1$ under DAKERA_ENCRYPTION_KEY, folds BM25 delta logs into their bases, snapshots the write-ahead log and places it where v0.11.108 reads it, flushes owed tiered cold writes, empties the L2 disk cache and Redis, and re-embeds gte-modernbert-base / modernbert-embed-base stores in v0.11's recipe. It is idempotent and starting v0.12 again on the converted data is safe.

It refuses (exit 78, nothing changed) next to a live server (one answering on the configuration's port, one holding the data root's lock, or a live marker in the S3 bucket), on an empty or mis-pointed data root, and for deployments on bge-m3, colbert-small / late-interaction scoring or the visual lane: move them to a v0.11 model under v0.12 first. The full procedure, with the notes on the knowledge graph, the model volume and containers, is in Rolling back to v0.11.

Observability

Prometheus metrics are at /metrics (unauthenticated, for internal scraping) and /v1/ops/metrics (global admin). Requests refused by the rate limit (429), authentication (401 / 403), the timeout (408), the body limit (413) and unmatched paths (path="unmatched") are counted in dakera_http_requests_total and dakera_http_request_duration_seconds.

New or newly meaningful in v0.12.0:

Area Metrics
Health dakera_component_degraded{component}, dakera_config_warnings
Models and memory dakera_model_loads_total{model,outcome}, dakera_model_load_duration_seconds{model}, dakera_memory_budget_reserved_bytes
Durability dakera_wal_write_failures_total, dakera_storage_write_stalls_total, dakera_backups_total{outcome}, dakera_memory_read_failures_total
Caches dakera_cache_redis_failed_writes_total, dakera_cache_redis_failed_publishes_total, dakera_cache_redis_failed_evictions_total
Indexes dakera_memory_index_builds_total, dakera_memory_index_invalidations_total, dakera_late_interaction_first_stage_total{stage}
Re-embed dakera_reembed_migration_*
Encryption dakera_encryption_values_resealed_total{kind}, dakera_encryption_values_unreadable, dakera_encryption_reseal_running, dakera_encryption_reseal_namespaces_remaining, dakera_encryption_deferred_writes_total, dakera_encryption_keys{state}
Cluster dakera_cluster_sync_outbox_entries, dakera_cluster_sync_failures_total, dakera_cluster_sync_outbox_dropped_total, dakera_cluster_sync_outbox_persist_failures_total, dakera_cluster_sync_refused_total, dakera_cluster_sync_superseded_total, dakera_cluster_sync_held_entries, dakera_cluster_replication_queue_dropped_total, dakera_cluster_reconcile_total, dakera_cluster_reconcile_namespaces_total, dakera_cluster_reconcile_pushed_namespaces_total, dakera_cluster_tombstones, dakera_cluster_tombstones_collected_total, dakera_cluster_clock_ahead_ms, dakera_cluster_clock_skew_refused_writes_total
RocksDB hot tier dakera_rocksdb_*

Alert rules and Grafana dashboards are published in dakera-deploy monitoring/ (Prometheus config, alerting rules and provisioned dashboards). The v0.12.0 versions, including monitoring/dakera.rules.yml and the rebuilt dashboard, arrive there with the dakera-deploy v0.12.0 release. The v0.12.0 rules come in two groups, which you can load through your Prometheus rule_files: signals (DakeraConfigWarnings, DakeraComponentDegraded, DakeraWalReplayDroppedEntries, DakeraMemoryReadFailures, DakeraRedisCacheFailures, DakeraClusterSyncHeldForV011Peers, DakeraWalWriteFailures, DakeraModelLoadFailures, DakeraBackupFailed, DakeraEncryptionValuesUnreadable) and service availability (DakeraDown, DakeraHigh5xxRate, DakeraShedding, DakeraMemoryBudgetRefusals, DakeraColdTierCircuitOpen, DakeraColdTierBacklog, DakeraRocksdbWritesStopped, DakeraClusterOutboxDropping, DakeraStorageWriteStalled, DakeraClusterClockSkew, DakeraCorruptRecords), each with a severity and a description of what to check.

Logs: set the level with RUST_LOG. The per-request INFO lines of the data path are DEBUG in v0.12; the audit log (target audit) still records every request (Security). A request without a valid key is a DEBUG line with at most one WARN per minute counting them; dakera_errors_total{type="auth"} counts every one. Log messages that used internal identifiers now use plain names (the keyring's messages say encryption keyring), so alerts that match log text need checking.

Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.