Operations
This runbook is for operators who run Dakera v0.12.0 in production. It covers health and readiness probes, checking a configuration before it runs, backups and restores, the background re-embed migration, maintenance mode, shutdown, going back to v0.11.108, and what to scrape and alert on.
Health, readiness and liveness
Three unauthenticated endpoints, all GET:
| Endpoint | Use it for | Answers |
|---|---|---|
/health/live |
Liveness probe | 200 {"alive": true, "version", "uptime_ as soon as the port is bound, even while models download |
/health/ready |
Readiness probe | 200 {"ready": true, "version", "checks": {...}} once the embedding model is loaded and storage answers; 503 otherwise |
/health |
Dashboards, humans, alerts | 200 with the status, degraded components and config_warnings below |
The port binds before the models load. Until the embedding model (and, with DAKERA_TIERED=1, the tiered engine) has loaded, /health/live answers 200; /health/ready and /health answer 503 with Retry-After: 5; every other route answers 503 with Retry-After: 5. A server that is still downloading a model is therefore not restarted by a liveness probe. While starting, the body names what the server is doing and the progress of every download:
{
"ready": false,
"version": "0.12.0",
"starting": true,
"reason": "starting",
"downloads": []
}
(/health answers {"service": "dakera", "status": "starting", "version", "reason", "downloads"}.) Point readiness at /health/ready and liveness at /health/live, and give the liveness probe a start period long enough for a model download (the image's HEALTHCHECK uses 600 s; the first success ends it). A probe that still points at /health sees 503 while the server starts.
/health fields
| Field | Meaning |
|---|---|
service, version, build_sha |
Identity (build_sha only in image builds) |
status |
healthy, or degraded when any component is listed in degraded, the node holds records it cannot read, its clock is too far ahead of its cluster peers, or its tiered cold tier is not taking writes |
degraded |
Array of {component, reason}; always present, empty when nothing is degraded. Reasons never contain file paths or credentials (the server log has them in full) |
config_warnings |
Array of {component, reason}; always present. Informational, never makes status degraded |
embed_migration |
Progress of the background re-embed migration (see below); absent until it has started |
unreadable_, corrupt_records |
Records this node skipped since start because a newer Dakera wrote them (non-zero after a rollback or in a mixed-version fleet) or because they do not decode |
clock_ahead_ms, advice |
Present while the node's clock is too far ahead of its cluster peers (writes are refused) or when a count above is non-zero |
Degraded components
| Group | Components | Meaning |
|---|---|---|
| Enabled by the server itself and running without | reranker, ml_classifier, embedding_, embedding_, tiered_engine, cold_tier |
A part the server turns on by default failed to load (cold_tier: the S3 circuit is open) |
| Configured explicitly, on a deployment that already holds data | redis_cache, disk_cache, backups, ml_classifier, cluster_gossip, knowledge_graph, cluster_outbox, cluster_auth |
The same mistake refuses the boot on a fresh install; an existing deployment starts without the component and lists it here |
| Found while running | wal_replay, cluster_, shared_store, speech_to_text, vision_model, entity_ |
A lossy write-ahead-log replay at startup; changes held for peers still on v0.11; another cluster node uses this node's S3 bucket; a media model that the memory budget refused to load |
Each is exported as dakera_component_degraded{component} except embedding_engine, tiered_engine and cold_tier.
config_warnings kinds
component |
Meaning |
|---|---|
env_name |
A DAKERA_* name no code reads (with the closest real name) |
env_value |
A value that cannot be honored; the setting runs on its default |
env_reading |
A boolean this version reads differently from v0.11.108 (the deployment keeps v0.11.108's reading) |
env_ignored |
A variable set where it has no effect (for example DAKERA_MODEL under DAKERA_TIERED) |
data_location |
Data kept at its pre-0.12 location while the new one is empty |
deployment |
The object store could not be listed at boot, so the node started as an existing deployment |
topology |
Another cluster node uses this node's S3 bucket |
backup_schedule |
A stored, enabled backup schedule that now runs |
An upgraded deployment starts with these as warnings; a fresh install with the same mistakes is refused (exit 78). dakera_config_warnings (gauge) counts the entries.
Verify
curl -s localhost:3000/health/ready # 200 when ready
curl -s localhost:3000/health | jq '{status, degraded, config_warnings, embed_migration}'
Check a configuration before it runs
docker run --rm <the deployment's env and volumes> ghcr.io/dakera-ai/dakera:0.12.0 --check-config
dakera --check-config runs the startup configuration check alone (same variable registry, same fresh-install-versus-existing-deployment decision, same policy) against the environment and data root it is given, prints its findings and starts nothing. Output lines: deployment: ... (fresh install, existing, or installed by v0.12), the unknown names and invalid values it found, configuration error: ..., warning: ..., and a final result: line. Exit code 0 means the server would start (warnings included; they will be in /health config_warnings), 78 means it would refuse. It does not check what only a running server can see: Redis reachability, the S3 bucket's contents, the stored backup schedule. Run it with the new image before an upgrade (Upgrading from v0.11.108).
DAKERA_CONFIG_LENIENT=1 turns refusals into warnings (and starts with DAKERA_WAL_SYNC=periodic:0 as everywrite).
Models
dakera models list | pull | prune shows, pre-fetches, bakes into an image and cleans the models with the server's own code; the images ship bge-large and the reranker already converted, so a default start downloads nothing. Flags, proxies, mirrors and air-gapped installs are in Models and Docker.
Backups
All backup routes need a global admin key, except download, upload and restore, which need a global super_admin key (a bundle carries every API key record).
| Route | Purpose |
|---|---|
POST / |
Create a backup: {"name", "backup_ |
GET /, GET / |
List and inspect (status inprogress, completed, failed, deleting; node_id of the node that took it) |
DELETE / |
Delete (an inprogress backup that no node is taking is accepted) |
GET /, POST / |
Move a bundle (an upload may inflate to at most 16 times DAKERA_) |
POST / |
{"backup_ |
GET / |
Restore progress |
GET/POST / |
{"enabled", "cron", "backup_ |
- Encryption and compression:
"compression": "zstd"compresses every data and part file;"encrypt": trueseals each file with AES-256-GCM under a key derived from the operator key (DAKERA_ENCRYPTION_KEY). Both are recorded in the manifest."compression": "lz4", andencryptwithoutDAKERA_ENCRYPTION_KEY, are400.incrementalis reported asfull. - Contents: records with their sidecars, attachments and knowledge-graph edges; text sealed (
_textwas plaintext in v0.11 backups). - Schedule: v0.11 stored a schedule and never ran it. v0.12 takes a backup at every slot of the
cronexpression (UTC; on the leader in a cluster) and appliesretention_daysandmax_backupsto scheduled backups only (the newest completed one is always kept). An enabled stored schedule on an upgraded deployment shows as abackup_schedulewarning; sendPOST /admin/backups/schedulewith{"enabled": false}before its first slot to keep v0.11's behavior. An invalid expression, or anencrypt/compressiona backup cannot honor, is400when set. - Interrupted backups: a backup a stop interrupted is marked
failedat the next start. - Restore: without
overwritethe backup's records are written over the current copies and records added since are kept (v0.11 behavior); with"overwrite": truethe result is the backup's exact point in time (it deletes first and is never refused by a quota half-way). Revoked API keys stay revoked, keyring records are restored only where missing, one restore runs at a time (409), and the live backup schedule is kept (the saved copy is_backups/{id}/schedule.jsonin the backup store). A restore does not pause serving: run it in a maintenance window. Memories whose text the restore changed get their sentence sub-memories and graph edges re-derived. - One bucket per node: a backup's files live in the backup store of the node that took it; a restore or download on a node that does not hold them is
409naming the owner. - Rolling back: a backup taken under v0.12 is restored by v0.12 and then downgraded (Rolling back to v0.11).
Metric: dakera_backups_total{outcome}; a shipped alert fires on failed backups.
curl -s -X POST localhost:3000/admin/backups -H "Authorization: Bearer $ADMIN_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "nightly", "backup_type": "full", "compression": "zstd", "encrypt": true}'
Re-embed migration
Memories stored before v0.12 were embedded with the query instruction; on the first start the server re-embeds every stored memory that may carry it in the background (resumable, at most 8 per second, yielding to ingest, on the leader only in a cluster, never over a newer user write). Expect roughly 35 minutes per 10,000 memory records with bge-large on conversational turns; recall is served throughout and mixes migrated and unmigrated vectors until it finishes. There is nothing to re-embed (state not_needed) for a model that embeds queries and documents identically, such as minilm, on the visual lane, or under late-interaction scoring. gte-modernbert-base and modernbert-embed-base stores are re-embedded too, because their embedding recipe was corrected.
Monitor it with /health embed_migration (state: pending, surveying, running, standby, waiting, complete, not_needed; remaining, reembedded, skipped, rate_per_sec, eta_secs), GET /admin/reembed/migration (adds reason, target, rate_limit_per_sec, cursor and per-namespace progress) and the dakera_reembed_migration_* metrics.
Maintenance mode
POST /admin/cluster/maintenance/enable with {"reason": "...", "reject_requests": true, "duration_minutes": 30} puts the node that receives the request in maintenance mode; GET /admin/cluster/maintenance shows it; POST /admin/cluster/maintenance/disable ends it (409 while maintenance jobs run, unless force). It is held in memory by that node alone: not replicated, and ended by a restart. With reject_requests, every request outside /admin, /v1/admin, /ops, /health, /metrics and /internal is answered 503 with Retry-After (the time to the scheduled end, or a minute); gRPC calls other than Health fail UNAVAILABLE.
Write-ahead log, TTL and checkpoints
- The write-ahead log (
DAKERA_WAL, default on;DAKERA_WAL_SYNCeverywriteby default) checkpoints everyDAKERA_WAL_CHECKPOINT_SECS(default 300). A checkpoint rotates and truncates the active segment, so deleted content stays in the log at most one checkpoint interval and in at most 10 snapshots. Concurrent writes share the log's fsync; a write is acknowledged only once durable. - A replay at startup that dropped entries the store refused is
wal_replaydegraded; write and fsync errors are returned to the client and counted indakera_wal_write_failures_total. - The TTL sweep (
DAKERA_TTL_SWEEP, default on;DAKERA_TTL_SWEEP_SECS, default 60) removes expired memories from vectors, full-text search, the graph and the counters; an expiring memory no longer forces an index rebuild.POST /admin/ttl/cleanupruns it on demand. - A write whose client gave up after it reached the log is still applied (and replicated). A storage backend that stops answering mid-write is waited for, logged at error every 30 s and counted in
dakera_storage_write_stalls_total.
Graceful shutdown
SIGTERM, Ctrl+C or POST /ops/shutdown (global admin; it really shuts down in v0.12, v0.11 answered and kept running) starts one shutdown for both REST and gRPC: both servers drain, then the final flushes run (access statistics, the sync outbox, full-text indexes, the cold tier). In a cluster the node drains its outbox once and announces a graceful leave. Media jobs start no new store step during shutdown, and stores already running are waited for up to 15 s; a job still transcribing or embedding stores nothing and must be started again.
Going back to v0.11.108
dakera downgrade converts a stopped deployment's data back to what v0.11.108 reads, with no v0.11 patch release.
- Stop v0.12 (a crash is covered). In a cluster, wait until every node's outbox is empty (
GET /admin/cluster/status,replication.outbox_entries), then stop every node. - Run it once with the server's environment and volumes, after the server has stopped:
docker run --rm <env and volumes> <v0.12 image> downgrade,docker compose run --rm <service> downgrade, or a Kubernetes Job withargs: ["downgrade"]. Do not run it as the server's own command in a rolling update. - Read the JSON report (it is alone on stdout; its log goes to stderr). Exit
0: the data is v0.11.108's. Exit1: something is not converted yet (values it could not open, or a step failed withcompleted: false); do not start v0.11.108, fix the cause and run it again. Exit78: it refused and changed nothing. - Apply what the report lists and start v0.11.108 on the same data and environment.
Report fields: namespaces_scanned, records_scanned, records_rewritten, contents_resealed_v1, texts_opened, values_unreadable, memories_reembedded, texts_reembedded, fulltext_indices_rewritten, fulltext_deltas_replayed, fulltext_records_removed, fulltext_unreadable, keyring_pass_records_removed, cold_writes_flushed, write_ahead_log, caches_cleared, environment_for_v011 (boolean variables to respell for v0.11.108), memory_policies_to_reapply (namespace to policy body; v0.11.108 never loads persisted policies, so PUT /v1/namespaces/{ns}/memory_policy each again), notes, errors and completed.
It converts $enc$v2$ values (key rotations since the upgrade included) to $enc$v1$ under DAKERA_ENCRYPTION_KEY, folds BM25 delta logs into their bases, snapshots the write-ahead log and places it where v0.11.108 reads it, flushes owed tiered cold writes, empties the L2 disk cache and Redis, and re-embeds gte-modernbert-base / modernbert-embed-base stores in v0.11's recipe. It is idempotent and starting v0.12 again on the converted data is safe.
It refuses (exit 78, nothing changed) next to a live server (one answering on the configuration's port, one holding the data root's lock, or a live marker in the S3 bucket), on an empty or mis-pointed data root, and for deployments on bge-m3, colbert-small / late-interaction scoring or the visual lane: move them to a v0.11 model under v0.12 first. The full procedure, with the notes on the knowledge graph, the model volume and containers, is in Rolling back to v0.11.
Observability
Prometheus metrics are at /metrics (unauthenticated, for internal scraping) and /v1/ops/metrics (global admin). Requests refused by the rate limit (429), authentication (401 / 403), the timeout (408), the body limit (413) and unmatched paths (path="unmatched") are counted in dakera_http_requests_total and dakera_http_request_duration_seconds.
New or newly meaningful in v0.12.0:
| Area | Metrics |
|---|---|
| Health | dakera_, dakera_ |
| Models and memory | dakera_, dakera_, dakera_ |
| Durability | dakera_, dakera_, dakera_, dakera_ |
| Caches | dakera_, dakera_, dakera_ |
| Indexes | dakera_, dakera_, dakera_ |
| Re-embed | dakera_ |
| Encryption | dakera_, dakera_, dakera_, dakera_, dakera_, dakera_ |
| Cluster | dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_, dakera_ |
| RocksDB hot tier | dakera_ |
Alert rules and Grafana dashboards are published in dakera-deploy monitoring/ (Prometheus config, alerting rules and provisioned dashboards). The v0.12.0 versions, including monitoring/dakera.rules.yml and the rebuilt dashboard, arrive there with the dakera-deploy v0.12.0 release. The v0.12.0 rules come in two groups, which you can load through your Prometheus rule_files: signals (DakeraConfigWarnings, DakeraComponentDegraded, DakeraWalReplayDroppedEntries, DakeraMemoryReadFailures, DakeraRedisCacheFailures, DakeraClusterSyncHeldForV011Peers, DakeraWalWriteFailures, DakeraModelLoadFailures, DakeraBackupFailed, DakeraEncryptionValuesUnreadable) and service availability (DakeraDown, DakeraHigh5xxRate, DakeraShedding, DakeraMemoryBudgetRefusals, DakeraColdTierCircuitOpen, DakeraColdTierBacklog, DakeraRocksdbWritesStopped, DakeraClusterOutboxDropping, DakeraStorageWriteStalled, DakeraClusterClockSkew, DakeraCorruptRecords), each with a severity and a description of what to check.
Logs: set the level with RUST_LOG. The per-request INFO lines of the data path are DEBUG in v0.12; the audit log (target audit) still records every request (Security). A request without a valid key is a DEBUG line with at most one WARN per minute counting them; dakera_errors_total{type="auth"} counts every one. Log messages that used internal identifiers now use plain names (the keyring's messages say encryption keyring), so alerts that match log text need checking.
Related
- Upgrading from v0.11.108: every behavior change and the first-start checklist
- High Availability: clusters, replication, rolling upgrades
- Security: keys, encryption at rest, permissions
- Configuration: every variable