NewDakera v0.12.0 is out: multilingual, multimodal and multi-vector memory, faster reranked recall, one-command rollbackSee what's new →

Multilingual search

Dakera v0.12.0 can store and recall memories in many languages. This page is for deployments whose agents store or query text that is not English: it covers the multilingual embedding model, full-text analysis per language, dates and query routing in seven languages, and the per-request lang field. Four independent settings make that work, and each one is opt-in, so a server that sets none of them behaves as v0.11.108 did (English):

Layer What it does Setting
Embeddings A multilingual model puts text from many languages in one vector space DAKERA_MODEL=bge-m3
Full-text (BM25) Stemming and stop words per language; character bigrams for Chinese, Japanese, Korean, Thai and similar scripts DAKERA_FULLTEXT_LANGUAGE, DAKERA_FULLTEXT_CJK_BIGRAMS
Query routing and dates Question patterns ("when", "how long ago") and date expressions in seven languages DAKERA_QUERY_LANG, per-request lang
Operations Rebuild an existing namespace's full-text index under a new analyzer POST /admin/fulltext/reindex

Use them together for a multilingual deployment, or separately: a German-only deployment needs only DAKERA_FULLTEXT_LANGUAGE=de (its query language follows it), while a deployment that mixes languages benefits most from bge-m3 plus DAKERA_FULLTEXT_LANGUAGE=multilingual.

Multilingual embeddings: bge-m3

DAKERA_MODEL=bge-m3 selects BAAI bge-m3, a multilingual model covering 100+ languages.

Property Value
Architecture XLM-RoBERTa, dense head (CLS token, L2-normalized)
Dimension 1024 (the same as the default bge-large)
Model window 8192 tokens
Served truncation 2048 tokens by default; DAKERA_MAX_SEQ_LENGTH sets it between 16 and 8192
Artifact INT8 ONNX, about 570 MB, pinned to a SHA-256 checksum
Prefixes None (the model needs no query or document instruction)
Backend CPU ONNX only

Why 2048 and not 8192: a full-attention encoder needs heads x tokens^2 x 4 bytes of attention scores per layer for one row. For bge-m3 (16 heads) that is about 4 GiB at 8192 tokens and about 256 MiB at 2048. Raise DAKERA_MAX_SEQ_LENGTH only if you have the memory and long texts.

Constraints

Choosing it on a new store

docker run -d -p 3000:3000 -v dakera-data:/data -v dakera-models:/app/models \
  -e DAKERA_ROOT_API_KEY=dk-mykey \
  -e DAKERA_MODEL=bge-m3 \
  -e DAKERA_FULLTEXT_LANGUAGE=multilingual \
  ghcr.io/dakera-ai/dakera:latest

Switching an existing store

The store records which model embedded it. Starting with another model is refused at startup (exit 1, both models named), and a request whose model differs from the server's model is 400, because bge-large and bge-m3 are both 1024-d and mixing them would trip no dimension check. To switch:

  1. Take a backup (POST /admin/backups).
  2. Start once with DAKERA_MODEL=bge-m3 and DAKERA_ALLOW_MODEL_CHANGE=1. This only acknowledges the change; it re-embeds nothing.
  3. Re-embed with an admin key:
curl -s -X POST localhost:3000/admin/namespaces/migrate-dimensions \
  -H "Authorization: Bearer $ADMIN_KEY" -H "Content-Type: application/json" \
  -d '{"target_dimension": 1024, "reembed_same_dimension": true}'

The request takes namespaces (default: every agent memory namespace), target_dimension (default 1024) and reembed_same_dimension (re-embed namespaces already at that dimension, needed here because both models are 1024-d). The response lists migrated, failed, already_current, per-namespace results (status: success, failed or skipped_already_current) and a job_id readable at GET /ops/jobs/{id}. The work runs on its own task, so a request that times out does not stop it. Vectors that clients upserted directly are not re-embedded, since Dakera never held their text.

Until every agent namespace has been re-embedded, recall mixes the two embedding spaces and GET /v1/capabilities reports "reembed_pending": true.

Re-embedding does not change the full-text analyzer. If the data is not English, also set DAKERA_FULLTEXT_LANGUAGE and rebuild each namespace's index (see below).

Full-text languages

Full-text (BM25) search runs on every memory. DAKERA_FULLTEXT_LANGUAGE chooses its analyzer for newly created indexes. Unset, empty or unrecognized means English with stemming, exactly as in v0.11.108. An unrecognized value refuses a fresh install and, on a deployment that already holds data, becomes a config_warnings entry on /health while English stays in effect.

Accepted values are an ISO 639-1 code or the English name, case-insensitive, with an optional region suffix (de, German, pt-BR, de_DE):

Language Code Language Code
Arabic ar Norwegian no (nb and nn are accepted)
Danish da Portuguese pt
Dutch nl Romanian ro
English en Russian ru
Finnish fi Spanish es
French fr Swedish sv
German de Tamil ta
Greek el Turkish tr
Hungarian hu
Italian it

These 18 languages use a Snowball stemmer and their own stop-word list (English keeps its existing list). English stop words are content words in other languages (die, was, an), which is why a German deployment must not keep the English list.

No stemming: zh, ja, ko, th (and chinese, japanese, korean, thai), none, off, multi and multilingual. These select an analyzer with no stemmer and no stop words. Use multilingual for a corpus that mixes languages, where any single stemmer would damage the others.

What the analyzer does

CJK and other unsegmented scripts: DAKERA_FULLTEXT_CJK_BIGRAMS

Chinese, Japanese, Korean and Thai are written without spaces, so whitespace tokenization turns a sentence into one very long token. In v0.11.108 that token exceeded the token length limit and was dropped: such text was not indexed at all.

With bigrams on, runs of unsegmented script are indexed as overlapping character bigrams (and as single characters). A query of two or more characters searches by bigrams, and a one-character query searches that character, so single-character words are findable. Scripts covered: Han, Hiragana, Katakana, Hangul, Thai, Lao, Khmer, Myanmar and Tibetan (Thai, Khmer, Myanmar and Tibetan clusters keep their combining marks together).

Setting Effect
unset On for every language except English; off for English
1, true, yes, on On (for example an English deployment that also stores Japanese notes)
0, false, no, off Off: the historical analyzer for that language. Dakera warns, because unsegmented scripts are then not indexed

A value that is not a boolean is ignored with a warning and the language default applies.

The analyzer is stored with the index

Each full-text index persists the analyzer it was built with, so an index and its queries always agree. Changing DAKERA_FULLTEXT_LANGUAGE or DAKERA_FULLTEXT_CJK_BIGRAMS therefore affects new indexes only. An existing namespace keeps its analyzer (the server logs that the restored index was built with a different one) until you rebuild it.

Rebuilding a namespace: POST /admin/fulltext/reindex

Requires a key with the global admin scope (not restricted to namespaces).

Request field Type Meaning
namespace string, optional One namespace. Omitted: every agent memory namespace (_dakera_agent_*)
rebuild bool, default false Re-analyze every memory under the current configuration instead of only adding the missing ones. Applied automatically when the namespace's index was built under another analyzer
curl -s -X POST localhost:3000/admin/fulltext/reindex \
  -H "Authorization: Bearer $ADMIN_KEY" -H "Content-Type: application/json" \
  -d '{"namespace": "_dakera_agent_support-de", "rebuild": true}'
{
  "namespaces_processed": 1,
  "total_indexed": 1520,
  "total_skipped": 0,
  "details": [
    {
      "namespace": "_dakera_agent_support-de",
      "vectors_scanned": 1520,
      "newly_indexed": 1520,
      "already_indexed": 0,
      "parse_failures": 0,
      "rebuilt": true,
      "analyzer_language": "de",
      "text_vectors_resealed": 0
    }
  ]
}

Query languages and dates

DAKERA_QUERY_LANG sets the language of the rule-based query handling:

Supported: en, de, fr, es, it, pt, nl. All seven use one calendar registry for month names, seasons and relative expressions. Values: one of those codes (or its English or native name), or auto.

Setting Behavior
unset The deployment language: DAKERA_FULLTEXT_LANGUAGE when that is one of the seven, else en. A default deployment is fixed to en and runs no language detector
a code Every query uses that language
auto The language is detected per query from function words and falls back to the deployment language when detection is unreliable
unrecognized Refuses a fresh install; on an existing deployment a config_warnings entry and the deployment language

Examples the date parser resolves (verified by the engine's tests): German "Worüber haben wir gestern gesprochen?" (yesterday), "vor zwei Wochen" (two weeks ago), "im Sommer 2022"; French "il y a deux semaines", "le mois dernier", "au mois d'août"; Spanish, Italian, Portuguese and Dutch have equivalent sets. Names that look like months or seasons are not treated as dates ("Frau Sommer", "Mai hat mir eine Nachricht geschickt").

Auto-detection is opt-in because a statistical detector can misjudge a short English query. Prefer a fixed DAKERA_QUERY_LANG or the per-request lang when you know the language.

A note on classifiers: Dakera's learned query classifier uses English prototypes. With an English-only embedding model, a query that is not English falls back to the rule-based classifier for its language. With bge-m3 the learned classifier applies to every language.

Per-request lang

lang accepts the same values (de, German, deutsch, pt-BR). Omitted or blank means the server-wide choice. Anything else is 400 naming the supported languages. It never translates or filters.

Route What lang selects
POST /v1/memory/recall Routing patterns and temporal expressions of the query
POST /v1/memory/search The same, for the query
POST /v1/memory/store How the content's event date, rule-based date entities and full-text date metadata are parsed; recorded on the memory
POST /v1/memories/store/batch The same, request-level, for every item
PUT /v1/memory/update/{id} The same; a lang that differs from the recorded one re-derives that data without re-embedding
POST /v1/extract Rule set and prompt language
POST /v1/memories/extract Rule set and prompt language

POST /v1/memories/recall/batch uses the server-wide language.

On a write, the canonical code is recorded in the memory's metadata as _dakera_lang. Sentence sub-memories, updates, imports, peer sync, consolidation and POST /admin/fulltext/reindex all resolve a memory's language through that record, so a memory is always parsed the way its write parsed it.

# Store a German memory; the date "vor drei Tagen" is parsed with German rules.
curl -s -X POST localhost:3000/v1/memory/store \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"agent_id": "support-de", "lang": "de",
       "content": "Vor drei Tagen hat der Kunde das Konto gekündigt.", "importance": 0.7}'
# 200 {"memory": {"id": "...", "metadata": {"_dakera_lang": "de", ...}, ...}, "embedding_time_ms": 41}

# Recall in German
curl -s -X POST localhost:3000/v1/memory/recall \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"agent_id": "support-de", "lang": "de", "query": "Wann hat der Kunde gekündigt?", "top_k": 5}'

Recall of Chinese, Japanese and Thai queries

A query in a space-free script has no spaces to count, so v0.11.108 counted it as one word and routed it as a keyword lookup, which weighs the dense lane at 0.25 instead of 0.50 or 0.75. v0.12 estimates the length of such queries from their characters (about two characters per word), recognizes the full-width question mark and routes them like their English equivalents.

Defects fixed in v0.12.0

These affected non-English text in v0.11.108 and need a rebuild of an existing index to take effect (POST /admin/fulltext/reindex with rebuild):

Verify it works

# What the server is configured for
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" \
  | jq '{default_model, fulltext_language, query_languages, reembed_pending,
         bge_m3: (.models[] | select(.name=="bge-m3") | {active, dimension, max_seq_length, effective_max_seq_length})}'

Expect default_model: "bge-m3", bge_m3.active: true, effective_max_seq_length: 2048 (or your DAKERA_MAX_SEQ_LENGTH), fulltext_language as configured ("none" for multilingual and the CJK choices), query_languages: ["en","de","fr","es","it","pt","nl"] and reembed_pending: false once any model switch has finished.

Then:

  1. Store two or three memories in the target language and recall with a paraphrase. With bge-m3 a query in another language than the stored text should still find it.
  2. For an existing namespace after a language change, run the reindex and check analyzer_language and rebuilt: true in the response.
  3. GET /health lists a config_warnings entry for an unrecognized DAKERA_FULLTEXT_LANGUAGE or DAKERA_QUERY_LANG on an upgraded deployment; an empty list means every value was honored.
  4. After a CJK reindex, a full-text search for a two-character word returns its documents.

Resource implications

See also: Configuration, Models and Docker, Multimodal memory, Search and ranking.

Stay sharp on agent memory
Benchmark results, SDK releases, and production patterns. Under 500 words per issue.
✓ You're in. First issue lands soon — watch for Dakera in your inbox.