Multilingual search
Dakera v0.12.0 can store and recall memories in many languages. This page is for deployments whose agents store or query text that is not English: it covers the multilingual embedding model, full-text analysis per language, dates and query routing in seven languages, and the per-request lang field. Four independent settings make that work, and each one is opt-in, so a server that sets none of them behaves as v0.11.108 did (English):
| Layer | What it does | Setting |
|---|---|---|
| Embeddings | A multilingual model puts text from many languages in one vector space | DAKERA_ |
| Full-text (BM25) | Stemming and stop words per language; character bigrams for Chinese, Japanese, Korean, Thai and similar scripts | DAKERA_, DAKERA_ |
| Query routing and dates | Question patterns ("when", "how long ago") and date expressions in seven languages | DAKERA_, per-request lang |
| Operations | Rebuild an existing namespace's full-text index under a new analyzer | POST / |
Use them together for a multilingual deployment, or separately: a German-only deployment needs only DAKERA_FULLTEXT_LANGUAGE=de (its query language follows it), while a deployment that mixes languages benefits most from bge-m3 plus DAKERA_FULLTEXT_LANGUAGE=multilingual.
Multilingual embeddings: bge-m3
DAKERA_MODEL=bge-m3 selects BAAI bge-m3, a multilingual model covering 100+ languages.
| Property | Value |
|---|---|
| Architecture | XLM-RoBERTa, dense head (CLS token, L2-normalized) |
| Dimension | 1024 (the same as the default bge-large) |
| Model window | 8192 tokens |
| Served truncation | 2048 tokens by default; DAKERA_ sets it between 16 and 8192 |
| Artifact | INT8 ONNX, about 570 MB, pinned to a SHA-256 checksum |
| Prefixes | None (the model needs no query or document instruction) |
| Backend | CPU ONNX only |
Why 2048 and not 8192: a full-attention encoder needs heads x tokens^2 x 4 bytes of attention scores per layer for one row. For bge-m3 (16 heads) that is about 4 GiB at 8192 tokens and about 256 MiB at 2048. Raise DAKERA_MAX_SEQ_LENGTH only if you have the memory and long texts.
Constraints
- CPU ONNX backend only. GPU mode (
DAKERA_USE_GPU=1) is refused for this model (its FP32 export is an external-data graph the loader cannot fetch), the Candle backend refuses it ("unsupported by the Candle backend"), and the static backend cannot use it (a different vocabulary). Dakera refuses at startup instead of producing wrong vectors. - Not in the image. The images ship
bge-largeand the reranker.bge-m3downloads on first start into the model cache, or ahead of time withdakera models pull bge-m3(see Models and Docker). The first load converts it once to the memory-mapped ORT form, which needs disk about the size of the model again. - Not with
DAKERA_TIERED=1. The tiered engine always embeds with the default model and ignoresDAKERA_MODEL. - No rollback to v0.11 while it is in use. v0.11.108 does not know
bge-m3;dakera downgraderefuses (exit 78, nothing changed) a store embedded with it. Move to a v0.11 model under v0.12 first (see Rolling back to v0.11).
Choosing it on a new store
docker run -d -p 3000:3000 -v dakera-data:/data -v dakera-models:/app/models \
-e DAKERA_ROOT_API_KEY=dk-mykey \
-e DAKERA_MODEL=bge-m3 \
-e DAKERA_FULLTEXT_LANGUAGE=multilingual \
ghcr.io/dakera-ai/dakera:latest
Switching an existing store
The store records which model embedded it. Starting with another model is refused at startup (exit 1, both models named), and a request whose model differs from the server's model is 400, because bge-large and bge-m3 are both 1024-d and mixing them would trip no dimension check. To switch:
- Take a backup (
POST /admin/backups). - Start once with
DAKERA_MODEL=bge-m3andDAKERA_ALLOW_MODEL_CHANGE=1. This only acknowledges the change; it re-embeds nothing. - Re-embed with an admin key:
curl -s -X POST localhost:3000/admin/namespaces/migrate-dimensions \
-H "Authorization: Bearer $ADMIN_KEY" -H "Content-Type: application/json" \
-d '{"target_dimension": 1024, "reembed_same_dimension": true}'
The request takes namespaces (default: every agent memory namespace), target_dimension (default 1024) and reembed_same_dimension (re-embed namespaces already at that dimension, needed here because both models are 1024-d). The response lists migrated, failed, already_current, per-namespace results (status: success, failed or skipped_already_current) and a job_id readable at GET /ops/jobs/{id}. The work runs on its own task, so a request that times out does not stop it. Vectors that clients upserted directly are not re-embedded, since Dakera never held their text.
Until every agent namespace has been re-embedded, recall mixes the two embedding spaces and GET /v1/capabilities reports "reembed_pending": true.
Re-embedding does not change the full-text analyzer. If the data is not English, also set DAKERA_FULLTEXT_LANGUAGE and rebuild each namespace's index (see below).
Full-text languages
Full-text (BM25) search runs on every memory. DAKERA_FULLTEXT_LANGUAGE chooses its analyzer for newly created indexes. Unset, empty or unrecognized means English with stemming, exactly as in v0.11.108. An unrecognized value refuses a fresh install and, on a deployment that already holds data, becomes a config_warnings entry on /health while English stays in effect.
Accepted values are an ISO 639-1 code or the English name, case-insensitive, with an optional region suffix (de, German, pt-BR, de_DE):
| Language | Code | Language | Code |
|---|---|---|---|
| Arabic | ar |
Norwegian | no (nb and nn are accepted) |
| Danish | da |
Portuguese | pt |
| Dutch | nl |
Romanian | ro |
| English | en |
Russian | ru |
| Finnish | fi |
Spanish | es |
| French | fr |
Swedish | sv |
| German | de |
Tamil | ta |
| Greek | el |
Turkish | tr |
| Hungarian | hu |
||
| Italian | it |
These 18 languages use a Snowball stemmer and their own stop-word list (English keeps its existing list). English stop words are content words in other languages (die, was, an), which is why a German deployment must not keep the English list.
No stemming: zh, ja, ko, th (and chinese, japanese, korean, thai), none, off, multi and multilingual. These select an analyzer with no stemmer and no stop words. Use multilingual for a corpus that mixes languages, where any single stemmer would damage the others.
What the analyzer does
- Text is Unicode NFKC-normalized first, so width and composition variants meet: half-width katakana, full-width digits and letters, decomposed kana.
- Combining marks stay inside the word, for every script. Thai tone marks, the Tamil pulli, the Devanagari virama and Arabic vowel marks no longer split a word. Optional Arabic vocalization and the tatweel stretch are dropped, so a vocalized and a plain spelling of a word are one term.
- Token length limits count characters, not bytes.
CJK and other unsegmented scripts: DAKERA_FULLTEXT_CJK_BIGRAMS
Chinese, Japanese, Korean and Thai are written without spaces, so whitespace tokenization turns a sentence into one very long token. In v0.11.108 that token exceeded the token length limit and was dropped: such text was not indexed at all.
With bigrams on, runs of unsegmented script are indexed as overlapping character bigrams (and as single characters). A query of two or more characters searches by bigrams, and a one-character query searches that character, so single-character words are findable. Scripts covered: Han, Hiragana, Katakana, Hangul, Thai, Lao, Khmer, Myanmar and Tibetan (Thai, Khmer, Myanmar and Tibetan clusters keep their combining marks together).
| Setting | Effect |
|---|---|
| unset | On for every language except English; off for English |
1, true, yes, on |
On (for example an English deployment that also stores Japanese notes) |
0, false, no, off |
Off: the historical analyzer for that language. Dakera warns, because unsegmented scripts are then not indexed |
A value that is not a boolean is ignored with a warning and the language default applies.
The analyzer is stored with the index
Each full-text index persists the analyzer it was built with, so an index and its queries always agree. Changing DAKERA_FULLTEXT_LANGUAGE or DAKERA_FULLTEXT_CJK_BIGRAMS therefore affects new indexes only. An existing namespace keeps its analyzer (the server logs that the restored index was built with a different one) until you rebuild it.
Rebuilding a namespace: POST /admin/fulltext/reindex
Requires a key with the global admin scope (not restricted to namespaces).
| Request field | Type | Meaning |
|---|---|---|
namespace |
string, optional | One namespace. Omitted: every agent memory namespace (_dakera_agent_*) |
rebuild |
bool, default false |
Re-analyze every memory under the current configuration instead of only adding the missing ones. Applied automatically when the namespace's index was built under another analyzer |
curl -s -X POST localhost:3000/admin/fulltext/reindex \
-H "Authorization: Bearer $ADMIN_KEY" -H "Content-Type: application/json" \
-d '{"namespace": "_dakera_agent_support-de", "rebuild": true}'
{
"namespaces_processed": 1,
"total_indexed": 1520,
"total_skipped": 0,
"details": [
{
"namespace": "_dakera_agent_support-de",
"vectors_scanned": 1520,
"newly_indexed": 1520,
"already_indexed": 0,
"parse_failures": 0,
"rebuilt": true,
"analyzer_language": "de",
"text_vectors_resealed": 0
}
]
}
- Without
rebuild, and when the analyzer is unchanged, only memories missing from the index are added (a backfill). parse_failurescounts vectors that are not memories or whose content could not be read; they are never indexed as ciphertext.text_vectors_resealedis non-zero only withrebuildand encryption at rest: text documents are rewritten sealed under the current key.- The call runs inside the request and reads every vector of each namespace it processes. On a large store run it one namespace at a time, in a quiet period.
- Memories are re-analyzed in the language they were written in (
_dakera_lang, below).
Query languages and dates
DAKERA_QUERY_LANG sets the language of the rule-based query handling:
- the patterns that classify a question ("when", "how long ago", "after she ...") and so choose how recall weighs dense and keyword search;
- temporal expressions in queries ("last week", "in March 2023");
- dates found in stored content and rule-based date entities.
Supported: en, de, fr, es, it, pt, nl. All seven use one calendar registry for month names, seasons and relative expressions. Values: one of those codes (or its English or native name), or auto.
| Setting | Behavior |
|---|---|
| unset | The deployment language: DAKERA_ when that is one of the seven, else en. A default deployment is fixed to en and runs no language detector |
| a code | Every query uses that language |
auto |
The language is detected per query from function words and falls back to the deployment language when detection is unreliable |
| unrecognized | Refuses a fresh install; on an existing deployment a config_warnings entry and the deployment language |
Examples the date parser resolves (verified by the engine's tests): German "Worüber haben wir gestern gesprochen?" (yesterday), "vor zwei Wochen" (two weeks ago), "im Sommer 2022"; French "il y a deux semaines", "le mois dernier", "au mois d'août"; Spanish, Italian, Portuguese and Dutch have equivalent sets. Names that look like months or seasons are not treated as dates ("Frau Sommer", "Mai hat mir eine Nachricht geschickt").
Auto-detection is opt-in because a statistical detector can misjudge a short English query. Prefer a fixed DAKERA_QUERY_LANG or the per-request lang when you know the language.
A note on classifiers: Dakera's learned query classifier uses English prototypes. With an English-only embedding model, a query that is not English falls back to the rule-based classifier for its language. With bge-m3 the learned classifier applies to every language.
Per-request lang
lang accepts the same values (de, German, deutsch, pt-BR). Omitted or blank means the server-wide choice. Anything else is 400 naming the supported languages. It never translates or filters.
| Route | What lang selects |
|---|---|
POST / |
Routing patterns and temporal expressions of the query |
POST / |
The same, for the query |
POST / |
How the content's event date, rule-based date entities and full-text date metadata are parsed; recorded on the memory |
POST / |
The same, request-level, for every item |
PUT / |
The same; a lang that differs from the recorded one re-derives that data without re-embedding |
POST / |
Rule set and prompt language |
POST / |
Rule set and prompt language |
POST /v1/memories/recall/batch uses the server-wide language.
On a write, the canonical code is recorded in the memory's metadata as _dakera_lang. Sentence sub-memories, updates, imports, peer sync, consolidation and POST /admin/fulltext/reindex all resolve a memory's language through that record, so a memory is always parsed the way its write parsed it.
# Store a German memory; the date "vor drei Tagen" is parsed with German rules.
curl -s -X POST localhost:3000/v1/memory/store \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"agent_id": "support-de", "lang": "de",
"content": "Vor drei Tagen hat der Kunde das Konto gekündigt.", "importance": 0.7}'
# 200 {"memory": {"id": "...", "metadata": {"_dakera_lang": "de", ...}, ...}, "embedding_time_ms": 41}
# Recall in German
curl -s -X POST localhost:3000/v1/memory/recall \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"agent_id": "support-de", "lang": "de", "query": "Wann hat der Kunde gekündigt?", "top_k": 5}'
Recall of Chinese, Japanese and Thai queries
A query in a space-free script has no spaces to count, so v0.11.108 counted it as one word and routed it as a keyword lookup, which weighs the dense lane at 0.25 instead of 0.50 or 0.75. v0.12 estimates the length of such queries from their characters (about two characters per word), recognizes the full-width question mark and routes them like their English equivalents.
Defects fixed in v0.12.0
These affected non-English text in v0.11.108 and need a rebuild of an existing index to take effect (POST /admin/fulltext/reindex with rebuild):
- Eleven accepted stemmer languages had an empty stop-word list.
- Token length limits counted bytes, so long Cyrillic, Greek and Arabic words (over 25 letters) and Tamil and Devanagari words (over 16) were dropped.
- Combining marks split Tamil and Devanagari words (and Thai, Arabic).
- Lao, Khmer, Myanmar and Tibetan text was dropped whole.
- Pseudo-relevance feedback injected non-English stop words into the query.
- CJK queries were routed as keyword queries.
Verify it works
# What the server is configured for
curl -s localhost:3000/v1/capabilities -H "Authorization: Bearer $KEY" \
| jq '{default_model, fulltext_language, query_languages, reembed_pending,
bge_m3: (.models[] | select(.name=="bge-m3") | {active, dimension, max_seq_length, effective_max_seq_length})}'
Expect default_model: "bge-m3", bge_m3.active: true, effective_max_seq_length: 2048 (or your DAKERA_MAX_SEQ_LENGTH), fulltext_language as configured ("none" for multilingual and the CJK choices), query_languages: ["en","de","fr","es","it","pt","nl"] and reembed_pending: false once any model switch has finished.
Then:
- Store two or three memories in the target language and recall with a paraphrase. With
bge-m3a query in another language than the stored text should still find it. - For an existing namespace after a language change, run the reindex and check
analyzer_languageandrebuilt: truein the response. GET /healthlists aconfig_warningsentry for an unrecognizedDAKERA_FULLTEXT_LANGUAGEorDAKERA_QUERY_LANGon an upgraded deployment; an empty list means every value was honored.- After a CJK reindex, a full-text search for a two-character word returns its documents.
Resource implications
bge-m3: about 570 MB download plus an ORT-format copy on disk; slower per token thanbge-large(it has about 568M parameters). Its attention memory grows with the square of the truncation length.- Stemming, stop words and bigrams add no model and no measurable memory: they change what is indexed. Bigrams index each character as well as each pair, so a CJK namespace's full-text index is larger than the same text would be unsegmented.
- Re-embedding a store runs as a job (
GET /ops/jobs/{id}) that embeds every memory again, so budget CPU time in proportion to the number of memories.
See also: Configuration, Models and Docker, Multimodal memory, Search and ranking.