1.5 KiB
1.5 KiB
ADR-0022: Per-Embedding-Model semantic_threshold
Status
Approved
Context
Semantic Boundary Detection cuts when adjacent-unit cosine similarity falls below a threshold. Cloud OpenAI and local Nomic produce different similarity distributions for the same Farsi text: the global SEMANTIC_THRESHOLD (config, historically ~0.3) rarely triggers cuts under Nomic, collapsing semantic into a single chunk. Operators need a higher Nomic default without changing OpenAI behavior, and a way to tune without editing .env and restarting.
Decision
- Each Embedding Model Registry entry has a
default_semantic_threshold(OpenAItext-embedding-3-small: 0.3; Nomicnomic-embed-text-v2-moe: **0.6`). - Admin may override the effective value per model id in SQLite (
semantic_threshold:{model_id}). - Process snapshots the Active Embedding Model and uses that model’s effective threshold for
semanticandsemantic_parent_child. - Admin API: list includes
semantic_threshold/default_semantic_threshold;PUT /admin/embedding-models/{id}/semantic-thresholdpersists overrides. - Changing the threshold does not rewrite existing corpora — re-process to apply.
Consequences
- Global
SEMANTIC_THRESHOLDremains a fallback only when a strategy is called without an explicit threshold (tests/legacy); production process path always passes the model-resolved value. - Operators must re-process after tuning; Experiments under different thresholds are not auto-invalidated.