Files
chatbot_v3/docs/adr/0005-reranking-model-and-sparse-analyzer-selection.md
Ali Zarinkolah 5c0a5938f8 feat(ingestion): add bounded, benchmark-aligned embedding execution
Why:
- Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into
  the inline upload path, with process-wide capacity/timeout/chunk-limit
  guards (ADR-0017).
- The BM25 analyzer and dense-model config are ported from the `emet`
  evaluation lab, which benchmarked them against the real Farsi corpus
  (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large
  at native 3072-dim), closing open items in ADR-0001/ADR-0005.

Changes:
- New: embedding ports, orchestration (embed_chunks), request-bounds
  helpers, and dense/sparse adapters (analyzers.py, bm25.py,
  openai_compatible.py).
- upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY
  (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413);
  every failure path still writes a terminal job row.
- Lifespan builds and warms both dense embedders at startup (fail-soft) and
  creates the sparse embedder and concurrency semaphore.
- httpx moves from dev to main dependencies (adapters use it directly).

Impact:
- Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0.
- New EMBEDDING_* env vars documented in .env.example; safe defaults.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 17:13:32 +03:30

225 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 0005. Reranking model and sparse (BM25) analyzer selection for the Farsi corpus
## Status
Proposed — the fusion/rerank *shape*, the reranker model, and (as of the
`emet` benchmark, see "Benchmark outcome" below) the **BM25 analyzer** are
decided. The commercial license status of the reranker remains open per the
follow-up items below.
## Context
ADR-0003 fixed the retrieval *shape* (dense_nomic + dense_openai + sparse
prefetch → RRF fusion → late-interaction rerank) but deliberately left two
things open: which model powers `late_interaction`, and the sparse-side
analyzer detail behind `bm25-fa-norm-stop` (ADR-0001). This ADR resolves
those, driven by the corpus being **Farsi (Persian)** — a morphologically
rich, low-resource language for most public embedding/rerank benchmarks and
for Qdrant's own hosted tooling.
Two things specific to Farsi drove this investigation rather than picking a
generic default:
- Qdrant's hosted `Qdrant/bm25` FastEmbed model's documented supported-
language list does not include Farsi (`fa`) stemming — using it as-is
would silently apply no stemming rule, or the wrong one, for this corpus.
- Farsi's morphology (verb conjugation, ezafe constructions, high-frequency
function words) makes pure BM25/keyword signals noisier than in English,
which raised the bar on how load-bearing the rerank stage needs to be —
treated here as close to mandatory for quality, not optional.
## Decision
### 1. Fusion: RRF by default, weighted RRF as a fallback
Confirms ADR-0003's fusion stage combines all three prefetch results
(`dense_nomic`, `dense_openai`, `sparse`) via **RRF** by default. If one
signal (typically sparse, on Farsi text) is observed to dominate the fused
ranking unexpectedly, **weighted RRF** is the fallback — not a switch to a
different fusion algorithm. This refines, not replaces, ADR-0003's fusion
decision.
### 2. Late-interaction reranker: jina-colbert-v2
Two multilingual late-interaction (ColBERT-style) options were evaluated
against the requirement of confirmed Farsi support:
| Model | Farsi support | License | Local hosting |
|---|---|---|---|
| **jina-colbert-v2** | Confirmed — `fa` explicitly listed among 89 supported languages | Conflicting: HF repo metadata says `cc-by-4.0`; Jina's own announcement states CC BY-NC-4.0 (non-commercial), commercial use via paid API/AWS/Azure only | Possible via PyLate/RAGatouille; requires `flash_attn` (CUDA GPU effectively required); reported loading issues via generic `transformers.AutoModel` |
| **LFM2-ColBERT-350M** | Not supported (8 languages: en, ar, zh, fr, de, ja, ko, es) | — | — |
| BGE-M3 (alternative, not adopted) | Strong Farsi performance in the FaMTEB benchmark | Apache-2.0 — unambiguous, commercial-friendly | CPU-runnable, no `flash_attn` dependency |
**jina-colbert-v2 is adopted** as the `late_interaction` reranker: it has
explicitly confirmed Farsi support among its 89 languages, which is the
primary requirement for this corpus. This means:
- A **CUDA GPU is now a required dependency** for the rerank stage
(`flash_attn`), not optional infrastructure — this changes the self-hosted
Docker deployment assumption from ADR-0001/0003, which had not committed
to GPU hosting.
- The **commercial license is unresolved** — HF repo metadata states
`cc-by-4.0` while Jina's own announcement states CC BY-NC-4.0
(non-commercial, with commercial use only via Jina's paid API/AWS/Azure
offering). This must be confirmed directly with Jina **before** this
project ships commercially on a self-hosted jina-colbert-v2 model; if
Jina confirms the non-commercial reading, self-hosting it commercially is
not an option and the fallback is Jina's paid hosted API or BGE-M3 as a
substitute reranker (same `late_interaction` vector shape, no schema
change needed either way).
- Loading it outside PyLate/RAGatouille (e.g. generic
`transformers.AutoModel`) has reported issues — plan to load it through
one of those two libraries, not a raw `transformers` call.
`late_interaction` continues to use `hnsw_config: m=0` (per ADR-0001,
now made explicit: HNSW indexing is disabled for this vector because it is
used only for reranking already-fetched candidates via MAX_SIM, never for
independent ANN search — the recommended pattern for rerank-only
multivectors) plus on-disk storage.
### 3. Sparse retrieval: custom BM25 pipeline, not Qdrant's hosted model
Confirms ADR-0001's choice: the project's own BM25 pipeline (Farsi
normalization/stopword/stemming computed outside Qdrant Cloud Inference) is
used instead of Qdrant's hosted `Qdrant/bm25` FastEmbed model, specifically
because that hosted model's documented language list omits Farsi. The
sparse vector is uploaded with `modifier="idf"`, the standard BM25 sparse
vector configuration.
Four analyzer variants were benchmarked, all sharing identical BM25 scoring
parameters (`k=1.2`, `b=0.75`) so any accuracy difference is attributable
entirely to the analyzer stage, not the ranking formula:
| Analyzer | Description |
|---|---|
| `bm25-raw` | no normalization |
| `bm25-fa-norm` | Farsi normalization only |
| `bm25-fa-norm-stop` | normalization + stopword removal — **current best performer** |
| `bm25-fa-norm-stem` | normalization + stemming |
`bm25-fa-norm-stop` is confirmed as the sparse analyzer named in ADR-0001,
consistent with Farsi's high density of function words (ezafe particles,
prepositions, common verbs) adding TF/IDF noise if left in.
### 3a. Benchmark outcome: `bm25-fa-norm-stop` confirmed, and where the BM25 math runs
The `emet` evaluation lab (`~/code/talie/emet`) ran the four-variant
comparison above against the real Farsi corpus and confirmed
**`bm25-fa-norm-stop`** as the winner. It is the only sparse variant promoted
into emet's hybrid matrix (`emet/hybrid.yaml`). This closes follow-up item 4
below.
The winning analyzer is a specific, reproducible artifact, ported into
`src/infrastructure/embedding/analyzers.py` and verified token-for-token
against emet's implementation. Its details are load-bearing:
- Unicode **NFC** (not NFKC), then ZWNJ → space, then Persian/Arabic-Indic
digits → ASCII, then `ي→ی ك→ک ة→ه ؤ→و إ→ا أ→ا`.
- Tokenizer `[^\W_]+`, which **keeps digits**. This matters for an insurance
corpus: policy numbers, dates, and amounts are exactly the terms lexical
retrieval should match, and the digit folding above means a query in ASCII
digits matches a document authored in Persian ones.
- A 51-entry stopword set (40 Persian/Arabic + 11 English, the corpus being
mixed-script). Deliberately not a full `hazm` list.
- No stemming — `fa_norm_stem` was the losing arm.
**The BM25 formula is split across two systems, deliberately.** The client
applies term-frequency saturation, including the `k`/`b` document-length
normalization; **IDF is supplied by Qdrant** via `modifier="idf"` on the
sparse vector field, computed from collection-wide statistics rather than
from a fixed client-side corpus.
That split is a correctness trap worth stating plainly: a `chunks` collection
created *without* `modifier="idf"` will score these vectors as saturated term
frequencies with no IDF weighting at all — no error, no warning, just
materially worse lexical retrieval. The collection bootstrap must set it.
Document and query encoding are asymmetric in exactly one term: documents
carry the `b` length normalization, queries do not (standard BM25 practice).
Both sides must therefore encode through the same implementation, which is
why the sparse port carries a `query` flag rather than leaving retrieval to
grow a second, silently divergent encoder.
Term → sparse-index mapping is `blake2b(token, digest_size=8) % (2**31 - 1)`,
a pure hash with no vocabulary table, so it needs no shared state and stays
identical across processes and between ingest and query time. Changing the
hash orphans every stored sparse vector: that is a re-ingestion, not a deploy.
### 4. BM25 parameters: keep `k=1.2`, `b=0.75`; tune analyzer, not formula
These are standard, well-validated defaults (Trotman, Puurula & Burgess,
2014) and are not the source of the observed analyzer-variant accuracy
differences — so formula tuning is deprioritized in favor of the analyzer
comparison and `b` sweep in the follow-ups below.
## Consequences
### Positive
- Resolves ADR-0003's two deferred decisions (reranker model, sparse
analyzer detail) with a corpus-specific rationale instead of a generic
default.
- jina-colbert-v2 is the only evaluated option with explicitly confirmed
Farsi support, directly matching this project's primary requirement.
- Isolating BM25 formula parameters from analyzer choice gives a clean,
defensible experimental basis for the `bm25-fa-norm-stop` selection.
### Negative
- Introduces a hard GPU dependency (`flash_attn`/CUDA) for the rerank stage
that ADR-0001/0003 hadn't assumed — self-hosted deployment now needs GPU
capacity, not just Docker on commodity hardware.
- Commercial license status is unresolved; shipping this commercially on a
self-hosted jina-colbert-v2 model without confirming licensing with Jina
is a legal risk, not just a technical one.
- Loading path is constrained to PyLate/RAGatouille due to reported
`transformers.AutoModel` issues — an extra library dependency and less
flexibility than a standard `transformers` load would give.
- `bm25-fa-norm-stop` is provisional until compared directly against
`bm25-fa-norm-stem` — Persian stemming can over-collapse distinct words
(irregular verb conjugation, Arabic-loanword plurals), so the current
"best performer" result could shift.
## Alternatives Considered
- **BGE-M3 as the reranker**: not adopted — Apache-2.0 license and
CPU-runnability are attractive and it remains the fallback if
jina-colbert-v2's license is confirmed non-commercial, but jina-colbert-v2
was prioritized for its explicit Farsi support.
- **LFM2-ColBERT-350M**: rejected outright — no Farsi support among its 8
supported languages.
- **Qdrant's hosted `Qdrant/bm25` FastEmbed model**: rejected — its
documented language support does not include Farsi stemming; using it
would risk silently wrong or absent stemming for this corpus.
- **Treating "RRF vs. rerank" as either/or**: rejected — RRF alone can't
resolve disagreement between two distinct dense embedding spaces plus a
noisy Farsi BM25 signal, so the two-stage prefetch-fusion-then-rerank
pipeline from ADR-0003 is kept, not replaced with either component alone.
## Follow-up / Open Items
1. Confirm jina-colbert-v2's commercial license status directly with Jina
before commercial deployment; fall back to Jina's paid hosted API or
BGE-M3 if the non-commercial reading is confirmed.
2. Provision GPU capacity for self-hosted jina-colbert-v2 (`flash_attn`
requires CUDA) as part of the deployment plan, not an afterthought.
3. Run an ablation: single dense model + sparse + rerank vs. the current
dual-dense-model + sparse + rerank setup, on real Farsi queries, to
justify (or drop) the second dense vector (`dense_openai`).
4. ~~Compare `bm25-fa-norm-stop` vs. `bm25-fa-norm-stem` in isolation~~ —
**done**, see "Benchmark outcome" above. `fa_norm_stop` won; stemming was
not adopted.
5. Sweep BM25 `b` (e.g. 0.5–0.9) for the winning analyzer, since document
length varies significantly across the corpus (short chat messages vs.
long articles) and `0.75` is a generic default, not corpus-tuned.
6. **Recalibrate `avg_len`.** The client-side `b` term needs an average
document length in *analyzer tokens*. The ported value (256.0) is emet's
own placeholder, and emet measured it over short Q&A records rather than
this service's ~400-token chunks, so it is very likely miscalibrated here.
Exposed as `EMBEDDING_SPARSE_AVG_LEN` so it can be corrected from real
corpus statistics without a code change.
7. **Re-benchmark the analyzer with diacritic stripping.** `fa_norm_stop`
does not remove harakat or tatweel, so `ســلام` and `سلام` are distinct
terms. `src/application/ingestion/normalization.py` already strips both
for chunk *content*; extending that to the analyzer is plausibly an
improvement but would deviate from the measured configuration, so it
belongs in an emet run rather than an unmeasured edit.