Files
chunking_strategies_evaluation/docs/final-chunking-strategy-decision.md
Mahdi Bazrafshan dc76ca67b8 docs: lock fixed_size ±3 in decision record and lld
Why:
- Close the open ±N follow-up from the strategy finalization memo.

Changes:
- Add addendum with ±0…±3 evidence and binding ±3/3 default.
- Sync LLD config table with new neighbor expansion defaults.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-22 13:27:18 +03:30

172 lines
8.7 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Chunking Strategy — Final Decision
**Date:** 17 August 2026 (family); **22 August 2026** (±N locked)
**Decision:** Adopt **`fixed_size`** as the chunking Strategy family with **Neighbor Expansion ±3/3**.
**Corpus Embedding Model:** `text-embedding-3-large`
**Evaluation set:** 10 Word documents (Decision Board / neighbor-sweep universe)
**Questions:** 201 per Candidate (nomic semantic: 195)
**One-line close:** **`fixed_size ±3`** is the stabilized Strategy Candidate (**9.040** mean composite vs **8.727** for semantic @ large). **Do not use ±0** as the default — it loses to semantic on mean.
Charts below are generated from `data/chunking_benchmark.db` with the same composite as the Decision Board:
```text
(0.3 × context relevance + 0.4 × answer similarity + 0.3 × faithfulness)
× (1 − hallucination rate)
```
Regenerate figures after new Experiments:
```bash
.venv/bin/python docs/assets/decision/generate_charts.py
```
---
## 1. Why `fixed_size` is the winner
Decision Board ranks by **mean composite**, not by “how many documents won.” Stage 1 auto-picks **`fixed_size ±3`** vs **`semantic @ text-embedding-3-large`**. Stage 2 mean composite is **+0.313** for `fixed_size`.
![Stage 1 mean composite by Strategy Candidate](assets/decision/stage1-mean-composite.svg)
| Candidate | Mean composite | Context | Similarity | Faithfulness | Hallucination | Docs |
|-----------|----------------|---------|------------|--------------|---------------|------|
| **fixed_size ±3** | **9.040** | 9.40 | 8.95 | 9.47 | 2.2% | 10/10 |
| fixed_size ±2 | 8.987 | 9.43 | 8.93 | 9.44 | 2.7% | 10/10 |
| fixed_size ±1 | 8.926 | 9.38 | 8.89 | 9.42 | 3.0% | 10/10 |
| **semantic @ text-embedding-3-large** | **8.727** | 9.38 | 8.83 | 9.30 | 4.8% | 10/10 |
| fixed_size ±0 | 8.606 | 9.18 | 8.79 | 9.25 | 4.9% | 10/10 |
| semantic @ nomic-embed-text-v2-moe | 8.551 | 9.19 | 8.65 | 9.14 | 4.7% | 10/10 |
`fixed_size ±1`, `±2`, and `±3` all beat the best semantic Candidate on **mean**. **`±0` (no Neighbor Expansion) does not** (8.606 vs 8.727). Shipping `fixed_size` without expansion would weaken this family decision.
---
## 2. Judge metrics (stage 2 winners)
`fixed_size ±3` leads on every judge metric, including **lower hallucination**.
![Mean judge metrics for stage 2 winners](assets/decision/stage2-metrics.svg)
---
## 3. Per-document showdown
Mean ranking still favors `fixed_size`. **Head-to-head at ±3: semantic 6 / `fixed_size` 4.** Semantic’s six wins are mostly **small**; `fixed_size`’s four wins are **larger**, especially **fire**.
![Per-document composite: fixed_size ±3 vs semantic @ large](assets/decision/stage2-per-document.svg)
![Signed margin per document (fixed_size − semantic)](assets/decision/stage2-margins.svg)
| Document | Questions | fixed_size ±3 | semantic @ large | Winner |
|----------|-----------|---------------|------------------|--------|
| fire.docx | 20 | **8.810** | 6.228 | **fixed_size** (+2.58) |
| moavenin.docx | 20 | **9.735** | 8.621 | **fixed_size** (+1.11) |
| general-havades-individuals.doc | 20 | **9.205** | 8.564 | **fixed_size** (+0.64) |
| havades.docx | 20 | **8.483** | 7.969 | **fixed_size** (+0.51) |
| bazresi.docx | 16 | 9.366 | **9.405** | semantic (−0.04) |
| life-time-individual.docx | 20 | 9.445 | **9.505** | semantic (−0.06) |
| Refah.docx | 20 | 9.680 | **9.815** | semantic (−0.14) |
| lifetime-compensation.docx | 15 | 8.761 | **9.129** | semantic (−0.37) |
| customer1.docx | 20 | 8.451 | **8.973** | semantic (−0.52) |
| website.docx | 30 | 8.466 | **9.058** | semantic (−0.59) |
Win-count would pick semantic. **Official product ranking (mean composite) picks `fixed_size`.** This memo follows the Decision Board rule (ADR-0026).
---
## 4. Full Candidate heatmap
Every cell is a newest single-strategy Experiment under `text-embedding-3-large`. Semantic @ large **collapses on fire**; `fixed_size ±1…±3` stay high across the set.
![Heatmap of composite scores by document and Candidate](assets/decision/heatmap-candidates.svg)
Best `fixed_size` ±N **by document** (does not change the family call):
| ±N | Documents where it is the best fixed_size cell |
|----|------------------------------------------------|
| ±3 | bazresi, fire, general-havades-individuals, havades, life-time-individual (5) |
| ±1 | customer1, moavenin, Refah (3) |
| ±2 | website (1) |
| ±0 | lifetime-compensation (1) |
---
## 5. Method
- Corpus Embedding Model locked to **`text-embedding-3-large`**
- Newest **single-strategy** Experiment per document × Candidate (Decision Board cells)
- Retrieval `top_k = 5`
- Stage 1: best `fixed_size` ±N vs best `semantic@Boundary`
- Stage 2: mean composite + per-document breakdown
- Experiments dated **9–10 August 2026**
- `logs/neighbor_sweep.log` recorded mid-run connection errors on some units; **SQLite now has a full 10×4 `fixed_size` grid** — treat the database as source of truth
---
## 6. Excluded strategies
No comparable **10-doc, single-strategy, `text-embedding-3-large`** Experiments exist for the three Strategies below. They are not Decision Board Candidates (ADR-0026).
| Strategy | Why it is not the winner |
|----------|--------------------------|
| **recursive** | Not on the Decision Board grid. Informal PDF / `text-embedding-3-small` multi-strategy runs (not comparable to this close-out) were mixed; recursive **does not get Neighbor Expansion** (expansion is `fixed_size` only). |
| **contextual_retrieval** | Extra LLM call per chunk at process time. Same informal PDF/`small` runs scored **below** recursive and `fixed_size`. |
| **semantic_parent_child** | Boundary-detection cost plus parent fetch at query. Informal PDF/`small` composites were **much worse** (~1.8–2.8). Fail-hard if Boundary embeds mismatch (ADR-0020). |
One leftover five-strategy Experiment on `customer1` under large is **invalid** (four Strategies scored 0 / errors) and was ignored.
---
## 7. Binding decision
**Use `fixed_size` for chunking** under **`text-embedding-3-large`.**
**Use Neighbor Expansion `neighbor_prev=3`, `neighbor_next=3`** (symmetric **±3/3**) for Query and Experiment defaults when running `fixed_size`.
**Do not use `semantic` as the default Strategy** on this evaluation universe.
**Do not default to ±0.** Without expansion, semantic @ large beats `fixed_size` on mean composite (8.727 vs 8.606).
Operational defaults are set in `src/core/config.py`, `.env.example`, and the Dashboard Query/Benchmarks forms. Override per request is still supported.
---
## 8. Addendum — why ±3 (22 August 2026)
Stage 1 ranked all `fixed_size` Candidates on the 10-doc grid:
| ±N | Mean composite | vs semantic @ large |
|----|----------------|---------------------|
| ±0 | 8.606 | **loses** (−0.121) |
| ±1 | 8.926 | wins (+0.199) |
| ±2 | 8.987 | wins (+0.260) |
| **±3** | **9.040** | **wins (+0.313)** |
**±3** is the stage 1 auto-pick and the highest mean composite. **±1** and **±2** are close; **±0** is ruled out.
Per-document best ±N varies (±3 wins on 5 docs, ±1 on 3, ±2 on 1, ±0 on 1). The **global** default is still **±3** because Decision Board ranks by mean composite across the full set, not by win-count among ±N levels.
**Caveats kept from stage 2:** semantic still wins head-to-head on 6/10 docs vs ±3 (website, customer1, etc.), but mean composite and judge metrics favor **`fixed_size ±3`**. Retrieval Inspect on outlier docs remains optional follow-up.
**Production RAG outside this repo** is not changed by this addendum — only benchmarker defaults and this record.
---
## 9. Evidence (stage 2 cells)
| Document | `fixed_size ±3` experiment id | `semantic @ large` experiment id |
|----------|-------------------------------|----------------------------------|
| bazresi.docx | `6fe08750ee324130b3d46d8b3bc95280` | `6263c4f9e63347cea9eb19f4d36309c0` |
| customer1.docx | `8c4d888d367b4c0299b2d19dd9654f80` | `0c4572e80902408796f0dd11776e8e55` |
| fire.docx | `230a3cf57ae74b679ecde83cfeb9b6f5` | `414162efc314406db082ba8537c2423f` |
| general-havades-individuals.doc | `7ad892cfcd6143109f77a7808c8dfbd9` | `14fa4f872373446c8b042fc8e84e6983` |
| havades.docx | `04bd45d836a64c3da2e68bb747c0c9b8` | `232238b4d3304243b82b02326ef64617` |
| life-time-individual.docx | `d682706f54824469b235c099bf5a3d3c` | `bc7d9f1fbe3540c097a5f8cc1a1c44bb` |
| moavenin.docx | `63f81cf0a632418aab2f93884518ed71` | `4e01960b5db14e81975a4fbc86c4e10c` |
| Refah.docx | `6304dba39f3f4c05935195c588a36a64` | `b5c630c45615451a93d97d80dee361df` |
| website.docx | `fb9d47750ca74761a3c82c57aa51030b` | `53256897cfb74dd7a6508e750ff1ba63` |
| lifetime-compensation.docx | `99736ffbc6da455b8b3ca3a0fb8e5ef7` | `e51c2b57e7364dfa97e5429f615d6e4e` |
HTML reports: `GET /benchmarks/{id}/report`. Dashboard: Decision Tab, Corpus = `text-embedding-3-large`.