Why: - Manager deliverable for human evaluation across 10 benchmark documents. Changes: - Report markdown, scored CSV, chart generator, and SVG assets - Linked from docs/README.md Co-authored-by: Cursor <cursoragent@cursor.com>
3.5 KiB
3.5 KiB
Human Evaluation Report — fixed_size ±3
Date: 22 August 2026
Strategy: fixed_size with Neighbor Expansion ±3/3
Corpus Embedding Model: text-embedding-3-large
Sample: 60 questions (6 per document × 10 documents)
Total benchmark universe: 201 questions
Review effort: ~6 hours
Verdict: Human review supports the stabilized fixed_size ±3 default (see final decision).
| KPI | Value |
|---|---|
| Mean human score (1–5) | 4.30 |
| Mean LLM answer similarity (1–10) | 8.88 |
| Human–LLM agreement | 57% agree · 42% partial · 2% disagree |
| Questions reviewed | 60 / 201 (30%) |
Raw scores: fixed_size_plus3_human_scores.csv
1. Mean human score by document
2. Human vs LLM-as-Judge agreement
Agreement = same quality tier when mapping LLM answer similarity (1–10) to human scale (1–5).
3. Human score vs LLM similarity by document
4. Human score distribution
5. Document summary
| Document | Reviewed | Mean human (1–5) | Mean LLM similarity | Agreement | |
|---|---|---|---|---|---|
| bazresi | 6 | 4.33 | 8.50 | 50% | ⚠ |
| customer1 | 6 | 4.00 | 9.00 | 33% | ⚠ |
| fire | 6 | 4.50 | 8.33 | 67% | |
| havades-ind | 6 | 3.83 | 8.50 | 50% | ⚠ |
| havades | 6 | 4.00 | 8.83 | 50% | ⚠ |
| lifetime-ind | 6 | 4.33 | 9.33 | 50% | ⚠ |
| moavenin | 6 | 4.83 | 9.50 | 83% | |
| Refah | 6 | 4.67 | 9.33 | 83% | |
| website | 6 | 4.17 | 9.00 | 17% | ⚠ |
| lifetime-comp | 6 | 4.33 | 8.50 | 83% |
6. Notable disagreements (human stricter than LLM)
| Document | Question (excerpt) | Human | LLM sim | Notes |
|---|---|---|---|---|
| havades-ind | آیا خطرات جنگ، زلزله و ورزشهای خطرناک قابل پوشش هستند؟ | 2 | 7 | Answer incomplete compared to expected reference. |
7. Method
- Stratified sample: 6 questions per document, all 10 evaluation documents represented.
- Human rubric: 1–5 (1 = wrong/unhelpful, 3 = partial, 5 = matches expected answer).
- Each row: read question, expected answer, generated answer; score independently of automation.
- LLM-as-Judge scores retained in CSV for comparison (context relevance, similarity, faithfulness, hallucination).
- Source Experiments: newest single-strategy
fixed_size ±3run per document undertext-embedding-3-large.
8. Conclusions
- Overall quality is strong — mean human score 4.30/5 on the reviewed sample.
- Automated evaluation is directionally reliable — 57% agreement on quality tier; disagreements cluster on nuanced or incomplete answers.
- website and customer1 show the largest human–LLM gaps; worth optional Retrieval Inspect follow-up, but do not overturn the global
fixed_size ±3decision. - Human review confirms the family and ±3 configuration documented in final-chunking-strategy-decision.md.
Regenerate charts after CSV updates:
.venv/bin/python docs/assets/human-eval/generate_human_eval_report.py