3.0 KiB
3.0 KiB
ADR 0005: Retrieval Evaluation Framework
Status
Accepted
Context
The current system has no metrics for retrieval quality. The previous classification accuracy was unknown, making improvement impossible. The new architecture needs measurement from day one to:
- Establish baseline performance
- Compare embedding models (OpenAI vs nomic)
- Compare reranking impact
- Detect regressions after changes
- Tune hyperparameters (K, N, thresholds)
Decision
Implement a multi-source test set combining:
1. Historical Query Labeling
- Sample real queries from production logs
- Manually label with correct chunk IDs and domain
- Focus on edge cases and ambiguous queries
- Target: ~50-100 labeled queries
2. Synthetic Q&A Generation
- Use LLM to generate questions from each chunk
- Question + source chunk = labeled test pair
- Ensures coverage across all domains and document types
- Target: ~100-200 synthetic pairs
3. LLM Auto-Labeling with Human Review
- Use LLM to judge relevance of retrieved chunks
- Flag low-confidence or disputed labels for human review
- Accelerates labeling while maintaining quality
- Target: ~50-100 auto-labeled with review
Metrics to Track
| Metric | Definition | Target |
|---|---|---|
| Precision@5 | Fraction of top 5 results that are relevant | > 0.80 |
| Precision@10 | Fraction of top 10 results that are relevant | > 0.70 |
| MRR | Mean Reciprocal Rank of first relevant result | > 0.85 |
| Recall@20 | Fraction of relevant docs found in top 20 | > 0.90 |
| Latency p50 | Median retrieval time | < 500ms |
| Latency p95 | 95th percentile retrieval time | < 1000ms |
| Domain Accuracy | Correct domain in top 5 results | > 0.95 |
Evaluation Pipeline
Test Query → Embed → Retrieve (K=20) → Rerank → Top N → Compare to Labels → Log Metrics
Comparison Modes
- Online vs Offline — Compare OpenAI + Cohere vs Nomic + BGE
- With vs Without Reranking — Measure reranking impact
- Before vs After Changes — Detect regressions
Consequences
Positive
- Measurable improvement — quantifiable baseline and progress
- Model comparison — data-driven decision on embedding/reranking choices
- Regression detection — catch quality degradation early
- Hyperparameter tuning — optimize K, N, thresholds with data
Negative
- Upfront investment — labeling requires human effort
- Maintenance — test set needs updates as corpus grows
- Storage — need to store labels and evaluation results
Neutral
- Synthetic questions may not match real query distribution
- LLM labeling may have bias — human review essential for edge cases
Implementation Notes
- Store test set in JSON/Parquet:
{query, relevant_chunk_ids, domain, source} - Build evaluation script that runs retrieval and computes metrics
- Log results to a simple database or file for trend analysis
- Run evaluation on every significant change
- Consider integrating with LlamaIndex's evaluation module