Why:
- Queries and Experiments must hit the Corpus Embedding Model's collections and optionally widen fixed_size context.
Changes:
- Resolve corpus model per request; apply Neighbor Expansion with Expansion Tree; persist and report expansion provenance.
Co-authored-by: Cursor <cursoragent@cursor.com>
Why:
- Make per-question strategy answers inspectable and show DOCX vs Text PDF source.
Changes:
- Side-by-side compare modal with HTML escaping; header source badge from filename.
Co-authored-by: Cursor <cursoragent@cursor.com>
Why:
- Users need to see each strategy's overall score at a glance in the report
- Technical report should show average metrics across all strategies
Changes:
- Each strategy card now shows its overall score below the metric bars
- Technical report aggregate table includes an AVERAGE row across all strategies
Impact:
- Managerial report: strategy cards now show overall score
- Technical report: aggregate table has a new AVERAGE row at the bottom
Why:
- Experiment list and detail responses showed raw IDs instead of filenames
- No way to delete experiments from the UI
Changes:
- Add document_filename field to ExperimentDetailResponse
- Enrich list endpoint with filenames and calculated best_strategy from aggregate_metrics
- Add DELETE /experiments/{id} endpoint
Impact:
- API responses now include document_filename for all experiment endpoints
- Frontend can display filenames instead of IDs
Why:
- Need request/response models for benchmark endpoints
- Need routes for creating and retrieving benchmarks
- Need view parameter for managerial vs technical report views
Changes:
- models.py: Added BenchmarkRequest, BenchmarkResponse, StrategyMetrics, ExperimentDetailResponse
- routes.py: Added POST /benchmarks, GET /benchmarks/{id}, GET /experiments, view parameter for reports
Why:
- Need LLM-as-Judge evaluation for automated scoring
- Need benchmark orchestration to run questions × strategies
- Need HTML report generation with two views (managerial/technical)
Changes:
- evaluation.py: LLM-as-Judge scoring on 4 metrics (context, similarity, faithfulness, hallucination)
- benchmark_service.py: Orchestration with per-strategy failure isolation
- report.py: Dual-view HTML reports with dark mode, charts, and strategy cards