ef8bab59628fb6097a7da734b977ea8b6b8d911c
Why: - Need LLM-as-Judge evaluation for automated scoring - Need benchmark orchestration to run questions × strategies - Need HTML report generation with two views (managerial/technical) Changes: - evaluation.py: LLM-as-Judge scoring on 4 metrics (context, similarity, faithfulness, hallucination) - benchmark_service.py: Orchestration with per-strategy failure isolation - report.py: Dual-view HTML reports with dark mode, charts, and strategy cards
chunking_strategies_evaluation
This repo is in order to reach multiple chunking strategies and evaulate them
Languages
Python
62.4%
HTML
37.6%