# Evaluation Metrics How scoring works in the benchmarking system. --- ## Overview The system uses **LLM-as-Judge** evaluation to automatically score answer quality across 4 metrics. --- ## Metrics ### 1. Context Relevance (1-10) **Question:** How relevant are the retrieved chunks to answering the question? | Score | Meaning | |-------|---------| | 9-10 | Chunks directly answer the question | | 7-8 | Chunks are highly relevant | | 5-6 | Chunks are somewhat relevant | | 3-4 | Chunks are partially relevant | | 1-2 | Chunks are not relevant | **What it measures:** Did the retrieval system find the right information? --- ### 2. Answer Similarity (1-10) **Question:** How similar is the generated answer to the expected answer? | Score | Meaning | |-------|---------| | 9-10 | Nearly identical to expected | | 7-8 | Covers most key points | | 5-6 | Covers some key points | | 3-4 | Partially correct | | 1-2 | Completely different | **What it measures:** Did the system produce the right answer? --- ### 3. Faithfulness (1-10) **Question:** Is the generated answer grounded in the retrieved context? | Score | Meaning | |-------|---------| | 9-10 | Entirely based on context | | 7-8 | Mostly based on context | | 5-6 | Partially based on context | | 3-4 | Some external knowledge used | | 1-2 | Mostly external knowledge | **What it measures:** Is the answer trustworthy? --- ### 4. Hallucination (boolean) **Question:** Did the LLM invent information not in the context? | Value | Meaning | |-------|---------| | false | Answer is grounded in context | | true | Answer contains invented information | **What it measures:** Is the answer fabricated? --- ## Scoring Process ### Step 1: Prepare Evaluation Context ``` Question: {user_question} Context: {retrieved_chunks} Expected Answer: {golden_answer} Generated Answer: {system_answer} ``` ### Step 2: Send to LLM-as-Judge ```python prompt = f""" Evaluate this RAG system output: Question: {question} Context: {context} Expected: {expected} Generated: {generated} Return JSON: {{ "context_relevance": <1-10>, "answer_similarity": <1-10>, "faithfulness": <1-10>, "hallucination": , "reasoning": "" }} """ ``` ### Step 3: Parse Response ```python scores = json.loads(llm_response) # Validate ranges for metric in ["context_relevance", "answer_similarity", "faithfulness"]: scores[metric] = max(1, min(10, scores[metric])) ``` --- ## Overall Score Calculation Each strategy gets an overall score weighted by importance: ```python overall = ( context_relevance * 0.3 + # 30% weight answer_similarity * 0.4 + # 40% weight faithfulness * 0.3 # 30% weight ) * (1 - hallucination_rate) # Hallucination penalty ``` ### Why These Weights? | Metric | Weight | Rationale | |--------|--------|-----------| | Answer Similarity | 40% | Most important - did we get the right answer? | | Context Relevance | 30% | Found the right information | | Faithfulness | 30% | Answer is trustworthy | | Hallucination | Penalty | Fabricated info is unacceptable | --- ## Aggregate Metrics Per strategy, we calculate: | Metric | Calculation | |--------|-------------| | avg_context_relevance | mean of all context_relevance scores | | avg_answer_similarity | mean of all answer_similarity scores | | avg_faithfulness | mean of all faithfulness scores | | hallucination_rate | count(hallucination=True) / total_questions | | total_questions | number of questions evaluated | | failed_questions | questions that errored | --- ## Example Output ```json { "recursive": { "avg_context_relevance": 8.5, "avg_answer_similarity": 7.8, "avg_faithfulness": 9.2, "hallucination_rate": 0.05, "total_questions": 21, "failed_questions": 0 } } ``` --- ## Interpretation Guide ### Good Scores | Metric | Target | Meaning | |--------|--------|---------| | Context Relevance | ≥ 8 | Retrieval is accurate | | Answer Similarity | ≥ 8 | Answers match expectations | | Faithfulness | ≥ 9 | Answers are trustworthy | | Hallucination Rate | ≤ 0.1 | Low fabrication rate | ### Warning Signs | Metric | Warning | Meaning | |--------|---------|---------| | Context Relevance | < 6 | Retrieval needs improvement | | Answer Similarity | < 6 | Answers are off-target | | Faithfulness | < 7 | Model is adding external knowledge | | Hallucination Rate | > 0.2 | High fabrication rate | --- ## Cost Considerations Each evaluation uses: | Component | Cost | |-----------|------| | LLM call (gpt-4o-mini) | ~$0.001 per evaluation | | Input tokens | ~500 per evaluation | | Output tokens | ~100 per evaluation | **Total for 21 questions × 5 strategies:** - 105 evaluations × $0.001 = **$0.105** --- ## Configuration Evaluation parameters in `src/benchmarking/evaluation.py`: ```python _EVALUATION_SYSTEM_PROMPT = "You are an expert evaluator..." _EVALUATION_USER_PROMPT = "Evaluate this RAG system output..." model = settings.llm_model # gpt-4o-mini temperature = 0.0 # Deterministic max_tokens = 500 # Response limit ```