Files
chunking_strategies_evaluation/docs
Mahdi Bazrafshan d3f8a9a5e5 docs: add project documentation and task tracking
Why:
- Need to track domain model, ADRs, task progress, and feature planning

Changes:
- CONTEXT.md: domain model with ADRs 0001-0014
- docs/tasks.md: updated task list with Phase 6 (Dashboard)
- docs/cant-do-yet.md: backend-ready but no UI features
- docs/out-of-scope-v1.md: intentionally excluded features

Impact:
- Project documentation centralized for reference
2026-07-29 17:55:01 +03:30
..

Documentation Index

Complete documentation for the RAG Chunking Benchmarker.


Quick Start

  1. New to the project? Start with Architecture Overview
  2. Want to understand strategies? Read Strategy Technical Details
  3. Need to use the API? Check API Reference
  4. Configuring the system? See Configuration Guide
  5. Understanding results? Read Evaluation Metrics
  6. Curious about data flow? See Data Flow

Documentation Files

File Purpose Audience
architecture.md System structure and design New team members
strategy-technical-details.md Deep dive into each strategy Engineers
api-reference.md All endpoints documented Developers
configuration.md Settings and environment variables DevOps
evaluation-metrics.md How scoring works Data scientists
data-flow.md How data moves through the system Engineers
chunking_strategies.md High-level strategy overview Everyone
phases.md Implementation phases Project managers
tasks.md Task tracking Developers

Architecture at a Glance

┌─────────────────────────────────────────────────────────────────┐
│                    RAG Chunking Benchmarker                     │
├─────────────────────────────────────────────────────────────────┤
│                                                                 │
│  Documents → Parser → Chunking → Embedding → Qdrant           │
│                                                                 │
│  Questions → Query Pipeline → LLM → Answers                    │
│                                                                 │
│  Answers → Evaluation → Metrics → Reports                      │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

Key Concepts

Chunking Strategies

Strategy Splitting Best For
fixed_size Token count Baseline comparison
recursive Structural boundaries Structured documents
semantic Similarity thresholds Topic shifts
contextual_retrieval LLM-enriched tokens Retrieval quality
semantic_parent_child Paragraph clusters Context needed

Evaluation Metrics

Metric What It Measures
Context Relevance Did we find the right information?
Answer Similarity Did we produce the right answer?
Faithfulness Is the answer trustworthy?
Hallucination Did we invent information?

API Endpoints

Endpoint Method Purpose
/documents GET/POST Manage documents
/documents/{id}/process POST Run chunking strategies
/queries POST Ask questions
/benchmarks POST Run comparisons
/benchmarks/{id}/report GET View results
/experiments GET List experiments

Configuration

All settings via .env file. See Configuration Guide for details.

Key settings:

  • OPENAI_API_KEY - Required for embeddings and LLM
  • CHUNK_SIZE - Target tokens per chunk (default: 512)
  • SEMANTIC_THRESHOLD - Similarity threshold (default: 0.5)

Further Reading

  • Phases - Implementation roadmap
  • Tasks - Task tracking
  • ADRs - Architectural Decision Records