Documentation Index
Complete documentation for the RAG Chunking Benchmarker.
Quick Start
- New to the project? Start with Architecture Overview
- Want to understand strategies? Read Strategy Technical Details
- Need to use the API? Check API Reference
- Configuring the system? See Configuration Guide
- Understanding results? Read Evaluation Metrics
- Curious about data flow? See Data Flow
Documentation Files
Architecture at a Glance
Key Concepts
Chunking Strategies
| Strategy |
Splitting |
Best For |
| fixed_size |
Token count |
Baseline comparison |
| recursive |
Structural boundaries |
Structured documents |
| semantic |
Similarity thresholds |
Topic shifts |
| contextual_retrieval |
LLM-enriched tokens |
Retrieval quality |
| semantic_parent_child |
Paragraph clusters |
Context needed |
Evaluation Metrics
| Metric |
What It Measures |
| Context Relevance |
Did we find the right information? |
| Answer Similarity |
Did we produce the right answer? |
| Faithfulness |
Is the answer trustworthy? |
| Hallucination |
Did we invent information? |
API Endpoints
| Endpoint |
Method |
Purpose |
/documents |
GET/POST |
Manage documents |
/documents/{id}/process |
POST |
Run chunking strategies |
/queries |
POST |
Ask questions |
/benchmarks |
POST |
Run comparisons |
/benchmarks/{id}/report |
GET |
View results |
/experiments |
GET |
List experiments |
Configuration
All settings via .env file. See Configuration Guide for details.
Key settings:
OPENAI_API_KEY - Required for embeddings and LLM
CHUNK_SIZE - Target tokens per chunk (default: 512)
SEMANTIC_THRESHOLD - Similarity threshold (default: 0.5)
Further Reading
- Phases - Implementation roadmap
- Tasks - Task tracking
- ADRs - Architectural Decision Records