Files
chunking_strategies_evaluation/docs/configuration.md
Mahdi Bazrafshan 754da323ff docs: add comprehensive project documentation
Why:
- Need documentation for team onboarding and reference
- Need technical details for strategy implementations
- Need API reference for developers

Changes:
- README.md: Documentation index and quick start guide
- api-reference.md: All 10 endpoints with examples
- architecture.md: System structure and design decisions
- configuration.md: All environment variables and parameters
- data-flow.md: How data moves through the system
- evaluation-metrics.md: How scoring works with weights
- strategy-technical-details.md: Deep dive into each strategy's implementation
2026-07-27 14:15:28 +03:30

4.1 KiB

Configuration Guide

All settings and environment variables for the RAG Chunking Benchmarker.


Environment Variables

Create a .env file in the project root:

# ── Database ──────────────────────────────────────────────────────

# SQLite database path
DATABASE_URL=sqlite:///./data/chunking_benchmark.db

# Qdrant vector database
QDRANT_URL=http://localhost:6333
QDRANT_API_KEY=

# ── OpenAI ────────────────────────────────────────────────────────

# API key for embeddings and LLM
OPENAI_API_KEY=sk-your-key-here

# ── Chunking Parameters ───────────────────────────────────────────

# Fixed-size strategy
CHUNK_SIZE=512
CHUNK_OVERLAP=50

# Semantic strategy
SEMANTIC_THRESHOLD=0.5
SEMANTIC_MIN_CHUNK_SIZE=5

# ── Models ────────────────────────────────────────────────────────

# Embedding model (used for all strategies)
EMBEDDING_MODEL=text-embedding-3-small

# LLM model (used for contextual retrieval and evaluation)
LLM_MODEL=gpt-4o-mini

Parameter Details

Chunking Parameters

Parameter Type Default Range Description
CHUNK_SIZE int 512 100-2000 Target tokens per chunk
CHUNK_OVERLAP int 50 0-200 Token overlap between chunks
SEMANTIC_THRESHOLD float 0.5 0.0-1.0 Cosine similarity threshold for chunk boundaries
SEMANTIC_MIN_CHUNK_SIZE int 5 1-20 Minimum sentences per semantic chunk

Model Parameters

Parameter Type Default Options Description
EMBEDDING_MODEL str text-embedding-3-small text-embedding-3-small, text-embedding-3-large OpenAI embedding model
LLM_MODEL str gpt-4o-mini gpt-4o-mini, gpt-4o Model for contextual retrieval and evaluation

Default Values in Code

# src/core/config.py
class Settings(BaseSettings):
    # Database
    database_url: str = "sqlite:///./data/chunking_benchmark.db"
    qdrant_url: str = "http://localhost:6333"
    qdrant_api_key: str = ""
    
    # OpenAI
    openai_api_key: str = ""
    
    # Chunking
    chunk_size: int = 512
    chunk_overlap: int = 50
    semantic_threshold: float = 0.5
    semantic_min_chunk_size: int = 5
    
    # Models
    embedding_model: str = "text-embedding-3-small"
    llm_model: str = "gpt-4o-mini"

Effect of Parameters

chunk_size

Value Chunks Quality Speed
256 More Higher context loss Faster
512 Balanced Balanced Balanced
1024 Fewer Lower context loss Slower

chunk_overlap

Value Context Preservation Redundancy
0 None None
50 Moderate Low
100 High High

semantic_threshold

Value Chunk Size Topic Sensitivity
0.3 Larger Less sensitive
0.5 Balanced Balanced
0.7 Smaller More sensitive

Environment-Specific Configs

Development

DATABASE_URL=sqlite:///./data/dev.db
QDRANT_URL=http://localhost:6333
LOG_LEVEL=DEBUG

Production

DATABASE_URL=postgresql://user:pass@localhost/dbname
QDRANT_URL=http://qdrant:6333
QDRANT_API_KEY=your-api-key
LOG_LEVEL=WARNING

Testing

DATABASE_URL=sqlite:///./data/test.db
QDRANT_URL=http://localhost:6333
CHUNK_SIZE=100

Validation

Settings are validated on startup. Invalid values will cause:

pydantic.ValidationError: 1 validation error for Settings

Common errors:

Error Cause Fix
Missing OPENAI_API_KEY No API key Add to .env
Invalid QDRANT_URL Wrong URL format Check URL
CHUNK_SIZE < 1 Too small Increase value