Files
chunking_strategies_evaluation/docs/configuration.md
Mahdi Bazrafshan 754da323ff docs: add comprehensive project documentation
Why:
- Need documentation for team onboarding and reference
- Need technical details for strategy implementations
- Need API reference for developers

Changes:
- README.md: Documentation index and quick start guide
- api-reference.md: All 10 endpoints with examples
- architecture.md: System structure and design decisions
- configuration.md: All environment variables and parameters
- data-flow.md: How data moves through the system
- evaluation-metrics.md: How scoring works with weights
- strategy-technical-details.md: Deep dive into each strategy's implementation
2026-07-27 14:15:28 +03:30

165 lines
4.1 KiB
Markdown

# Configuration Guide
All settings and environment variables for the RAG Chunking Benchmarker.
---
## Environment Variables
Create a `.env` file in the project root:
```env
# ── Database ──────────────────────────────────────────────────────
# SQLite database path
DATABASE_URL=sqlite:///./data/chunking_benchmark.db
# Qdrant vector database
QDRANT_URL=http://localhost:6333
QDRANT_API_KEY=
# ── OpenAI ────────────────────────────────────────────────────────
# API key for embeddings and LLM
OPENAI_API_KEY=sk-your-key-here
# ── Chunking Parameters ───────────────────────────────────────────
# Fixed-size strategy
CHUNK_SIZE=512
CHUNK_OVERLAP=50
# Semantic strategy
SEMANTIC_THRESHOLD=0.5
SEMANTIC_MIN_CHUNK_SIZE=5
# ── Models ────────────────────────────────────────────────────────
# Embedding model (used for all strategies)
EMBEDDING_MODEL=text-embedding-3-small
# LLM model (used for contextual retrieval and evaluation)
LLM_MODEL=gpt-4o-mini
```
---
## Parameter Details
### Chunking Parameters
| Parameter | Type | Default | Range | Description |
|-----------|------|---------|-------|-------------|
| `CHUNK_SIZE` | int | 512 | 100-2000 | Target tokens per chunk |
| `CHUNK_OVERLAP` | int | 50 | 0-200 | Token overlap between chunks |
| `SEMANTIC_THRESHOLD` | float | 0.5 | 0.0-1.0 | Cosine similarity threshold for chunk boundaries |
| `SEMANTIC_MIN_CHUNK_SIZE` | int | 5 | 1-20 | Minimum sentences per semantic chunk |
### Model Parameters
| Parameter | Type | Default | Options | Description |
|-----------|------|---------|---------|-------------|
| `EMBEDDING_MODEL` | str | text-embedding-3-small | text-embedding-3-small, text-embedding-3-large | OpenAI embedding model |
| `LLM_MODEL` | str | gpt-4o-mini | gpt-4o-mini, gpt-4o | Model for contextual retrieval and evaluation |
---
## Default Values in Code
```python
# src/core/config.py
class Settings(BaseSettings):
# Database
database_url: str = "sqlite:///./data/chunking_benchmark.db"
qdrant_url: str = "http://localhost:6333"
qdrant_api_key: str = ""
# OpenAI
openai_api_key: str = ""
# Chunking
chunk_size: int = 512
chunk_overlap: int = 50
semantic_threshold: float = 0.5
semantic_min_chunk_size: int = 5
# Models
embedding_model: str = "text-embedding-3-small"
llm_model: str = "gpt-4o-mini"
```
---
## Effect of Parameters
### chunk_size
| Value | Chunks | Quality | Speed |
|-------|--------|---------|-------|
| 256 | More | Higher context loss | Faster |
| 512 | Balanced | Balanced | Balanced |
| 1024 | Fewer | Lower context loss | Slower |
### chunk_overlap
| Value | Context Preservation | Redundancy |
|-------|---------------------|------------|
| 0 | None | None |
| 50 | Moderate | Low |
| 100 | High | High |
### semantic_threshold
| Value | Chunk Size | Topic Sensitivity |
|-------|------------|-------------------|
| 0.3 | Larger | Less sensitive |
| 0.5 | Balanced | Balanced |
| 0.7 | Smaller | More sensitive |
---
## Environment-Specific Configs
### Development
```env
DATABASE_URL=sqlite:///./data/dev.db
QDRANT_URL=http://localhost:6333
LOG_LEVEL=DEBUG
```
### Production
```env
DATABASE_URL=postgresql://user:pass@localhost/dbname
QDRANT_URL=http://qdrant:6333
QDRANT_API_KEY=your-api-key
LOG_LEVEL=WARNING
```
### Testing
```env
DATABASE_URL=sqlite:///./data/test.db
QDRANT_URL=http://localhost:6333
CHUNK_SIZE=100
```
---
## Validation
Settings are validated on startup. Invalid values will cause:
```
pydantic.ValidationError: 1 validation error for Settings
```
Common errors:
| Error | Cause | Fix |
|-------|-------|-----|
| Missing OPENAI_API_KEY | No API key | Add to .env |
| Invalid QDRANT_URL | Wrong URL format | Check URL |
| CHUNK_SIZE < 1 | Too small | Increase value |