Why: - Need documentation for team onboarding and reference - Need technical details for strategy implementations - Need API reference for developers Changes: - README.md: Documentation index and quick start guide - api-reference.md: All 10 endpoints with examples - architecture.md: System structure and design decisions - configuration.md: All environment variables and parameters - data-flow.md: How data moves through the system - evaluation-metrics.md: How scoring works with weights - strategy-technical-details.md: Deep dive into each strategy's implementation
8.6 KiB
8.6 KiB
Data Flow
How data moves through the RAG Chunking Benchmarker system.
Overview
┌─────────────────────────────────────────────────────────────────┐
│ DATA FLOW │
├─────────────────────────────────────────────────────────────────┤
│ │
│ INGESTION │
│ ───────── │
│ .docx/.doc → Parser → DocumentTree + Markdown → SQLite │
│ │
│ PROCESSING │
│ ────────── │
│ Document → Strategy.chunk() → Chunks → Embed → Qdrant │
│ │
│ QUERYING │
│ ───────── │
│ Question → Embed → Qdrant Search → Chunks → LLM → Answer │
│ │
│ BENCHMARKING │
│ ─────────── │
│ Questions → (Query × Strategies) → Evaluation → Report │
│ │
└─────────────────────────────────────────────────────────────────┘
1. Document Ingestion
Flow
User uploads .docx/.doc
↓
FastAPI receives file
↓
Parser extracts text blocks
↓
Builds DocumentTree (hierarchy)
↓
Generates markdown (flat text)
↓
Stores in SQLite
Data Structures
Input: Binary file (.docx/.doc)
Output:
ParseResult(
tree=DocumentTree, # Hierarchical structure
markdown=str, # Flat text for chunking
plain_text=str, # Raw text
paragraph_count=int # Number of text blocks
)
Storage
SQLite:
INSERT INTO documents (id, filename, parsed_text, document_tree, created_at)
VALUES (?, ?, ?, ?, ?)
2. Document Processing
Flow
User selects strategies
↓
For each strategy:
↓
Load DocumentTree + markdown from SQLite
↓
Strategy.chunk() → list[Chunk]
↓
Embed chunks via OpenAI
↓
Store in Qdrant collection
↓
Update chunk_counts in SQLite
Data Structures
Input:
{
"document_id": "doc-abc123",
"strategies": ["recursive", "fixed_size", "semantic"]
}
Processing:
# For each strategy
chunks = strategy.chunk(
doc_name="insurance.docx",
tree=document_tree,
markdown=markdown_text
)
embeddings = embed_texts([chunk.text for chunk in chunks])
qdr.upsert_chunks(chunks, embeddings)
Output:
{
"strategies_completed": [
{"strategy": "recursive", "chunks_produced": 218}
]
}
Storage
Qdrant:
{
"id": "uuid5(chunk_id)",
"vector": [0.023, -0.156, ...], # 1536 dimensions
"payload": {
"document_name": "insurance.docx",
"chunk_id": "recursive_doc_000045",
"strategy_name": "recursive",
"text": "chunk content...",
"token_count": 146,
"character_count": 205
}
}
3. Query Pipeline
Flow
User asks question
↓
Embed question via OpenAI
↓
Search Qdrant for similar chunks
↓
Retrieve top-K chunks
↓
Build context from chunks
↓
Generate answer via LLM
↓
Store query result in SQLite
Data Structures
Input:
{
"document_id": "doc-abc123",
"strategy": "recursive",
"question": "What are the main topics?",
"top_k": 5
}
Processing:
# 1. Embed question
question_embedding = embed_single(question)
# 2. Vector search
hits = qdr.search(
strategy="recursive",
query_vector=question_embedding,
top_k=5
)
# 3. Build context
context = "\n".join([hit.text for hit in hits])
# 4. Generate answer
answer = llm.generate(
system="Answer based on context...",
user=f"Context: {context}\nQuestion: {question}"
)
Output:
{
"answer": "The document covers insurance regulations...",
"retrieved_chunks": [...],
"latency_breakdown": {...},
"token_usage": {...}
}
Storage
SQLite:
INSERT INTO queries (id, document_id, strategy_name, question, answer,
retrieved_chunks, latency_breakdown, token_usage)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
4. Benchmark Pipeline
Flow
Load questions from JSON
↓
For each question × strategy:
↓
Run query pipeline
↓
Evaluate with LLM-as-Judge
↓
Store per-question results
↓
Aggregate metrics per strategy
↓
Store experiment in SQLite
↓
Generate HTML report
Data Structures
Input:
{
"document_id": "doc-abc123",
"strategies": ["recursive", "fixed_size"],
"questions_file": "files/questions.json"
}
Processing:
# For each question
for question in questions:
for strategy in strategies:
# Query
result = run_query(question, strategy)
# Evaluate
scores = evaluate_single(
question=question,
context=result.retrieved_chunks,
expected=question.expected_answer,
generated=result.answer
)
# Store
per_question_results.append({...})
Output:
{
"experiment_id": "exp-abc123",
"aggregate_metrics": {
"recursive": {
"avg_context_relevance": 8.5,
"avg_answer_similarity": 7.8,
"avg_faithfulness": 9.2,
"hallucination_rate": 0.05
}
},
"best_strategy": "recursive"
}
5. Report Generation
Flow
Load experiment from SQLite
↓
Calculate rankings
↓
Generate HTML (managerial or technical view)
↓
Return HTML response
Views
| View | Focus | Content |
|---|---|---|
| Managerial | Decision-making | KPIs, winner, recommendations |
| Technical | Observability | Detailed scores, token usage, per-question |
Data Persistence
SQLite Tables
| Table | Purpose | Key Fields |
|---|---|---|
| documents | Uploaded documents | id, filename, parsed_text, document_tree |
| queries | Query results | id, document_id, question, answer |
| experiments | Benchmark results | id, document_id, aggregate_metrics |
Qdrant Collections
| Collection | Purpose | Fields |
|---|---|---|
| recursive_collection | Recursive strategy vectors | chunk_id, text, scores |
| fixed_size_collection | Fixed-size strategy vectors | chunk_id, text, scores |
| semantic_collection | Semantic strategy vectors | chunk_id, text, scores |
| contextual_retrieval_collection | Contextual strategy vectors | chunk_id, text, scores |
| semantic_parent_child_collection | Parent-child strategy vectors | chunk_id, text, parent_id |
Data Transformation
Document → Chunks
Input: Full document text (10,000 tokens)
Output: 50-300 chunks (100-500 tokens each)
Transformation:
- Tokenization
- Boundary detection
- Metadata enrichment
Chunks → Vectors
Input: Chunk text
Output: 1536-dimensional vector
Transformation:
- Text → Tokens → Embedding API → Vector
Question → Answer
Input: Question string
Output: Answer string + metadata
Transformation:
- Question → Embedding → Vector Search → Chunks → LLM → Answer
Performance Characteristics
| Operation | Time | Tokens | Cost |
|---|---|---|---|
| Document upload | ~2s | 0 | Free |
| Process (1 strategy) | ~5s | N/A | ~$0.001 (embeddings) |
| Process (5 strategies) | ~25s | N/A | ~$0.005 |
| Query | ~2s | ~1500 | ~$0.001 |
| Benchmark (21 questions) | ~120s | ~200,000 | ~$0.05 |