docs(rag): add domain model and original research input

This commit is contained in:
2026-07-21 10:10:06 +03:30
parent 6c9cebaaa0
commit 642aec93f2
15 changed files with 1147 additions and 0 deletions

View File

@@ -0,0 +1,62 @@
# ADR 0002: Chunking Strategy
## Status
Accepted
## Context
The current system chunks CSV files row-by-row, treating each Q&A pair as a single chunk. The migration expands support to PDF, Word, Excel, and text files with predominantly unstructured content (FAQs, guidance, mixed text).
Chunk quality directly impacts retrieval precision. Arbitrary splits (mid-sentence, mid-thought) produce incoherent chunks that degrade both embedding quality and answer generation.
## Decision
Adopt a **hybrid chunking strategy** with markdown as the intermediate format:
### 1. Markdown-First Pipeline
All documents are converted to markdown before chunking. This provides:
- Consistent processing regardless of source format
- Preserved structure (headers, lists, tables)
- Easier debugging and inspection
### 2. Type-Specific Chunking
| Source Type | Chunking Method | Rationale |
|-------------|-----------------|-----------|
| CSV | Row-by-row | Already natural Q&A pairs |
| Excel | Row-by-row | Preserves tabular context per record |
| PDF | Convert to markdown → semantic chunking | Unstructured prose needs coherent boundaries |
| Word (.docx) | Convert to markdown → semantic chunking | Same as PDF |
| Text files | Convert to markdown → semantic chunking | Same as PDF |
| Structured documents (policies with headers) | Markdown section-aware chunking | Headers define natural boundaries |
### 3. Semantic Chunking Parameters
- **Unit:** Sentences (not tokens)
- **Grouping:** Consecutive sentences grouped until semantic similarity drops below threshold
- **Overlap:** 1-2 sentences at boundaries to preserve context
- **Max chunk size:** Determined by embedding model limits (typically 512-8192 tokens depending on model)
## Consequences
### Positive
- **Coherent chunks** — complete thoughts, no mid-sentence splits
- **Consistent pipeline** — one chunking logic for all unstructured content
- **Debuggable** — markdown is human-readable, easy to inspect chunk quality
- **Flexible** — can adjust semantic threshold per document type if needed
### Negative
- **Processing overhead** — conversion to markdown adds a step
- **Variable chunk sizes** — harder to predict storage and latency
- **Semantic chunking requires embedding model** — adds computation during ingestion
### Neutral
- Need to select semantic similarity threshold (tuning required)
- Need to handle edge cases: very long sentences, tables in markdown, code blocks
## Implementation Notes
- Use `unstructured` or `pandoc` for document → markdown conversion
- Use LlamaIndex's `SemanticSplitter` or custom implementation with embedding model
- Consider `MarkdownElementNodeParser` for section-aware chunking when headers are present
- Store original source file path and chunk index in metadata for citation