docs(rag): add domain model and original research input
This commit is contained in:
62
chunking/docs/adr/0002-chunking-strategy.md
Normal file
62
chunking/docs/adr/0002-chunking-strategy.md
Normal file
@@ -0,0 +1,62 @@
|
||||
# ADR 0002: Chunking Strategy
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The current system chunks CSV files row-by-row, treating each Q&A pair as a single chunk. The migration expands support to PDF, Word, Excel, and text files with predominantly unstructured content (FAQs, guidance, mixed text).
|
||||
|
||||
Chunk quality directly impacts retrieval precision. Arbitrary splits (mid-sentence, mid-thought) produce incoherent chunks that degrade both embedding quality and answer generation.
|
||||
|
||||
## Decision
|
||||
|
||||
Adopt a **hybrid chunking strategy** with markdown as the intermediate format:
|
||||
|
||||
### 1. Markdown-First Pipeline
|
||||
All documents are converted to markdown before chunking. This provides:
|
||||
- Consistent processing regardless of source format
|
||||
- Preserved structure (headers, lists, tables)
|
||||
- Easier debugging and inspection
|
||||
|
||||
### 2. Type-Specific Chunking
|
||||
|
||||
| Source Type | Chunking Method | Rationale |
|
||||
|-------------|-----------------|-----------|
|
||||
| CSV | Row-by-row | Already natural Q&A pairs |
|
||||
| Excel | Row-by-row | Preserves tabular context per record |
|
||||
| PDF | Convert to markdown → semantic chunking | Unstructured prose needs coherent boundaries |
|
||||
| Word (.docx) | Convert to markdown → semantic chunking | Same as PDF |
|
||||
| Text files | Convert to markdown → semantic chunking | Same as PDF |
|
||||
| Structured documents (policies with headers) | Markdown section-aware chunking | Headers define natural boundaries |
|
||||
|
||||
### 3. Semantic Chunking Parameters
|
||||
- **Unit:** Sentences (not tokens)
|
||||
- **Grouping:** Consecutive sentences grouped until semantic similarity drops below threshold
|
||||
- **Overlap:** 1-2 sentences at boundaries to preserve context
|
||||
- **Max chunk size:** Determined by embedding model limits (typically 512-8192 tokens depending on model)
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
- **Coherent chunks** — complete thoughts, no mid-sentence splits
|
||||
- **Consistent pipeline** — one chunking logic for all unstructured content
|
||||
- **Debuggable** — markdown is human-readable, easy to inspect chunk quality
|
||||
- **Flexible** — can adjust semantic threshold per document type if needed
|
||||
|
||||
### Negative
|
||||
- **Processing overhead** — conversion to markdown adds a step
|
||||
- **Variable chunk sizes** — harder to predict storage and latency
|
||||
- **Semantic chunking requires embedding model** — adds computation during ingestion
|
||||
|
||||
### Neutral
|
||||
- Need to select semantic similarity threshold (tuning required)
|
||||
- Need to handle edge cases: very long sentences, tables in markdown, code blocks
|
||||
|
||||
## Implementation Notes
|
||||
|
||||
- Use `unstructured` or `pandoc` for document → markdown conversion
|
||||
- Use LlamaIndex's `SemanticSplitter` or custom implementation with embedding model
|
||||
- Consider `MarkdownElementNodeParser` for section-aware chunking when headers are present
|
||||
- Store original source file path and chunk index in metadata for citation
|
||||
Reference in New Issue
Block a user