Files
chunking_strategies_evaluation/docs/strategy-technical-details.md
Mahdi Bazrafshan 754da323ff docs: add comprehensive project documentation
Why:
- Need documentation for team onboarding and reference
- Need technical details for strategy implementations
- Need API reference for developers

Changes:
- README.md: Documentation index and quick start guide
- api-reference.md: All 10 endpoints with examples
- architecture.md: System structure and design decisions
- configuration.md: All environment variables and parameters
- data-flow.md: How data moves through the system
- evaluation-metrics.md: How scoring works with weights
- strategy-technical-details.md: Deep dive into each strategy's implementation
2026-07-27 14:15:28 +03:30

8.9 KiB

Strategy Technical Details

Deep technical explanation of each chunking strategy's implementation.


1. Fixed-Size Chunking

File: src/chunking/strategies/fixed_size.py

How It Works

Splits text into fixed-size token chunks with overlap.

Splitting Unit

Token-level using tiktoken (cl100k_base encoding).

Algorithm

1. Encode full markdown → token array
2. Take first chunk_size tokens → chunk 1
3. Slide forward by (chunk_size - overlap) tokens
4. Repeat until end of text

Parameters

Parameter Config Key Default Description
Chunk size chunk_size 512 Tokens per chunk
Overlap chunk_overlap 50 Token overlap between chunks

Code Logic

tokens = _encoder.encode(text)
start = 0
while start < len(tokens):
    end = min(start + chunk_size, len(tokens))
    chunk_tokens = tokens[start:end]
    chunks.append(_encoder.decode(chunk_tokens))
    start = end - overlap

Characteristics

  • Splitting boundary: Arbitrary (mid-sentence possible)
  • Structure awareness: None
  • Speed: Fastest (no API calls)
  • Best for: Baseline comparison, simple documents

2. Recursive Chunking

File: src/chunking/strategies/recursive.py

How It Works

Cascade splitting using a separator hierarchy. Tries meaningful boundaries first, falls back to less meaningful ones.

Splitting Unit

Character-level with regex patterns.

Separator Cascade (Priority Order)

Priority Separator Regex Pattern Description
1 Header ^(#{1,6})\s+ Markdown headers
2 Double newline \n\n Paragraph breaks
3 Single newline \n Line breaks
4 Sentence [.!?]\s+ Sentence endings
5 Space \s+ Word boundaries

Algorithm

1. Try splitting by highest-priority separator
2. If parts still too big → recurse with next separator
3. Merge small parts back up to target size
4. If single part exceeds target → hard split by character count

Code Logic

for sep_name, pattern in _SEPARATORS:
    parts = pattern.split(text)
    if len(parts) <= 1:
        continue  # this separator didn't split
    
    # Merge parts back up to target_size
    current = ""
    for part in parts:
        candidate = (current + " " + part).strip()
        if len(candidate) <= target_size:
            current = candidate
        else:
            chunks.append(current)
            current = part
    
    return chunks  # Return on first successful split

Characteristics

  • Splitting boundary: Semantic (headers, paragraphs, sentences)
  • Structure awareness: High (respects markdown structure)
  • Speed: Fast (no API calls)
  • Best for: Structured documents with clear hierarchy

3. Semantic Chunking

File: src/chunking/strategies/semantic.py

How It Works

Groups sentences by semantic similarity. When similarity drops, starts a new chunk.

Splitting Unit

Sentence-level with similarity-based boundaries.

Algorithm

1. Split markdown into sentences
2. Embed each sentence via OpenAI (text-embedding-3-small)
3. Compute cosine similarity between adjacent sentences
4. When similarity < SEMANTIC_THRESHOLD → chunk boundary
5. Enforce SEMANTIC_MIN_CHUNK_SIZE (minimum sentences per chunk)

Parameters

Parameter Config Key Default Description
Threshold semantic_threshold 0.5 Similarity threshold for boundary
Min size semantic_min_chunk_size 5 Minimum sentences per chunk

Code Logic

sentences = split_sentences(markdown)
embeddings = embed_texts(sentences)  # OpenAI API call

current_group = [sentences[0]]
for i in range(1, len(sentences)):
    sim = cosine_similarity(embeddings[i-1], embeddings[i])
    
    if sim < threshold and len(current_group) >= min_size:
        # Topic shift → new chunk
        chunks.append(" ".join(current_group))
        current_group = [sentences[i]]
    else:
        current_group.append(sentences[i])

Characteristics

  • Splitting boundary: Semantic (topic changes)
  • Structure awareness: Medium (sentence-level)
  • Speed: Medium (requires embeddings)
  • Best for: Documents with topic shifts, no clear structure

4. Contextual Retrieval

File: src/chunking/strategies/contextual_retrieval.py

How It Works

Based on Anthropic's 2024 research. Prepends a short context summary to each chunk before embedding.

Splitting Unit

Token-level (same as fixed-size).

Algorithm

1. Split text using fixed-size token splitting
2. For each chunk:
   a. Find surrounding text (500 chars before/after)
   b. Send to LLM with context prompt
   c. LLM generates 1-2 sentence context prefix
   d. Enriched chunk = context + original text
3. Embed enriched chunks (not original)

LLM Prompt

System: You are a document analysis assistant. Given a section of text 
from a document, write a short context prefix (1-2 sentences) that 
would help someone find this section later via search.

User: 
Preceding text: {preceding_500_chars}
This section: {chunk_text}
Following text: {following_500_chars}

Code Logic

for chunk_text in raw_chunks:
    # Find surrounding context
    preceding = full_text[pos-500:pos]
    following = full_text[pos+len:pos+len+500]
    
    # Generate context prefix
    context = llm.generate(
        system=CONTEXT_PROMPT,
        user=f"Preceding: {preceding}\nSection: {chunk_text}\nFollowing: {following}"
    )
    
    # Enrich chunk
    enriched = f"{context}\n\n{chunk_text}"
    chunks.append(enriched)

Characteristics

  • Splitting boundary: Token-based (like fixed-size)
  • Structure awareness: None (relies on LLM for context)
  • Speed: Slowest (1 LLM call per chunk)
  • Best for: Improving retrieval quality, unstructured documents

5. Semantic Parent-Child

File: src/chunking/strategies/semantic_parent_child.py

How It Works

Groups paragraphs into semantic clusters. Each cluster is a parent; each paragraph is a child.

Splitting Unit

Paragraph-level with semantic clustering.

Algorithm

1. Split markdown into paragraphs
2. Embed each paragraph
3. Cluster consecutive paragraphs by similarity
4. Each cluster = parent chunk (full cluster text)
5. Each paragraph = child chunk (linked to parent)

Query-Time Behavior

1. Vector search finds matching child paragraph
2. Use parent_id to fetch parent cluster
3. Return both child + parent to LLM

Parameters

Parameter Config Key Default Description
Threshold semantic_threshold 0.5 Similarity threshold for clustering

Code Logic

paragraphs = split_paragraphs(markdown)
embeddings = embed_texts(paragraphs)

# Cluster paragraphs
clusters = [[0]]
for i in range(1, len(paragraphs)):
    sim = cosine_similarity(embeddings[i-1], embeddings[i])
    if sim >= threshold:
        clusters[-1].append(i)  # Same cluster
    else:
        clusters.append([i])    # New cluster

# Create parent-child chunks
for cluster in clusters:
    parent_text = "\n\n".join(paragraphs[i] for i in cluster)
    parent_id = make_chunk_id(...)
    
    for para_idx in cluster:
        chunks.append(build_chunk(
            text=paragraphs[para_idx],
            parent_id=parent_id
        ))

Characteristics

  • Splitting boundary: Semantic (paragraph clusters)
  • Structure awareness: Medium (paragraph-level)
  • Speed: Medium (requires embeddings)
  • Best for: Documents needing context, flat structure

Comparison Table

Strategy Split Unit Boundary Logic API Calls Speed Best For
fixed_size Token Arbitrary 0 Fast Baseline
recursive Character Structural 0 Fast Structured docs
semantic Sentence Similarity 1 (embeddings) Medium Topic shifts
contextual_retrieval Token LLM-generated N (1 per chunk) Slow Retrieval quality
semantic_parent_child Paragraph Similarity 1 (embeddings) Medium Context needed

Configuration Reference

All strategy parameters are in src/core/config.py:

class Settings(BaseSettings):
    chunk_size: int = 512
    chunk_overlap: int = 50
    semantic_threshold: float = 0.5
    semantic_min_chunk_size: int = 5
    embedding_model: str = "text-embedding-3-small"
    llm_model: str = "gpt-4o-mini"

ADR References

ADR Strategy Decision
ADR 0012 semantic Sentence-level with min chunk size
ADR 0015 recursive Direct API, no LangChain/LlamaIndex
ADR 0011 contextual_retrieval Embed enriched text, not original
ADR 0003 all Per-strategy failure isolation