From 0488a7e0d06d6fccb29d462715edd598418cce53 Mon Sep 17 00:00:00 2001 From: Mahdi Bazrafshan Date: Tue, 21 Jul 2026 10:14:42 +0330 Subject: [PATCH] docs(rag): add top-level readme for chunking directory --- chunking/README.md | 92 ++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 92 insertions(+) create mode 100644 chunking/README.md diff --git a/chunking/README.md b/chunking/README.md new file mode 100644 index 0000000..cf2e8ee --- /dev/null +++ b/chunking/README.md @@ -0,0 +1,92 @@ +# Insurance RAG System — Architecture & Design + +Architecture decision records (ADRs) and domain model for migrating an insurance company's RAG chatbot from ChromaDB to Qdrant. + +## Problem + +The current system uses ~30 ChromaDB collections, one per insurance domain (car, health, fire, life, travel...). A collection classifier routes each query to exactly one collection before retrieval — but insurance domains are semantically similar, and misclassification produces irrelevant results with no recovery path. The classifier has no measurable accuracy, and its blind spots are unknown. + +The goal is to redesign the retrieval architecture so the system can handle PDFs, Word docs, CSVs, Excel files, and text — not just CSV Q&A pairs — while remaining fully operational during internet blackouts of up to 90 days. + +## Constraints + +| Constraint | Implication | +|------------|-------------| +| Internet blackouts (up to 90 days) | All components must have local fallbacks. Cloud APIs are optional enhancements, not dependencies. | +| Bilingual content (Persian + English) | Embedding models, rerankers, and LLMs must handle both languages. | +| Enterprise-scale, on-premises | No reliance on managed cloud services. Full data sovereignty. | + +## Architecture at a Glance + +``` +User Query + → Embed (OpenAI online / Nomic offline) + → Hybrid Search (dense + BM25 sparse) in Qdrant + → Rerank (Cohere online / BGE offline) + → Top-N chunks + metadata + → LLM (Qwen 2.5 local / cloud) + → Answer + citations +``` + +**Core decisions:** + +| Area | Choice | Why | +|------|--------|-----| +| Collection strategy | Single collection + metadata filtering | Eliminates the classification failure mode entirely | +| Chunking | Hybrid: semantic for unstructured, row-by-row for tabular | Markdown-first pipeline, coherent chunks | +| Embeddings | Dual: OpenAI `text-embedding-3-large` + Nomic `nomic-embed-text` | Seamless online/offline switching at query time | +| Reranking | Dual: Cohere Rerank + BGE Reranker v2-m3 | Cross-encoder precision, multilingual support | +| Vector DB | Qdrant (Docker, on-prem) | Filtered search, named vectors, sparse support | +| Search | Dense + sparse (BM25) with RRF fusion | Exact term matching + semantic similarity | +| Inference | Ollama + Qwen 2.5 on RTX 3090 | Local bilingual LLM, ~20GB VRAM fits one GPU | +| Observability | Langfuse (self-hosted) | Full request tracing, cost tracking, quality scoring | + +## Project Structure + +``` +chunking/ +├── README.md ← You are here +├── CONTEXT.md ← Domain model, glossary, constraints +├── raw/ +│ └── 001.md ← Original research request +└── docs/ + ├── README.md ← ADR index with architecture overview + └── adr/ + ├── 0001-single-collection-with-metadata.md + ├── 0002-chunking-strategy.md + ├── 0003-dual-embedding-strategy.md + ├── 0004-reranking-strategy.md + ├── 0005-retrieval-evaluation-framework.md + ├── 0006-qdrant-deployment.md + ├── 0007-hybrid-search.md + ├── 0008-metadata-schema.md + ├── 0009-ingestion-pipeline.md + ├── 0010-query-pipeline.md + ├── 0011-inference-infrastructure.md + └── 0012-observability-auth.md +``` + +## How to Read This + +1. **Start with `CONTEXT.md`** — understand the domain, the glossary, and why the current architecture fails. +2. **Browse `docs/README.md`** — get the full architecture overview and a table of all 12 decisions. +3. **Read individual ADRs** as needed — each follows the Context / Decision / Consequences format with alternatives considered. + +## Tech Stack + +| Component | Tool | Deployment | +|-----------|------|------------| +| Vector database | Qdrant | Docker, on-prem | +| Primary embeddings | OpenAI `text-embedding-3-large` | Cloud API | +| Offline embeddings | Nomic `nomic-embed-text` | Ollama / sentence-transformers | +| Primary reranker | Cohere Rerank | Cloud API | +| Offline reranker | BGE Reranker v2-m3 | Local PyTorch | +| LLM (offline) | Qwen 2.5 (14B) | Ollama on RTX 3090 | +| Hybrid search | BM25 sparse + dense vectors | Qdrant native | +| Observability | Langfuse | Self-hosted Docker | +| Ingestion | FastAPI + Celery + Redis | On-prem | +| Chunking | Markdown-first pipeline | Unstructured / pandoc | + +## License + +Internal research document. Not for distribution.