docs(rag): add domain model and original research input

This commit is contained in:
2026-07-21 10:10:06 +03:30
parent 6c9cebaaa0
commit 642aec93f2
15 changed files with 1147 additions and 0 deletions

View File

@@ -0,0 +1,52 @@
# ADR 0001: Single Collection with Metadata Filtering
## Status
Accepted
## Context
The current architecture uses ~30 collections in ChromaDB, each representing an insurance domain. The system must classify incoming queries to a single collection before retrieval. Misclassification leads to irrelevant results with no recovery path.
The current hybrid classifier (keyword-first cascade with similarity + LLM) has:
- No measurable accuracy metrics
- No visibility into keyword coverage
- Degraded performance on ambiguous queries
- Similarity score domination when signals conflict
Insurance domains are semantically similar — many questions could plausibly belong to multiple domains (e.g., "claim denial appeal" could be car, health, fire, etc.).
## Decision
Migrate to a **single collection** in Qdrant with rich metadata on each chunk:
- `domain`: car | health | fire | life | travel | ...
- `document_type`: policy | claim | faq | terms | ...
- `language`: fa | en
- `source_file`: filename
- `chunk_index`: position in source document
Queries will use metadata filtering to narrow search space when domain is known, or search the entire collection when ambiguous.
## Consequences
### Positive
- **Eliminates classification failure mode** — no single point of failure at the routing stage
- **Graceful degradation** — ambiguous queries search broader, still retrieve relevant results
- **Simpler pipeline** — one collection to manage, one index to optimize
- **Qdrant excels at filtered search** — metadata filtering happens before vector search, maintaining performance
- **Easier to add new domains** — just add documents with new domain metadata, no collection creation
### Negative
- **Larger index to search** — all domains in one collection
- **Requires disciplined metadata** — incorrect or missing metadata will cause retrieval failures
- **Potential for cross-domain noise** — irrelevant domains could appear in results if filtering isn't precise
### Neutral
- Reranking becomes more important to sort results across domains
- Embedding quality becomes more critical — must distinguish semantically similar domains
## Alternatives Considered
1. **Keep 30 collections + improve classifier** — adds metrics but retains single point of failure
2. **Hybrid: few collections (3-5) + metadata** — reduces but doesn't eliminate classification problem
3. **Multi-collection search with reranking** — more robust but slower and more complex