3.2 KiB
3.2 KiB
ADR 0010: Query Pipeline
Status
Accepted
Context
The query pipeline transforms user questions into answers using the RAG architecture. It must handle:
- Dual embedding models (online/offline)
- Hybrid search (dense + sparse)
- Dual reranking (online/offline)
- Bilingual generation (Persian/English)
- Citation and traceability for admin monitoring
Decision
Architecture
User Query
↓
Query Embedder (OpenAI or Nomic)
↓
Hybrid Search (dense + sparse vectors)
↓
Reranker (Cohere or BGE)
↓
Top-N Chunks (N=5-10)
↓
Context Builder (chunks + metadata + history)
↓
LLM Generator (Qwen 2.5 or cloud LLM)
↓
Answer + Citations + Metadata
Context for LLM
The LLM receives:
- Retrieved chunks — text content of top-N results
- Metadata — domain, source_file, document_type for each chunk
- Conversation history — previous turns for multi-turn context
Prompt template:
You are an insurance assistant. Answer the user's question based on the provided context.
Context:
[Domain: {domain}] [Source: {source_file}]
{chunk_text}
...
Conversation history:
{history}
Question: {query}
Answer the question using only the provided context. Cite sources when relevant.
Admin Layer Metadata
API response includes:
{
"answer": "Your car insurance policy covers...",
"citations": [
{"source_file": "car_faq.pdf", "chunk_index": 5, "domain": "car"}
],
"retrieval_metadata": {
"embedding_model": "nomic",
"reranker": "bge",
"latency_ms": 450,
"total_chunks_retrieved": 20,
"chunks_after_rerank": 5
},
"confidence": "high"
}
Query-Time Model Selection
def get_query_pipeline():
if is_internet_available():
return {
"embedder": "openai",
"reranker": "cohere",
"llm": "cloud" # or Qwen if preferred
}
else:
return {
"embedder": "nomic",
"reranker": "bge",
"llm": "qwen"
}
Citations in Response
LLM is instructed to cite sources:
- "According to
car_faq.pdf, your policy covers..." - "As stated in the
terms_and_conditions.pdfdocument..."
Citations use source_file metadata, enabling admins to trace answers to exact documents.
Consequences
Positive
- Full context — LLM has all information needed for accurate answers
- Traceability — admin layer can audit which files informed each answer
- Multi-turn support — conversation history enables follow-up questions
- Model flexibility — automatic fallback during offline periods
Negative
- Token usage — full context increases generation cost
- Latency — more context means longer LLM inference time
- Context window limits — may need to truncate for very long conversations
Neutral
- Need to implement conversation history storage
- May need to tune how much history to include
Implementation Notes
- Use LlamaIndex or LangChain for pipeline orchestration
- Store conversation history in Redis or database
- Implement rate limiting on query endpoint
- Log all retrieval metadata for analysis
- Consider streaming responses for better UX