--- last_updated: 2026-07-20 tags: [embedding, decision-report, farsi, insurance, rag, jira, recommendations] source: import-knowledge --- # Embedding Model Decision Report ## Context We are building an insurance domain RAG pipeline using LangGraph + Qdrant. The only language is Farsi (Persian). We need embedding models that produce high-quality Farsi retrieval (top-K ranking of correct insurance document chunks). Current state: cloud = OpenAI text-embedding-3-large, local = Nomic Embed v1.5 via Ollama. Hardware: 1× NVIDIA RTX 3090 (24 GB VRAM). **Critical discovery**: Nomic Embed v1.5 is English-only and does NOT support Farsi. This must be replaced. --- ## Decision 1: Local Embedding Model ### Option 1 — BAAI/bge-m3 ⭐ RECOMMENDED - **Ranking**: 1st - **Rationale**: Triple-mode retrieval (dense + sparse + ColBERT) from a single model, native support for 100+ languages including Farsi, only 1.1 GB VRAM on a 24 GB card, MIT license, and available on Ollama/TEI. Sparse vector support enables hybrid search in Qdrant without a separate sparse encoder — directly matches our LangGraph + Qdrant architecture. - **Pros**: - Supports Farsi (100+ languages) - Triple-mode: dense + sparse + ColBERT multi-vector - Hybrid search ready — no separate sparse encoder needed - 8,192 token context (can handle long insurance policy chunks) - Only 1.1 GB VRAM — leaves headroom for other workloads - MIT license — no commercial restrictions - Available on Ollama, TEI, sentence-transformers - Active development by BAAI (Beijing Academy of AI) - **Cons**: - No published Farsi-specific retrieval benchmark scores - Slightly less community adoption than multilingual-E5 - Triple-mode adds architectural complexity if only dense mode is used - **Verdict**: Best overall fit. Triple-mode retrieval is uniquely suited to Qdrant hybrid search. The sparse mode is a built-in BM25 alternative. Fits easily on RTX 3090 with room to spare. ### Option 2 — Alibaba-NLP/gte-multilingual-base - **Ranking**: 2nd - **Rationale**: Lightest model with confirmed Farsi support (610 MB VRAM, 70+ languages), Apache-2.0 license, 8,192 token context. Good choice if BGE-M3 proves too complex or if minimal resource footprint is the priority. - **Pros**: - Explicitly lists Farsi in supported languages - Only 610 MB VRAM — lightest meaningful option - 8,192 token context - Benchmarked on MIRACL and PL-MTEB - Apache-2.0 license - Simpler architecture (dense-only) — easier to deploy and debug - **Cons**: - No ColBERT/sparse support — dense-only retrieval - No published isolated Farsi retrieval scores - Alibaba NLP ecosystem less mature than HuggingFace ecosystem - **Verdict**: Best lightweight alternative. Choose over BGE-M3 if simplicity matters more than hybrid search capability. ### Option 3 — intfloat/multilingual-e5-large - **Ranking**: 3rd - **Rationale**: Proven multilingual baseline used in production worldwide. 560M parameters, 100 languages incl. Farsi, MIT license. - **Pros**: - Most widely used multilingual embedding model - Strong community support and documentation - 100 languages including Farsi - MIT license, available on Ollama - **Cons**: - 512 token max context — requires aggressive chunking for insurance documents - No ColBERT/sparse support - Requires `"query: "/"passage: "` prefixes (inference quirk) - **Verdict**: Safe fallback. The 512-token limit is a real constraint for insurance policy chunks. ### Option 4 — intfloat/e5-mistral-7b-instruct - **Ranking**: 4th - **Rationale**: Highest raw MTEB performance. LLM-based embedding with 4096 dimensions. Fits on RTX 3090 (14.2 GB) but tight. - **Cons**: 14.2 GB VRAM, Farsi support uncertain (English-only model card), 512 token limit. - **Verdict**: Only if empirical testing shows major gains over lighter models. ### Final Recommendation (Local) **BAAI/bge-m3** — deploy via Ollama or TEI. Triple-mode retrieval, 8K context, 1.1 GB VRAM, confirmed Farsi support. Fallback: gte-multilingual-base if simplicity is preferred. --- ## Decision 2: Cloud Embedding API ### Option 1 — Cohere embed-v4.0 ⭐ RECOMMENDED (if budget allows) - **Ranking**: 1st - **Rationale**: Strongest Farsi support (100+ languages), 512K context (embed full documents), MIRACL-Farsi scores. Enterprise pricing. - **Pros**: 512K context, strong MIRACL scores, configurable dimensions, v2 API with input_type - **Cons**: Enterprise-only pricing (~$4–5/hr), no public per-token pricing - **Verdict**: Best quality for Farsi if budget permits. ### Option 2 — Voyage AI voyage-4-large - **Ranking**: 2nd - **Rationale**: $0.12/1M tokens, 32K context, 200M free tokens. Best cost/quality balance. - **Pros**: 200M tokens free, 4× OpenAI context, competitive pricing - **Cons**: No Farsi-specific scores, smaller company - **Verdict**: Risk-free to test. Best starting point. ### Option 3 — OpenAI text-embedding-3-large *(current baseline)* - **Ranking**: 3rd - **Rationale**: Already integrated, proven at scale. But 8K context, no Farsi optimization, data leaves company. - **Verdict**: Keep as fallback. ### Option 4 — Jina AI jina-embeddings-v3 - **Ranking**: 4th - **Rationale**: Cheapest (~$0.02/1M tokens), 89 langs incl. Farsi. But non-commercial local license, small company. - **Verdict**: Budget option. Verify licensing. ### Final Recommendation (Cloud) **Cohere embed-v4.0** if budget allows. Otherwise **Voyage voyage-4-large** with 200M free tokens. Keep **OpenAI as fallback**. On-premises concern means local BGE-M3 may eventually replace cloud for sensitive data. --- ## Decision 3: Local Serving Platform ### Option 1 — Ollama ⭐ RECOMMENDED - **Ranking**: 1st - **Rationale**: Already deployed, BGE-M3 available. Zero migration effort. - **Pros**: No new infrastructure, built-in model management, OpenAI-compatible API - **Cons**: Not all models available, lower throughput than TEI - **Verdict**: Default choice. Fastest path from current state. ### Option 2 — HuggingFace TEI - **Ranking**: 2nd - **Rationale**: Purpose-built for embeddings, best throughput, Docker-based. - **Pros**: Highest throughput, native HF support, OpenAI-compatible API - **Cons**: New service to deploy, requires Docker - **Verdict**: Best for production-grade throughput (>100 queries/sec). ### Option 3 — sentence-transformers (direct Python) - **Ranking**: 3rd - **Rationale**: Maximum flexibility, no server overhead. Good for dev/test. - **Cons**: Python GIL limits concurrency, not production-grade - **Verdict**: Use for development only. ### Final Recommendation (Serving Platform) **Ollama** for immediate migration. Migrate to **TEI** if throughput becomes a bottleneck. --- ## Acceptance Criteria - [ ] Local baseline replaced: Nomic Embed → BGE-M3 - [ ] Retrieval quality validated: Eval script built with 100-row Q&A CSV (Recall@5 + MRR) - [ ] Cloud comparison completed: ≥3 cloud APIs tested with free-tier credits - [ ] Final model selection documented with rationale and cost projection - [ ] MIRACL-Farsi benchmark downloaded for secondary validation - [ ] Qdrant re-indexed with new embeddings