docs(research): add farsi embedding model research knowledge base

Body:
        Why:
        - Insurance RAG pipeline needs Farsi-optimized embeddings
        - Current local baseline (Nomic Embed) is English-only

        Changes:
        - 8 cloud model profiles with curl + python test commands
        - 5 local model profiles with VRAM, serving platform notes
        - Persian benchmark landscape (MIRACL-Fa, mMARCO, FaMTE, MTEB)
        - Decision report with ranked options and recommendations
        - Full comparison table (15 models × 10 columns)

        Impact:
        - Knowledge base for embedding model selection
        - Top picks: BGE-M3 (local), Cohere embed-v4.0 (cloud)
        - Critical: Nomic Embed must be replaced (English-only)
This commit is contained in:
2026-07-20 15:00:11 +03:30
parent 72d8a3ebe4
commit 6c9cebaaa0
13 changed files with 1296 additions and 0 deletions

View File

@@ -0,0 +1,71 @@
---
last_updated: 2026-07-20
tags: [embedding, gte, alibaba, local, farsi, insurance, rag, lightweight]
source: import-knowledge
---
# Alibaba-NLP/gte-multilingual-base
## Overview
- **Type**: Local
- **Provider**: Alibaba NLP
- **Parameters**: ~305M
- **VRAM (fp16)**: ~610 MB
- **Embedding dimensions**: 768
- **Max input tokens**: 8,192
- **Language coverage**: 70+ languages **explicitly including Farsi** (fa)
- **Serving platform**: sentence-transformers, TEI
- **License**: Apache-2.0
- **Model card**: https://huggingface.co/Alibaba-NLP/gte-multilingual-base
## Why Consider It
Lightest model with **confirmed Farsi support**. Only 610 MB VRAM. 8K context. Apache-2.0 license. Benchmarked on MIRACL and PL-MTEB. Simpler than BGE-M3 (dense-only, no ColBERT/sparse overhead).
## Pros
- Explicitly lists Farsi in supported languages
- Only 610 MB VRAM — lightest meaningful option with Farsi
- 8,192 token context
- Benchmarked on MIRACL and PL-MTEB
- Apache-2.0 license
- Simpler architecture (dense-only) — easier to deploy and debug
## Cons
- No ColBERT/sparse support — dense-only retrieval
- No published isolated Farsi retrieval scores
- Alibaba NLP ecosystem less mature than HuggingFace ecosystem
- Not available on Ollama (may need manual GGUF conversion)
## API Test — Python
```python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Alibaba-NLP/gte-multilingual-base")
# Farsi insurance query
query = "شرایط لغو بیمه نامه چیست?"
embedding = model.encode(query)
print(f"Dimensions: {len(embedding)}")
# Expected: Dimensions: 768
```
## API Test — TEI (Docker)
```bash
# Start TEI server
docker run --gpus all -p 8080:80 -v $PWD/data:/data ghcr.io/huggingface/text-embeddings-inference:latest --model-id Alibaba-NLP/gte-multilingual-base
# Test
curl -s http://localhost:8080/embed -H "Content-Type: application/json" -d '{"inputs": "بیمه نامه شخص ثالث چیست؟"}' | python3 -c "import sys,json; d=json.load(sys.stdin); print(f'dims={len(d[0])}')"
```
## TODO
- [ ] Benchmark on MIRACL-Farsi subset
- [ ] Compare with BGE-M3 on Farsi retrieval quality
- [ ] Test GGUF conversion for Ollama compatibility
- [ ] Measure inference latency on RTX 3090