feat(ingestion): add bounded, benchmark-aligned embedding execution

Why:
- Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into
  the inline upload path, with process-wide capacity/timeout/chunk-limit
  guards (ADR-0017).
- The BM25 analyzer and dense-model config are ported from the `emet`
  evaluation lab, which benchmarked them against the real Farsi corpus
  (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large
  at native 3072-dim), closing open items in ADR-0001/ADR-0005.

Changes:
- New: embedding ports, orchestration (embed_chunks), request-bounds
  helpers, and dense/sparse adapters (analyzers.py, bm25.py,
  openai_compatible.py).
- upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY
  (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413);
  every failure path still writes a terminal job row.
- Lifespan builds and warms both dense embedders at startup (fail-soft) and
  creates the sparse embedder and concurrency semaphore.
- httpx moves from dev to main dependencies (adapters use it directly).

Impact:
- Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0.
- New EMBEDDING_* env vars documented in .env.example; safe defaults.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-19 17:13:32 +03:30
parent aa6d595424
commit 5c0a5938f8
33 changed files with 2455 additions and 536 deletions

View File

@@ -1,7 +1,11 @@
"""Hand-written fakes for narrow application-owned ports (ADR-0016)."""
import asyncio
from collections.abc import Sequence
from dataclasses import dataclass, field
from src.application.ingestion.models import SparseVector
@dataclass
class FakeObjectStorage:
@@ -15,3 +19,42 @@ class FakeObjectStorage:
self.fail_next = False
raise OSError("simulated object storage failure")
self.objects[key] = data
@dataclass
class FakeDenseEmbedder:
"""A scripted `DenseEmbedder`. Returns a fixed-dimension zero vector per
text by default; `fail_next` simulates one batch failure.
"""
name: str
dimensions: int = 4
calls: list[list[str]] = field(default_factory=list)
fail_next: bool = False
delay_seconds: float = 0.0
"""Simulates a slow provider call, e.g. to exercise timeout handling."""
async def embed_batch(self, texts: Sequence[str]) -> list[list[float]]:
self.calls.append(list(texts))
if self.delay_seconds:
await asyncio.sleep(self.delay_seconds)
if self.fail_next:
self.fail_next = False
raise RuntimeError("simulated embedder failure")
return [[0.0] * self.dimensions for _ in texts]
@dataclass
class FakeSparseEmbedder:
"""A scripted `SparseEmbedder`. Returns an empty sparse vector per text."""
name: str = "sparse"
calls: list[list[str]] = field(default_factory=list)
fail_next: bool = False
def embed_batch(self, texts: Sequence[str], *, query: bool = False) -> list[SparseVector]:
self.calls.append(list(texts))
if self.fail_next:
self.fail_next = False
raise RuntimeError("simulated embedder failure")
return [SparseVector(indices=[], values=[]) for _ in texts]