feat(ingestion): add bounded, benchmark-aligned embedding execution
Why: - Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into the inline upload path, with process-wide capacity/timeout/chunk-limit guards (ADR-0017). - The BM25 analyzer and dense-model config are ported from the `emet` evaluation lab, which benchmarked them against the real Farsi corpus (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large at native 3072-dim), closing open items in ADR-0001/ADR-0005. Changes: - New: embedding ports, orchestration (embed_chunks), request-bounds helpers, and dense/sparse adapters (analyzers.py, bm25.py, openai_compatible.py). - upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413); every failure path still writes a terminal job row. - Lifespan builds and warms both dense embedders at startup (fail-soft) and creates the sparse embedder and concurrency semaphore. - httpx moves from dev to main dependencies (adapters use it directly). Impact: - Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0. - New EMBEDDING_* env vars documented in .env.example; safe defaults. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
49
src/application/ports/embedding.py
Normal file
49
src/application/ports/embedding.py
Normal file
@@ -0,0 +1,49 @@
|
||||
"""Embedding ports (ADR-0001, ADR-0017).
|
||||
|
||||
`src/infrastructure/embedding/` holds the production adapters; tests use
|
||||
scripted fakes (ADR-0016). Application code depends on these Protocols, not
|
||||
on `httpx`/provider SDKs directly.
|
||||
"""
|
||||
|
||||
from collections.abc import Sequence
|
||||
from typing import Protocol
|
||||
|
||||
from src.application.ingestion.models import SparseVector
|
||||
|
||||
|
||||
class DenseEmbedder(Protocol):
|
||||
"""One named dense vector's embedding client (`dense_nomic`/`dense_openai`).
|
||||
|
||||
`embed_batch` is a single batched network call — callers own concurrency
|
||||
bounding (ADR-0017's `embed_concurrency` semaphore), not this Protocol.
|
||||
"""
|
||||
|
||||
name: str
|
||||
|
||||
async def embed_batch(self, texts: Sequence[str]) -> list[list[float]]:
|
||||
"""Return one vector per input text, same order. Raises `EmbedderError`
|
||||
(see `src/application/ingestion/errors.py`) on transport/response
|
||||
failure.
|
||||
"""
|
||||
...
|
||||
|
||||
|
||||
class SparseEmbedder(Protocol):
|
||||
"""The `sparse` (BM25) vector's embedding client.
|
||||
|
||||
Blocking/CPU-bound (ADR-0017): callers offload it via
|
||||
`anyio.to_thread.run_sync` with the ingestion `CapacityLimiter`, not call
|
||||
it directly from an `async def`.
|
||||
"""
|
||||
|
||||
name: str
|
||||
|
||||
def embed_batch(self, texts: Sequence[str], *, query: bool = False) -> list[SparseVector]:
|
||||
"""Return one sparse vector per input text, same order.
|
||||
|
||||
`query=True` selects the query-side weighting, which omits document
|
||||
length normalization. Ingestion always passes `False`; the flag exists
|
||||
so retrieval (ADR-0003) encodes queries through this same port rather
|
||||
than growing a second, silently divergent implementation.
|
||||
"""
|
||||
...
|
||||
Reference in New Issue
Block a user