Files
chatbot_v3/docs/adr/0001-ingestion-pipeline-and-collection-schema.md
Ali Zarinkolah 5c0a5938f8 feat(ingestion): add bounded, benchmark-aligned embedding execution
Why:
- Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into
  the inline upload path, with process-wide capacity/timeout/chunk-limit
  guards (ADR-0017).
- The BM25 analyzer and dense-model config are ported from the `emet`
  evaluation lab, which benchmarked them against the real Farsi corpus
  (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large
  at native 3072-dim), closing open items in ADR-0001/ADR-0005.

Changes:
- New: embedding ports, orchestration (embed_chunks), request-bounds
  helpers, and dense/sparse adapters (analyzers.py, bm25.py,
  openai_compatible.py).
- upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY
  (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413);
  every failure path still writes a terminal job row.
- Lifespan builds and warms both dense embedders at startup (fail-soft) and
  creates the sparse embedder and concurrency semaphore.
- httpx moves from dev to main dependencies (adapters use it directly).

Impact:
- Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0.
- New EMBEDDING_* env vars documented in .env.example; safe defaults.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 17:13:32 +03:30

16 KiB
Raw Blame History

0001. Ingestion pipeline and Qdrant collection schema

Status

Accepted

Context

FastAPI needs to accept uploaded documents (.docx, .csv), preprocess them, split them into chunks, embed the chunks, and store them in Qdrant. This collection is shared by two other pipelines that will be defined in later ADRs: direct CRUD/search over chunks (0002) and hybrid retrieval for an AI agent (0003). The schema decided here — vector configuration and payload fields — is load-bearing for both, so it needs to be right before either is built on top of it.

Two Qdrant behaviors make this schema decision urgent rather than deferrable:

  • Sparse vectors cannot be added to an existing unnamed-dense-vector collection — doing so requires recreating the collection. If we start dense-only and add sparse/late-interaction vectors later, that's a disruptive migration.
  • Co-locating large late-interaction (multivector) representations with dense vectors in the same segment degrades all queries, including plain dense search, unless the multivector is stored on disk.

We've decided (see project discussion) that:

  • Qdrant runs self-hosted via Docker.
  • Tenancy is a single shared collection, partitioned by payload, not collection-per-tenant — Qdrant's own multitenancy guidance is explicit that per-tenant collections "do not scale past a few hundred and waste resources," while payload partitioning scales to 10k+ tenants.
  • The dense embedding model is now decided — nomic-embed-text-v2-moe, see 0004. Sparse and late-interaction models are not yet chosen; those remain swappable via FastEmbed- compatible config.
  • The exact payload field list beyond the baseline below is still open for discussion; this ADR proposes a starting schema, not a final one.

Decision

Collection

One collection, e.g. chunks, shared by all tenants and domains.

Vectors (named vectors, defined at creation time)

Name Type Purpose Notes
dense_nomic dense vector primary semantic similarity (multilingual, incl. Persian) nomic-embed-text-v2-moe, 768-dim (0004)
dense_openai dense vector second semantic signal text-embedding-3-large at its native 3072 dimensions — the dimensions param is deliberately left unset (see below)
sparse sparse vector lexical/keyword-sensitive retrieval bm25-fa-norm-stop — Qdrant FastEmbed's BM25 sparse encoder configured for Persian (stopword removal + normalization), not a separately trained model
late_interaction multivector reserved for late-interaction rerank (0003) jina-colbert-v2 (0005), comparator: max_sim, hnsw_config: m=0 (rerank-only, never independently ANN-searched), stored on disk

All four are defined now, even though late_interaction won't be populated until the agent retrieval work in ADR-0003 lands, specifically to avoid the forced-recreation problem described above. late_interaction is configured with on-disk storage so its larger footprint doesn't degrade the dense/sparse query latency. Two dense vectors are provisioned deliberately — dense_nomic and dense_openai are two independent semantic signals, both prefetched and fused at query time (ADR-0003), not a primary/fallback pair.

Dense model endpoints and dimensions (resolved by the emet benchmark)

Both dense models are reached over the same OpenAI-compatible /embeddings API, so one adapter (src/infrastructure/embedding/openai_compatible.py) serves both named vectors with different configuration:

Named vector Model Endpoint Dimensions
dense_nomic nomic-embed-text-v2-moe self-hosted Ollama OpenAI-compat shim 768 (verified against the live endpoint)
dense_openai text-embedding-3-large OpenAI hosted API 3072 (native; dimensions unset)

dense_openai's dimension was previously listed as an open dependency. It is now pinned to the native 3072, because that is the configuration the emet lab benchmarked — it never passed a dimensions argument. Setting it later would truncate via Matryoshka and is a re-embedding migration, not a config tweak, exactly as the negative consequence below warns.

Two operational notes about the self-hosted embedder, both learned by measurement rather than assumption:

  • Cold load exceeds 150s, far beyond INGESTION_TIMEOUT_SECONDS, so an idle-then-upload would return 504. Mitigated on both ends: Ollama's keep_alive keeps the model resident, and the FastAPI lifespan warms each dense embedder at startup (fail-soft — a down embedder must not block boot).
  • Once warm it is fast: ~0.30s for one input and ~0.34s for a batch of 16. Batching is therefore nearly free, which is what keeps ADR-0017's inline ingestion viable.

Multitenancy / indexing config

  • HNSW: m: 0 (disable the global index) + payload_m: 16, per Qdrant's multitenant collection guidance.
  • Payload index on tenant_id: keyword index with is_tenant: true — this co-locates a tenant's vectors on disk for sequential reads.
  • Payload index on domain: keyword index (secondary partition dimension within a tenant).
  • Payload index on file_id: keyword index, used by CRUD lookups in ADR-0002 (e.g. "delete all chunks belonging to this file").
  • Payload index on order_id: float index (Qdrant's order_by and Range filter conditions only support numeric/datetime payloads, not keyword/string — see the order_id format note below). Used for order_by when listing/scrolling a file's chunks in sequence (ADR-0002) and for previous/next lookups.
  • Payload index on previous_chunk_id / next_chunk_id: keyword index, used for O(1) adjacency retrieval (see below).

Payload schema

This schema is now decided for the fields below. Additional document-context fields (e.g. page/row position, section heading, effective/expiration dates for insurance policy documents) are deliberately deferred to a future ADR that will accompany the docx/csv chunking-strategy work — that decision involves format-specific tradeoffs not yet made.

Field Type Purpose
tenant_id keyword, is_tenant: true index tenant isolation
domain keyword logical partition within a tenant (e.g. fire, car for insurance lines)
file_id keyword groups chunks back to their source file (formerly referred to as document_id — standardized on file_id)
chunk_id keyword stable identifier for a single chunk
content_type keyword classification of the chunk's content; exact value set (e.g. paragraph, table_row, heading) to be finalized alongside the chunking-strategy ADR
source_filename keyword original uploaded filename
source_type keyword (docx | csv) which parser produced this chunk
order_id float (see below) chunk's display position within the file; mutable so the backend can reorder/insert chunks
chunk_index integer chunk's original ingestion ordinal — immutable, used to derive the deterministic point ID below (kept separate from order_id precisely because order_id can change)
previous_chunk_id keyword, nullable chunk_id of the preceding chunk in display order (null for the first chunk in a file) — O(1) adjacency pointer for context-window expansion in ADR-0003
next_chunk_id keyword, nullable chunk_id of the following chunk in display order (null for the last chunk in a file) — same purpose as previous_chunk_id
content text (full-text indexed) the chunk text itself, also used for keyword search in ADR-0002
is_active boolean soft-delete / visibility flag — inactive chunks are excluded from CRUD listing and agent retrieval but retained for audit
deleted_at datetime, nullable set when is_active transitions to false; distinguishes "deactivated" from "never active"
created_at datetime ingestion timestamp
updated_at datetime last modification timestamp
created_by keyword user/service that created the chunk
updated_by keyword user/service that last modified the chunk
version integer optimistic-concurrency counter, used in ADR-0002
content_hash keyword hash of the chunk's raw text; lets re-ingestion detect unchanged content and skip re-embedding it
embedding_model_version keyword identifies which embedding model(s) produced this chunk's vectors; needed to know which chunks require re-embedding after a future model swap

order_id format

order_id is a fractional float key (e.g. 1.0, 2.0, 3.0, ...), not a plain sequential integer and not a lexicographic string. We are not implementing chunk insertion/reordering yet, but choosing this format now means that when that feature is added, inserting a chunk between two existing ones (e.g. assigning it 1.5, then 1.25 for a subsequent insert in the same gap) only touches that one chunk's payload — it never requires renumbering every subsequent chunk, the same benefit a sortable string would give.

A string key was considered first but does not work in Qdrant: the Range filter condition (gt/gte/lt/lte) only supports float/integer payloads (datetime gets its own separate DatetimeRange condition), and order_by on scroll requires a payload index that supports Range filtering — so order_by is likewise limited to numeric/datetime fields. A keyword/string order_id could only be sorted client-side after fetching every chunk for a file, and couldn't support a targeted previous/next query at all. A float key gets native Range/order_by support instead.

Known limitation: repeatedly inserting into the exact same gap (~50+ times between the same two neighbors) runs into floating-point precision limits. Mitigate with an occasional rebalance job that respaces a file's order_id values (e.g. back to 1000, 2000, 3000, ...); this is standard for any fractional-indexing scheme and not expected to be hit in normal use.

Previous/next chunk retrieval

Two ways to get a chunk's neighbors, both viable given this schema:

  1. Pointer fields (primary, recommended for the agent path): read previous_chunk_id/next_chunk_id off the retrieved chunk and fetch those chunk IDs in a single batch "retrieve points by ID" call. This is the intended pattern for ADR-0003's context-window expansion, since it runs on every retrieved chunk and a single batch-get is cheaper than a filtered query per chunk.
  2. Range query (fallback, only viable because order_id is numeric): filter file_id = X + order_id < current (Range lt) + order_by desc
    • limit 1 for the previous chunk; mirror with gt/ascending for the next chunk. Useful if the pointer fields are ever missing/stale, or for ad-hoc debugging.

Pointer fields are only as correct as the mutation logic that maintains them — see ADR-0002 for how reorder/insert/delete operations keep previous_chunk_id/next_chunk_id in sync.

Ingestion flow

  1. FastAPI upload endpoint receives a .docx or .csv file.
  2. Parse: python-docx for .docx, pandas for .csv.
  3. Preprocess: clean/normalize extracted text.
  4. Chunk: split into chunks (size/overlap are tunable config, not fixed by this ADR).
  5. Embed each chunk into all four vectors: dense_nomic, dense_openai, sparse, and late_interaction. dense_nomic uses nomic-embed-text-v2-moe with the chunk text prefixed search_document: (0004); dense_openai uses the OpenAI large embedding model; sparse uses bm25-fa-norm-stop; late_interaction uses jina-colbert-v2 (0005) — its document-side multivector is computed and stored at ingestion time here, while the query-side multivector is computed per-query in ADR-0003 for the MAX_SIM rerank comparison.
  6. Assign a deterministic point ID — UUIDv5 derived from file_id + chunk_index (the immutable ingestion ordinal, not the mutable order_id) — so re-ingesting the same file upserts existing chunks instead of creating duplicates, and re-ordering chunks later never changes their IDs. Initialize order_id from chunk_index at ingestion time (e.g. chunk_index 0, 1, 2 → order_id 1.0, 2.0, 3.0), and set previous_chunk_id/next_chunk_id to each chunk's immediate ingestion-order neighbor (null at the two ends of the file).
  7. Batch upsert into Qdrant: 64–256 points per request, 2–4 parallel upload streams, per Qdrant's bulk-upload guidance.

Consequences

Positive

  • Defining all four named vectors up front avoids a future forced collection recreation when late-interaction rerank is added in ADR-0003.
  • Payload-partitioned multitenancy scales to large tenant counts without per-tenant collection sprawl, and shares infrastructure/config across ADR-0002 and ADR-0003.
  • Deterministic point IDs make ingestion idempotent — safe to re-run on the same document.
  • Numeric order_id is natively supported by Qdrant's Range/order_by, enabling both ordered listing (ADR-0002) and previous/next chunk lookups (ADR-0003's context-window expansion) without client-side sorting.

Negative

  • The collection carries four named vectors and every chunk is embedded into all of them at ingestion (including late_interaction via jina-colbert-v2), rather than only the vectors actually queried at launch — more ingestion-time compute than a leaner initial cut.
  • Document-context payload fields (page/row position, section heading, effective/expiration dates, etc.) are still deferred to the chunking-strategy ADR; adding them later means an additive payload migration, though it won't touch the fields already decided here.
  • Two dense vectors means every chunk is embedded twice (nomic + OpenAI) at ingestion time and both are queried at retrieval time — roughly double the dense embedding cost/latency of a single-dense-vector design, plus an external network dependency on OpenAI's API in the ingestion path.
  • dense_openai's exact output dimension is still an open dependency — resolved: pinned to the native 3072 (see "Dense model endpoints and dimensions" above). The warning still stands for any future change: re-dimensioning is a re-embedding migration, not a config tweak.
  • The sparse vector must be created with modifier="idf". The client computes only BM25's term-frequency saturation; without that modifier Qdrant applies no IDF at all and lexical retrieval silently degrades (ADR-0005).
  • jina-colbert-v2 (0005) adds a hard GPU dependency to ingestion (not just query time, since the document-side multivector is computed here) and its commercial license is still unconfirmed.
  • previous_chunk_id/next_chunk_id are denormalized pointers — every reorder/insert/delete must update the affected neighbors' payloads too (see ADR-0002), or the pointers go stale. Fractional float order_id also needs an (infrequent) rebalance job as a long-term maintenance task.

Alternatives Considered

  • Collection per domain: rejected. Domains would need materially different vector configs or strict data-boundary requirements to justify this; neither applies here, and it fragments tenant-scaling benefits.
  • Collection per tenant: rejected outright per Qdrant's own guidance — doesn't scale past a few hundred tenants.
  • Start dense-only, add sparse/late-interaction later: rejected — Qdrant requires recreating the collection to add sparse vectors to an unnamed dense-only collection, which is a disruptive migration we can avoid by deciding the full vector shape now.