Why: - Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into the inline upload path, with process-wide capacity/timeout/chunk-limit guards (ADR-0017). - The BM25 analyzer and dense-model config are ported from the `emet` evaluation lab, which benchmarked them against the real Farsi corpus (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large at native 3072-dim), closing open items in ADR-0001/ADR-0005. Changes: - New: embedding ports, orchestration (embed_chunks), request-bounds helpers, and dense/sparse adapters (analyzers.py, bm25.py, openai_compatible.py). - upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413); every failure path still writes a terminal job row. - Lifespan builds and warms both dense embedders at startup (fail-soft) and creates the sparse embedder and concurrency semaphore. - httpx moves from dev to main dependencies (adapters use it directly). Impact: - Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0. - New EMBEDDING_* env vars documented in .env.example; safe defaults. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
16 KiB
0001. Ingestion pipeline and Qdrant collection schema
Status
Accepted
Context
FastAPI needs to accept uploaded documents (.docx, .csv), preprocess them,
split them into chunks, embed the chunks, and store them in Qdrant. This
collection is shared by two other pipelines that will be defined in later
ADRs: direct CRUD/search over chunks (0002)
and hybrid retrieval for an AI agent
(0003). The schema decided here — vector
configuration and payload fields — is load-bearing for both, so it needs to
be right before either is built on top of it.
Two Qdrant behaviors make this schema decision urgent rather than deferrable:
- Sparse vectors cannot be added to an existing unnamed-dense-vector collection — doing so requires recreating the collection. If we start dense-only and add sparse/late-interaction vectors later, that's a disruptive migration.
- Co-locating large late-interaction (multivector) representations with dense vectors in the same segment degrades all queries, including plain dense search, unless the multivector is stored on disk.
We've decided (see project discussion) that:
- Qdrant runs self-hosted via Docker.
- Tenancy is a single shared collection, partitioned by payload, not collection-per-tenant — Qdrant's own multitenancy guidance is explicit that per-tenant collections "do not scale past a few hundred and waste resources," while payload partitioning scales to 10k+ tenants.
- The dense embedding model is now decided —
nomic-embed-text-v2-moe, see 0004. Sparse and late-interaction models are not yet chosen; those remain swappable via FastEmbed- compatible config. - The exact payload field list beyond the baseline below is still open for discussion; this ADR proposes a starting schema, not a final one.
Decision
Collection
One collection, e.g. chunks, shared by all tenants and domains.
Vectors (named vectors, defined at creation time)
| Name | Type | Purpose | Notes |
|---|---|---|---|
dense_nomic |
dense vector | primary semantic similarity (multilingual, incl. Persian) | nomic-embed-text-v2-moe, 768-dim (0004) |
dense_openai |
dense vector | second semantic signal | text-embedding-3-large at its native 3072 dimensions — the dimensions param is deliberately left unset (see below) |
sparse |
sparse vector | lexical/keyword-sensitive retrieval | bm25-fa-norm-stop — Qdrant FastEmbed's BM25 sparse encoder configured for Persian (stopword removal + normalization), not a separately trained model |
late_interaction |
multivector | reserved for late-interaction rerank (0003) | jina-colbert-v2 (0005), comparator: max_sim, hnsw_config: m=0 (rerank-only, never independently ANN-searched), stored on disk |
All four are defined now, even though late_interaction won't be
populated until the agent retrieval work in ADR-0003 lands, specifically to
avoid the forced-recreation problem described above. late_interaction is
configured with on-disk storage so its larger footprint doesn't degrade the
dense/sparse query latency. Two dense vectors are provisioned deliberately —
dense_nomic and dense_openai are two independent semantic signals, both
prefetched and fused at query time (ADR-0003), not a primary/fallback pair.
Dense model endpoints and dimensions (resolved by the emet benchmark)
Both dense models are reached over the same OpenAI-compatible
/embeddings API, so one adapter
(src/infrastructure/embedding/openai_compatible.py) serves both named
vectors with different configuration:
| Named vector | Model | Endpoint | Dimensions |
|---|---|---|---|
dense_nomic |
nomic-embed-text-v2-moe |
self-hosted Ollama OpenAI-compat shim | 768 (verified against the live endpoint) |
dense_openai |
text-embedding-3-large |
OpenAI hosted API | 3072 (native; dimensions unset) |
dense_openai's dimension was previously listed as an open dependency. It is
now pinned to the native 3072, because that is the configuration the emet
lab benchmarked — it never passed a dimensions argument. Setting it later
would truncate via Matryoshka and is a re-embedding migration, not a config
tweak, exactly as the negative consequence below warns.
Two operational notes about the self-hosted embedder, both learned by measurement rather than assumption:
- Cold load exceeds 150s, far beyond
INGESTION_TIMEOUT_SECONDS, so an idle-then-upload would return504. Mitigated on both ends: Ollama'skeep_alivekeeps the model resident, and the FastAPI lifespan warms each dense embedder at startup (fail-soft — a down embedder must not block boot). - Once warm it is fast: ~0.30s for one input and ~0.34s for a batch of 16. Batching is therefore nearly free, which is what keeps ADR-0017's inline ingestion viable.
Multitenancy / indexing config
- HNSW:
m: 0(disable the global index) +payload_m: 16, per Qdrant's multitenant collection guidance. - Payload index on
tenant_id: keyword index withis_tenant: true— this co-locates a tenant's vectors on disk for sequential reads. - Payload index on
domain: keyword index (secondary partition dimension within a tenant). - Payload index on
file_id: keyword index, used by CRUD lookups in ADR-0002 (e.g. "delete all chunks belonging to this file"). - Payload index on
order_id: float index (Qdrant'sorder_byandRangefilter conditions only support numeric/datetime payloads, not keyword/string — see theorder_idformat note below). Used fororder_bywhen listing/scrolling a file's chunks in sequence (ADR-0002) and for previous/next lookups. - Payload index on
previous_chunk_id/next_chunk_id: keyword index, used for O(1) adjacency retrieval (see below).
Payload schema
This schema is now decided for the fields below. Additional document-context fields (e.g. page/row position, section heading, effective/expiration dates for insurance policy documents) are deliberately deferred to a future ADR that will accompany the docx/csv chunking-strategy work — that decision involves format-specific tradeoffs not yet made.
| Field | Type | Purpose |
|---|---|---|
tenant_id |
keyword, is_tenant: true index |
tenant isolation |
domain |
keyword | logical partition within a tenant (e.g. fire, car for insurance lines) |
file_id |
keyword | groups chunks back to their source file (formerly referred to as document_id — standardized on file_id) |
chunk_id |
keyword | stable identifier for a single chunk |
content_type |
keyword | classification of the chunk's content; exact value set (e.g. paragraph, table_row, heading) to be finalized alongside the chunking-strategy ADR |
source_filename |
keyword | original uploaded filename |
source_type |
keyword (docx | csv) |
which parser produced this chunk |
order_id |
float (see below) | chunk's display position within the file; mutable so the backend can reorder/insert chunks |
chunk_index |
integer | chunk's original ingestion ordinal — immutable, used to derive the deterministic point ID below (kept separate from order_id precisely because order_id can change) |
previous_chunk_id |
keyword, nullable | chunk_id of the preceding chunk in display order (null for the first chunk in a file) — O(1) adjacency pointer for context-window expansion in ADR-0003 |
next_chunk_id |
keyword, nullable | chunk_id of the following chunk in display order (null for the last chunk in a file) — same purpose as previous_chunk_id |
content |
text (full-text indexed) | the chunk text itself, also used for keyword search in ADR-0002 |
is_active |
boolean | soft-delete / visibility flag — inactive chunks are excluded from CRUD listing and agent retrieval but retained for audit |
deleted_at |
datetime, nullable | set when is_active transitions to false; distinguishes "deactivated" from "never active" |
created_at |
datetime | ingestion timestamp |
updated_at |
datetime | last modification timestamp |
created_by |
keyword | user/service that created the chunk |
updated_by |
keyword | user/service that last modified the chunk |
version |
integer | optimistic-concurrency counter, used in ADR-0002 |
content_hash |
keyword | hash of the chunk's raw text; lets re-ingestion detect unchanged content and skip re-embedding it |
embedding_model_version |
keyword | identifies which embedding model(s) produced this chunk's vectors; needed to know which chunks require re-embedding after a future model swap |
order_id format
order_id is a fractional float key (e.g. 1.0, 2.0, 3.0, ...),
not a plain sequential integer and not a lexicographic string. We are not
implementing chunk insertion/reordering yet, but choosing this format now
means that when that feature is added, inserting a chunk between two
existing ones (e.g. assigning it 1.5, then 1.25 for a subsequent insert
in the same gap) only touches that one chunk's payload — it never requires
renumbering every subsequent chunk, the same benefit a sortable string would
give.
A string key was considered first but does not work in Qdrant: the
Range filter condition (gt/gte/lt/lte) only supports float/integer
payloads (datetime gets its own separate DatetimeRange condition), and
order_by on scroll requires a payload index that supports Range
filtering — so order_by is likewise limited to numeric/datetime fields.
A keyword/string order_id could only be sorted client-side after fetching
every chunk for a file, and couldn't support a targeted previous/next query
at all. A float key gets native Range/order_by support instead.
Known limitation: repeatedly inserting into the exact same gap (~50+
times between the same two neighbors) runs into floating-point precision
limits. Mitigate with an occasional rebalance job that respaces a file's
order_id values (e.g. back to 1000, 2000, 3000, ...); this is standard
for any fractional-indexing scheme and not expected to be hit in normal use.
Previous/next chunk retrieval
Two ways to get a chunk's neighbors, both viable given this schema:
- Pointer fields (primary, recommended for the agent path): read
previous_chunk_id/next_chunk_idoff the retrieved chunk and fetch those chunk IDs in a single batch "retrieve points by ID" call. This is the intended pattern for ADR-0003's context-window expansion, since it runs on every retrieved chunk and a single batch-get is cheaper than a filtered query per chunk. - Range query (fallback, only viable because
order_idis numeric): filterfile_id = X+order_id < current(Rangelt) +order_by desclimit 1for the previous chunk; mirror withgt/ascending for the next chunk. Useful if the pointer fields are ever missing/stale, or for ad-hoc debugging.
Pointer fields are only as correct as the mutation logic that maintains
them — see ADR-0002 for how reorder/insert/delete operations keep
previous_chunk_id/next_chunk_id in sync.
Ingestion flow
- FastAPI upload endpoint receives a
.docxor.csvfile. - Parse:
python-docxfor.docx,pandasfor.csv. - Preprocess: clean/normalize extracted text.
- Chunk: split into chunks (size/overlap are tunable config, not fixed by this ADR).
- Embed each chunk into all four vectors:
dense_nomic,dense_openai,sparse, andlate_interaction.dense_nomicusesnomic-embed-text-v2-moewith the chunk text prefixedsearch_document:(0004);dense_openaiuses the OpenAI large embedding model;sparseusesbm25-fa-norm-stop;late_interactionusesjina-colbert-v2(0005) — its document-side multivector is computed and stored at ingestion time here, while the query-side multivector is computed per-query in ADR-0003 for the MAX_SIM rerank comparison. - Assign a deterministic point ID — UUIDv5 derived from
file_id+chunk_index(the immutable ingestion ordinal, not the mutableorder_id) — so re-ingesting the same file upserts existing chunks instead of creating duplicates, and re-ordering chunks later never changes their IDs. Initializeorder_idfromchunk_indexat ingestion time (e.g.chunk_index0, 1, 2 →order_id1.0,2.0,3.0), and setprevious_chunk_id/next_chunk_idto each chunk's immediate ingestion-order neighbor (nullat the two ends of the file). - Batch upsert into Qdrant: 64–256 points per request, 2–4 parallel upload streams, per Qdrant's bulk-upload guidance.
Consequences
Positive
- Defining all four named vectors up front avoids a future forced collection recreation when late-interaction rerank is added in ADR-0003.
- Payload-partitioned multitenancy scales to large tenant counts without per-tenant collection sprawl, and shares infrastructure/config across ADR-0002 and ADR-0003.
- Deterministic point IDs make ingestion idempotent — safe to re-run on the same document.
- Numeric
order_idis natively supported by Qdrant'sRange/order_by, enabling both ordered listing (ADR-0002) and previous/next chunk lookups (ADR-0003's context-window expansion) without client-side sorting.
Negative
- The collection carries four named vectors and every chunk is embedded
into all of them at ingestion (including
late_interactionviajina-colbert-v2), rather than only the vectors actually queried at launch — more ingestion-time compute than a leaner initial cut. - Document-context payload fields (page/row position, section heading, effective/expiration dates, etc.) are still deferred to the chunking-strategy ADR; adding them later means an additive payload migration, though it won't touch the fields already decided here.
- Two dense vectors means every chunk is embedded twice (nomic + OpenAI) at ingestion time and both are queried at retrieval time — roughly double the dense embedding cost/latency of a single-dense-vector design, plus an external network dependency on OpenAI's API in the ingestion path.
— resolved: pinned to the native 3072 (see "Dense model endpoints and dimensions" above). The warning still stands for any future change: re-dimensioning is a re-embedding migration, not a config tweak.dense_openai's exact output dimension is still an open dependency- The
sparsevector must be created withmodifier="idf". The client computes only BM25's term-frequency saturation; without that modifier Qdrant applies no IDF at all and lexical retrieval silently degrades (ADR-0005). jina-colbert-v2(0005) adds a hard GPU dependency to ingestion (not just query time, since the document-side multivector is computed here) and its commercial license is still unconfirmed.previous_chunk_id/next_chunk_idare denormalized pointers — every reorder/insert/delete must update the affected neighbors' payloads too (see ADR-0002), or the pointers go stale. Fractional floatorder_idalso needs an (infrequent) rebalance job as a long-term maintenance task.
Alternatives Considered
- Collection per domain: rejected. Domains would need materially different vector configs or strict data-boundary requirements to justify this; neither applies here, and it fragments tenant-scaling benefits.
- Collection per tenant: rejected outright per Qdrant's own guidance — doesn't scale past a few hundred tenants.
- Start dense-only, add sparse/late-interaction later: rejected — Qdrant requires recreating the collection to add sparse vectors to an unnamed dense-only collection, which is a disruptive migration we can avoid by deciding the full vector shape now.