# 0001. Ingestion pipeline and Qdrant collection schema ## Status Accepted ## Context FastAPI needs to accept uploaded documents (`.docx`, `.csv`), preprocess them, split them into chunks, embed the chunks, and store them in Qdrant. This collection is shared by two other pipelines that will be defined in later ADRs: direct CRUD/search over chunks ([0002](0002-chunk-crud-and-search-api.md)) and hybrid retrieval for an AI agent ([0003](0003-agent-hybrid-retrieval.md)). The schema decided here — vector configuration and payload fields — is load-bearing for both, so it needs to be right before either is built on top of it. Two Qdrant behaviors make this schema decision urgent rather than deferrable: - **Sparse vectors cannot be added to an existing unnamed-dense-vector collection** — doing so requires recreating the collection. If we start dense-only and add sparse/late-interaction vectors later, that's a disruptive migration. - **Co-locating large late-interaction (multivector) representations with dense vectors in the same segment degrades *all* queries**, including plain dense search, unless the multivector is stored on disk. We've decided (see project discussion) that: - Qdrant runs self-hosted via Docker. - Tenancy is a **single shared collection**, partitioned by payload, not collection-per-tenant — Qdrant's own multitenancy guidance is explicit that per-tenant collections "do not scale past a few hundred and waste resources," while payload partitioning scales to 10k+ tenants. - The dense embedding model is now decided — `nomic-embed-text-v2-moe`, see [0004](0004-docx-csv-chunking-strategy.md). Sparse and late-interaction models are **not yet chosen**; those remain swappable via FastEmbed- compatible config. - The exact payload field list beyond the baseline below is still open for discussion; this ADR proposes a starting schema, not a final one. ## Decision ### Collection One collection, e.g. `chunks`, shared by all tenants and domains. ### Vectors (named vectors, defined at creation time) | Name | Type | Purpose | Notes | |---|---|---|---| | `dense_nomic` | dense vector | primary semantic similarity (multilingual, incl. Persian) | `nomic-embed-text-v2-moe`, 768-dim ([0004](0004-docx-csv-chunking-strategy.md)) | | `dense_openai` | dense vector | second semantic signal | `text-embedding-3-large` at its **native 3072 dimensions** — the `dimensions` param is deliberately left unset (see below) | | `sparse` | sparse vector | lexical/keyword-sensitive retrieval | `bm25-fa-norm-stop` — Qdrant FastEmbed's BM25 sparse encoder configured for Persian (stopword removal + normalization), not a separately trained model | | `late_interaction` | multivector | reserved for late-interaction rerank ([0003](0003-agent-hybrid-retrieval.md)) | `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)), `comparator: max_sim`, `hnsw_config: m=0` (rerank-only, never independently ANN-searched), stored **on disk** | All four are defined **now**, even though `late_interaction` won't be populated until the agent retrieval work in ADR-0003 lands, specifically to avoid the forced-recreation problem described above. `late_interaction` is configured with on-disk storage so its larger footprint doesn't degrade the dense/sparse query latency. Two dense vectors are provisioned deliberately — `dense_nomic` and `dense_openai` are two independent semantic signals, both prefetched and fused at query time (ADR-0003), not a primary/fallback pair. #### Dense model endpoints and dimensions (resolved by the `emet` benchmark) Both dense models are reached over the **same OpenAI-compatible `/embeddings` API**, so one adapter (`src/infrastructure/embedding/openai_compatible.py`) serves both named vectors with different configuration: | Named vector | Model | Endpoint | Dimensions | |---|---|---|---| | `dense_nomic` | `nomic-embed-text-v2-moe` | self-hosted Ollama OpenAI-compat shim | **768** (verified against the live endpoint) | | `dense_openai` | `text-embedding-3-large` | OpenAI hosted API | **3072** (native; `dimensions` unset) | `dense_openai`'s dimension was previously listed as an open dependency. It is now pinned to the native 3072, because that is the configuration the `emet` lab benchmarked — it never passed a `dimensions` argument. Setting it later would truncate via Matryoshka and is a **re-embedding migration, not a config tweak**, exactly as the negative consequence below warns. Two operational notes about the self-hosted embedder, both learned by measurement rather than assumption: - **Cold load exceeds 150s**, far beyond `INGESTION_TIMEOUT_SECONDS`, so an idle-then-upload would return `504`. Mitigated on both ends: Ollama's `keep_alive` keeps the model resident, and the FastAPI lifespan warms each dense embedder at startup (fail-soft — a down embedder must not block boot). - **Once warm it is fast**: ~0.30s for one input and ~0.34s for a batch of 16. Batching is therefore nearly free, which is what keeps ADR-0017's inline ingestion viable. ### Multitenancy / indexing config - HNSW: `m: 0` (disable the global index) + `payload_m: 16`, per Qdrant's multitenant collection guidance. - Payload index on `tenant_id`: keyword index with `is_tenant: true` — this co-locates a tenant's vectors on disk for sequential reads. - Payload index on `domain`: keyword index (secondary partition dimension within a tenant). - Payload index on `file_id`: keyword index, used by CRUD lookups in ADR-0002 (e.g. "delete all chunks belonging to this file"). - Payload index on `order_id`: **float** index (Qdrant's `order_by` and `Range` filter conditions only support numeric/datetime payloads, not keyword/string — see the `order_id` format note below). Used for `order_by` when listing/scrolling a file's chunks in sequence (ADR-0002) and for previous/next lookups. - Payload index on `previous_chunk_id` / `next_chunk_id`: keyword index, used for O(1) adjacency retrieval (see below). ### Payload schema This schema is now decided for the fields below. Additional document-context fields (e.g. page/row position, section heading, effective/expiration dates for insurance policy documents) are deliberately **deferred to a future ADR** that will accompany the docx/csv chunking-strategy work — that decision involves format-specific tradeoffs not yet made. | Field | Type | Purpose | |---|---|---| | `tenant_id` | keyword, `is_tenant: true` index | tenant isolation | | `domain` | keyword | logical partition within a tenant (e.g. `fire`, `car` for insurance lines) | | `file_id` | keyword | groups chunks back to their source file (formerly referred to as `document_id` — standardized on `file_id`) | | `chunk_id` | keyword | stable identifier for a single chunk | | `content_type` | keyword | classification of the chunk's content; exact value set (e.g. `paragraph`, `table_row`, `heading`) to be finalized alongside the chunking-strategy ADR | | `source_filename` | keyword | original uploaded filename | | `source_type` | keyword (`docx` \| `csv`) | which parser produced this chunk | | `order_id` | float (see below) | chunk's *display* position within the file; mutable so the backend can reorder/insert chunks | | `chunk_index` | integer | chunk's *original ingestion* ordinal — immutable, used to derive the deterministic point ID below (kept separate from `order_id` precisely because `order_id` can change) | | `previous_chunk_id` | keyword, nullable | `chunk_id` of the preceding chunk in display order (`null` for the first chunk in a file) — O(1) adjacency pointer for context-window expansion in ADR-0003 | | `next_chunk_id` | keyword, nullable | `chunk_id` of the following chunk in display order (`null` for the last chunk in a file) — same purpose as `previous_chunk_id` | | `content` | text (full-text indexed) | the chunk text itself, also used for keyword search in ADR-0002 | | `is_active` | boolean | soft-delete / visibility flag — inactive chunks are excluded from CRUD listing and agent retrieval but retained for audit | | `deleted_at` | datetime, nullable | set when `is_active` transitions to `false`; distinguishes "deactivated" from "never active" | | `created_at` | datetime | ingestion timestamp | | `updated_at` | datetime | last modification timestamp | | `created_by` | keyword | user/service that created the chunk | | `updated_by` | keyword | user/service that last modified the chunk | | `version` | integer | optimistic-concurrency counter, used in ADR-0002 | | `content_hash` | keyword | hash of the chunk's raw text; lets re-ingestion detect unchanged content and skip re-embedding it | | `embedding_model_version` | keyword | identifies which embedding model(s) produced this chunk's vectors; needed to know which chunks require re-embedding after a future model swap | #### `order_id` format `order_id` is a **fractional float key** (e.g. `1.0`, `2.0`, `3.0`, ...), not a plain sequential integer and not a lexicographic string. We are not implementing chunk insertion/reordering yet, but choosing this format now means that when that feature is added, inserting a chunk between two existing ones (e.g. assigning it `1.5`, then `1.25` for a subsequent insert in the same gap) only touches that one chunk's payload — it never requires renumbering every subsequent chunk, the same benefit a sortable string would give. A string key was considered first but **does not work in Qdrant**: the `Range` filter condition (`gt`/`gte`/`lt`/`lte`) only supports float/integer payloads (datetime gets its own separate `DatetimeRange` condition), and `order_by` on scroll requires a payload index that supports `Range` filtering — so `order_by` is likewise limited to numeric/datetime fields. A keyword/string `order_id` could only be sorted client-side after fetching every chunk for a file, and couldn't support a targeted previous/next query at all. A float key gets native `Range`/`order_by` support instead. **Known limitation**: repeatedly inserting into the exact same gap (~50+ times between the same two neighbors) runs into floating-point precision limits. Mitigate with an occasional rebalance job that respaces a file's `order_id` values (e.g. back to `1000, 2000, 3000, ...`); this is standard for any fractional-indexing scheme and not expected to be hit in normal use. #### Previous/next chunk retrieval Two ways to get a chunk's neighbors, both viable given this schema: 1. **Pointer fields (primary, recommended for the agent path)**: read `previous_chunk_id`/`next_chunk_id` off the retrieved chunk and fetch those chunk IDs in a single batch "retrieve points by ID" call. This is the intended pattern for ADR-0003's context-window expansion, since it runs on every retrieved chunk and a single batch-get is cheaper than a filtered query per chunk. 2. **Range query (fallback, only viable because `order_id` is numeric)**: filter `file_id = X` + `order_id < current` (Range `lt`) + `order_by desc` + `limit 1` for the previous chunk; mirror with `gt`/ascending for the next chunk. Useful if the pointer fields are ever missing/stale, or for ad-hoc debugging. Pointer fields are only as correct as the mutation logic that maintains them — see ADR-0002 for how reorder/insert/delete operations keep `previous_chunk_id`/`next_chunk_id` in sync. ### Ingestion flow 1. FastAPI upload endpoint receives a `.docx` or `.csv` file. 2. Parse: `python-docx` for `.docx`, `pandas` for `.csv`. 3. Preprocess: clean/normalize extracted text. 4. Chunk: split into chunks (size/overlap are tunable config, not fixed by this ADR). 5. Embed each chunk into all four vectors: `dense_nomic`, `dense_openai`, `sparse`, and `late_interaction`. `dense_nomic` uses `nomic-embed-text-v2-moe` with the chunk text prefixed `search_document: ` ([0004](0004-docx-csv-chunking-strategy.md)); `dense_openai` uses the OpenAI large embedding model; `sparse` uses `bm25-fa-norm-stop`; `late_interaction` uses `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)) — its document-side multivector is computed and stored at ingestion time here, while the query-side multivector is computed per-query in ADR-0003 for the MAX_SIM rerank comparison. 6. Assign a **deterministic point ID** — UUIDv5 derived from `file_id` + `chunk_index` (the immutable ingestion ordinal, not the mutable `order_id`) — so re-ingesting the same file upserts existing chunks instead of creating duplicates, and re-ordering chunks later never changes their IDs. Initialize `order_id` from `chunk_index` at ingestion time (e.g. `chunk_index` 0, 1, 2 → `order_id` `1.0`, `2.0`, `3.0`), and set `previous_chunk_id`/`next_chunk_id` to each chunk's immediate ingestion-order neighbor (`null` at the two ends of the file). 7. Batch upsert into Qdrant: 64–256 points per request, 2–4 parallel upload streams, per Qdrant's bulk-upload guidance. ## Consequences ### Positive - Defining all four named vectors up front avoids a future forced collection recreation when late-interaction rerank is added in ADR-0003. - Payload-partitioned multitenancy scales to large tenant counts without per-tenant collection sprawl, and shares infrastructure/config across ADR-0002 and ADR-0003. - Deterministic point IDs make ingestion idempotent — safe to re-run on the same document. - Numeric `order_id` is natively supported by Qdrant's `Range`/`order_by`, enabling both ordered listing (ADR-0002) and previous/next chunk lookups (ADR-0003's context-window expansion) without client-side sorting. ### Negative - The collection carries four named vectors and every chunk is embedded into all of them at ingestion (including `late_interaction` via `jina-colbert-v2`), rather than only the vectors actually queried at launch — more ingestion-time compute than a leaner initial cut. - Document-context payload fields (page/row position, section heading, effective/expiration dates, etc.) are still deferred to the chunking-strategy ADR; adding them later means an additive payload migration, though it won't touch the fields already decided here. - Two dense vectors means every chunk is embedded twice (nomic + OpenAI) at ingestion time and both are queried at retrieval time — roughly double the dense embedding cost/latency of a single-dense-vector design, plus an external network dependency on OpenAI's API in the ingestion path. - ~~`dense_openai`'s exact output dimension is still an open dependency~~ — **resolved**: pinned to the native 3072 (see "Dense model endpoints and dimensions" above). The warning still stands for any future change: re-dimensioning is a re-embedding migration, not a config tweak. - The `sparse` vector must be created with `modifier="idf"`. The client computes only BM25's term-frequency saturation; without that modifier Qdrant applies no IDF at all and lexical retrieval silently degrades (ADR-0005). - `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)) adds a hard GPU dependency to ingestion (not just query time, since the document-side multivector is computed here) and its commercial license is still unconfirmed. - `previous_chunk_id`/`next_chunk_id` are denormalized pointers — every reorder/insert/delete must update the affected neighbors' payloads too (see ADR-0002), or the pointers go stale. Fractional float `order_id` also needs an (infrequent) rebalance job as a long-term maintenance task. ## Alternatives Considered - **Collection per domain**: rejected. Domains would need materially different vector configs or strict data-boundary requirements to justify this; neither applies here, and it fragments tenant-scaling benefits. - **Collection per tenant**: rejected outright per Qdrant's own guidance — doesn't scale past a few hundred tenants. - **Start dense-only, add sparse/late-interaction later**: rejected — Qdrant requires recreating the collection to add sparse vectors to an unnamed dense-only collection, which is a disruptive migration we can avoid by deciding the full vector shape now.