Files
chatbot_v3/docs/adr/0001-ingestion-pipeline-and-collection-schema.md
Ali Zarinkolah 5c0a5938f8 feat(ingestion): add bounded, benchmark-aligned embedding execution
Why:
- Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into
  the inline upload path, with process-wide capacity/timeout/chunk-limit
  guards (ADR-0017).
- The BM25 analyzer and dense-model config are ported from the `emet`
  evaluation lab, which benchmarked them against the real Farsi corpus
  (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large
  at native 3072-dim), closing open items in ADR-0001/ADR-0005.

Changes:
- New: embedding ports, orchestration (embed_chunks), request-bounds
  helpers, and dense/sparse adapters (analyzers.py, bm25.py,
  openai_compatible.py).
- upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY
  (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413);
  every failure path still writes a terminal job row.
- Lifespan builds and warms both dense embedders at startup (fail-soft) and
  creates the sparse embedder and concurrency semaphore.
- httpx moves from dev to main dependencies (adapters use it directly).

Impact:
- Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0.
- New EMBEDDING_* env vars documented in .env.example; safe defaults.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 17:13:32 +03:30

272 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 0001. Ingestion pipeline and Qdrant collection schema
## Status
Accepted
## Context
FastAPI needs to accept uploaded documents (`.docx`, `.csv`), preprocess them,
split them into chunks, embed the chunks, and store them in Qdrant. This
collection is shared by two other pipelines that will be defined in later
ADRs: direct CRUD/search over chunks ([0002](0002-chunk-crud-and-search-api.md))
and hybrid retrieval for an AI agent
([0003](0003-agent-hybrid-retrieval.md)). The schema decided here — vector
configuration and payload fields — is load-bearing for both, so it needs to
be right before either is built on top of it.
Two Qdrant behaviors make this schema decision urgent rather than deferrable:
- **Sparse vectors cannot be added to an existing unnamed-dense-vector
collection** — doing so requires recreating the collection. If we start
dense-only and add sparse/late-interaction vectors later, that's a
disruptive migration.
- **Co-locating large late-interaction (multivector) representations with
dense vectors in the same segment degrades *all* queries**, including
plain dense search, unless the multivector is stored on disk.
We've decided (see project discussion) that:
- Qdrant runs self-hosted via Docker.
- Tenancy is a **single shared collection**, partitioned by payload, not
collection-per-tenant — Qdrant's own multitenancy guidance is explicit that
per-tenant collections "do not scale past a few hundred and waste
resources," while payload partitioning scales to 10k+ tenants.
- The dense embedding model is now decided — `nomic-embed-text-v2-moe`, see
[0004](0004-docx-csv-chunking-strategy.md). Sparse and late-interaction
models are **not yet chosen**; those remain swappable via FastEmbed-
compatible config.
- The exact payload field list beyond the baseline below is still open for
discussion; this ADR proposes a starting schema, not a final one.
## Decision
### Collection
One collection, e.g. `chunks`, shared by all tenants and domains.
### Vectors (named vectors, defined at creation time)
| Name | Type | Purpose | Notes |
|---|---|---|---|
| `dense_nomic` | dense vector | primary semantic similarity (multilingual, incl. Persian) | `nomic-embed-text-v2-moe`, 768-dim ([0004](0004-docx-csv-chunking-strategy.md)) |
| `dense_openai` | dense vector | second semantic signal | `text-embedding-3-large` at its **native 3072 dimensions** — the `dimensions` param is deliberately left unset (see below) |
| `sparse` | sparse vector | lexical/keyword-sensitive retrieval | `bm25-fa-norm-stop` — Qdrant FastEmbed's BM25 sparse encoder configured for Persian (stopword removal + normalization), not a separately trained model |
| `late_interaction` | multivector | reserved for late-interaction rerank ([0003](0003-agent-hybrid-retrieval.md)) | `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)), `comparator: max_sim`, `hnsw_config: m=0` (rerank-only, never independently ANN-searched), stored **on disk** |
All four are defined **now**, even though `late_interaction` won't be
populated until the agent retrieval work in ADR-0003 lands, specifically to
avoid the forced-recreation problem described above. `late_interaction` is
configured with on-disk storage so its larger footprint doesn't degrade the
dense/sparse query latency. Two dense vectors are provisioned deliberately —
`dense_nomic` and `dense_openai` are two independent semantic signals, both
prefetched and fused at query time (ADR-0003), not a primary/fallback pair.
#### Dense model endpoints and dimensions (resolved by the `emet` benchmark)
Both dense models are reached over the **same OpenAI-compatible
`/embeddings` API**, so one adapter
(`src/infrastructure/embedding/openai_compatible.py`) serves both named
vectors with different configuration:
| Named vector | Model | Endpoint | Dimensions |
|---|---|---|---|
| `dense_nomic` | `nomic-embed-text-v2-moe` | self-hosted Ollama OpenAI-compat shim | **768** (verified against the live endpoint) |
| `dense_openai` | `text-embedding-3-large` | OpenAI hosted API | **3072** (native; `dimensions` unset) |
`dense_openai`'s dimension was previously listed as an open dependency. It is
now pinned to the native 3072, because that is the configuration the `emet`
lab benchmarked — it never passed a `dimensions` argument. Setting it later
would truncate via Matryoshka and is a **re-embedding migration, not a config
tweak**, exactly as the negative consequence below warns.
Two operational notes about the self-hosted embedder, both learned by
measurement rather than assumption:
- **Cold load exceeds 150s**, far beyond `INGESTION_TIMEOUT_SECONDS`, so an
idle-then-upload would return `504`. Mitigated on both ends: Ollama's
`keep_alive` keeps the model resident, and the FastAPI lifespan warms each
dense embedder at startup (fail-soft — a down embedder must not block boot).
- **Once warm it is fast**: ~0.30s for one input and ~0.34s for a batch of 16.
Batching is therefore nearly free, which is what keeps ADR-0017's inline
ingestion viable.
### Multitenancy / indexing config
- HNSW: `m: 0` (disable the global index) + `payload_m: 16`, per Qdrant's
multitenant collection guidance.
- Payload index on `tenant_id`: keyword index with `is_tenant: true` — this
co-locates a tenant's vectors on disk for sequential reads.
- Payload index on `domain`: keyword index (secondary partition dimension
within a tenant).
- Payload index on `file_id`: keyword index, used by CRUD lookups in
ADR-0002 (e.g. "delete all chunks belonging to this file").
- Payload index on `order_id`: **float** index (Qdrant's `order_by` and
`Range` filter conditions only support numeric/datetime payloads, not
keyword/string — see the `order_id` format note below). Used for
`order_by` when listing/scrolling a file's chunks in sequence (ADR-0002)
and for previous/next lookups.
- Payload index on `previous_chunk_id` / `next_chunk_id`: keyword index,
used for O(1) adjacency retrieval (see below).
### Payload schema
This schema is now decided for the fields below. Additional document-context
fields (e.g. page/row position, section heading, effective/expiration dates
for insurance policy documents) are deliberately **deferred to a future ADR**
that will accompany the docx/csv chunking-strategy work — that decision
involves format-specific tradeoffs not yet made.
| Field | Type | Purpose |
|---|---|---|
| `tenant_id` | keyword, `is_tenant: true` index | tenant isolation |
| `domain` | keyword | logical partition within a tenant (e.g. `fire`, `car` for insurance lines) |
| `file_id` | keyword | groups chunks back to their source file (formerly referred to as `document_id` — standardized on `file_id`) |
| `chunk_id` | keyword | stable identifier for a single chunk |
| `content_type` | keyword | classification of the chunk's content; exact value set (e.g. `paragraph`, `table_row`, `heading`) to be finalized alongside the chunking-strategy ADR |
| `source_filename` | keyword | original uploaded filename |
| `source_type` | keyword (`docx` \| `csv`) | which parser produced this chunk |
| `order_id` | float (see below) | chunk's *display* position within the file; mutable so the backend can reorder/insert chunks |
| `chunk_index` | integer | chunk's *original ingestion* ordinal — immutable, used to derive the deterministic point ID below (kept separate from `order_id` precisely because `order_id` can change) |
| `previous_chunk_id` | keyword, nullable | `chunk_id` of the preceding chunk in display order (`null` for the first chunk in a file) — O(1) adjacency pointer for context-window expansion in ADR-0003 |
| `next_chunk_id` | keyword, nullable | `chunk_id` of the following chunk in display order (`null` for the last chunk in a file) — same purpose as `previous_chunk_id` |
| `content` | text (full-text indexed) | the chunk text itself, also used for keyword search in ADR-0002 |
| `is_active` | boolean | soft-delete / visibility flag — inactive chunks are excluded from CRUD listing and agent retrieval but retained for audit |
| `deleted_at` | datetime, nullable | set when `is_active` transitions to `false`; distinguishes "deactivated" from "never active" |
| `created_at` | datetime | ingestion timestamp |
| `updated_at` | datetime | last modification timestamp |
| `created_by` | keyword | user/service that created the chunk |
| `updated_by` | keyword | user/service that last modified the chunk |
| `version` | integer | optimistic-concurrency counter, used in ADR-0002 |
| `content_hash` | keyword | hash of the chunk's raw text; lets re-ingestion detect unchanged content and skip re-embedding it |
| `embedding_model_version` | keyword | identifies which embedding model(s) produced this chunk's vectors; needed to know which chunks require re-embedding after a future model swap |
#### `order_id` format
`order_id` is a **fractional float key** (e.g. `1.0`, `2.0`, `3.0`, ...),
not a plain sequential integer and not a lexicographic string. We are not
implementing chunk insertion/reordering yet, but choosing this format now
means that when that feature is added, inserting a chunk between two
existing ones (e.g. assigning it `1.5`, then `1.25` for a subsequent insert
in the same gap) only touches that one chunk's payload — it never requires
renumbering every subsequent chunk, the same benefit a sortable string would
give.
A string key was considered first but **does not work in Qdrant**: the
`Range` filter condition (`gt`/`gte`/`lt`/`lte`) only supports float/integer
payloads (datetime gets its own separate `DatetimeRange` condition), and
`order_by` on scroll requires a payload index that supports `Range`
filtering — so `order_by` is likewise limited to numeric/datetime fields.
A keyword/string `order_id` could only be sorted client-side after fetching
every chunk for a file, and couldn't support a targeted previous/next query
at all. A float key gets native `Range`/`order_by` support instead.
**Known limitation**: repeatedly inserting into the exact same gap (~50+
times between the same two neighbors) runs into floating-point precision
limits. Mitigate with an occasional rebalance job that respaces a file's
`order_id` values (e.g. back to `1000, 2000, 3000, ...`); this is standard
for any fractional-indexing scheme and not expected to be hit in normal use.
#### Previous/next chunk retrieval
Two ways to get a chunk's neighbors, both viable given this schema:
1. **Pointer fields (primary, recommended for the agent path)**: read
`previous_chunk_id`/`next_chunk_id` off the retrieved chunk and fetch
those chunk IDs in a single batch "retrieve points by ID" call. This is
the intended pattern for ADR-0003's context-window expansion, since it
runs on every retrieved chunk and a single batch-get is cheaper than a
filtered query per chunk.
2. **Range query (fallback, only viable because `order_id` is numeric)**:
filter `file_id = X` + `order_id < current` (Range `lt`) + `order_by desc`
+ `limit 1` for the previous chunk; mirror with `gt`/ascending for the
next chunk. Useful if the pointer fields are ever missing/stale, or for
ad-hoc debugging.
Pointer fields are only as correct as the mutation logic that maintains
them — see ADR-0002 for how reorder/insert/delete operations keep
`previous_chunk_id`/`next_chunk_id` in sync.
### Ingestion flow
1. FastAPI upload endpoint receives a `.docx` or `.csv` file.
2. Parse: `python-docx` for `.docx`, `pandas` for `.csv`.
3. Preprocess: clean/normalize extracted text.
4. Chunk: split into chunks (size/overlap are tunable config, not fixed by
this ADR).
5. Embed each chunk into all four vectors: `dense_nomic`, `dense_openai`,
`sparse`, and `late_interaction`. `dense_nomic` uses
`nomic-embed-text-v2-moe` with the chunk text prefixed
`search_document: ` ([0004](0004-docx-csv-chunking-strategy.md));
`dense_openai` uses the OpenAI large embedding model; `sparse` uses
`bm25-fa-norm-stop`; `late_interaction` uses `jina-colbert-v2`
([0005](0005-reranking-model-and-sparse-analyzer-selection.md)) — its
document-side multivector is computed and stored at ingestion time here,
while the query-side multivector is computed per-query in ADR-0003 for
the MAX_SIM rerank comparison.
6. Assign a **deterministic point ID** — UUIDv5 derived from
`file_id` + `chunk_index` (the immutable ingestion ordinal, not the
mutable `order_id`) — so re-ingesting the same file upserts existing
chunks instead of creating duplicates, and re-ordering chunks later never
changes their IDs. Initialize `order_id` from `chunk_index` at ingestion
time (e.g. `chunk_index` 0, 1, 2 → `order_id` `1.0`, `2.0`, `3.0`), and set
`previous_chunk_id`/`next_chunk_id` to each chunk's immediate ingestion-order
neighbor (`null` at the two ends of the file).
7. Batch upsert into Qdrant: 64–256 points per request, 2–4 parallel upload
streams, per Qdrant's bulk-upload guidance.
## Consequences
### Positive
- Defining all four named vectors up front avoids a future forced
collection recreation when late-interaction rerank is added in ADR-0003.
- Payload-partitioned multitenancy scales to large tenant counts without
per-tenant collection sprawl, and shares infrastructure/config across
ADR-0002 and ADR-0003.
- Deterministic point IDs make ingestion idempotent — safe to re-run on the
same document.
- Numeric `order_id` is natively supported by Qdrant's `Range`/`order_by`,
enabling both ordered listing (ADR-0002) and previous/next chunk lookups
(ADR-0003's context-window expansion) without client-side sorting.
### Negative
- The collection carries four named vectors and every chunk is embedded
into all of them at ingestion (including `late_interaction` via
`jina-colbert-v2`), rather than only the vectors actually queried at
launch — more ingestion-time compute than a leaner initial cut.
- Document-context payload fields (page/row position, section heading,
effective/expiration dates, etc.) are still deferred to the
chunking-strategy ADR; adding them later means an additive payload
migration, though it won't touch the fields already decided here.
- Two dense vectors means every chunk is embedded twice (nomic + OpenAI) at
ingestion time and both are queried at retrieval time — roughly double
the dense embedding cost/latency of a single-dense-vector design, plus an
external network dependency on OpenAI's API in the ingestion path.
- ~~`dense_openai`'s exact output dimension is still an open dependency~~ —
**resolved**: pinned to the native 3072 (see "Dense model endpoints and
dimensions" above). The warning still stands for any future change:
re-dimensioning is a re-embedding migration, not a config tweak.
- The `sparse` vector must be created with `modifier="idf"`. The client
computes only BM25's term-frequency saturation; without that modifier
Qdrant applies no IDF at all and lexical retrieval silently degrades
(ADR-0005).
- `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md))
adds a hard GPU dependency to ingestion (not just query time, since the
document-side multivector is computed here) and its commercial license is
still unconfirmed.
- `previous_chunk_id`/`next_chunk_id` are denormalized pointers — every
reorder/insert/delete must update the affected neighbors' payloads too
(see ADR-0002), or the pointers go stale. Fractional float `order_id` also
needs an (infrequent) rebalance job as a long-term maintenance task.
## Alternatives Considered
- **Collection per domain**: rejected. Domains would need materially
different vector configs or strict data-boundary requirements to justify
this; neither applies here, and it fragments tenant-scaling benefits.
- **Collection per tenant**: rejected outright per Qdrant's own guidance —
doesn't scale past a few hundred tenants.
- **Start dense-only, add sparse/late-interaction later**: rejected — Qdrant
requires recreating the collection to add sparse vectors to an unnamed
dense-only collection, which is a disruptive migration we can avoid by
deciding the full vector shape now.