diff --git a/docs/adr/0002-chunk-crud-and-search-api.md b/docs/adr/0002-chunk-crud-and-search-api.md index 1a1d549..5dcd328 100644 --- a/docs/adr/0002-chunk-crud-and-search-api.md +++ b/docs/adr/0002-chunk-crud-and-search-api.md @@ -105,6 +105,62 @@ Qdrant's `update_filter`, giving an optimistic-concurrency-style guard against races between a concurrent ingestion re-run (ADR-0001) and a manual edit through this API. +### Re-embedding on content edit + +`PUT /points/{point_id}` can change `content`, which leaves the stored +vectors stale unless they're recomputed. When `content` changes, the point +is **re-embedded inline**, reusing the same async embedding ports and +batch/semaphore bounds ingestion uses ([0017](0017-synchronous-ingestion-in-the-request-path.md)), +for parity between the two write paths. When `content` is unchanged, the +edit applies only the supplied vector/payload fields and skips re-embedding +entirely. Failure modes on this path reuse ingestion's status codes: `502` +on embedder failure, `504` if the edit's embedding step exceeds the same +timeout budget class as ingestion. The `version` guard (`update_filter`) +still applies to the write — re-embedding happens before the guarded write, +not instead of it, so a stale-version edit still fails with `409` rather +than re-embedding for nothing. + +Rejected alternatives: requiring the caller to supply vectors when content +changes (pushes model knowledge onto the client, and is easy to get subtly +wrong); marking the point stale for background re-embedding later (needs +background work, which ADR-0017 currently rules out for this slice). + +### `order_id` gap exhaustion + +Repeatedly inserting into the same gap between two neighbors eventually +exhausts float precision (ADR-0001's known limitation). This slice does +**not** ship a renormalize endpoint. Instead, any operation that assigns a +new fractional `order_id` between two neighbors (insert, reorder) computes +the resulting gap and: + +- logs a structured warning (`points.order_id.gap_low`) with `file_id` and + the two neighbor point IDs once the gap falls under a defined safety + threshold, so the condition is observable before it becomes uninsertable; +- **rejects** the write with `409` and a distinct error code if the + computed gap is no longer representable (would collapse to one of the two + neighbor values), instead of silently applying an imprecise value. + +Recovering from an exhausted gap is a manual data-fix operation covered by +the operator runbook, not an endpoint this slice builds — deferring the +renormalize primitive is acceptable, silently producing an unrepresentable +gap is not. + +### `POST /points/batch` semantics + +Batch requests are **all-or-nothing**, capped at **100 operations per +request**. The service layer validates every operation's `version` +precondition before applying any of them; if any operation's precondition +fails, the whole request is rejected with `409` and nothing is applied — no +partially-applied batch ever reaches Qdrant. This follows directly from the +`version`-guard rule above applied at the batch level, and from the +pointer-relinking rule (a reorder/insert/delete's neighbor updates must land +in the same `points/batch` call, and a partial relink is a defect): partial +application of a batch is exactly the failure mode that would produce a +stale pointer chain. The 100-operation cap is independent of ADR-0001's +64–256-point bulk-ingestion batch sizing — that number is about upload +throughput; this one bounds an admin/manual edit request to something that +comfortably finishes inside a normal request timeout. + ## Consequences ### Positive @@ -122,6 +178,15 @@ edit through this API. - Optimistic concurrency via `version` requires every writer (ingestion and this API) to consistently read-check-write; a writer that skips this can silently clobber concurrent edits. +- Inline re-embedding puts embedder latency and `502`/`504` failure modes on + an admin content edit, not just on ingestion — an edit that only intended + to fix a typo pays the same embedding cost as a fresh chunk. +- Deferring the `order_id` renormalize endpoint means a file whose gaps are + genuinely exhausted has no automated recovery in this slice; an operator + must intervene by hand until that endpoint exists. +- All-or-nothing batch semantics mean one stale operation in a 100-operation + batch fails the entire request, even when the other 99 operations are + independent and would have succeeded on their own. ## Alternatives Considered @@ -132,3 +197,16 @@ edit through this API. - **Client-supplied `tenant_id` in request body**: rejected — trusting client input for the isolation boundary is a direct multitenancy security risk; it must come from server-side auth context. +- **Caller-supplied vectors on content edit**: rejected — pushes embedding + model knowledge onto the client and makes it easy to silently desync + vectors from content. +- **Mark-stale-and-re-embed-later on content edit**: rejected for this + slice — needs background work, which ADR-0017 currently rules out. +- **Renormalize `order_id` automatically within this slice**: rejected — + nothing in current scope has hit gap exhaustion; building the primitive + now is speculative. Revisit if the logged warning starts firing in + practice. +- **Partial-success batch semantics (per-operation status)**: rejected — + a partially-applied batch is exactly the failure mode that leaves the + pointer chain (`previous_chunk_id`/`next_chunk_id`) inconsistent, which + this ADR treats as a defect, not a degraded-but-acceptable outcome. diff --git a/docs/adr/0008-rest-api-and-fastapi-boundary.md b/docs/adr/0008-rest-api-and-fastapi-boundary.md index 7712e37..7c52c3f 100644 --- a/docs/adr/0008-rest-api-and-fastapi-boundary.md +++ b/docs/adr/0008-rest-api-and-fastapi-boundary.md @@ -164,16 +164,20 @@ metadata such as `domain`. File validation is server-side: - Derive `tenant_id`, `created_by`, and `updated_by` from `AuthContext`, not from form fields. -Ingestion may be slow because it parses, chunks, embeds, and writes many -Qdrant points. The REST contract is job-shaped even if the first -implementation runs inline: +Ingestion parses, chunks, embeds, and writes many Qdrant points. +[ADR-0017](0017-synchronous-ingestion-in-the-request-path.md) supersedes the +job-shaped contract originally specified here: ingestion runs inline and the +response is terminal. ```text -202 Accepted -> { file_id, ingestion_job_id, status: "queued" | "running" } +201 Created -> { file_id, ingestion_job_id, status: "succeeded", chunks_indexed } ``` -A durable worker/job queue can be added later without changing the API -contract. +`ingestion_job_id` is retained so the attempt stays inspectable via +`GET /v1/files/{file_id}`, and so a future move back to a queued `202 Accepted` +contract (ADR-0014) is additive for clients that already read it. Ingestion +failures are HTTP failures: `400` unparseable, `413` too large, `502` embedder +failure, `503` at capacity, `504` past the ingestion timeout. ### Point endpoints replace the older `/chunks` sketches @@ -228,7 +232,7 @@ Important status codes: | Status | Use | |---|---| -| `202 Accepted` | Ingestion accepted as a job. | +| `201 Created` | Ingestion completed inline (ADR-0017). | | `400 Bad Request` | Invalid domain/filter combinations or unsupported file type. | | `401 Unauthorized` | Missing/invalid API key. | | `403 Forbidden` | Valid key without required scope. | @@ -252,8 +256,8 @@ and Qdrant operations can be correlated. records — while preserving the chunk payload schema underneath. - `/threads/{thread_id}/runs` remains compatible with the LangGraph thread/run model already chosen in ADR-0007. -- Job-shaped file ingestion lets the first implementation be simple while - keeping room for a durable worker without breaking clients. +- Inline file ingestion (ADR-0017) gives callers a terminal result in one + request, with failures surfaced as ordinary HTTP errors. - Router-level dependencies and typed FastAPI dependencies keep auth, tenant resolution, sessions, and scopes reusable instead of repeated per endpoint. @@ -266,8 +270,9 @@ and Qdrant operations can be correlated. public product API. - API-key auth in Postgres adds a database lookup to every request unless short-lived caching is introduced. Caching must preserve revocation semantics. -- A job-shaped ingestion contract needs a job status store even if the initial - implementation processes inline. +- Inline ingestion ties the upload's duration to proxy/client timeouts, and + moving back to a queued `202` contract later is a breaking change for clients + (see ADR-0017's trigger list). The `ingestion_jobs` store is kept either way. - The REST layer now depends on the tenant/API-key, ingestion-job, audit, and usage tables defined in [ADR-0009](0009-postgres-sqlalchemy-alembic-schema.md). @@ -288,10 +293,12 @@ and Qdrant operations can be correlated. - **Expose one generic `/v1/qdrant/*` proxy**: rejected. It would leak Qdrant's full API surface, bypass tenant/scoping rules too easily, and couple clients to storage operations the service should hide. -- **Synchronous file ingestion only**: rejected as the contract. It is simpler - to implement, but embedding and late-interaction vector generation can be - slow enough to exceed HTTP timeouts. The job-shaped response gives the - implementation room to evolve. +- **Synchronous file ingestion only**: originally rejected here on the grounds + that embedding and late-interaction vector generation can exceed HTTP timeouts; + **adopted** by ADR-0017 for the first slice. Dense embedding is async network + I/O that batches and runs concurrently, and late-interaction vectors are not + populated at ingest yet — which is what made the original objection decisive + and is now the named trigger for reverting to a job-shaped contract. - **Create threads with `POST /v1/threads`**: rejected for now. The main backend owns conversation/session records, and LangGraph can create a checkpoint sequence on the first run for a `thread_id`. A create endpoint diff --git a/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md b/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md index ac1a954..d0196a1 100644 --- a/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md +++ b/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md @@ -47,6 +47,7 @@ Use SQLAlchemy 2.x ORM models with typed mappings: class Base(DeclarativeBase): pass + class Tenant(Base): __tablename__ = "tenants" @@ -208,9 +209,12 @@ on the `content_hash` policy. #### `ingestion_jobs` -One attempt to parse/chunk/embed/upsert a source file. This table is required -even if the first implementation processes inline, because ADR-0008's file -upload contract is job-shaped. +One attempt to parse/chunk/embed/upsert a source file. Under +[ADR-0017](0017-synchronous-ingestion-in-the-request-path.md) that attempt runs +inline in the upload request, so a row is written `running` before the work and +updated to a terminal status after it — the table is a durable record of the +attempt, not a queue. It is what makes failures inspectable, re-ingestion +idempotent, and a later move back to queued dispatch (ADR-0014) additive. | Column | Notes | |---|---| diff --git a/docs/adr/0012-application-resource-lifetime-and-dependency-ownership.md b/docs/adr/0012-application-resource-lifetime-and-dependency-ownership.md index 875fd73..9acf426 100644 --- a/docs/adr/0012-application-resource-lifetime-and-dependency-ownership.md +++ b/docs/adr/0012-application-resource-lifetime-and-dependency-ownership.md @@ -165,6 +165,7 @@ Avoid hidden global access: # Do not do this. session = SessionLocal() + async def create_user(data: CreateUserRequest) -> User: session.add(User(email=data.email)) await session.commit() @@ -206,8 +207,7 @@ async def get_point( qdrant: QdrantClient, auth: AuthContext, point_id: str, -) -> PointResponse: - ... +) -> PointResponse: ... ``` Do not have lower layers import mutable resource singletons. Explicit parameters diff --git a/docs/adr/0014-durable-job-dispatch-with-rabbitmq.md b/docs/adr/0014-durable-job-dispatch-with-rabbitmq.md index 34163fe..d4d0149 100644 --- a/docs/adr/0014-durable-job-dispatch-with-rabbitmq.md +++ b/docs/adr/0014-durable-job-dispatch-with-rabbitmq.md @@ -2,7 +2,14 @@ ## Status -Proposed +Superseded by +[0017](0017-synchronous-ingestion-in-the-request-path.md) + +ADR-0017 defers this decision rather than rejecting it: ingestion currently runs +inline in the `POST /v1/files` request, with no broker, no outbox, and no queue. +This ADR remains the intended design for when ingestion becomes slow enough to +need one — ADR-0017 lists the triggers, and names an in-process Postgres-claimed +runner as the likely intermediate step before adopting a broker. ## Context diff --git a/docs/adr/0015-modular-monolith-package-architecture.md b/docs/adr/0015-modular-monolith-package-architecture.md index c6171bd..73dda16 100644 --- a/docs/adr/0015-modular-monolith-package-architecture.md +++ b/docs/adr/0015-modular-monolith-package-architecture.md @@ -4,6 +4,13 @@ Proposed +> Amended by [ADR-0017](0017-synchronous-ingestion-in-the-request-path.md): +> there is no broker, outbox publisher, queue, or separate worker process for +> now. `messaging/`, `workers/`, and `infrastructure/rabbitmq/` are part of the +> target shape but are **not created yet**; ingestion runs inline in the request, +> so the FastAPI route is the only entry adapter and it calls +> `application/ingestion/` directly. Every layering rule below applies unchanged. + ## Context The repository currently contains only a small FastAPI-oriented scaffold under @@ -69,17 +76,18 @@ src/ │ │ ├── models/ │ │ ├── repositories/ │ │ ├── database.py -│ │ └── outbox.py +│ │ └── outbox.py # ADR-0017: not yet; no outbox today │ ├── qdrant/ │ ├── minio/ -│ ├── rabbitmq/ +│ ├── embedding/ # async embedding adapters (batched, bounded) +│ ├── rabbitmq/ # ADR-0017: not yet; ADR-0014 target only │ ├── langgraph/ │ └── observability/ -├── messaging/ +├── messaging/ # ADR-0017: not yet; ADR-0014 target only │ ├── events.py │ ├── subjects.py │ └── outbox_publisher.py -└── workers/ +└── workers/ # ADR-0017: not yet; ADR-0014 target only ├── ingestion.py └── maintenance.py ``` @@ -103,7 +111,6 @@ tests/ ├── integration/ │ ├── postgres/ │ ├── minio/ -│ ├── rabbitmq/ │ └── qdrant/ └── e2e/ ``` @@ -193,8 +200,10 @@ bootstrap. Graph nodes call application services, particularly - `qdrant/` owns Qdrant client lifecycle, collection/bootstrap helpers, low-level point operations, and hybrid retrieval adapter mechanics. - `minio/` implements object-storage operations against MinIO/S3-compatible APIs. +- `embedding/` implements the async dense/sparse embedding adapters, including + provider batching and the concurrency semaphore from ADR-0017. - `rabbitmq/` owns the RabbitMQ connection/channel lifecycle plus low-level - publish and consumer adapters (aio-pika). + publish and consumer adapters (aio-pika). Not created while ADR-0017 stands. - `langgraph/` configures the concrete Postgres-backed LangGraph persistence adapters. - `observability/` configures structlog and Langfuse integrations. @@ -204,6 +213,12 @@ must not create mutable external clients at import time. ### Messaging and workers +Under ADR-0017 neither package exists yet: ingestion runs inline in the request, +so the FastAPI route plays the entry-adapter role described here and obeys the +same rule — it binds logging context and invokes an application service, and +holds no parsing/chunking/Qdrant business logic itself. The rest of this section +describes the shape both packages take when ADR-0014 is adopted. + `messaging/` contains versioned event schemas, stable routing-key names, and the outbox-publisher orchestration. The outbox publisher coordinates Postgres outbox records with the RabbitMQ adapter; it does not become a second source of job diff --git a/docs/adr/0016-testing-strategy-and-quality-gates.md b/docs/adr/0016-testing-strategy-and-quality-gates.md index ee64afb..8b12c33 100644 --- a/docs/adr/0016-testing-strategy-and-quality-gates.md +++ b/docs/adr/0016-testing-strategy-and-quality-gates.md @@ -4,17 +4,24 @@ Proposed +> Amended by [ADR-0017](0017-synchronous-ingestion-in-the-request-path.md): +> there is no broker, outbox, or queue, so the `rabbitmq` marker and +> `tests/integration/rabbitmq/` are not carried until ADR-0014 is adopted. +> Ingestion is exercised through the upload request itself, which now returns a +> terminal result. Every reliability invariant below still applies — retrying an +> upload stands in for redelivery. + ## Context The project has ADRs for tenant-scoped ingestion, explicit resource ownership, -MinIO object storage, transactional outbox dispatch, RabbitMQ workers, +MinIO object storage, inline request-path ingestion, Qdrant indexing, and a modular monolith. It has no test runner, test fixtures, or executable test suite yet. The first CSV ingestion slice has correctness properties that cannot be left to manual testing: Alembic is the only schema-management path; tenant identity is -trusted server-side context; a file, job, and outbox event commit atomically; -workers are safe under at-least-once delivery; and generated Qdrant points are +trusted server-side context; the source file and its job row commit +atomically; job execution is safe under at-least-once semantics; and generated Qdrant points are idempotent and tenant-filtered. ADR-0015 already reserves a test layout by boundary, while ADR-0012 requires explicit dependencies and resource lifetimes that should make tests practical without import-time client patching. @@ -23,8 +30,11 @@ Tests need to give fast feedback during implementation without replacing integration coverage with mocks or making routine development depend on Docker, provider credentials, live models, or Langfuse availability. -ADR-0014's transactional-outbox decision controls ingestion dispatch. The upload -path records durable dispatch intent; a separate outbox publisher publishes it. +ADR-0017 controls ingestion: `POST /v1/files` parses, chunks, embeds, and +indexes inline, then returns a terminal `201`. `ingestion_jobs` records the +attempt. There is no outbox, queue, or publisher to test — but the request's +bounds (size, timeout, capacity) and its two-transaction shape are testable +contracts. ## Decision @@ -42,8 +52,8 @@ ADR-0012. `FastAPI.TestClient` is not the default project test client. Register these markers: - one primary boundary marker per test: `unit`, `integration`, or `e2e`; -- `postgres`, `minio`, `rabbitmq`, or `qdrant` for the real service used by an - integration test; +- `postgres`, `minio`, or `qdrant` for the real service used by an integration + test (a `rabbitmq` marker returns with ADR-0014); - `slow` only where a test materially exceeds the normal integration feedback target; - `live_provider` for an opt-in, credential-gated external-provider smoke test. @@ -68,7 +78,6 @@ tests/ ├── integration/ │ ├── postgres/ │ ├── minio/ -│ ├── rabbitmq/ │ └── qdrant/ └── e2e/ ``` @@ -84,7 +93,7 @@ tests/ tests. Hand-written fakes and spies implement narrow application-owned ports, not -MinIO, RabbitMQ, Qdrant, or model SDK-shaped interfaces. Scripted model, embedder, +MinIO, Qdrant, or model SDK-shaped interfaces. Scripted model, embedder, retrieval, clock, and UUID fakes make normal test runs deterministic. ### Apply pragmatic TDD @@ -106,7 +115,7 @@ real-adapter integration test before declaring that boundary complete. ### Use disposable real infrastructure in integration tests Use Testcontainers as the standard automated integration-test resource mechanism -for Postgres, MinIO, RabbitMQ, and Qdrant. +for Postgres, MinIO, and Qdrant. - Tests never connect to a developer's local services or Langfuse-owned storage and credentials. @@ -119,8 +128,8 @@ for Postgres, MinIO, RabbitMQ, and Qdrant. adapters through their normal constructors. Docker Compose remains the mechanism for manual local validation and a later, -serialized operational smoke test where web, outbox-publisher, and worker run as -independent processes. It is not the default pytest fixture mechanism. +serialized operational smoke test of the running web process, which performs +ingestion inline under ADR-0017. It is not the default pytest fixture mechanism. ### Treat invariants as reusable contracts @@ -128,21 +137,25 @@ Test the following requirements at the applicable application, adapter, and E2E boundaries: - Tenant identity comes from server-side authenticated context. Request payloads, - query parameters, object metadata, and broker messages cannot override it. + query parameters, object metadata, and job payloads cannot override it. - Cross-tenant access does not disclose tenant-owned data. Public routes normally return `404` for inaccessible resources. - Alembic creates the schema from an empty database. Tests never use `Base.metadata.create_all()`, and FastAPI startup performs readiness checks only, never DDL. -- The upload transaction records `source_files`, a queued `ingestion_jobs` row, - and an unpublished `outbox_events` row atomically. The HTTP route does not - directly publish the ingestion event. -- Broker messages contain durable identifiers and correlation metadata only. The - worker reloads job and source-file records from Postgres before tenant-scoped - side effects. -- Worker acknowledgement follows durable progress or terminal-state persistence. - Duplicate publication and redelivery do not regress terminal jobs, inflate - counters, or create duplicate logical chunks. +- The first upload transaction records `source_files` and a `running` + `ingestion_jobs` row atomically, and commits before any parse/embed work; no + session or transaction is held open across that work. +- Every terminating path — success, parse failure, embedder failure, timeout — + writes a terminal job status and its `ingestion_job_events` row. A job is + never left in `running` by a handled failure. +- Each bound maps to its status code: oversized upload `413`, capacity `503`, + timeout `504`, embedder failure `502`. +- Retrying an upload does not regress terminal jobs, inflate counters, or create + duplicate logical chunks; identical content is recognized rather than + re-ingested. +- Embedding is batched and concurrency-bounded rather than serial per chunk, and + blocking work is executed off the event loop under an explicit limiter. - MinIO keys are server-derived internal paths. Qdrant reads and mutations use a server-derived tenant filter, deterministic point IDs, and upsert semantics. @@ -176,8 +189,8 @@ ratcheting threshold rather than encouraging low-value coverage. - Unit tests provide fast, deterministic TDD feedback for core application behavior. - Real-service tests cover the behaviors least safe to simulate: Alembic - migrations, object storage, RabbitMQ acknowledgements/redelivery, and Qdrant - filtering/upserts. + migrations, object storage, transaction boundaries under real sessions, and + Qdrant filtering/upserts. - Explicit fakes reinforce the dependency direction and resource ownership rules from ADR-0012 and ADR-0015. - The ingestion path has concrete tenant-isolation and reliability contracts, @@ -197,8 +210,8 @@ ratcheting threshold rather than encouraging low-value coverage. ## Alternatives Considered - **Mock all external SDKs**: rejected. Mocks cannot prove migrations, real - RabbitMQ acknowledgement/redelivery behavior, MinIO semantics, or Qdrant - tenant filtering. + transaction/connection behavior, MinIO semantics, or Qdrant tenant + filtering. - **Use full-stack Compose tests only**: rejected. They are slow and opaque for the default development loop and make failures difficult to localize. - **Run all integration containers on every pytest invocation**: rejected. Test diff --git a/docs/adr/0017-synchronous-ingestion-in-the-request-path.md b/docs/adr/0017-synchronous-ingestion-in-the-request-path.md new file mode 100644 index 0000000..ef59aba --- /dev/null +++ b/docs/adr/0017-synchronous-ingestion-in-the-request-path.md @@ -0,0 +1,252 @@ +# 0017. Synchronous ingestion in the request path + +## Status + +Proposed + +## Context + +ADR-0014 chose RabbitMQ plus a transactional outbox for durable job dispatch, +with three process entrypoints (web, outbox publisher, ingestion worker). +ADR-0008 made `POST /v1/files` job-shaped — `202 Accepted` with an +`ingestion_job_id` and a non-terminal status — explicitly so that a durable +queue could be added later without breaking clients. + +That contract was chosen on the assumption that ingestion is slow enough to +exceed HTTP timeouts. For the first slice, that assumption does not hold, and +paying for it costs a broker, an outbox table, a publisher process, a worker +process, and a client that must poll for completion. + +What ingestion actually does per file, and how expensive each stage is: + +| Stage | Kind of work | Cost | +|---|---|---| +| Parse (`python-docx`, `csv`) | Blocking CPU, sync SDK | Modest; seconds at worst for a large document | +| Chunk (fixed-size, ADR-0004) | Blocking CPU, pure Python | Negligible | +| `dense_nomic`, `dense_openai` | **Async network I/O** | Dominant by wall-clock, but batchable and concurrent | +| `sparse` (own BM25 pipeline, ADR-0005) | Blocking CPU | Modest, GIL-bound | +| Qdrant upsert | Async network I/O | Modest, batchable | + +The dense embedding stage is the largest term, but it is I/O, not computation +this process performs: both embedders take arrays of inputs, and independent +batches can be in flight at once. A few hundred chunks is a small number of +batched requests, issued concurrently — not a serial per-chunk round trip. + +`late_interaction` (jina-colbert-v2) is the one genuinely heavy ingest-time +stage: a document-side multivector per chunk from a local GPU model requiring +`flash_attn`. ADR-0001 defines the vector now to avoid collection recreation but +does not require it to be populated, and plan 001 puts the reranker and GPU +deployment explicitly out of scope for the first slice. It is therefore **not +computed during synchronous ingestion**, and enabling it is a trigger to revisit +this decision rather than a cost this decision has to carry. + +## Decision + +**`POST /v1/files` performs ingestion inline and returns a terminal result.** No +broker, no outbox, no queue, no worker process, and no polling by the client. + +```text +POST /v1/files + -> authenticate, resolve tenant, validate the upload + -> 201 Created { file_id, ingestion_job_id, status: "succeeded", chunks_indexed } +``` + +ADR-0008's job-shaped `202 Accepted` contract is superseded by this ADR; +ADR-0014 remains `Superseded by 0017` and is the design to adopt if the triggers +below are hit. `ingestion_jobs` is kept exactly as ADR-0009 defines it — it is +now a record of an ingestion *attempt* rather than a queue entry, and it is what +makes failures inspectable and re-ingestion idempotent. + +### Request shape + +The request runs in three phases, and the phase boundaries matter: + +```text +1. txn A (short): insert source_files + insert ingestion_jobs(status='running', started_at=now()) + commit -- release the connection +2. no transaction: store bytes in MinIO + parse + chunk (threads) + embed dense + sparse (async, batched, bounded) + upsert Qdrant points (deterministic ids) +3. txn B (short): update ingestion_jobs -> succeeded/failed + counters + append ingestion_job_events + commit +``` + +**Never hold a Postgres session or transaction open across phase 2.** A slow +upload would otherwise pin a pool connection — and an idle-in-transaction row +lock — for the entire ingestion. Take a session, commit, release, and take a +fresh one for phase 3. This is ADR-0012's "one session per unit of work" applied +to a request that contains two units. + +Committing the job row *before* the work means a request that dies mid-ingestion +still leaves durable evidence: a `running` job that never reached a terminal +state, discoverable by age. + +### Embedding is concurrent and batched, not serial + +Chunks are embedded through the async embedding ports, batched per provider +limits, with independent batches in flight concurrently and bounded by a +semaphore: + +```python +limiter = asyncio.Semaphore(settings.ingestion.embed_concurrency) + +async def embed_batch(batch: Sequence[str]) -> list[Vector]: + async with limiter: + return await embedder.embed(batch) + +batches = chunk_into_batches(texts, size=settings.ingestion.embed_batch_size) +vectors = flatten(await asyncio.gather(*(embed_batch(b) for b in batches))) +``` + +Rules: + +- **Batch before you parallelize.** Both dense embedders accept arrays; sending + one request per chunk wastes far more time than concurrency can recover. +- **Bound concurrency with a semaphore**, never an unbounded `gather` over every + batch. The limit exists for the provider's rate limits and for the self-hosted + nomic server's capacity — past saturation, extra concurrency just moves the + queue somewhere you cannot see it. +- Retry `429`/transient failures with backoff *inside* the request's overall + timeout budget, not beyond it. +- The two dense embedders are themselves independent and run concurrently with + each other. + +### Blocking work runs on threads + +Parsing (`python-docx`, `csv`), chunking, hashing, the sparse BM25 pipeline, and +the synchronous `minio` SDK are blocking. They run via +`anyio.to_thread.run_sync`, with an explicit `anyio.CapacityLimiter` so ingestion +threads cannot exhaust the pool Starlette uses for sync route handlers and +dependencies: + +```python +chunks = await anyio.to_thread.run_sync(parse_and_chunk, raw_bytes, limiter=ingestion_limiter) +``` + +Calling any of them directly from an `async def` service is a defect: one large +`python-docx` parse would block every concurrent request in the process. + +### The request is bounded, and says so when it cannot finish + +Synchronous ingestion means the client's timeout is now the system's deadline. +Three bounds, all configured, all enforced server-side: + +- `INGESTION_MAX_UPLOAD_SIZE_MB` (and a max chunk count) reject work that is + obviously too large *before* any of it starts — a `413`, not a timeout. +- `INGESTION_TIMEOUT_SECONDS` bounds the whole of phase 2. On expiry the job is + marked `failed` with a timeout error code in phase 3, and the response is + `504`. A timeout must never leave the job stuck in `running`. +- `INGESTION_MAX_CONCURRENCY` bounds how many ingestions run in the process at + once. Over the limit, the request is rejected with `503` and `Retry-After` + rather than queued behind an unbounded wait — a queue that the client is + blocked on is the worst of both designs. + +Document the deployment consequence: proxy, load balancer, and client read +timeouts must all exceed `INGESTION_TIMEOUT_SECONDS`, or the client will give up +on work that is still succeeding. + +### Re-running an ingestion stays safe + +The idempotency requirements survive, because a client that times out will +retry, and phase 2 has no transaction protecting it: + +- Point ids are deterministic from `file_id` + `chunk_index` (ADR-0001), so a + retried upload upserts rather than duplicates. +- Identical uploads are recognized by `(tenant_id, domain, content_sha256)` and + return the existing file/job rather than re-ingesting (plan 001). +- `tenant_id` comes from `AuthContext`, never from the request body. +- A terminal job is never transitioned back to `running`. +- Qdrant points from a failed attempt do not replace the previous successful + index; replacement happens only after a successful attempt. + +### Failures are HTTP failures + +There is no dead-letter queue and no retry loop. A failed ingestion marks the +job `failed` with an error code and a terminating `ingestion_job_events` row, +and returns the corresponding status code (`400` for an unparseable file, `413` +too large, `503` at capacity, `504` on timeout, `502` for an embedder failure). +The client decides whether to retry, which is the correct owner of that decision +when the client is synchronous. + +### `GET /v1/files/{file_id}` stays + +It reports the stored job status. It is no longer a polling mechanism for the +upload, but it remains how an operator inspects a past ingestion, and how a +client that lost its connection mid-upload discovers what happened. + +### When to revisit + +Return to a queue — ADR-0014's design, or an in-process Postgres-claimed runner +as an intermediate step — when any of these becomes true: + +- `late_interaction` document vectors are populated at ingest (GPU, per-chunk + model inference: this alone is likely sufficient); +- typical ingestion approaches `INGESTION_TIMEOUT_SECONDS`, or `504`/`503` rates + stop being negligible; +- file types arrive that are large or slow enough to be unbounded (bulk XLSX, + scanned PDFs with OCR); +- ingestion load starts degrading chat/retrieval latency in the same process; +- ingestion needs to be retried automatically rather than by the caller. + +Track p95 ingestion duration, `503`/`504` counts, and the count of `running` +jobs older than the timeout. Those are the trigger metrics. + +## Consequences + +### Positive +- The client gets its answer in one call. No polling, no job-status endpoint in + the happy path, no "queued" state to explain to a frontend. +- The first slice needs Postgres, MinIO, and Qdrant only — no broker, no outbox + table, no publisher or worker process, no exchange/queue/DLX provisioning. +- Failures reach the caller directly, with a real status code and message, + instead of being discovered later by polling a job row. +- One process to run, deploy, and reason about; `docker compose up` is the whole + application. +- Batched, bounded-concurrent embedding is a property worth having regardless of + where ingestion runs — it carries over unchanged into a queued design. + +### Negative +- The request's duration is now a product constraint. Proxy/client timeouts + become deployment configuration that can silently break uploads. +- Ingestion competes with chat/retrieval for CPU, threads, and connections in + the same process, and a burst of uploads degrades API latency for everyone. +- No automatic retry: a transient embedder failure surfaces as a `502` and + depends on the caller to retry. +- A client disconnect does not cancel the work, and leaves a job that can sit in + `running` until observed by age. +- Capacity rejection (`503`) is a worse experience than queueing for callers who + would rather wait — accepted deliberately, because a blocked client waiting on + a hidden queue is worse still. +- Reversing this decision changes the HTTP contract (`201` with a terminal status + becomes `202` with a pending one), so clients would have to change. This is the + real cost of the decision, and the reason the trigger list above is explicit. + +## Alternatives Considered + +- **Job-shaped `202 Accepted` with a queue (ADR-0008 as written, ADR-0014)**: + rejected for now. It is the right design once ingestion is genuinely slow or + needs automatic retry, and the trigger list says when to adopt it. Building it + first would add a broker, an outbox, and two processes to make a + seconds-long operation asynchronous. +- **In-process Postgres-claimed job runner with `202`**: rejected as the current + step, but it is the natural intermediate — it keeps the single process and the + single deployment while decoupling the request from the work. It is the first + thing to reach for when the triggers fire, before adopting a broker. +- **Serial per-chunk embedding**: rejected. It is the version of synchronous + ingestion that genuinely would exceed HTTP timeouts, and it is avoidable with + batching plus bounded concurrency. +- **Unbounded `asyncio.gather` over all batches**: rejected. It converts a large + upload into a rate-limit burst against the provider and an unbounded memory + spike locally; the semaphore is what makes the concurrency safe. +- **Run the blocking stages on the event loop directly**: rejected. A single + large `python-docx` parse would stall every concurrent request in the process. +- **`BackgroundTasks` after returning `201`**: rejected as the worst of both — + the client is told the work succeeded before it has, with no durable record + and no retry if the process restarts. +- **Populate `late_interaction` during synchronous ingestion**: rejected for this + slice, consistent with plan 001's scope. Per-chunk GPU model inference is the + stage that makes ingestion unbounded; adding it is a trigger to revisit this + ADR, not something to absorb into a request. diff --git a/docs/plans/001-ingestion-vertical-slice.md b/docs/plans/001-ingestion-vertical-slice.md index 43e7d2f..1b2b6e0 100644 --- a/docs/plans/001-ingestion-vertical-slice.md +++ b/docs/plans/001-ingestion-vertical-slice.md @@ -4,9 +4,8 @@ This plan turns the accepted architectural direction in the ADRs into the first working product slice: a tenant-scoped CSV upload is stored in MinIO, represented -by durable Postgres records, dispatched through RabbitMQ using a -transactional outbox, processed by a separate worker, and indexed as Qdrant -points. +by durable Postgres records, parsed/chunked/embedded inline in the request +(ADR-0017), and indexed as Qdrant points before the response returns. This is an implementation plan, not an Architecture Decision Record. ADRs explain why major technologies and boundaries were chosen; this document defines the @@ -19,10 +18,9 @@ The first vertical slice uses these responsibilities: | System | Responsibility | |---|---| | FastAPI | HTTP boundary, validation, authentication, tenant derivation, and job creation. | -| Postgres | Tenant/auth data, source-file metadata, ingestion job state/progress, audit, and transactional outbox events. | +| Postgres | Tenant/auth data, source-file metadata, ingestion attempt records/progress, and audit. | | MinIO | Private source-file bytes and retained derived ingestion blobs. | -| RabbitMQ | Durable delivery of ingestion and maintenance work. | -| Ingestion worker | Parsing, chunking, embedding, ingestion-generated Chunk/Point CRUD, and job status updates. | +| Ingestion service | Parsing, chunking, embedding, ingestion-generated Chunk/Point CRUD, and job status updates — inline in the request. | | Qdrant | Tenant-filtered generated chunks and their vectors/payloads. | The controlling ADRs are: @@ -35,9 +33,11 @@ The controlling ADRs are: resource lifetime, dependency injection, and explicit transaction ownership. - [ADR-0013](../adr/0013-s3-compatible-object-storage-with-minio.md): MinIO object storage boundary. -- [ADR-0014](../adr/0014-durable-job-dispatch-with-rabbitmq.md): RabbitMQ, - transactional outbox, separate workers, and worker-owned ingestion Chunk/Point - CRUD. +- [ADR-0017](../adr/0017-synchronous-ingestion-in-the-request-path.md): + inline ingestion, batched/bounded-concurrent embedding, + `anyio.to_thread.run_sync` for blocking work, and request bounds. It supersedes + [ADR-0014](../adr/0014-durable-job-dispatch-with-rabbitmq.md), which remains + the design to adopt when a broker becomes necessary. The cited ADRs are currently proposed. Treat them as the implementation baseline only after the project owner accepts them; code should not silently diverge from @@ -50,13 +50,14 @@ them. - `POST /v1/files` for authenticated tenant-scoped **CSV** upload. - File validation, size limits, content hashing, and streaming upload to MinIO. - Alembic-managed Postgres schema for the minimal tenant/auth, source file, - ingestion job, job event, and outbox records needed by this slice. -- Transactional outbox publication of `ingestion.job.created` to RabbitMQ. -- A separate ingestion worker process with a durable RabbitMQ consumer. + ingestion job, and job event records needed by this slice. +- Inline ingestion in `POST /v1/files`, with batched/bounded-concurrent + embedding, thread-offloaded parsing, and enforced size/timeout/capacity + bounds. - CSV parsing and deterministic chunk creation. - Tenant-filtered Qdrant point upserts using deterministic point identifiers. - Job status/progress persistence and `GET /v1/files/{file_id}` status lookup. -- Structured correlation logging at HTTP, outbox, and worker ingress. +- Structured correlation logging at HTTP and ingestion-stage boundaries. - Automated tests for the critical state transitions, redelivery, and tenant boundaries. @@ -70,31 +71,40 @@ them. ingestion. - Full tenant erasure and hard-deletion workflow. - A public download API or presigned object URLs. -- Exactly-once end-to-end processing. The worker must instead be safe under - at-least-once delivery. +- A message broker, transactional outbox, job queue, and separate + worker/publisher processes (ADR-0014, deferred by ADR-0017). +- `late_interaction` (jina-colbert-v2) document vectors at ingest — populating + them is ADR-0017's primary trigger to move ingestion back off the request. +- Exactly-once end-to-end processing. Job execution must instead be safe when a + job runs more than once. ## Required invariants The implementation must preserve these rules from the ADRs: 1. `tenant_id` is derived from trusted authentication context; it is never - accepted from the upload body, query parameters, MinIO metadata, or a broker - message as authority. + accepted from the upload body, query parameters, MinIO metadata, or a + dispatch payload as authority. 2. MinIO stores bytes; Postgres stores metadata, lifecycle state, job progress, - audit records, and dispatch intent. -3. Broker messages contain stable IDs and correlation metadata only. They never + audit records, and the queue. +3. Dispatch payloads contain stable IDs and correlation metadata only. They never contain file bytes, extracted text, chunks, embeddings, secrets, raw prompts, or raw model output. -4. The worker reloads the job and source-file records from Postgres before doing - tenant-scoped work. -5. The worker, not the HTTP publisher, performs parsing, chunking, embedding, and - generated Qdrant Chunk/Point CRUD. +4. Ingestion works from the persisted job and source-file records, not from + request-supplied values. +5. Parsing, chunking, embedding, and generated Qdrant Chunk/Point CRUD live in + the ingestion application service, not in the route handler. Blocking work + runs through `anyio.to_thread.run_sync`, never directly on the event loop. 6. Qdrant reads and mutations are tenant-filtered. Ingestion-generated point IDs are deterministic so retrying a job does not create duplicate logical chunks. -7. The worker acknowledges a RabbitMQ message only after it has durably persisted - the applicable Postgres progress/final state. -8. Application clients are built at FastAPI or worker-process startup and closed - at shutdown. No mutable clients are opened as import-time globals. +7. No Postgres session or transaction is held open across parse/embed/upsert. + The job row is committed `running` before the work and updated to a terminal + status after it, in a second short transaction. +8. Application clients are built at FastAPI startup and closed at shutdown. No + mutable clients are opened as import-time globals. +9. Embedding is batched per provider limits and concurrency-bounded by a + semaphore; blocking work (parse, chunk, BM25, `minio`) runs through + `anyio.to_thread.run_sync` with an explicit `CapacityLimiter`. ## Required decisions before implementing the affected phase @@ -120,22 +130,25 @@ rather than becoming an accidental repository behavior. For the first release, file deletion should be soft and job-shaped: 1. mark the source file as deletion requested/soft deleted in Postgres; -2. write a maintenance outbox event; -3. have a worker soft-delete the related Qdrant points; +2. soft-delete the related Qdrant points inline, recording the attempt; 4. retain the MinIO object until an explicit retention or hard-erasure workflow. Hard deletion requires a later retention/erasure implementation covering MinIO, Qdrant, and the relevant Postgres data. -### Broker operations +### Ingestion bounds and operations Before deploying an environment, define and document: -- RabbitMQ exchange, queue, binding, and dead-letter-exchange configuration; -- queue name, prefetch, manual-ack policy, and DLX retry/backoff; -- worker concurrency and resource limits; -- outbox polling/publish interval and stuck-event alert threshold; -- how failed jobs are inspected, retried, and cancelled. +- `INGESTION_MAX_CONCURRENCY` and the thread-pool capacity limiter, and their + relation to the process's CPU/memory budget; +- `INGESTION_TIMEOUT_SECONDS`, and the proxy/load-balancer/client read timeouts + that must exceed it; +- `INGESTION_EMBED_BATCH_SIZE` and `INGESTION_EMBED_CONCURRENCY`, sized to the + provider's rate limits and the self-hosted embedder's capacity; +- alert thresholds for p95 ingestion duration, `503`/`504` rates, and jobs left + in `running` past the timeout; +- how failed ingestions are inspected and retried. These are deployment/runbook settings, not new ADRs unless they change the reliability guarantee or system boundary. @@ -144,15 +157,14 @@ reliability guarantee or system boundary. ### Phase 1: Foundation and local dependencies -1. Add typed configuration in `src/config.py` for Postgres, MinIO, RabbitMQ, - Qdrant, application limits, and logging. +1. Add typed configuration in `src/config.py` for Postgres, MinIO, Qdrant, + ingestion bounds, application limits, and logging. 2. Populate `.env.example` with non-secret local-development configuration. -3. Add application Docker Compose services for Postgres, MinIO, RabbitMQ, and - Qdrant. Keep application MinIO buckets/credentials separate from Langfuse - infrastructure. +3. Add application Docker Compose services for Postgres, MinIO, and Qdrant. Keep + application MinIO buckets/credentials separate from Langfuse infrastructure. 4. Add direct Python dependencies and lock them with `uv`: - SQLAlchemy async/Postgres driver, MinIO/S3 client, aio-pika, Qdrant client, - and structured logging dependencies chosen by ADR-0011. + SQLAlchemy async/Postgres driver, MinIO/S3 client, Qdrant client, and + structured logging dependencies chosen by ADR-0011. 5. Add the pytest foundation from ADR-0016: async test configuration, boundary markers, and support for dependency-injected fakes. Add a lifespan smoke test before creating external clients. @@ -166,11 +178,10 @@ run without Docker or live providers. ### Phase 2: Database, migrations, and domain contracts 1. Define SQLAlchemy models and Alembic migrations for the minimum required - tables: `tenants`, `api_keys`, `source_files`, `ingestion_jobs`, - `ingestion_job_events`, and `outbox_events`. -2. Define Pydantic request/response/message schemas, including a versioned - `ingestion.job.created` message containing `event_id`, `tenant_id`, `file_id`, - `ingestion_job_id`, `request_id`, and `api_key_id`. + tables: `tenants`, `api_keys`, `source_files`, `ingestion_jobs`, and + `ingestion_job_events`, per ADR-0009. +2. Define Pydantic request/response schemas, including the terminal upload + response (`file_id`, `ingestion_job_id`, `status`, `chunks_indexed`). 3. Implement explicit repositories/services with a request/job-lifetime `AsyncSession`; routes/services own commit/rollback boundaries as specified by ADR-0012. @@ -188,75 +199,75 @@ reads/writes and valid job transitions. 2. Implement `POST /v1/files` for CSV only, including streaming-size controls, file-type validation, SHA-256 calculation, and a private MinIO upload using an internal object key. -3. In one Postgres transaction, persist `source_files`, create - `ingestion_jobs(status='queued')`, and write the corresponding unpublished - `outbox_events` row. -4. Return `202 Accepted` with `file_id`, `ingestion_job_id`, and `queued` status. +3. In one short Postgres transaction, persist `source_files` and create + `ingestion_jobs(status='running')`, then commit and release the connection + before any parse/embed work. +4. Return `201 Created` with `file_id`, `ingestion_job_id`, terminal status, and + `chunks_indexed` once ingestion completes. 5. Implement `GET /v1/files/{file_id}` with tenant filtering and a public status response that does not expose raw storage credentials or internal artifacts. 6. Add cleanup/compensation handling for a MinIO upload that succeeds while the database transaction fails. 7. Add unit/API tests for trusted tenant derivation, CSV validation, idempotency, - `202 Accepted`, and tenant-scoped status. Add MinIO adapter integration tests + the terminal `201 Created` response, and tenant-scoped status. Add MinIO adapter integration tests for server-derived private object paths and compensation behavior. -**Exit criteria:** an authenticated CSV upload creates a private object, a queued -job, and an unpublished outbox event; a tenant cannot retrieve another tenant's -file status; the HTTP path does not publish directly to RabbitMQ. +**Exit criteria:** an authenticated CSV upload creates a private object and a +`running` job row committed before any ingestion work; a tenant cannot retrieve +another tenant's file status. -### Phase 4: RabbitMQ and outbox publisher +### Phase 4: Bounded execution primitives -1. Provision/verify the application-owned RabbitMQ topic exchange, queue, and - binding configuration through deployment/bootstrap code rather than route - startup side effects. -2. Implement an outbox-publisher process that safely claims unpublished events, - publishes them with the stable event id using publisher confirms, and records - success/failure attempts. -3. RabbitMQ has no broker-native message de-duplication; retain idempotency in - all consumers as the sole guard against duplicate processing. -4. Add monitoring/logging for publish attempts, unpublished-event age, and - repeated failures. -5. Add unit tests for event claiming and retryable failures, then Testcontainers - RabbitMQ integration tests for durable publication, restart/retry, duplicate - publication, and message metadata. +1. Add the async embedding ports and adapters in `src/infrastructure/embedding/`, + with per-provider batching and an `asyncio.Semaphore` bounding in-flight + batches. +2. Route blocking work (parse, chunk, BM25, `minio`) through + `anyio.to_thread.run_sync` with an explicit `CapacityLimiter` created at + startup, so ingestion cannot exhaust Starlette's thread pool. +3. Enforce the request bounds: `INGESTION_MAX_CONCURRENCY` (`503` + `Retry-After` + when exceeded), `INGESTION_TIMEOUT_SECONDS` around the whole work phase + (`504`), and the size/chunk-count ceiling (`413`) checked before work starts. +4. Guarantee a bounded failure always writes a terminal job status — a timeout + must never leave a job stuck in `running`. +5. Add unit tests for batching/concurrency limits, timeout-to-terminal-status, + and capacity rejection, using scripted embedder fakes. -**Exit criteria:** an outbox event becomes a durable RabbitMQ message after a -publisher restart; retrying publication cannot create duplicate application work. +**Exit criteria:** embedding a few hundred chunks issues batched, concurrent +requests rather than serial ones; exceeding any bound produces the right status +code and a terminal job row. -### Phase 5: Ingestion worker and Qdrant Chunk/Point CRUD +### Phase 5: Ingestion execution and Qdrant Chunk/Point CRUD -1. Create a separate worker entrypoint with its own application-lifetime database, - MinIO, RabbitMQ, Qdrant, model, and logging clients. -2. Consume `ingestion.job.created`; reload and validate Postgres records before - fetching the MinIO object. -3. Transition the job from `queued` to `running` conditionally, append progress - events, parse CSV, create deterministic chunks, and upsert tenant-scoped - Qdrant points. -4. Mark the job `succeeded` with counters or `failed` with a safe error summary; - acknowledge only after final/progress state is persisted. -5. Make a repeated delivery of the same `ingestion_job_id` safe: no duplicate - logical chunks, no incorrect counters, and no transition from a terminal state - back to `running`. +1. Implement the ingestion service called by the route, using the + application-lifetime database, MinIO, Qdrant, model, and logging clients. +2. Validate the persisted records before fetching the MinIO object. +3. Append progress events, parse CSV, create deterministic chunks, embed them, + and upsert tenant-scoped Qdrant points — without holding a Postgres session + open across the work. +4. In a second short transaction, mark the job `succeeded` with counters or + `failed` with a safe error summary, then return the terminal response. +5. Make a retried upload safe: no duplicate logical chunks, no incorrect + counters, and no transition from a terminal state back to `running`. 6. Add unit tests for deterministic CSV chunks, point IDs, and terminal job - transitions. Add Testcontainers Qdrant and RabbitMQ integration tests for - tenant-filtered upserts, acknowledgement after durable state, redelivery, and + transitions. Add Testcontainers Qdrant and Postgres integration tests for + tenant-filtered upserts, terminal state persistence, retrying an upload, and parser/Qdrant failure handling. -**Exit criteria:** a successful upload reaches `succeeded`, and its points are -retrievable only under the owning tenant's Qdrant filter. Forced worker failure -and message redelivery produce a correct final job state. +**Exit criteria:** a successful upload returns `201` with a terminal status, and +its points are retrievable only under the owning tenant's Qdrant filter. A forced +failure mid-ingestion produces a `failed` job and the right HTTP status, and +retrying the upload produces a correct final state without duplicate chunks. ### Phase 6: Operations, integration tests, and documentation -1. Add an operator runbook covering local startup, migrations, MinIO bucket setup, - RabbitMQ exchange/queue setup, web/worker/outbox commands, and job replay. -2. Add a serialized Compose-based operational smoke test where the web process, - outbox publisher, and worker run independently for upload through indexed - points. Testcontainers remains the standard pytest mechanism for individual - adapter integration tests. -3. Add end-to-end tests for duplicate upload, duplicate RabbitMQ delivery, outbox - publisher crash/restart, worker crash/restart, tenant isolation, and failed - parser/Qdrant behavior. +1. Add an operator runbook covering local startup, migrations, MinIO bucket + setup, the run command, ingestion-bound tuning, the proxy/client timeout + requirement, and how to retry a failed ingestion. +2. Add a serialized Compose-based operational smoke test covering upload through + indexed points against the running web process. Testcontainers remains the + standard pytest mechanism for individual adapter integration tests. +3. Add end-to-end tests for duplicate upload, retrying a failed upload, tenant + isolation, capacity/timeout rejection, and failed parser/Qdrant behavior. 4. Add health/readiness checks that distinguish process health from dependency readiness. 5. Update the README with local-start instructions and links to ADRs, this plan, @@ -274,10 +285,11 @@ covered by automated tests: ```text POST /v1/files (authenticated CSV upload) -> raw bytes stored privately in MinIO - -> source file, queued job, and outbox event committed in Postgres - -> outbox publisher writes a durable RabbitMQ message - -> ingestion worker processes and indexes deterministic Qdrant points + -> source file and running job committed in Postgres, connection released + -> parse/chunk on threads, embed in bounded concurrent batches + -> deterministic Qdrant points upserted -> Postgres records progress and terminal job status + -> 201 Created returns the terminal result in the same request -> GET /v1/files/{file_id} reports that status within the owning tenant only ```