Why:
- resolve_auth_context() runs on every authenticated request and logged
nothing; four distinct rejection reasons (malformed/unknown/inactive/
expired key, inactive tenant) were all invisible.
- The domain-allowlist rejection in upload_source_file() happens before any
ingestion_jobs row exists, so it wasn't covered by the job-level
ingestion.job.failed event either -- a rejected upload left zero trace.
- Four of upload_source_file()'s five failure branches (parse_failed,
chunk_limit_exceeded, embedding_failed, index_failed) called
_mark_job_failed(), which wrote to Postgres but never logged; only
storage_failed and timeout had an ad-hoc logger.warning duplicated at their
own call sites.
Changes:
- auth/service.py: auth.succeeded / auth.failed (with a reason field per
rejection type), matching ADR-0011's own event catalog.
- domains/service.py: domain.rejected on the allowlist check;
domain.created / domain.updated / domain.status_changed on the three
mutations.
- files/upload.py: centralized failure logging inside _mark_job_failed
(every failure branch already calls it, so logging there once closes all
five branches instead of duplicating a log call at each site) as
ingestion.job.failed; added ingestion.job.started; renamed the ad-hoc
files.upload.succeeded to ingestion.job.completed for catalog consistency.
Impact:
- None to request/response behavior -- log events only.
Why:
- Domain values are denormalized into every Qdrant point payload. Without
validation, an unregistered or typo'd domain (e.g. "fier" for "fire")
silently creates a new partition that retrieval never queries — the file
ends up invisible rather than rejected. Tenants also need independently
sized domain sets (one may run 14 insurance lines, another 6), which rules
out an enum.
Changes:
- tenant_domains table (migration 41335d162de8) + repository, unique on
(tenant_id, domain).
- src/application/domains/: ensure_domain_allowed() is the strict-allowlist
check now run inside upload_source_file()'s first transaction, before any
MinIO object, job row, or Qdrant point is written.
- /v1/domains (list/create/patch/disable/enable) gated on its own
domains:read/domains:write scopes, deliberately separate from files:write
so an upload key cannot create partitions. domain itself is immutable
(denormalized into every point payload); only display_name is editable.
Disable blocks new uploads without touching already-indexed points.
Impact:
- BREAKING: POST /v1/files now rejects any domain without an active
tenant_domains row (400, unknown_domain). A domain must be created via
POST /v1/domains before the first upload to it.
Why:
- The chunks collection is now created by an explicit deployment step
(qdrant_bootstrap), not at startup, which means a process can boot against
a healthy Qdrant that has no collection at all. /readyz's previous check
only called get_collections(), so it reported ready in that state — the
misconfiguration stayed invisible until the first upload failed with a 502
after already paying for the MinIO write and embedding round trips.
Changes:
- ping_qdrant() now checks collection_exists(collection) instead of just
reachability.
Why:
- POST /v1/files was reporting chunks_indexed=0/points_created=0 unconditionally
— chunks were parsed and embedded but never written to Qdrant, so nothing
was actually searchable after upload.
Changes:
- upload_source_file() now calls index_chunks() after embedding, inside the
same INGESTION_TIMEOUT_SECONDS window, and marks the job failed
(error_code=index_failed, 502) if it raises.
- Job counters (points_created, points_soft_deleted) and the response's
chunks_indexed now reflect the real indexing result instead of a hardcoded
zero.
- Wired PointStorage through AppResources/lifespan/the files router.
Impact:
- A successful upload is now searchable in Qdrant by the time 201 returns.
Why:
- Ingested chunks need to become searchable Qdrant points before the upload
response returns, with tenant/domain isolation and a safe re-ingestion
story per ADR-0001/0017.
Changes:
- src/application/points/: index_chunks() is the sole entry point, owning
payload construction, batched/bounded-concurrency upserts
(upsert_concurrency semaphore), and a soft-delete sweep for points a
shorter re-ingestion leaves behind. The sweep runs only after every upsert
in the attempt succeeds, so a failed attempt can leave a stale prefix but
never removes content from a working index.
- PointStorage port (application/ports/) + QdrantPointStorage adapter
(infrastructure/qdrant/points.py), keeping the qdrant_client SDK out of
application code per ADR-0015.
- FakePointStorage test double for exercising the ordering/idempotency
guarantees without a real Qdrant.
Why:
- The chunks collection needs four named vectors (dense_nomic, dense_openai,
sparse, late_interaction) and payload indexes defined at creation time per
ADR-0001; sparse/multivector fields cannot be added to an existing
collection without recreating it, so schema drift here is expensive.
- Creating it at FastAPI startup would mirror the DDL-at-boot anti-pattern
ADR-0009 already rejects for Postgres and ADR-0012 rejects for LangGraph's
setup(), so it is a deployment step instead.
Changes:
- src/infrastructure/qdrant/collection.py: ensure_chunks_collection(),
idempotent and schema-verifying (raises on dimension/modifier mismatch
rather than silently accepting a misconfigured collection).
- src/cli/qdrant_bootstrap.py: the operator entry point
(python -m src.cli.qdrant_bootstrap).
- QdrantSettings gains collection/upsert_batch_size/upsert_concurrency.
Impact:
- Deployments must run the new bootstrap command before the first upload;
see ADR-0001's new "Collection provisioning" section.
Why:
- Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into
the inline upload path, with process-wide capacity/timeout/chunk-limit
guards (ADR-0017).
- The BM25 analyzer and dense-model config are ported from the `emet`
evaluation lab, which benchmarked them against the real Farsi corpus
(bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large
at native 3072-dim), closing open items in ADR-0001/ADR-0005.
Changes:
- New: embedding ports, orchestration (embed_chunks), request-bounds
helpers, and dense/sparse adapters (analyzers.py, bm25.py,
openai_compatible.py).
- upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY
(503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413);
every failure path still writes a terminal job row.
- Lifespan builds and warms both dense embedders at startup (fail-soft) and
creates the sparse embedder and concurrency semaphore.
- httpx moves from dev to main dependencies (adapters use it directly).
Impact:
- Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0.
- New EMBEDDING_* env vars documented in .env.example; safe defaults.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Why:
- Code under test (auth resolution, the two-phase upload) opens more than
one session per operation; the existing fixture only exposed one
rolled-back session.
Changes:
- Add a db_sessionmaker fixture sharing one outer transaction.
- Pin loop_scope="session" -- without it, a second async test against the
session-scoped Postgres container fails with "Event loop is closed."
Testcontainers reports the container host as `localhost`, which resolves
to ::1 before 127.0.0.1 on this machine. Docker publishes the mapped port
on IPv4 only, and the IPv6 SYN is dropped rather than refused, so asyncpg
blocked on the first address until its connect timeout instead of falling
back to the second -- the migration fixture hung rather than failing.
127.0.0.1 connects in 0.07s where localhost timed out at 15s; the
integration suite now runs in 3.6s, inside the 10s pytest-timeout.
alembic/env.py previously always overwrote sqlalchemy.url from
Settings().postgres.dsn, which made it impossible for a test fixture to
point Alembic at a Testcontainers-managed database. Now env.py only
sets it when unset, and a new integration suite runs `alembic upgrade
head` against a real Postgres container per ADR-0016 (no create_all()).