Compare commits

...

44 Commits

Author SHA1 Message Date
0b1932f716 docs(claude): record plan 002 phase 3 in the project status
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 17:13:05 +03:30
73bdac0da2 test(points): cover soft delete, relinking, and the file sweep
Why:
- The failure modes worth testing are races, and they are cheap to force
  against the fake and expensive to observe anywhere else.

Changes:
- Unit tests for both boundaries, the repeat delete, a missing neighbour, the
  traversal property after several deletes, and — via a repository that bumps a
  rival's version before each apply — both the partial-apply repair and the
  unconvergent 409.
- One new shared contract scenario (a multi-point batch applies every patch,
  including nulling a pointer) so it runs against the fake and real Qdrant.
- HTTP tests against real Postgres and Qdrant for relinking, soft-not-hard
  delete, cross-tenant 404, and scope enforcement on both routes.
- create_source_file factory for tests addressing a file without uploading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 17:12:56 +03:30
b58f4630f3 feat(points): add soft delete with neighbour relinking
Why:
- Plan 002 Phase 3. ADR-0002 makes delete soft by default and treats a partial
  neighbour relink as a defect, since ADR-0003's context-window expansion walks
  the previous/next pointer chain.

Changes:
- relinking.py computes the patches still missing between the state just read
  and the desired end state, so a normal delete, a repeat delete, and recovery
  from a half-applied batch are one path.
- deletion.py re-plans and re-applies up to three times, verifying by read-back,
  because Qdrant has no multi-point transaction and reports success for a
  filtered set_payload that matched nothing; exhausting the retries raises
  PointVersionConflictError (409).
- DELETE /v1/points/{point_id} and DELETE /v1/files/{file_id}, both on
  points:write. The file route sweeps points first, then marks the source_files
  row soft_deleted in its own short transaction.
- Log events carry ADR-0011 duration_ms plus rounds.

Impact:
- A deleted file's source_files row leaves 'active', so re-uploading the same
  bytes now re-ingests instead of matching the duplicate path.
- No migration; no point is ever removed from Qdrant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 17:12:43 +03:30
b25c15fefa docs(claude): record plan 002 phase 2 in the project status
Why:
- The status section described plan 002 as Phase 1 only. The /v1/points read
  routes, the query service, and the points scopes had shipped but were listed
  under "not built yet", so a fresh session would start from a wrong map of the
  codebase.

Changes:
- Record the five read routes, the points:read gating, and the 404-not-403
  mapping.
- Call out the three route-level rules that fail quietly when broken: /count and
  /search must precede /{point_id}, file_id is required on the listing, and the
  search query is Persian-folded before matching.
- Point at tests/support/point_contract.py as the place new repository behaviour
  belongs, since it runs against both the fake and real Qdrant.
- Narrow the "not built yet" list to Phases 3-6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 16:04:59 +03:30
ba7921dd4e test(points): cover the read paths and the query service
Why:
- The isolation and pagination guarantees are the ones that fail silently, so
  they need tests that would actually notice.

Changes:
- API tests over real Postgres and Qdrant together, because the invariant worth
  testing spans both: the tenant Postgres derived is the only one Qdrant is ever
  queried with.
- Pagination holds when a point is inserted behind the cursor mid-listing --
  the defect an offset cursor would have.
- A route-order guard, since /{point_id} declared first turns count into a 422
  and nothing else in the suite would catch it.
- Unit tests for the query service against the fake, including the Arabic to
  Persian letterform fold and raising rather than returning None.

Impact:
- Suite goes to 325 passed, 3 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 15:09:31 +03:30
802bae429d docs(claude): record plan 002 phase 1 in the project status 2026-08-22 15:09:18 +03:30
43932b6562 docs(adr): record query normalization and file-scoped pagination
Why:
- Both are contracts callers depend on, not implementation details, and neither
  was written down. This repo treats the ADR as the source of truth rather than
  letting code diverge from it silently.

Changes:
- Record that the search query is folded the same way ingested content was, and
  why the alternative fails in the worst available way: an exact-looking query
  returning nothing, with no error and nothing in the logs to distinguish it
  from a genuine miss.
- Record that listing requires file_id and paginates by order_id value, and why
  an offset cursor repeats an already-served row under a concurrent insert.
- State that results carry no relevance score and no ranked order, so callers
  cannot read array position as relevance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 15:09:07 +03:30
3b9434faf4 feat(points): add the /v1/points read paths
Why:
- Ingestion writes points in bulk but nothing could read one back. Plan 002
  Phase 2 opens the read surface an admin frontend needs.

Changes:
- GET /v1/points/{point_id}, /v1/points?file_id=..., /v1/points/count,
  /v1/points/search, and /v1/files/{file_id}/points, all under points:read --
  the scope follows the data, so an upload key does not become a way to read
  every chunk of every file.
- The keyword query is Persian-normalized before matching, because ingestion
  letter-folds content at ingest and an unfolded query would return an empty
  result set silently rather than an error.
- file_id is required on the listing: the cursor is an order_id value and
  order_id is only unique within one file.
- PointNotFoundError maps to 404, never 403, so a cross-tenant point id is
  indistinguishable from a nonexistent one.
- Route order is load-bearing: /count and /search precede /{point_id}, or
  "count" is parsed as a UUID and fails 422.

Impact:
- Requires the content/is_active/chunk_index payload indexes, so a deployed
  environment needs qdrant_bootstrap re-run before search works.
- Keyword search returns no relevance score and no ranked order; callers must
  not read array position as relevance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 15:08:54 +03:30
5251990444 feat(tenant): grant points scopes on provisioned keys
Why:
- points:read and points:write are named in ADR-0009 but were absent from
  DEFAULT_SCOPES, so a provisioned key could not reach the read paths that
  follow. The runbook documented only files:write and domains:write, which
  understated what a default key can now do.

Changes:
- Add points:read and points:write to DEFAULT_SCOPES.
- Document the full scope table in the runbook, calling out that points:read
  grants the text of every chunk of every file -- so an upload-only key gets
  files:write alone.

Impact:
- Keys issued before this change keep their existing scopes; provisioning does
  not backfill. Reissue or widen an existing key explicitly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 15:08:39 +03:30
062fdd7ac1 refactor(test): share the Qdrant point-seeding helper
Why:
- The helper turning a SeedSpec into an upsertable ChunkPoint lived inside one
  integration test. A second suite now needs to seed real Qdrant the same way,
  and a copy would let the two drift.

Changes:
- Move `_chunk_point` into tests/support/point_contract.py as `chunk_point_for`,
  beside the `build_point` read model it derives its payload from.

Impact:
- Pure move. No behaviour change; enables the API suite that follows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 15:08:25 +03:30
923ac8e5d6 test(points): hold the fake and the Qdrant adapter to one shared contract
Why:
- Two parallel test files let a fake drift more permissive than the store it
  stands in for, so unit tests stay green while production diverges. Plan 002
  Phase 1's exit criterion is precisely that the two agree.

Changes:
- One scenario suite in tests/support/point_contract.py, run against
  FakePointRepository (unit) and QdrantPointRepository (integration). A
  divergence fails one of the two runs rather than hiding.
- The fake models the behaviours services branch on: the implied is_active read
  filter, value-based cursor pagination, and a stale version guard that matches
  nothing rather than raising -- the no-op Qdrant's filtered set_payload actually
  has, and the reason a service must read back to know its write landed.
- Patched points are re-validated rather than model_copy'd, so the fake holds a
  datetime where a read from real Qdrant returns one.
- The seeded corpus gives each tenant its own file: point IDs derive from
  file_id plus chunk_index alone, so two tenants in one file would collide on a
  single ID and the fixture would assert an impossible state.

Impact:
- 15 scenarios pass against both implementations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 13:09:50 +03:30
4da30f9983 feat(points): add the point read/edit port and its Qdrant adapter
Why:
- PointStorage is deliberately the two bulk operations ingestion performs. Reads,
  single-point edits, and keyword search have a different caller, a different
  failure vocabulary, and a different tenant-filter obligation, so they get their
  own port rather than accreting onto the ingestion one.

Changes:
- tenant_id is a required keyword argument on every port method, making a
  forgotten tenant filter a type error rather than a review question.
- Reads go through scroll with a HasIdCondition, not retrieve: retrieve takes no
  filter and would push the tenant check into Python after Qdrant already
  answered -- the shape ADR-0002's isolation rule exists to prevent.
- Ordered listing paginates by order_id value, not offset. Qdrant returns no page
  offset under order_by, and an offset cursor skips or repeats rows when a
  concurrent insert shifts positions underneath the reader.
- Point.from_payload takes a Mapping, not a dict: dict is invariant in its value
  type, so the SDK's concrete vector union is not a dict[str, object].
- Request schemas forbid extra keys and omit server-owned fields, so a client
  sending tenant_id or version gets 422 rather than having it silently ignored.

Impact:
- No route uses this yet; the /v1/points surface is Phase 2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 13:09:37 +03:30
5e935e5895 feat(qdrant): index content, is_active, and chunk_index on the chunks collection
Why:
- ADR-0002's keyword search needs a full-text index on content, which
  collection.py deliberately deferred to plan 002. is_active and chunk_index
  were unindexed while ingestion was the only reader; every /v1/points read path
  filters on them.

Changes:
- content gets a TEXT index with the multilingual tokenizer, which segments
  Persian correctly where the word tokenizer mishandles ZWNJ-joined compounds.
  No stemmer or stopword list: content is already letter-folded by
  normalize_persian_text at ingest, and the ranked Farsi lexical path is the
  benchmarked BM25 sparse vector, not this index.
- Tests assert content is TEXT rather than KEYWORD -- a keyword index would only
  match an entire chunk verbatim, which never happens and fails silently.
- Adds a test that a missing index is added to an already-live collection.

Impact:
- Requires re-running `python -m src.cli.qdrant_bootstrap`. Payload indexes are
  additive, so no collection rebuild and no re-embedding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 13:09:26 +03:30
ac3810182d docs(adr): record the re-ingestion rule and close plan 002's open decisions
Why:
- ADR-0002 already answered three of the four questions plan 002 listed as
  "decisions needed"; the fourth -- what happens to a manually edited point when
  its file is re-uploaded -- was left for a Phase 6 test to force. Deciding it in
  code rather than in the ADR would invert this repo's rule.

Changes:
- ADR-0002 gains "Re-ingestion versus manual edits": the new file wins,
  surviving points are overwritten with an incremented version, absent points
  are flagged inactive rather than removed, and manually created points sit past
  the ingested chunk_index range so the existing sweep covers them.
- Plan 002's stale decisions section becomes a pointer table; its audit scope is
  pinned to both ADR-0009 tables.

Impact:
- Clobbered edits are recoverable from point_audit_events, not from Qdrant: the
  deterministic point ID cannot hold both versions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 13:09:15 +03:30
Ali Zarinkolah
3d9269e54f docs(ops): add the operator runbook and record Phase 6 as complete
Why:
- the slice had no operator documentation: no tuning guidance, no failure
  procedure, no statement of the proxy timeout requirement.

Changes:
- add docs/runbook.md: startup, deployment steps, provisioning, the INGESTION_*
  tuning table with each bound's status code, the proxy read-timeout rule,
  /healthz vs /readyz, failure investigation by real event name plus job/event
  SQL, retry semantics, and alert thresholds
- link it from the README and note provisioning there
- mark plan 001 Phase 6 done and refresh CLAUDE.md's status paragraph

Impact:
- documentation only

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:26:58 +03:30
Ali Zarinkolah
7c1fe79f1c test(e2e): add the Compose smoke test of the running web process
Why:
- ADR-0016 reserves Compose for a serialized smoke test of the running web
  process. Nothing else exercises the deployment steps, a real HTTP server, or
  the real logging configuration, which every in-process test no-ops.

Changes:
- scripts/smoke.sh brings up Compose, runs both bootstrap steps, provisions a
  throwaway tenant, starts uvicorn, and drives the test against it
- the test skips unless SMOKE_BASE_URL is set, so `uv run pytest` never invokes
  Compose; it asserts the ADR-0011 JSON log sink

Impact:
- a pre-release gate, not a per-PR one

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:26:47 +03:30
Ali Zarinkolah
1b873e5a6f test(e2e): cover the ingestion slice against real Postgres, MinIO, and Qdrant
Why:
- the slice's reliability invariants (ADR-0016) had no end-to-end coverage.

Changes:
- 13 Testcontainers-based tests: duplicate upload, retry after a failed job,
  cross-tenant 404, unregistered domain, capacity 503, timeout 504, parse 400,
  real Qdrant 502, missing scope 403, and both readiness states
- only the dense embedders are faked (ADR-0016 bars live providers); they
  return the pinned 768/3072 dimensions
- the capacity test uses a committing sessionmaker, since the shared-connection
  fixture cannot serve concurrent sessions

Impact:
- runs in the default `uv run pytest`; needs Docker, like every integration test

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:26:36 +03:30
Ali Zarinkolah
133f565704 feat(tenant): add operator provisioning for tenants, API keys, and domains
Why:
- nothing over HTTP could create the first tenant: every /v1 route needs an API
  key, and a key cannot exist before its tenant. The service was unusable by a
  human without hand-written SQL.

Changes:
- add `provision_tenant`, owning tenant reuse-or-create, key generation and
  hashing, and domain registration in one transaction
- expose it as `python -m src.cli.provision_tenant`, alongside
  `alembic upgrade head` and `qdrant_bootstrap`
- add `tenants.get_by_slug`/`create` and `api_keys.create`
- log `tenant.provisioned` / `api_key.provisioned` with the key prefix only

Impact:
- a third deployment step; the plaintext key is printed once and never logged
  or stored (ADR-0011, ADR-0009)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:26:25 +03:30
Ali Zarinkolah
c7a5b69c0a fix(api): use the non-deprecated 422 status constant
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:26:12 +03:30
Ali Zarinkolah
2c72688440 refactor(test): share container fixtures across the integration and e2e suites
Why:
- tests/e2e/ cannot reach fixtures defined in a per-boundary conftest, and
  `pytest_plugins` is only honoured in the root conftest.

Changes:
- move the Postgres/MinIO/Qdrant container fixtures into
  tests/support/containers.py and register it as a root plugin
- fold MinIO bucket creation into `minio_settings`; an autouse fixture in a
  globally registered plugin would pull a container into unit runs
- add a `postgres_settings` fixture so a component can be built from it directly

Impact:
- no behavior change; `pytest -m unit` still needs no Docker

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:26:02 +03:30
Ali Zarinkolah
012b44d5f2 feat(observability): backfill logging for upload, auth, and domain services
Why:
- resolve_auth_context() runs on every authenticated request and logged
  nothing; four distinct rejection reasons (malformed/unknown/inactive/
  expired key, inactive tenant) were all invisible.
- The domain-allowlist rejection in upload_source_file() happens before any
  ingestion_jobs row exists, so it wasn't covered by the job-level
  ingestion.job.failed event either -- a rejected upload left zero trace.
- Four of upload_source_file()'s five failure branches (parse_failed,
  chunk_limit_exceeded, embedding_failed, index_failed) called
  _mark_job_failed(), which wrote to Postgres but never logged; only
  storage_failed and timeout had an ad-hoc logger.warning duplicated at their
  own call sites.

Changes:
- auth/service.py: auth.succeeded / auth.failed (with a reason field per
  rejection type), matching ADR-0011's own event catalog.
- domains/service.py: domain.rejected on the allowlist check;
  domain.created / domain.updated / domain.status_changed on the three
  mutations.
- files/upload.py: centralized failure logging inside _mark_job_failed
  (every failure branch already calls it, so logging there once closes all
  five branches instead of duplicating a log call at each site) as
  ingestion.job.failed; added ingestion.job.started; renamed the ad-hoc
  files.upload.succeeded to ingestion.job.completed for catalog consistency.

Impact:
- None to request/response behavior -- log events only.
2026-08-20 19:24:39 +03:30
Ali Zarinkolah
ac779dec7e fix(tests): isolate structlog global state from lifespan-triggering fixtures
Why:
- Any test using the client/api_client fixtures runs the app's real lifespan,
  which calls the production configure_logging() -- setting
  cache_logger_on_first_use=True (ADR-0011). That permanently monkeypatches
  the .bind method on whichever module-level
  logger = structlog.get_logger(__name__) instance is used first.
  structlog.reset_defaults() only resets *global* config, not that
  per-instance mutation, so once triggered, structlog.testing.capture_logs()
  silently stops intercepting events in every test that runs afterward in the
  same pytest process -- order-dependent flakiness with no useful failure
  message (assertions just see an empty list).

Changes:
- Added two autouse fixtures: one no-ops configure_logging for tests that
  spin up the app via LifespanManager (they test HTTP behavior, not logging
  output, so they don't need the real thing), one resets structlog defaults
  after every test as defense in depth.

Impact:
- Test-only; makes capture_logs()-based assertions reliable regardless of
  test execution order.
2026-08-20 19:21:56 +03:30
Ali Zarinkolah
9e8987968c feat(observability): add dual local logging sinks and static environment context
Why:
- Wanted human-readable console output while developing locally, without
  losing a machine-parseable log for later grepping/parsing. A single
  renderer chosen by a flag can't do both at once.
- ADR-0011 had no way to correlate an issue with a specific deployment
  (build/region/instance) independent of any one request.

Changes:
- configure_logging() now builds two independent handlers: console (always
  on, colored unless LOG_JSON_FORMAT=true) and an optional rotating JSON file
  (LOG_FILE_PATH, unset by default) -- the same structlog event fans out to
  both, so call sites are unaffected.
- A static structlog processor binds env/service_version onto every event.
  Deliberately not a contextvar: RequestIdMiddleware's clear_contextvars()
  would wipe a value bound there before the first request.
- New settings: APP_SERVICE_VERSION, LOG_FILE_PATH/LOG_FILE_MAX_BYTES/
  LOG_FILE_BACKUP_COUNT.
- ADR-0011 amended with both decisions ("console and file are independent
  sinks locally"; "bind process-level environment context once at startup").

Impact:
- configure_logging() signature changed to (logging_settings, app_settings);
  both call sites (lifespan, qdrant_bootstrap CLI) updated.
2026-08-20 19:20:27 +03:30
Ali Zarinkolah
e9e83b3a26 feat(tenant): add tenant_domains allowlist and /v1/domains management API
Why:
- Domain values are denormalized into every Qdrant point payload. Without
  validation, an unregistered or typo'd domain (e.g. "fier" for "fire")
  silently creates a new partition that retrieval never queries — the file
  ends up invisible rather than rejected. Tenants also need independently
  sized domain sets (one may run 14 insurance lines, another 6), which rules
  out an enum.

Changes:
- tenant_domains table (migration 41335d162de8) + repository, unique on
  (tenant_id, domain).
- src/application/domains/: ensure_domain_allowed() is the strict-allowlist
  check now run inside upload_source_file()'s first transaction, before any
  MinIO object, job row, or Qdrant point is written.
- /v1/domains (list/create/patch/disable/enable) gated on its own
  domains:read/domains:write scopes, deliberately separate from files:write
  so an upload key cannot create partitions. domain itself is immutable
  (denormalized into every point payload); only display_name is editable.
  Disable blocks new uploads without touching already-indexed points.

Impact:
- BREAKING: POST /v1/files now rejects any domain without an active
  tenant_domains row (400, unknown_domain). A domain must be created via
  POST /v1/domains before the first upload to it.
2026-08-20 18:20:24 +03:30
Ali Zarinkolah
fa933b08ff fix(qdrant): verify collection existence in the readiness check
Why:
- The chunks collection is now created by an explicit deployment step
  (qdrant_bootstrap), not at startup, which means a process can boot against
  a healthy Qdrant that has no collection at all. /readyz's previous check
  only called get_collections(), so it reported ready in that state — the
  misconfiguration stayed invisible until the first upload failed with a 502
  after already paying for the MinIO write and embedding round trips.

Changes:
- ping_qdrant() now checks collection_exists(collection) instead of just
  reachability.
2026-08-20 18:17:56 +03:30
Ali Zarinkolah
cc915f0f1a feat(ingestion): index embedded chunks into Qdrant on upload
Why:
- POST /v1/files was reporting chunks_indexed=0/points_created=0 unconditionally
  — chunks were parsed and embedded but never written to Qdrant, so nothing
  was actually searchable after upload.

Changes:
- upload_source_file() now calls index_chunks() after embedding, inside the
  same INGESTION_TIMEOUT_SECONDS window, and marks the job failed
  (error_code=index_failed, 502) if it raises.
- Job counters (points_created, points_soft_deleted) and the response's
  chunks_indexed now reflect the real indexing result instead of a hardcoded
  zero.
- Wired PointStorage through AppResources/lifespan/the files router.

Impact:
- A successful upload is now searchable in Qdrant by the time 201 returns.
2026-08-20 18:17:39 +03:30
Ali Zarinkolah
d00d436e5c feat(qdrant): add tenant-scoped point storage for ingestion
Why:
- Ingested chunks need to become searchable Qdrant points before the upload
  response returns, with tenant/domain isolation and a safe re-ingestion
  story per ADR-0001/0017.

Changes:
- src/application/points/: index_chunks() is the sole entry point, owning
  payload construction, batched/bounded-concurrency upserts
  (upsert_concurrency semaphore), and a soft-delete sweep for points a
  shorter re-ingestion leaves behind. The sweep runs only after every upsert
  in the attempt succeeds, so a failed attempt can leave a stale prefix but
  never removes content from a working index.
- PointStorage port (application/ports/) + QdrantPointStorage adapter
  (infrastructure/qdrant/points.py), keeping the qdrant_client SDK out of
  application code per ADR-0015.
- FakePointStorage test double for exercising the ordering/idempotency
  guarantees without a real Qdrant.
2026-08-20 18:16:35 +03:30
Ali Zarinkolah
58ca6109d1 feat(embedding): expose model_version on dense and sparse embedder ports
Why:
- ADR-0001's Qdrant payload records embedding_model_version so a future model
  swap can identify which chunks need re-embedding. The embedder is what
  knows which model produced its vectors, so it reports this rather than the
  call site reconstructing it from settings.

Changes:
- DenseEmbedder/SparseEmbedder protocols gain a model_version: str attribute.
- OpenAICompatibleEmbedder reports its configured model; Bm25SparseEmbedder
  reports its analyzer (bm25-<analyzer>).
2026-08-20 18:15:38 +03:30
Ali Zarinkolah
5e0addcc55 feat(qdrant): provision the chunks collection as an explicit deployment step
Why:
- The chunks collection needs four named vectors (dense_nomic, dense_openai,
  sparse, late_interaction) and payload indexes defined at creation time per
  ADR-0001; sparse/multivector fields cannot be added to an existing
  collection without recreating it, so schema drift here is expensive.
- Creating it at FastAPI startup would mirror the DDL-at-boot anti-pattern
  ADR-0009 already rejects for Postgres and ADR-0012 rejects for LangGraph's
  setup(), so it is a deployment step instead.

Changes:
- src/infrastructure/qdrant/collection.py: ensure_chunks_collection(),
  idempotent and schema-verifying (raises on dimension/modifier mismatch
  rather than silently accepting a misconfigured collection).
- src/cli/qdrant_bootstrap.py: the operator entry point
  (python -m src.cli.qdrant_bootstrap).
- QdrantSettings gains collection/upsert_batch_size/upsert_concurrency.

Impact:
- Deployments must run the new bootstrap command before the first upload;
  see ADR-0001's new "Collection provisioning" section.
2026-08-20 18:15:21 +03:30
Ali Zarinkolah
e8fb41af87 test(ci): scope pytest timeout to test bodies, not fixture setup
Why:
- Testcontainers' session-scoped container startup (~25s on a cold Docker
  cache) was charged against the global 10s pytest-timeout budget, causing
  every integration test to fail regardless of its own runtime.

Changes:
- Set timeout_func_only = true so the budget applies to the test function
  only, not fixture setup.
2026-08-20 18:14:45 +03:30
9a4b173b95 fix(config): cascade .env file loading to nested settings classes
Why:
- No setting in this app ever actually read from .env: only the outer
  Settings declared env_file=".env", and pydantic-settings does not cascade
  that to nested BaseSettings classes. Every previously-correct local value
  was coincidence (.env.example defaults matching class defaults). Found by
  testing EMBEDDING_OPENAI_API_KEY against the live OpenAI API.

Changes:
- Every nested settings class now declares env_file=".env" itself.
- Settings.__init__/EmbeddingSettings.__init__ explicitly thread an
  _env_file override to every nested constructor, so overriding it (as
  tests do) reaches the whole tree, not just the outer class.
- env_ignore_empty=True everywhere, since the fix surfaced a second bug:
  a blank env var (e.g. EMBEDDING_OPENAI_DIMENSIONS=) failed to parse as
  int | None instead of falling back to the field default.

Impact:
- Real deployments setting env vars directly (Docker Compose) are
  unaffected. Local .env-file development now actually works.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 17:14:56 +03:30
5c0a5938f8 feat(ingestion): add bounded, benchmark-aligned embedding execution
Why:
- Plan 001 Phase 4 needs batched, concurrency-bounded embedding wired into
  the inline upload path, with process-wide capacity/timeout/chunk-limit
  guards (ADR-0017).
- The BM25 analyzer and dense-model config are ported from the `emet`
  evaluation lab, which benchmarked them against the real Farsi corpus
  (bm25-fa-norm-stop; nomic-embed-text-v2-moe at 768-dim; text-embedding-3-large
  at native 3072-dim), closing open items in ADR-0001/ADR-0005.

Changes:
- New: embedding ports, orchestration (embed_chunks), request-bounds
  helpers, and dense/sparse adapters (analyzers.py, bm25.py,
  openai_compatible.py).
- upload.py now parses/chunks/embeds inline behind INGESTION_MAX_CONCURRENCY
  (503), INGESTION_TIMEOUT_SECONDS (504), and the chunk-count ceiling (413);
  every failure path still writes a terminal job row.
- Lifespan builds and warms both dense embedders at startup (fail-soft) and
  creates the sparse embedder and concurrency semaphore.
- httpx moves from dev to main dependencies (adapters use it directly).

Impact:
- Qdrant point upserts are still Phase 5 -- chunks_indexed stays 0.
- New EMBEDDING_* env vars documented in .env.example; safe defaults.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 17:13:32 +03:30
aa6d595424 chore(claude): add Explore subagent definition 2026-08-19 15:01:21 +03:30
07b50d6987 docs(claude): record the deep-module principle for future development 2026-08-19 15:01:10 +03:30
9858e27c2d test(integration): add coverage for repositories, auth resolution, and the two-phase upload 2026-08-19 15:01:00 +03:30
e70ad13b10 test(postgres): support multi-session integration tests
Why:
- Code under test (auth resolution, the two-phase upload) opens more than
  one session per operation; the existing fixture only exposed one
  rolled-back session.

Changes:
- Add a db_sessionmaker fixture sharing one outer transaction.
- Pin loop_scope="session" -- without it, a second async test against the
  session-scoped Postgres container fails with "Event loop is closed."
2026-08-19 15:00:50 +03:30
3bced65926 feat(api): add POST/GET /v1/files with auth, error envelope, and request-id middleware
Why:
- Wires the ADR-0008 error envelope, per-request correlation id, and the
  /v1/files routes into the app.

Changes:
- Extend AppResources/lifespan with the ingestion CapacityLimiter and
  ObjectStorage adapter.

Impact:
- /v1 now exposes routes for the first time.
2026-08-19 15:00:39 +03:30
c9cf7b368b feat(files): add source-file upload with MinIO storage and Postgres repositories
Why:
- Implements plan 001 Phase 3's upload orchestration.

Changes:
- ObjectStorage port and MinIO adapter, thread-offloaded per ADR-0017.
- Tenant-scoped repositories for api_keys, source_files, ingestion_jobs.
- upload_source_file implementing the two-transaction shape with
  (tenant_id, domain, content_sha256) idempotency.

Impact:
- This phase stores bytes only -- chunks_indexed is always 0 until
  Phase 4/5 add parsing/embedding.
2026-08-19 15:00:27 +03:30
e97ce6e5f3 feat(auth): add API-key authentication and tenant resolution 2026-08-19 15:00:13 +03:30
3803d9c79a feat(config): move the upload-size ceiling under ingestion settings
Why:
- ADR-0017 names the bound INGESTION_MAX_UPLOAD_SIZE_MB; the code had it
  as APP_MAX_UPLOAD_SIZE_MB. Reconciled to match the ADR, grouped with the
  other three ADR-0017 bounds on IngestionSettings.
2026-08-19 15:00:02 +03:30
94684d97ae refactor(ingestion): give the pipeline package a single async entry point
Why:
- The package exposed 8 modules directly, pushing source-type dispatch and
  the ADR-0017 thread-offload obligation onto every caller.

Changes:
- Add parse_and_chunk_document as the sole public entry point.
- Demote the individual parsers to internal/test-only.
2026-08-19 14:59:51 +03:30
7753651dd6 fix(tests): pin the Postgres container URL to IPv4 so asyncpg can connect
Testcontainers reports the container host as `localhost`, which resolves
to ::1 before 127.0.0.1 on this machine. Docker publishes the mapped port
on IPv4 only, and the IPv6 SYN is dropped rather than refused, so asyncpg
blocked on the first address until its connect timeout instead of falling
back to the second -- the migration fixture hung rather than failing.

127.0.0.1 connects in 0.07s where localhost timed out at 15s; the
integration suite now runs in 3.6s, inside the 10s pytest-timeout.
2026-08-18 10:40:32 +03:30
5cdfb70085 feat(ingestion): add DOCX/CSV/XLSX parsing and fixed-size chunking (ADR-0018)
Adds src/application/ingestion/ -- Persian normalization, DOCX body
walk with structural data/layout table classification, CSV/XLSX row
rendering, and fixed-size token chunking (cl100k_base, 400/60/512) --
as pure functions per ADR-0015, tested against real production
documents (asia_data_sample, kept out of the repo). ADR-0018 records
where this diverges from ADR-0004 (fixed-size default, no invented
headings/tree, structural table classification, header-provable
labeling only). Plan 001's scope line is corrected from CSV-only to
DOCX/XLSX/CSV, and CLAUDE.md's stale project-status paragraph is
updated to match current implementation state.
2026-08-18 10:22:17 +03:30
80ed5b1577 test(postgres): allow Testcontainers to override the migration sqlalchemy.url
alembic/env.py previously always overwrote sqlalchemy.url from
Settings().postgres.dsn, which made it impossible for a test fixture to
point Alembic at a Testcontainers-managed database. Now env.py only
sets it when unset, and a new integration suite runs `alembic upgrade
head` against a real Postgres container per ADR-0016 (no create_all()).
2026-08-18 10:21:56 +03:30
158 changed files with 15895 additions and 443 deletions

View File

@@ -0,0 +1,8 @@
---
name: Explore
description: Fast, read-only codebase search
model: sonnet
effort: low
tools: Read, Grep, Glob, Bash, WebFetch, WebSearch
maxTurns: 20
---

View File

@@ -9,12 +9,19 @@
# Application # Application
APP_ENV=local APP_ENV=local
APP_MAX_UPLOAD_SIZE_MB=25
APP_READINESS_CHECK_TIMEOUT_SECONDS=2.0 APP_READINESS_CHECK_TIMEOUT_SECONDS=2.0
# Set by CI/CD at build/deploy time; never computed at runtime.
APP_SERVICE_VERSION=dev
# Logging # Logging
LOG_LEVEL=INFO LOG_LEVEL=INFO
LOG_JSON_FORMAT=false LOG_JSON_FORMAT=false
# Optional second sink, always JSON regardless of LOG_JSON_FORMAT. Local dev
# only -- leave unset in production, where stdout/stderr collection is
# preferred over an in-container log file.
# LOG_FILE_PATH=logs/app.log
LOG_FILE_MAX_BYTES=10485760
LOG_FILE_BACKUP_COUNT=5
# Postgres (application database, separate from Langfuse's Postgres) # Postgres (application database, separate from Langfuse's Postgres)
# Use 127.0.0.1 rather than localhost: some environments resolve localhost to # Use 127.0.0.1 rather than localhost: some environments resolve localhost to
@@ -37,6 +44,7 @@ MINIO_BUCKET=chatbot-source-files
INGESTION_MAX_CONCURRENCY=4 INGESTION_MAX_CONCURRENCY=4
INGESTION_THREAD_POOL_SIZE=8 INGESTION_THREAD_POOL_SIZE=8
INGESTION_TIMEOUT_SECONDS=120.0 INGESTION_TIMEOUT_SECONDS=120.0
INGESTION_MAX_UPLOAD_SIZE_MB=25
INGESTION_MAX_CHUNKS_PER_FILE=5000 INGESTION_MAX_CHUNKS_PER_FILE=5000
INGESTION_EMBED_BATCH_SIZE=128 INGESTION_EMBED_BATCH_SIZE=128
INGESTION_EMBED_CONCURRENCY=4 INGESTION_EMBED_CONCURRENCY=4
@@ -44,3 +52,59 @@ INGESTION_EMBED_CONCURRENCY=4
# Qdrant # Qdrant
QDRANT_URL=http://127.0.0.1:6343 QDRANT_URL=http://127.0.0.1:6343
QDRANT_API_KEY= QDRANT_API_KEY=
QDRANT_COLLECTION=chunks
QDRANT_UPSERT_BATCH_SIZE=128
QDRANT_UPSERT_CONCURRENCY=4
# Dense embedders (ADR-0001). Both speak an OpenAI-compatible /embeddings
# endpoint, so one adapter serves both. Models and endpoints are the ones the
# `emet` evaluation lab benchmarked as winners on the Farsi corpus.
#
# dense_nomic runs behind Ollama's OpenAI-compat shim, which accepts any
# non-empty API key. KEEP_ALIVE holds the model resident: a cold load of
# nomic-embed-text-v2-moe takes >150s, well past INGESTION_TIMEOUT_SECONDS,
# so an idle-then-upload would otherwise 504.
EMBEDDING_NOMIC_BASE_URL=http://192.168.10.10:11435/v1
EMBEDDING_NOMIC_MODEL=nomic-embed-text-v2-moe
EMBEDDING_NOMIC_API_KEY=sk-not-set
EMBEDDING_NOMIC_KEEP_ALIVE=30m
EMBEDDING_NOMIC_TIMEOUT_SECONDS=30.0
# Empty = emet parity. The model card specifies `search_document: ` (ADR-0004),
# but the benchmark ran without it and the prefix shifts the vector a lot
# (cosine 0.57 on identical text) — so if you set this, the query side must
# send `search_query: ` to match, or retrieval gets worse rather than better.
EMBEDDING_NOMIC_DOCUMENT_PREFIX=
# Leave DIMENSIONS empty for text-embedding-3-large's native 3072, which is
# what was benchmarked. Setting it truncates via Matryoshka and is a
# re-embedding migration, not a config tweak.
EMBEDDING_OPENAI_BASE_URL=https://api.openai.com/v1
EMBEDDING_OPENAI_MODEL=text-embedding-3-large
EMBEDDING_OPENAI_API_KEY=
EMBEDDING_OPENAI_DIMENSIONS=
EMBEDDING_OPENAI_DOCUMENT_PREFIX=
EMBEDDING_OPENAI_TIMEOUT_SECONDS=30.0
# Sparse BM25 (ADR-0001, ADR-0005): the benchmarked `bm25-fa-norm-stop`.
# k/b saturation is applied client-side; IDF comes from Qdrant's
# modifier="idf" on the sparse vector field. AVG_LEN is the average document
# length in analyzer tokens — emet's placeholder, worth recalibrating from
# real corpus statistics.
EMBEDDING_SPARSE_ANALYZER=fa_norm_stop
EMBEDDING_SPARSE_K=1.2
EMBEDDING_SPARSE_B=0.75
EMBEDDING_SPARSE_AVG_LEN=256.0
# Parsing and chunking (ADR-0018).
# max_chunk_tokens is nomic-embed-text-v2-moe's sequence length; text past it
# is silently truncated by the model, so the cap is enforced before embedding.
# chunk_size sits under it to leave room for the `search_document: ` prefix.
CHUNKING_STRATEGY=fixed_size
CHUNKING_CHUNK_SIZE=400
CHUNKING_CHUNK_OVERLAP=60
CHUNKING_MAX_CHUNK_TOKENS=512
CHUNKING_ENCODING_NAME=cl100k_base
# tiktoken downloads its vocabulary on first use; point this at a
# pre-populated directory for offline/air-gapped deployments.
# TIKTOKEN_CACHE_DIR=

195
CLAUDE.md
View File

@@ -4,12 +4,135 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
## Project status ## Project status
This repo is currently ADR-driven and mostly pre-implementation: `src/` contains This repo is ADR-driven and early in implementation. Working today: the FastAPI
only an empty `main.py`/`config.py` scaffold and empty `api/routers`, app factory and lifespan wiring (`src/bootstrap/`), `/healthz` and `/readyz`,
`api/dependencies`, `db`, and `schemas` directories. Architecture decisions live structlog config, Postgres/MinIO/Qdrant clients (`src/infrastructure/`), five
in `docs/adr/` (17 ADRs plus the 0000 template; 0001–0004 are `Accepted`, SQLAlchemy models with one Alembic migration, document parsing plus fixed-size
0014 is `Superseded by 0017`, and the rest — 0005–0013 and 0015–0017 — are chunking (`src/application/ingestion/`), API-key auth, `POST`/`GET /v1/files`
`Proposed`). Implementation plans live in `docs/plans/`: with durable two-phase job creation (`src/application/files/`), and bounded
inline embedding: `dense_nomic`/`dense_openai` adapters over an
OpenAI-compatible HTTP client and a `bm25-fa-norm-stop` sparse adapter
(`src/infrastructure/embedding/`), wired into the upload path behind
`INGESTION_MAX_CONCURRENCY` (`503`), `INGESTION_TIMEOUT_SECONDS` (`504`), and
the chunk-count ceiling (`413`). The embedding configuration is **ported from
the `emet` evaluation lab** (`~/code/talie/emet`), which benchmarked these
models and analyzers on the real Farsi corpus — the analyzer and BM25 weights
are verified token-for-token against it, so treat them as a measured artifact
and re-benchmark rather than tune them in place (ADR-0005). Also working: the
`chunks` collection bootstrap (`src/infrastructure/qdrant/collection.py`, run as
a deployment step via `uv run python -m src.cli.qdrant_bootstrap` — never at
startup) and tenant-scoped point upserts (`src/application/points/` behind the
`PointStorage` port), so an upload is searchable by the time `201` returns.
Also working: `tenant_domains` plus `/v1/domains` (`src/application/domains/`),
a strict per-tenant allowlist — `POST /v1/files` rejects an unregistered or
disabled `domain` with `400` before anything is written, and domain management
sits behind its own `domains:read`/`domains:write` scopes, never `files:write`.
Also working: the operator runbook (`docs/runbook.md`), tenant/API-key/domain
provisioning (`uv run python -m src.cli.provision_tenant` — the third deployment
step, since nothing over HTTP can create the first tenant), a Testcontainers
e2e suite in the default pytest run (`tests/e2e/test_ingestion_slice.py`:
duplicate upload, retry after failure, tenant isolation, capacity, timeout,
parse and Qdrant failure), and the one Compose-based test — `scripts/smoke.sh`
driving `tests/e2e/test_compose_smoke.py` against a real uvicorn process, which
skips itself unless `SMOKE_BASE_URL` is set. That maps to plan 001 Phases 1-6
done.
Plan 002 (`/v1/points` CRUD and keyword search) is **Phases 1-3 done**. Phase 1
landed the `PointRepository` port (`src/application/ports/point_repository.py`)
with its `Point` read model (`src/application/points/point.py`), the Qdrant adapter
(`src/infrastructure/qdrant/point_repository.py`), request/response schemas
(`src/api/schemas/points.py`), and lifespan wiring. This port is **separate from
`PointStorage`**, which stays exactly the two bulk operations ingestion
performs — reads, single-point edits, and keyword search have a different caller
and a different tenant-filter obligation, so do not accrete them onto the
ingestion port. `tenant_id` is a required keyword argument on every
`PointRepository` method by design; keep it that way, because it is what turns a
forgotten tenant filter into a type error. The `chunks` collection also gained
full-text `content`, `is_active`, and `chunk_index` payload indexes, so a
deployed environment needs `qdrant_bootstrap` re-run (indexes are additive — no
rebuild, no re-embedding).
Two adapter mechanics there are load-bearing and easy to "simplify" into bugs:
reads go through `scroll` with a `HasIdCondition` rather than `retrieve` (which
takes no filter, and would move the tenant check to *after* Qdrant answered),
and ordered listing paginates by `order_id` value rather than offset (Qdrant
returns no page offset under `order_by`, and an offset cursor skips or repeats
rows under a concurrent insert).
Phase 2 added the **read routes**: `GET /v1/points/{point_id}`,
`GET /v1/points?file_id=...`, `GET /v1/points/count`, `GET /v1/points/search`,
and `GET /v1/files/{file_id}/points`, over `src/application/points/queries.py`
(`src/api/routers/points.py`). All are gated on `points:read`, which — with
`points:write` — is now in `DEFAULT_SCOPES`; `GET /v1/files/{file_id}/points`
uses `points:read` rather than `files:write`, so the scope follows the data
rather than the URL prefix. `PointNotFoundError` maps to `404` in
`src/api/errors.py`, never `403`. Three route-level rules are load-bearing:
`/count` and `/search` are declared **before** `/{point_id}` (FastAPI matches in
declaration order, so reordering them makes `/v1/points/count` a `422`),
`file_id` is **required** on the listing (the cursor is an `order_id` value and
`order_id` is unique only within one file), and `search_points` folds the query
with `normalize_persian_text` before matching, because ingestion letter-folds
content and an unfolded Arabic-keyboard query would return an empty result set
silently rather than erroring (ADR-0002).
Phase 3 added **soft delete**: `DELETE /v1/points/{point_id}` and
`DELETE /v1/files/{file_id}`, over `src/application/points/deletion.py` (with
the pure relinking primitive in `src/application/points/relinking.py`) and
`src/application/files/deletion.py`. Both are gated on `points:write` — the
file route included, since the data it destroys is points. Nothing is ever
removed from Qdrant.
Four rules there are load-bearing, and three of them look like complications
until the concurrency is taken seriously:
- `patches_for_removal` computes **what is still missing between the state just
read and the desired end state**, not "the patches a delete implies". That is
what makes a normal delete, a second delete of an already-inactive point (a
no-op success, never `404`), and recovery from a half-applied batch one code
path. Rewriting it as a straight-line "deactivate, patch prev, patch next"
breaks all three.
- Qdrant has no multi-point transaction and reports success for a filtered
`set_payload` that matched nothing, so a batch whose second operation loses a
version race applies its first anyway. `soft_delete_point` therefore re-plans
and re-applies up to three times, verifying by read-back, and only then raises
`PointVersionConflictError` (`409`). A single-shot delete would be able to
leave a stale pointer, which ADR-0002 calls a defect.
- A soft-deleted point **keeps its own** `previous_chunk_id`/`next_chunk_id`;
only the surviving neighbours are rewritten. Those pointers are unreachable
rather than stale, they are the only record of where the point sat, and the
retry re-plans from them. The whole-file sweep follows from the same rule:
every point leaves at once, so no survivor can dangle and no pointer is
touched at all.
- `DELETE /v1/files/{file_id}` marks the `source_files` row `soft_deleted`
**after** the point sweep, in its own short transaction (no session is held
across the Qdrant work). Order matters: a half-finished sweep leaves the row
`active` and a retried `DELETE` finishes it, and retiring the row is what
makes a later re-upload of the same bytes re-ingest instead of matching
`find_active_by_content_hash` and returning a file whose points are gone.
Audit rows are still Phase 4/6 work; Phase 3 emits log events only
(`points.soft_deleted`, `files.soft_deleted`, `points.relink.neighbour_missing`,
and the two `*.conflict` warnings). The completion and conflict events carry
ADR-0011's `duration_ms` plus `rounds`, and the pair is what makes them
diagnostic: relinking itself is O(1) (that is what the adjacency pointers buy),
so a single-point delete costs a fixed ~5 Qdrant round trips and a `rounds`
above 1 means contention, not a slow store. The whole-file sweep is the one
whose cost scales — two round trips per 100-point page.
Also worth knowing before touching the points tests: `tests/support/point_contract.py`
holds **one** scenario suite run against both `FakePointRepository` (unit) and
`QdrantPointRepository` (integration), so new repository behaviour belongs there
rather than in one of the two runners — that is what keeps the fake from drifting
more permissive than the real store.
Not built yet: plan 002 Phases 4-6 — create/replace/patch, reorder and batch,
the `api_request_logs`/`point_audit_events` tables, and the runbook section on
inspecting and repairing a file's pointer chain — and `src/agent/`.
Architecture decisions live in `docs/adr/` (18 ADRs plus the 0000 template;
0001–0004 are `Accepted` — 0004 amended by 0018; 0014 is `Superseded by 0017`;
the rest — 0005–0013 and 0015–0018 — are `Proposed`). Implementation plans live
in `docs/plans/`:
`001-ingestion-vertical-slice.md` and `001-ingestion-vertical-slice.md` and
`002-point-crud-and-keyword-search.md`. **Read the relevant `002-point-crud-and-keyword-search.md`. **Read the relevant
ADR(s) before implementing anything** — the ADRs are the source of truth for ADR(s) before implementing anything** — the ADRs are the source of truth for
@@ -106,6 +229,31 @@ MinIO/Qdrant/SQLAlchemy client-construction code. Use ports only for
external side effects/persistence — not around pure local functions. external side effects/persistence — not around pure local functions.
(ADR-0015) (ADR-0015)
### Prefer deep modules over shallow ones
When a package exposes several small pure functions that a caller must
compose correctly every time (right dispatch, right order, right
thread/async offload), give it one entry point that owns that composition,
and keep the small functions internal — exported only where their own unit
tests need them. A shallow interface (one whose surface is nearly as complex
as its implementation) pushes a correctness obligation onto every call site;
a deep one absorbs it once. Apply the deletion test when unsure: if deleting
the wrapper would concentrate the composition logic back into every caller
rather than just relocate it, the wrapper is worth having.
Worked example: `src/application/ingestion/` exposes `parse_and_chunk_document`
as its only caller-facing entry point. It dispatches on source type and owns
the `anyio.to_thread.run_sync` + `CapacityLimiter` offload ADR-0017 requires;
`parse_docx`/`parse_csv`/`parse_xlsx`/`chunk_document` stay in the package,
exported mainly for their own tests, not for outside callers to reach for
directly. `src/application/points/` follows the same shape: `index_chunks` is
the only caller-facing entry point, owning payload construction, batching,
the `upsert_concurrency` semaphore, and the ordering rule that the soft-delete
sweep runs only after every upsert succeeds; `build_chunk_payload` stays
internal. Follow this pattern in `application/` as new packages are added
there — `retrieval/`, `threads/` — rather than exposing their internals as the
primary surface.
### Resource lifetime rules (ADR-0012) ### Resource lifetime rules (ADR-0012)
- Application-lifetime objects (SQLAlchemy engine/sessionmaker, Qdrant client, - Application-lifetime objects (SQLAlchemy engine/sessionmaker, Qdrant client,
@@ -183,6 +331,11 @@ content_sha256)` idempotency, no terminal job returning to `running`.
### Postgres conventions (ADR-0009) ### Postgres conventions (ADR-0009)
`domain` is never free-form: it must match an `active` `tenant_domains` row for
the authenticated tenant (ADR-0009). Domain sets are per-tenant and vary in
size. The key itself is immutable — it is denormalized into every Qdrant point
payload and into `source_files`, so renaming it is a migration, not an edit.
UUID primary keys (app-generated), `timestamptz` for all timestamps, UUID primary keys (app-generated), `timestamptz` for all timestamps,
`Numeric(18, 8)` for money (never floats), `JSONB` for flexible metadata but `Numeric(18, 8)` for money (never floats), `JSONB` for flexible metadata but
typed/indexed columns for query-critical fields, string status columns with typed/indexed columns for query-critical fields, string status columns with
@@ -199,7 +352,24 @@ Postgres remains system of record for tenants, API keys, audit, jobs,
`graph_runs`, `llm_calls`/`llm_pricing`. Correlate the two via `request_id`, `graph_runs`, `llm_calls`/`llm_pricing`. Correlate the two via `request_id`,
`tenant_id`, `thread_id`, `run_id`. Use `structlog` with stable event names `tenant_id`, `thread_id`, `run_id`. Use `structlog` with stable event names
and structured fields (`logger.info("graph.run.completed", ...)`), not and structured fields (`logger.info("graph.run.completed", ...)`), not
interpolated prose; JSON logs by default in production. interpolated prose; JSON logs by default in production, plus an optional
local-only JSON file sink independent of the console renderer (`LOG_FILE_PATH`).
**Add logging in the same change that adds the code, not as a follow-up.**
When you add a new service-level entry point (an `application/` function a
route calls directly, an ingestion phase, a mutation) or a new failure branch
inside one, add its `logger.*` event in that same diff, using ADR-0011's
level/event-naming table. Deferring it means re-deriving the failure modes and
field names later from code that no longer has them in working memory — as
happened with `src/application/files/upload.py`, where four failure branches
(`parse_failed`, `chunk_limit_exceeded`, `embedding_failed`, `index_failed`)
shipped with no log event and had to be retrofitted.
This does not mean logging every function. Pure functions, models, schemas,
and repositories (`infrastructure/postgres/repositories/`) stay silent by
convention — the caller that turns their result into a business-meaningful
outcome (job succeeded, upload rejected, domain disabled) is where the event
belongs, not the row-level function underneath it.
## Testing (ADR-0016) ## Testing (ADR-0016)
@@ -213,9 +383,14 @@ interpolated prose; JSON logs by default in production.
- Test naming: `test_<unit>_<scenario>_<outcome>`, Arrange–Act–Assert. - Test naming: `test_<unit>_<scenario>_<outcome>`, Arrange–Act–Assert.
- Layout mirrors architecture: `tests/unit/{application,agent}`, - Layout mirrors architecture: `tests/unit/{application,agent}`,
`tests/integration/{postgres,minio,qdrant}`, `tests/e2e/`. `tests/integration/{postgres,minio,qdrant}`, `tests/e2e/`.
- Integration tests use **Testcontainers** (never a developer's local - Integration **and e2e** tests use **Testcontainers** (never a developer's
services or Langfuse-owned storage/credentials) — this is the standard local services or Langfuse-owned storage/credentials) — this is the standard
automated mechanism, not Docker Compose. Isolate data per test via unique automated mechanism, not Docker Compose. Compose is reserved for exactly one
thing: the serialized operational smoke test of the *running web process*
(`scripts/smoke.sh`), which is gated out of `uv run pytest`. Shared container
fixtures live in `tests/support/containers.py`, registered from the root
`tests/conftest.py` via `pytest_plugins` (a non-root conftest cannot declare
it). Isolate data per test via unique
keys/queue/collection names; parallel integration execution is disabled keys/queue/collection names; parallel integration execution is disabled
until fixture isolation is proven safe. until fixture isolation is proven safe.
- Pytest never calls a live/paid model provider in routine runs — that's - Pytest never calls a live/paid model provider in routine runs — that's

View File

@@ -2,6 +2,49 @@
Architecture decisions live in [`docs/adr`](docs/adr). The first implementation Architecture decisions live in [`docs/adr`](docs/adr). The first implementation
milestone is documented in the [ingestion vertical-slice plan](docs/plans/001-ingestion-vertical-slice.md). milestone is documented in the [ingestion vertical-slice plan](docs/plans/001-ingestion-vertical-slice.md).
Day-to-day operation — tuning the ingestion bounds, the proxy timeout
requirement, and how to investigate or retry a failed upload — is the
[operator runbook](docs/runbook.md).
## Provisioning the datastores
Both schema steps run as explicit deployment steps. The application performs no
DDL at startup — not for Postgres (ADR-0009) and not for Qdrant (ADR-0001,
"Collection provisioning").
```bash
docker compose up -d # Postgres, MinIO, Qdrant
uv run alembic upgrade head # Postgres schema
uv run python -m src.cli.qdrant_bootstrap # the `chunks` collection
uv run fastapi dev src/main.py
```
Nothing over HTTP can create the first tenant — every `/v1` route needs an API
key, and a key cannot exist before its tenant. One command issues both, plus any
domains, printing the key once (only its hash is stored):
```bash
uv run python -m src.cli.provision_tenant --slug acme --domain fire
```
Before a tenant can upload, its domains must be registered — `POST /v1/files`
rejects an unregistered or disabled `domain` with `400`. The calling backend
manages them over `/v1/domains` using a key with the `domains:write` scope:
```bash
curl -X POST http://localhost:8000/v1/domains \
-H "Authorization: Bearer $API_KEY" \
-H 'Content-Type: application/json' \
-d '{"domain": "fire", "display_name": "Fire insurance"}'
```
Both bootstrap commands are idempotent and safe to re-run. `qdrant_bootstrap` verifies an
existing collection against the pinned schema and exits non-zero on a mismatch,
rather than leaving a silently degraded sparse index in place.
`./scripts/smoke.sh` verifies the whole path — Compose up, both deployment
steps, provisioning, an upload through the running web process to indexed Qdrant
points. See the [runbook](docs/runbook.md#12-verifying-a-deployment).
## Local Langfuse ## Local Langfuse

View File

@@ -84,9 +84,9 @@ path_separator = os
# output_encoding = utf-8 # output_encoding = utf-8
# database URL. This is consumed by the user-maintained env.py script only. # database URL. This is consumed by the user-maintained env.py script only.
# other means of configuring database URLs may be customized within the env.py # Left unset here: env.py falls back to Settings().postgres.dsn (ADR-0009),
# file. # and test fixtures may override it programmatically before invoking Alembic.
sqlalchemy.url = driver://user:pass@localhost/dbname # sqlalchemy.url =
[post_write_hooks] [post_write_hooks]

View File

@@ -19,10 +19,13 @@ if config.config_file_name is not None:
fileConfig(config.config_file_name) fileConfig(config.config_file_name)
# Application models' MetaData, used for 'autogenerate' support. The database # Application models' MetaData, used for 'autogenerate' support. The database
# URL is likewise sourced from application settings, not alembic.ini, so both # URL is likewise sourced from application settings by default, so both
# migrations and the app read the same env-derived configuration (ADR-0009). # migrations and the app read the same env-derived configuration (ADR-0009) —
# unless a caller (e.g. a test fixture pointing at a Testcontainers database)
# has already set sqlalchemy.url on this Config before invoking Alembic.
target_metadata = Base.metadata target_metadata = Base.metadata
config.set_main_option("sqlalchemy.url", Settings().postgres.dsn) if not config.get_main_option("sqlalchemy.url"):
config.set_main_option("sqlalchemy.url", Settings().postgres.dsn)
def run_migrations_offline() -> None: def run_migrations_offline() -> None:

View File

@@ -0,0 +1,48 @@
"""create tenant_domains
Revision ID: 41335d162de8
Revises: bfc6c81c2542
Create Date: 2026-08-20 17:48:29.443293
"""
from typing import Sequence, Union
from alembic import op
import sqlalchemy as sa
from sqlalchemy.dialects import postgresql
# revision identifiers, used by Alembic.
revision: str = '41335d162de8'
down_revision: Union[str, Sequence[str], None] = 'bfc6c81c2542'
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
"""Upgrade schema."""
# ### commands auto generated by Alembic - please adjust! ###
op.create_table('tenant_domains',
sa.Column('id', sa.Uuid(), nullable=False),
sa.Column('tenant_id', sa.Uuid(), nullable=False),
sa.Column('domain', sa.String(length=80), nullable=False),
sa.Column('display_name', sa.String(length=200), nullable=False),
sa.Column('status', sa.String(length=20), server_default='active', nullable=False),
sa.Column('metadata', postgresql.JSONB(astext_type=sa.Text()), server_default='{}', nullable=False),
sa.Column('created_at', sa.DateTime(timezone=True), server_default=sa.text('now()'), nullable=False),
sa.Column('updated_at', sa.DateTime(timezone=True), server_default=sa.text('now()'), nullable=False),
sa.Column('disabled_at', sa.DateTime(timezone=True), nullable=True),
sa.CheckConstraint("status IN ('active', 'disabled')", name='ck_tenant_domains_status'),
sa.ForeignKeyConstraint(['tenant_id'], ['tenants.id'], ondelete='CASCADE'),
sa.PrimaryKeyConstraint('id'),
sa.UniqueConstraint('tenant_id', 'domain', name='uq_tenant_domains_tenant_id_domain')
)
op.create_index(op.f('ix_tenant_domains_tenant_id'), 'tenant_domains', ['tenant_id'], unique=False)
# ### end Alembic commands ###
def downgrade() -> None:
"""Downgrade schema."""
# ### commands auto generated by Alembic - please adjust! ###
op.drop_index(op.f('ix_tenant_domains_tenant_id'), table_name='tenant_domains')
op.drop_table('tenant_domains')
# ### end Alembic commands ###

View File

@@ -49,7 +49,7 @@ One collection, e.g. `chunks`, shared by all tenants and domains.
| Name | Type | Purpose | Notes | | Name | Type | Purpose | Notes |
|---|---|---|---| |---|---|---|---|
| `dense_nomic` | dense vector | primary semantic similarity (multilingual, incl. Persian) | `nomic-embed-text-v2-moe`, 768-dim ([0004](0004-docx-csv-chunking-strategy.md)) | | `dense_nomic` | dense vector | primary semantic similarity (multilingual, incl. Persian) | `nomic-embed-text-v2-moe`, 768-dim ([0004](0004-docx-csv-chunking-strategy.md)) |
| `dense_openai` | dense vector | second semantic signal | OpenAI large embedding model (e.g. `text-embedding-3-large`), dimension per OpenAI's `dimensions` param (TBD — full 3072 vs. a truncated size) | | `dense_openai` | dense vector | second semantic signal | `text-embedding-3-large` at its **native 3072 dimensions** — the `dimensions` param is deliberately left unset (see below) |
| `sparse` | sparse vector | lexical/keyword-sensitive retrieval | `bm25-fa-norm-stop` — Qdrant FastEmbed's BM25 sparse encoder configured for Persian (stopword removal + normalization), not a separately trained model | | `sparse` | sparse vector | lexical/keyword-sensitive retrieval | `bm25-fa-norm-stop` — Qdrant FastEmbed's BM25 sparse encoder configured for Persian (stopword removal + normalization), not a separately trained model |
| `late_interaction` | multivector | reserved for late-interaction rerank ([0003](0003-agent-hybrid-retrieval.md)) | `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)), `comparator: max_sim`, `hnsw_config: m=0` (rerank-only, never independently ANN-searched), stored **on disk** | | `late_interaction` | multivector | reserved for late-interaction rerank ([0003](0003-agent-hybrid-retrieval.md)) | `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)), `comparator: max_sim`, `hnsw_config: m=0` (rerank-only, never independently ANN-searched), stored **on disk** |
@@ -61,6 +61,35 @@ dense/sparse query latency. Two dense vectors are provisioned deliberately —
`dense_nomic` and `dense_openai` are two independent semantic signals, both `dense_nomic` and `dense_openai` are two independent semantic signals, both
prefetched and fused at query time (ADR-0003), not a primary/fallback pair. prefetched and fused at query time (ADR-0003), not a primary/fallback pair.
#### Dense model endpoints and dimensions (resolved by the `emet` benchmark)
Both dense models are reached over the **same OpenAI-compatible
`/embeddings` API**, so one adapter
(`src/infrastructure/embedding/openai_compatible.py`) serves both named
vectors with different configuration:
| Named vector | Model | Endpoint | Dimensions |
|---|---|---|---|
| `dense_nomic` | `nomic-embed-text-v2-moe` | self-hosted Ollama OpenAI-compat shim | **768** (verified against the live endpoint) |
| `dense_openai` | `text-embedding-3-large` | OpenAI hosted API | **3072** (native; `dimensions` unset) |
`dense_openai`'s dimension was previously listed as an open dependency. It is
now pinned to the native 3072, because that is the configuration the `emet`
lab benchmarked — it never passed a `dimensions` argument. Setting it later
would truncate via Matryoshka and is a **re-embedding migration, not a config
tweak**, exactly as the negative consequence below warns.
Two operational notes about the self-hosted embedder, both learned by
measurement rather than assumption:
- **Cold load exceeds 150s**, far beyond `INGESTION_TIMEOUT_SECONDS`, so an
idle-then-upload would return `504`. Mitigated on both ends: Ollama's
`keep_alive` keeps the model resident, and the FastAPI lifespan warms each
dense embedder at startup (fail-soft — a down embedder must not block boot).
- **Once warm it is fast**: ~0.30s for one input and ~0.34s for a batch of 16.
Batching is therefore nearly free, which is what keeps ADR-0017's inline
ingestion viable.
### Multitenancy / indexing config ### Multitenancy / indexing config
- HNSW: `m: 0` (disable the global index) + `payload_m: 16`, per Qdrant's - HNSW: `m: 0` (disable the global index) + `payload_m: 16`, per Qdrant's
@@ -79,6 +108,38 @@ prefetched and fused at query time (ADR-0003), not a primary/fallback pair.
- Payload index on `previous_chunk_id` / `next_chunk_id`: keyword index, - Payload index on `previous_chunk_id` / `next_chunk_id`: keyword index,
used for O(1) adjacency retrieval (see below). used for O(1) adjacency retrieval (see below).
### Collection provisioning
The collection is created by an explicit **deployment step**, not by application
startup and not lazily on first write:
uv run python -m src.cli.qdrant_bootstrap
Creating a collection is DDL, and this project already keeps DDL out of the boot
and request paths: [0009](0009-postgres-sqlalchemy-alembic-schema.md) requires
Alembic for Postgres schema and forbids `create_all()` at startup, and
[0012](0012-application-resource-lifetime-and-dependency-ownership.md) makes
LangGraph's `.setup()` a deployment step for the same reason. Neither ADR named
Qdrant explicitly; this section closes that gap rather than letting the placement
be decided by whichever code happened to need it first.
Doing it in the FastAPI lifespan was rejected: it couples process boot to Qdrant
being reachable (which is `/readyz`'s job, not boot's), races across replicas,
and turns a misconfigured collection into a silent skip. Doing it lazily on first
upsert was rejected for putting DDL on a user request and hiding the
misconfiguration until traffic arrives.
`ensure_chunks_collection` is idempotent and **verifying**: against an existing
collection it compares the dense dimensions and the sparse `modifier` to the
pinned values and fails loudly on divergence. That check is the point of making
the step explicit — both properties degrade silently in production if wrong (a
missing `modifier="idf"` produces no error, just unweighted lexical retrieval).
Payload indexes are (re)created on every run, since unlike vector configuration
they can be added to a live collection. The full-text index on `content` is
therefore deferred to the keyword-search work in
[0002](0002-chunk-crud-and-search-api.md), not created here.
### Payload schema ### Payload schema
This schema is now decided for the fields below. Additional document-context This schema is now decided for the fields below. Additional document-context
@@ -95,7 +156,7 @@ involves format-specific tradeoffs not yet made.
| `chunk_id` | keyword | stable identifier for a single chunk | | `chunk_id` | keyword | stable identifier for a single chunk |
| `content_type` | keyword | classification of the chunk's content; exact value set (e.g. `paragraph`, `table_row`, `heading`) to be finalized alongside the chunking-strategy ADR | | `content_type` | keyword | classification of the chunk's content; exact value set (e.g. `paragraph`, `table_row`, `heading`) to be finalized alongside the chunking-strategy ADR |
| `source_filename` | keyword | original uploaded filename | | `source_filename` | keyword | original uploaded filename |
| `source_type` | keyword (`docx` \| `csv`) | which parser produced this chunk | | `source_type` | keyword (`docx` \| `xlsx` \| `csv`) | which parser produced this chunk — `xlsx` added by [0018](0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md) |
| `order_id` | float (see below) | chunk's *display* position within the file; mutable so the backend can reorder/insert chunks | | `order_id` | float (see below) | chunk's *display* position within the file; mutable so the backend can reorder/insert chunks |
| `chunk_index` | integer | chunk's *original ingestion* ordinal — immutable, used to derive the deterministic point ID below (kept separate from `order_id` precisely because `order_id` can change) | | `chunk_index` | integer | chunk's *original ingestion* ordinal — immutable, used to derive the deterministic point ID below (kept separate from `order_id` precisely because `order_id` can change) |
| `previous_chunk_id` | keyword, nullable | `chunk_id` of the preceding chunk in display order (`null` for the first chunk in a file) — O(1) adjacency pointer for context-window expansion in ADR-0003 | | `previous_chunk_id` | keyword, nullable | `chunk_id` of the preceding chunk in display order (`null` for the first chunk in a file) — O(1) adjacency pointer for context-window expansion in ADR-0003 |
@@ -107,7 +168,7 @@ involves format-specific tradeoffs not yet made.
| `updated_at` | datetime | last modification timestamp | | `updated_at` | datetime | last modification timestamp |
| `created_by` | keyword | user/service that created the chunk | | `created_by` | keyword | user/service that created the chunk |
| `updated_by` | keyword | user/service that last modified the chunk | | `updated_by` | keyword | user/service that last modified the chunk |
| `version` | integer | optimistic-concurrency counter, used in ADR-0002 | | `version` | integer | optimistic-concurrency counter, used in ADR-0002. Ingestion currently writes `1` unconditionally: the read-check-write that makes the guard meaningful costs one read per point and belongs with the `/v1/points` write paths, so plan 002 owns it. Safe while ingestion is the only writer of a file's points; it would clobber a concurrent manual edit's counter once `/v1/points` ships. |
| `content_hash` | keyword | hash of the chunk's raw text; lets re-ingestion detect unchanged content and skip re-embedding it | | `content_hash` | keyword | hash of the chunk's raw text; lets re-ingestion detect unchanged content and skip re-embedding it |
| `embedding_model_version` | keyword | identifies which embedding model(s) produced this chunk's vectors; needed to know which chunks require re-embedding after a future model swap | | `embedding_model_version` | keyword | identifies which embedding model(s) produced this chunk's vectors; needed to know which chunks require re-embedding after a future model swap |
@@ -212,9 +273,14 @@ them — see ADR-0002 for how reorder/insert/delete operations keep
ingestion time and both are queried at retrieval time — roughly double ingestion time and both are queried at retrieval time — roughly double
the dense embedding cost/latency of a single-dense-vector design, plus an the dense embedding cost/latency of a single-dense-vector design, plus an
external network dependency on OpenAI's API in the ingestion path. external network dependency on OpenAI's API in the ingestion path.
- `dense_openai`'s exact output dimension is still an open dependency that - ~~`dense_openai`'s exact output dimension is still an open dependency~~ —
should be pinned before ingestion is implemented — changing it later is a **resolved**: pinned to the native 3072 (see "Dense model endpoints and
re-embedding migration, not a config tweak. dimensions" above). The warning still stands for any future change:
re-dimensioning is a re-embedding migration, not a config tweak.
- The `sparse` vector must be created with `modifier="idf"`. The client
computes only BM25's term-frequency saturation; without that modifier
Qdrant applies no IDF at all and lexical retrieval silently degrades
(ADR-0005).
- `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md)) - `jina-colbert-v2` ([0005](0005-reranking-model-and-sparse-analyzer-selection.md))
adds a hard GPU dependency to ingestion (not just query time, since the adds a hard GPU dependency to ingestion (not just query time, since the
document-side multivector is computed here) and its commercial license is document-side multivector is computed here) and its commercial license is

View File

@@ -75,6 +75,51 @@ retrieval used by the AI agent in ADR-0003; the two "search" concepts serve
different callers (a human/admin managing chunks vs. an agent retrieving different callers (a human/admin managing chunks vs. an agent retrieving
context) and should not be conflated in the API or in future discussion. context) and should not be conflated in the API or in future discussion.
Two properties follow from the index being a *filter*: results carry no
relevance score, and their order is unspecified. The API therefore returns
neither a score field nor a ranked list, and callers must not read the array
order as relevance. A caller that wants ranking wants ADR-0003's path.
#### The query is normalized the way ingested content was
`normalize_persian_text` (ADR-0018) folds Arabic letterforms to their Persian
equivalents — U+064A to U+06CC, U+0643 to U+06A9 — on every text block before
chunking, so stored `content` is uniformly Persian-formed. A query string is
not chunk content and never passes through that path, so a term typed on an
Arabic keyboard reaches the index as a different codepoint sequence than the
document it should match.
The service therefore applies the same folding to the query before matching.
Without it the endpoint fails in the worst available way: an exact-looking
query returns an empty result set, with no error, no warning, and nothing in
the logs to distinguish "no such term" from "the term is spelled with the
other yeh". Note this is a *query-side* transformation only — it changes what
is compared, never what is stored.
This does not extend to stemming or synonyms. Qdrant's full-text index offers
neither, and adding a Farsi analyzer here would duplicate the benchmarked BM25
sparse pipeline (ADR-0005) in a code path that is not benchmarked against
anything.
#### Listing is scoped to one file, and paginates by `order_id`
`GET /points?file_id=...` requires `file_id` rather than treating it as one
optional filter among several, and its pagination cursor is an `order_id`
value rather than an offset. Both follow from `order_id` being per-file:
- A cursor is only meaningful against a totally ordered key. `order_id` orders
points within one file and says nothing across files, so an unscoped listing
has no stable sort to paginate along.
- An offset cursor is wrong even within one file. Insert, reorder, and delete
all shift positions, so a page-two request issued after a concurrent insert
ahead of the cursor would repeat a row already returned — silently. Ranging
on `order_id > cursor` is unaffected: the reader has passed that value, and a
point inserted behind it was already served.
The second point depends on `order_id` being unique within a file, which the
gap-exhaustion rule below preserves by rejecting a reorder whose computed gap
would collapse onto a neighbour value.
### Delete is soft by default ### Delete is soft by default
`DELETE /points/{point_id}` and `DELETE /points?file_id=...` set `DELETE /points/{point_id}` and `DELETE /points?file_id=...` set
@@ -105,6 +150,41 @@ Qdrant's `update_filter`, giving an optimistic-concurrency-style guard
against races between a concurrent ingestion re-run (ADR-0001) and a manual against races between a concurrent ingestion re-run (ADR-0001) and a manual
edit through this API. edit through this API.
### Re-ingestion versus manual edits
A file can be re-uploaded after someone has hand-edited one of its points
through this API. **The newly ingested file wins.** Ingestion is authoritative
for the content of the file it ingested; a manual edit is a correction that
survives only until the source document is replaced.
Concretely:
- A point that still exists in the new version (same `file_id` +
`chunk_index`, hence the same deterministic point ID) is **overwritten in
place**. Ingestion performs a read-check-write so `version` is incremented
from whatever the manual edit left it at, rather than reset to `1`.
- A point from the previous ingestion that is **absent** from the new version
is flagged `is_active: false` with `deleted_at` set. It is never removed
from Qdrant — the soft-delete rule above applies to re-ingestion exactly as
it applies to `DELETE`.
- A manually created point (`POST /points`) is assigned a `chunk_index` past
the ingested range, so the same sweep deactivates it on the next upload of
its file. This is the intended consequence of "the new file wins", not an
accident of the sweep's bounds.
Because the point ID is derived from the immutable `chunk_index`, an
overwritten point cannot hold both the manual edit and the new file's content.
The clobbered content is therefore recorded in `point_audit_events`
(ADR-0009) as a `reingest_overwrite` operation carrying `before_version`, so
the edit is recoverable from the audit trail even though it is no longer a
live point.
Rejected alternative: preserving manual edits by having ingestion skip points
with `version > 1`. It breaks the guarantee that a successful upload leaves
Qdrant matching the uploaded document, and it needs a second, separate rule
for edited points that no longer exist in the new version — two divergent
notions of authority over one file.
### Re-embedding on content edit ### Re-embedding on content edit
`PUT /points/{point_id}` can change `content`, which leaves the stored `PUT /points/{point_id}` can change `content`, which leaves the stored

View File

@@ -4,6 +4,23 @@
Accepted Accepted
> Amended by
> [ADR-0018](0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md):
> v1 ships **fixed-size** chunking rather than the semantic-aware default
> below, and defers `qa_pair` detection and image captioning. Docx tables do
> become `table_row` units as this ADR specifies, but only when they are data:
> a table holding a cell larger than one chunk, or a cell containing nested
> tables, is treated as page layout and its cells are chunked as prose.
> Row labels are applied only when row 0 is provably a header, and cells are
> joined unlabeled otherwise. No document tree is built, and no heading is
> inferred from text. `.doc` is rejected with `415` pending an out-of-process
> conversion service rather than shelling out to LibreOffice. ADR-0018 also
> adds a Persian normalization step at parse time and fixes the chunk-size
> numbers this ADR left open. The spreadsheet row-to-chunk rules, the
> embedding model, the task-prefix invariant, and the `content_type` value set
> below all apply unchanged — except that `.csv` is read with the standard
> library `csv` module rather than `pandas`.
## Context ## Context
ADR-0001 deferred two things to "when we start the docx/csv chunking work": ADR-0001 deferred two things to "when we start the docx/csv chunking work":
@@ -45,6 +62,23 @@ from its model card: 768-dim output, Matryoshka-truncatable down to 256;
every embedded string — `search_document: ` at ingestion time, `search_query: ` every embedded string — `search_document: ` at ingestion time, `search_query: `
on the agent's query side (ADR-0003). on the agent's query side (ADR-0003).
> **Amendment — the task prefix is currently not applied.** The `emet`
> benchmark that selected this model ran *without* any prefix: its Ollama
> deployment's template is a bare `{{ .Prompt }}` passthrough that injects
> nothing, which was verified directly against the running endpoint. The
> prefix is not cosmetic — embedding the same Persian text with and without
> `search_document: ` yields a cosine of only **0.5741** — so applying it at
> ingest while the query side omits `search_query: ` would make retrieval
> *worse* than using neither.
>
> Implementation therefore defaults `EMBEDDING_NOMIC_DOCUMENT_PREFIX` to
> empty, matching the measured configuration, and exposes it as config so the
> prefixed variant is a one-line experiment rather than a code change. The
> model card remains the reason to expect prefixing to help; what is missing
> is evidence on *this* corpus. Turning it on is a paired change — ingest and
> query must move together — and should be settled by an emet run that
> measures the pair, not by an unmeasured edit here.
## Decision ## Decision
### Parsing order: structural extraction before chunking ### Parsing order: structural extraction before chunking

View File

@@ -2,9 +2,10 @@
## Status ## Status
Proposed — the fusion/rerank *shape* and reranker model are decided; the Proposed — the fusion/rerank *shape*, the reranker model, and (as of the
final BM25 analyzer and the commercial license status of the reranker are `emet` benchmark, see "Benchmark outcome" below) the **BM25 analyzer** are
still open per the follow-up items below. decided. The commercial license status of the reranker remains open per the
follow-up items below.
## Context ## Context
@@ -100,6 +101,50 @@ entirely to the analyzer stage, not the ranking formula:
consistent with Farsi's high density of function words (ezafe particles, consistent with Farsi's high density of function words (ezafe particles,
prepositions, common verbs) adding TF/IDF noise if left in. prepositions, common verbs) adding TF/IDF noise if left in.
### 3a. Benchmark outcome: `bm25-fa-norm-stop` confirmed, and where the BM25 math runs
The `emet` evaluation lab (`~/code/talie/emet`) ran the four-variant
comparison above against the real Farsi corpus and confirmed
**`bm25-fa-norm-stop`** as the winner. It is the only sparse variant promoted
into emet's hybrid matrix (`emet/hybrid.yaml`). This closes follow-up item 4
below.
The winning analyzer is a specific, reproducible artifact, ported into
`src/infrastructure/embedding/analyzers.py` and verified token-for-token
against emet's implementation. Its details are load-bearing:
- Unicode **NFC** (not NFKC), then ZWNJ → space, then Persian/Arabic-Indic
digits → ASCII, then `ي→ی ك→ک ة→ه ؤ→و إ→ا أ→ا`.
- Tokenizer `[^\W_]+`, which **keeps digits**. This matters for an insurance
corpus: policy numbers, dates, and amounts are exactly the terms lexical
retrieval should match, and the digit folding above means a query in ASCII
digits matches a document authored in Persian ones.
- A 51-entry stopword set (40 Persian/Arabic + 11 English, the corpus being
mixed-script). Deliberately not a full `hazm` list.
- No stemming — `fa_norm_stem` was the losing arm.
**The BM25 formula is split across two systems, deliberately.** The client
applies term-frequency saturation, including the `k`/`b` document-length
normalization; **IDF is supplied by Qdrant** via `modifier="idf"` on the
sparse vector field, computed from collection-wide statistics rather than
from a fixed client-side corpus.
That split is a correctness trap worth stating plainly: a `chunks` collection
created *without* `modifier="idf"` will score these vectors as saturated term
frequencies with no IDF weighting at all — no error, no warning, just
materially worse lexical retrieval. The collection bootstrap must set it.
Document and query encoding are asymmetric in exactly one term: documents
carry the `b` length normalization, queries do not (standard BM25 practice).
Both sides must therefore encode through the same implementation, which is
why the sparse port carries a `query` flag rather than leaving retrieval to
grow a second, silently divergent encoder.
Term → sparse-index mapping is `blake2b(token, digest_size=8) % (2**31 - 1)`,
a pure hash with no vocabulary table, so it needs no shared state and stays
identical across processes and between ingest and query time. Changing the
hash orphans every stored sparse vector: that is a re-ingestion, not a deploy.
### 4. BM25 parameters: keep `k=1.2`, `b=0.75`; tune analyzer, not formula ### 4. BM25 parameters: keep `k=1.2`, `b=0.75`; tune analyzer, not formula
These are standard, well-validated defaults (Trotman, Puurula & Burgess, These are standard, well-validated defaults (Trotman, Puurula & Burgess,
@@ -159,8 +204,21 @@ comparison and `b` sweep in the follow-ups below.
3. Run an ablation: single dense model + sparse + rerank vs. the current 3. Run an ablation: single dense model + sparse + rerank vs. the current
dual-dense-model + sparse + rerank setup, on real Farsi queries, to dual-dense-model + sparse + rerank setup, on real Farsi queries, to
justify (or drop) the second dense vector (`dense_openai`). justify (or drop) the second dense vector (`dense_openai`).
4. Compare `bm25-fa-norm-stop` vs. `bm25-fa-norm-stem` in isolation to 4. ~~Compare `bm25-fa-norm-stop` vs. `bm25-fa-norm-stem` in isolation~~ —
determine whether gains come from stopword removal, stemming, or both. **done**, see "Benchmark outcome" above. `fa_norm_stop` won; stemming was
not adopted.
5. Sweep BM25 `b` (e.g. 0.5–0.9) for the winning analyzer, since document 5. Sweep BM25 `b` (e.g. 0.5–0.9) for the winning analyzer, since document
length varies significantly across the corpus (short chat messages vs. length varies significantly across the corpus (short chat messages vs.
long articles) and `0.75` is a generic default, not corpus-tuned. long articles) and `0.75` is a generic default, not corpus-tuned.
6. **Recalibrate `avg_len`.** The client-side `b` term needs an average
document length in *analyzer tokens*. The ported value (256.0) is emet's
own placeholder, and emet measured it over short Q&A records rather than
this service's ~400-token chunks, so it is very likely miscalibrated here.
Exposed as `EMBEDDING_SPARSE_AVG_LEN` so it can be corrected from real
corpus statistics without a code change.
7. **Re-benchmark the analyzer with diacritic stripping.** `fa_norm_stop`
does not remove harakat or tatweel, so `ســلام` and `سلام` are distinct
terms. `src/application/ingestion/normalization.py` already strips both
for chunk *content*; extending that to the analyzer is plausibly an
improvement but would deviate from the measured configuration, so it
belongs in an emet run rather than an unmeasured edit.

View File

@@ -98,13 +98,16 @@ One row per customer/tenant.
| `slug` | Stable short name, unique, human-readable. | | `slug` | Stable short name, unique, human-readable. |
| `name` | Display name. | | `name` | Display name. |
| `status` | `active` \| `suspended` \| `deleted`. Suspended tenants authenticate to a clear error but cannot run work. | | `status` | `active` \| `suspended` \| `deleted`. Suspended tenants authenticate to a clear error but cannot run work. |
| `settings` | JSONB for tenant-level feature flags/limits (max upload size, enabled file types, allowed domains, etc.). | | `settings` | JSONB for tenant-level feature flags/limits (max upload size, enabled file types, etc.). Allowed domains were previously listed here as well; they live in `tenant_domains` instead, per this ADR's own rule that query-critical fields get typed columns — `domain` is validated on every upload and filtered on every query. |
| `created_at`, `updated_at`, `deleted_at` | Audit/soft-delete timestamps. | | `created_at`, `updated_at`, `deleted_at` | Audit/soft-delete timestamps. |
#### `tenant_domains` #### `tenant_domains`
Optional but recommended. Validates the `domain` values used throughout Qdrant **Required.** (Previously "optional but recommended"; implemented and made
payloads (`car`, `fire`, etc.) per tenant. mandatory alongside `/v1/domains`.) Validates the `domain` values used
throughout Qdrant payloads (`car`, `fire`, etc.) per tenant. Domain sets are
per-tenant and differ in size — one tenant may run 14 insurance lines and
another 6 — so this is data, not an enum.
| Column | Notes | | Column | Notes |
|---|---| |---|---|
@@ -116,7 +119,37 @@ payloads (`car`, `fire`, etc.) per tenant.
| `metadata` | JSONB for domain-specific ingestion/retrieval settings. | | `metadata` | JSONB for domain-specific ingestion/retrieval settings. |
This prevents arbitrary caller-supplied domains from silently creating new This prevents arbitrary caller-supplied domains from silently creating new
partitions in Qdrant. partitions in Qdrant. The failure it guards against is quiet: a typo such as
`fier` for `fire` produces no error anywhere — the file is stored, parsed,
embedded, and indexed into a partition retrieval never queries, so it is
invisible rather than failed.
##### Enforcement and management
- **Strict allowlist.** `POST /v1/files` rejects a domain with no `active` row
for the tenant (`400`, error code `unknown_domain`). There is no auto-create
on first use: that would record the typo rather than prevent it. The check
runs inside the upload's first transaction, before any MinIO object, job row,
or Qdrant point is written.
- **Managed over the API, not by an operator.** `/v1/domains` (list, create,
update, disable, enable) is the surface the calling backend uses. Domains are
created by an explicit, scoped call rather than as a side effect of an upload
— that distinction, not who makes the call, is what "strict" means here.
- **Its own scope.** `domains:read`/`domains:write`, deliberately separate from
`files:write`. Folding domain creation into the upload scope would let an
upload key create partitions again, which is the exact hole this closes.
`api_keys.scopes` is already a free JSONB list, so this needs no schema change.
- **`tenant_id` stays derived from the API key.** One key per tenant; nothing
request-suppliable. A platform key acting across tenants would need a real
actor model and is not adopted.
- **`domain` is immutable; `display_name` is not.** The key is denormalized into
every Qdrant point payload and into `source_files`, so renaming it means
rewriting all of them — a migration, not a `PATCH`. The update schema
therefore has no `domain` field.
- **Disable is not delete.** `status='disabled'` blocks new uploads and hides
the domain from listings, leaving already-indexed points intact and
retrievable. Actual removal needs the retention/erasure workflow this ADR and
plan 001 defer.
#### `api_keys` #### `api_keys`

View File

@@ -70,17 +70,30 @@ logger.info(
Do not build log messages by interpolating operational metadata into prose. Do not build log messages by interpolating operational metadata into prose.
Prefer fields over long strings because fields are queryable. Prefer fields over long strings because fields are queryable.
### Emit JSON logs by default in production ### Emit JSON logs by default in production; console and file are independent sinks locally
Production logs are JSON on stdout so process managers, container runtimes, and Production logs are JSON on stdout so process managers, container runtimes, and
log collectors can ingest them directly. Local development may use a colored log collectors can ingest them directly. This does not change.
console renderer controlled by configuration.
File logging is optional and mainly for local development. If enabled, it must Locally, stdout and an optional file are two **independent, simultaneous**
use explicit rotation settings such as `maxBytes` and `backupCount`. Do not rely handlers on the same logger, not a single renderer chosen by a flag — the same
on a default `RotatingFileHandler` with no rotation parameters. In containerized structlog event fans out to both:
production, stdout/stderr collection is preferred over writing `logs/app.log`
inside the application container. - **Console handler**: always on, `structlog.dev.ConsoleRenderer(colors=True)`.
This is what a developer reads while the process runs, so it stays
human-readable regardless of whether file logging is also enabled.
- **File handler**: off by default, enabled by setting `LOG_FILE_PATH`. Always
renders JSON (`structlog.processors.JSONRenderer()`), independent of the
console handler's renderer, so a saved log is machine-parseable even though
the terminal output next to it is not. Must use explicit rotation
(`RotatingFileHandler` with `maxBytes`/`backupCount` — never an unrotated
handler).
In containerized production, stdout/stderr collection remains preferred over
writing `logs/app.log` inside the application container, so `LOG_FILE_PATH` is
expected to be unset there; the file handler exists for local development,
where reading a colored terminal *and* keeping a JSON trail to grep/parse later
are both useful at once.
### Configure stdlib and structlog together ### Configure stdlib and structlog together
@@ -188,6 +201,36 @@ Notes:
- `structlog.contextvars.merge_contextvars` ensures request-bound fields appear - `structlog.contextvars.merge_contextvars` ensures request-bound fields appear
on both structlog and stdlib logs processed through the formatter. on both structlog and stdlib logs processed through the formatter.
### Bind process-level environment context once at startup
Deployment identity — which build is running, in which environment, on which
instance — answers a different question than request correlation: "is this
issue specific to one deployment / one region / one instance?" rather than "is
this issue specific to one request?" It does not vary per request, so it must
not go through `structlog.contextvars`, which `RequestIdMiddleware` clears on
every request; a value bound there before the first request would be wiped the
moment that middleware runs.
Instead, add a static structlog **processor** — a plain closure over values read
once at `configure_logging()` time — so it runs on every event regardless of
request context:
```python
def _bind_environment(settings: AppLimitSettings):
def processor(logger, method_name, event_dict):
event_dict["env"] = settings.env
event_dict["service_version"] = settings.service_version
return event_dict
return processor
```
`service_version` should be the deployed commit SHA or release tag (e.g. from a
`GIT_SHA`/`APP_VERSION` build-time env var — not computed at runtime by
shelling out to `git`). This makes "is this only happening on the new
deployment?" answerable directly from logs, without cross-referencing a
separate deployment record.
### Bind request context with contextvars ### Bind request context with contextvars
At FastAPI ingress, clear stale context, bind request identifiers, and return the At FastAPI ingress, clear stale context, bind request identifiers, and return the

View File

@@ -159,8 +159,25 @@ retry, and phase 2 has no transaction protecting it:
return the existing file/job rather than re-ingesting (plan 001). return the existing file/job rather than re-ingesting (plan 001).
- `tenant_id` comes from `AuthContext`, never from the request body. - `tenant_id` comes from `AuthContext`, never from the request body.
- A terminal job is never transitioned back to `running`. - A terminal job is never transitioned back to `running`.
- Qdrant points from a failed attempt do not replace the previous successful - A failed attempt never *removes* content from a working index. The
index; replacement happens only after a successful attempt. soft-delete sweep that retires a shortened file's leftover points runs only
after every upsert in the attempt has succeeded.
This is deliberately weaker than "replacement happens only after a successful
attempt", which an earlier revision of this ADR claimed. That guarantee is not
achievable alongside ADR-0001's deterministic point ids: those ids are exactly
what makes a retry idempotent, and they also mean a re-ingestion overwrites
points **in place**, so a crash partway through leaves a prefix updated and the
remainder still on the old content. Buying literal atomicity would mean
generation-suffixed ids and an activation flip, which contradicts ADR-0001 and
ADR-0002's stable point ids. Staging the new points as `is_active=false` and
flipping them on success is strictly worse — the in-place overwrite would
deactivate the previously live points, silently emptying a working index if the
attempt were interrupted.
What holds instead: the index is never emptied, never partially deleted, and a
retry converges — deterministic ids rewrite every point and the sweep re-runs,
reaching the exact correct state.
### Failures are HTTP failures ### Failures are HTTP failures

View File

@@ -0,0 +1,306 @@
# 0018. DOCX and spreadsheet parsing with fixed-size chunking
## Status
Proposed
## Context
ADR-0004 specified the full parsing and chunking design: structural extraction
before chunking (table rows and Q&A pairs kept atomic, prose chunked
separately), semantic-aware chunking as the default, image captioning via a
vision API, and legacy `.doc` conversion through headless LibreOffice. Plan
001's first vertical slice needs a working parser now, and that full design is
substantially more work than the slice can absorb. This ADR records what v1
actually ships and why it differs, so the code does not silently contradict an
Accepted ADR.
Four forces shaped the decision:
**A working extractor already exists.** The `chunking_strategies_evaluation`
repository — the harness the user built to compare chunking strategies against
this same Farsi corpus — contains a DOCX extractor that walks the document body
in reading order and handles the "prose lives inside table cells" pattern
common in Farsi documents exported from older Word versions. Roughly 200 lines
of it are production-quality; the rest is evaluation scaffolding (five
competing strategies, an LLM-as-judge benchmark, a dashboard, an HTML report
generator). Porting it is cheaper and better-tested against real documents than
writing a parser from scratch.
**Semantic chunking does not fit the inline request.** ADR-0004 chose
semantic-aware chunking as the default on the strength of the user's own
offline accuracy comparison, in which fixed-size was the "viable, simpler
runner-up". That comparison measured retrieval accuracy, not ingestion cost.
Semantic boundary detection requires embedding every sentence *before* chunk
boundaries can be decided — a second network round-trip pass inside the request
budget ADR-0017 bounds with `INGESTION_TIMEOUT_SECONDS`. ADR-0017's own cost
table already assumes the cheaper strategy, listing "Chunk (fixed-size,
ADR-0004) | Blocking CPU, pure Python | Negligible". ADR-0004 and ADR-0017 are
therefore already in tension, and this ADR resolves it toward ADR-0017 for v1.
Fixed-size is acceptable specifically because ADR-0001 gives every chunk
`previous_chunk_id`/`next_chunk_id` pointers: a chunk boundary that cuts a
thought in half is recoverable by expanding to neighbors at retrieval time
(ADR-0003).
**Nothing normalizes the text the dense embedders see.** ADR-0005 resolved the
sparse analyzer as `bm25-fa-norm-stop` — "normalization + stopword removal",
computed in our own BM25 pipeline outside Qdrant. That covers the sparse vector
only. Persian text authored on mixed Arabic/Persian keyboards contains both
`ک` (U+06A9) and `ك` (U+0643), both `ی` (U+06CC) and `ي` (U+064A); these are
distinct codepoints and therefore distinct tokens to `nomic-embed-text-v2-moe`
and `text-embedding-3-large` alike, so the same Persian word can embed two
different ways depending on which key the author pressed. Word documents in
this corpus reliably contain both forms.
**The corpus is spreadsheets more than it is CSVs.** ADR-0004 inspected the
real sample and found the tabular files are entirely `.xlsx`; no `.csv` exists
in practice. Plan 001 and the `source_files.source_type` CHECK constraint both
name `csv`. The row-to-chunk mapping is identical either way — only the reader
differs — so v1 reads both rather than forcing a manual export step that would
silently drop the merged-cell values ADR-0004 warns about.
## Decision
### 1. Formats
v1 ingests `.docx`, `.csv`, and `.xlsx`.
`.doc` is rejected with `415 Unsupported Media Type`. ADR-0004 specified
conversion via `soffice --headless --convert-to docx`; that is a subprocess
with a multi-second startup cost running inside the inline request ADR-0017
defines, and it adds a system binary to the container image. Conversion is
deferred to an out-of-process HTTP conversion service, tracked in the backlog.
`source_files.source_type` continues to allow `doc` so the row can be recorded
once conversion lands.
### 2. Structural units before chunking
Every source file is decomposed into ordered **structural units** before any
chunking runs, as ADR-0004 requires. There are two kinds:
- `PARAGRAPH` — a run of flowing prose. Consecutive paragraphs accumulate into
one unit rather than one unit each, so the splitter sees flowing text instead
of a series of one-sentence fragments ("the whole remaining run of
paragraphs", per ADR-0004).
- `TABLE_ROW` — one row of a data table. Atomic; split only when a single row
exceeds the model's sequence length.
Walk `doc.element.body` children in document order — not `doc.paragraphs` and
`doc.tables` separately — so tables interleaved with paragraphs keep their
position. This is ADR-0004's rule, unchanged.
**No headings are invented.** A real `Heading N` Word style becomes a `#`
prefix on its paragraph; where a document declares none, none appear. The
evaluation repository upgrades paragraphs to headings by text pattern (`^بخش`,
`^\d+[-.]\d+`, leading `*`); those patterns are tuned to a Farsi regulatory
corpus, not an insurance one, and a wrongly-detected heading silently reshapes
the document in a way that is hard to notice downstream. They are not adopted.
There is also **no document tree**. An earlier draft built a `Document >
Section > Article > Paragraph` hierarchy from heading styles. Not one document
in the sample corpus carries a single `Heading` style, so that tree was flat in
every real case, and nothing consumed it — chunking works from the unit list,
and ADR-0001's payload has no tree field. It is not built.
### 3. Data tables against layout tables
A docx table is either data or page furniture, and the two need opposite
treatment. The classification is **structural, never a reading of content**: a
data cell is by definition small enough to be a chunk, so a table containing a
cell that alone exceeds `chunk_size`, or a cell containing nested tables, is a
layout container. In the sample corpus this separates by two orders of
magnitude — 24 to 171 tokens for the largest cell of each data table, against
54,007 tokens for a cell holding an entire sub-document across 738 paragraphs
and 5 nested tables.
- **Data table** → one `TABLE_ROW` unit per row, rendered by the same code that
renders spreadsheet rows.
- **Layout table** → its cells are prose, recursed into and folded into the
surrounding prose block.
### 4. Table rows are labeled only when a header is provable
Applied to docx tables and spreadsheets alike:
- A header is **row 0 or nothing**. Never scan further down for a
header-shaped row. Scanning discarded every row above the match and then
labeled the rest from a data row, turning a 30-row compensation table into
chunks reading `80: 70`.
- Leading rows that are structurally a merged banner — fewer than two populated
cells, or one value repeated across the row — are skipped first. That is a
fact about the merge, not a guess about meaning.
- Row 0 is accepted as a header only when it is **inconsistent with the column
beneath it**: a text label above a numeric column, or a short label above much
longer values. This is the test `csv.Sniffer.has_header` uses; it is a
property of the table rather than a pattern borrowed from one document.
- When no header is provable, cells are joined with `" | "` — unlabeled, but
never mislabeled. Losing a label is recoverable at retrieval time; labeling
every row from a data row is not.
- A header merged vertically across two rows resolves to the same text in the
row below it; that duplicate is skipped rather than emitted as data.
- A cell merged across columns is reported once per grid position it spans;
those repeats are collapsed.
is a pipeline invariant
### 5. Persian normalization is a pipeline invariant
Every extracted text block is normalized before chunking, for all formats:
- Arabic to Persian letter folding: `ك`→`ک`, `ي`→`ی`, `ى`→`ی`, `أ`/`إ`→`ا`
- `unicodedata.normalize("NFKC")`
- Removal of harakat (diacritics) and tatweel
- `¬` → space, then collapse runs of whitespace
Digits and punctuation are **not** rewritten. Persian digits (`۱۲۳`) and
Persian punctuation (`؛`, `٬`) are left as authored, because chunk `content` is
what citations render back to the user and Western digits inside Persian prose
read as wrong.
Normalization runs **per text block, before the markdown is assembled** — the
whitespace-collapse step maps `\n` to a space, so applying it to an assembled
document would flatten every heading and paragraph onto a single line.
This complements rather than replaces ADR-0005's sparse-side normalization,
which additionally removes stopwords and is specific to the BM25 vector. It
also stabilizes the text that feeds `content_hash`.
### 6. Spreadsheet handling
ADR-0004's rules stand, under the header discipline of section 4: each cell is
rendered as `"{column_header}: {cell_value}"`, merged cell ranges are
forward-filled before rendering, and sheets with no non-empty data rows are
skipped.
Forward-filling merges is not cosmetic. openpyxl stores a merged range's value
only in its top-left cell, so in the branch directory the province is present
on the first branch of each province and absent from every other one. Filling
the range makes each row-chunk self-contained — a branch carries its province
even though the source cell is blank.
One correction: `.csv` is read with the standard library's `csv` module, not
`pandas` as ADR-0004 states. Adding pandas for delimiter handling and row
iteration is not warranted. `.xlsx` uses `openpyxl`, as ADR-0004 assumed.
There is no sheet-shape sniffing. A two-column Q&A sheet and a branch-directory
sheet go through the same generic renderer; a Q&A row renders as
`question: …\nanswer: …` and a branch row as `branch_name: …\ncity: …`, both
self-describing without a schema heuristic that could misfire.
### 7. Chunking
The `fixed_size` strategy, over tokens counted with tiktoken `cl100k_base`:
| Setting | Value |
|---|---|
| `chunk_size` | 400 tokens |
| `chunk_overlap` | 60 tokens |
| `max_chunk_tokens` | 512 (hard cap) |
These numbers are recorded here because they exist in no ADR today — ADR-0001
explicitly left "size/overlap are tunable config, not fixed by this ADR" open,
and ADR-0004 gives only the 512 ceiling.
`cl100k_base` is a deliberate proxy. `text-embedding-3-large` has an 8191-token
window and never binds; `nomic-embed-text-v2-moe`'s 512-token sequence length
is the only real constraint. cl100k tokenizes Persian inefficiently while
nomic's multilingual tokenizer does not, so a cl100k count reliably
*over-estimates* the nomic count — measuring with cl100k and capping at 512 is
safe in the conservative direction, without shipping a second tokenizer and its
model download into the ingestion path. The 400/512 gap leaves headroom for the
mandatory `search_document: ` task prefix (ADR-0004) and any heading text
carried into a chunk.
Spreadsheet rows are atomic and bypass the splitter. A row that exceeds
`max_chunk_tokens` falls through the fixed-size splitter in place, emitting
several ordered chunks, rather than being silently truncated by the embedding
model.
### 8. `content_type` values emitted
ADR-0004's four-value set is unchanged. v1 emits `paragraph` (DOCX prose and
flattened tables) and `table_row` (spreadsheet rows). `qa_pair` and
`image_caption` remain defined but are not produced.
### 9. Deferred, not rejected
Semantic-aware chunking; DOCX `table_row` and `qa_pair` structural detection;
embedded-image captioning; `.doc` conversion; Matryoshka-256 truncation. Each
remains ADR-0004's stated intent; this ADR only records that v1 does not ship
them.
## Consequences
### Positive
- Plan 001's ingestion slice is unblocked with a parser already proven against
this specific Farsi corpus, rather than one written speculatively.
- Chunking stays pure, synchronous, and cheap — it fits ADR-0017's inline
request budget with no network round-trip, and runs safely under
`anyio.to_thread.run_sync`.
- Persian normalization closes a real defect that would otherwise degrade both
dense vectors silently, with no error and no obvious symptom.
- Concrete chunk-size numbers and their rationale are now recorded somewhere
other than a config default, so a later change is a visible decision.
- Q&A sheets, branch directories, and docx contact tables are all served by one
renderer with no schema sniffing, so a new sheet shape needs no new code.
- Refusing to label a table whose header is unprovable means the parser degrades
to unlabeled rows instead of producing confidently wrong ones, which is the
failure mode that is hard to notice downstream.
### Negative
- **v1 ships the strategy the user's own comparison ranked second.** Retrieval
accuracy is expected to be measurably lower than semantic chunking would
give. Neighbor expansion via ADR-0001's pointers is the mitigation, and it is
unproven at this scale.
- A table whose header cannot be proven — short text over short text, which is
genuinely ambiguous — produces unlabeled `" | "` rows. A reader or model can
still see the values but not which column each belongs to.
- The layout-table rule keys on `chunk_size`, so changing that setting silently
changes which tables are treated as data. The observed margin is two orders of
magnitude, so this is unlikely to flip in practice, but it is a coupling.
- A DOCX that alternates literal `سوال:`/`پاسخ:` paragraphs loses question/answer
atomicity; a boundary can fall between a question and its answer, because
`qa_pair` detection is deferred.
- cl100k is a proxy for nomic's tokenizer. The relationship is safe in the
conservative direction for Persian, but a document in another language could
in principle tokenize the other way; the 512 assertion is what catches it.
- Normalizing stored `content` means the text served in citations is not
byte-identical to the source document. Letter folding was chosen over full
normalization specifically to keep this difference invisible to a reader.
- `.doc` files are rejected outright rather than converted, so any legacy
document must be re-saved by hand until the conversion service lands.
## Alternatives Considered
- **Implement ADR-0004 in full now** (DOCX table-row detection with header
inference, Q&A pair heuristics, image captioning, semantic boundary
detection): rejected for v1 as roughly triple the work, none of which exists
in the evaluation repository to port, and which would block the first
ingestion slice on parser research.
- **Semantic chunking inside the inline request**: rejected — it adds a
per-sentence embedding pass to a request already bounded by
`INGESTION_TIMEOUT_SECONDS`, and ADR-0017 chose inline ingestion on the
assumption that chunking is negligible. Revisit when ingestion moves back off
the request path.
- **The `nomic-embed-text-v2-moe` tokenizer** for exact chunk sizing: rejected
— it requires `transformers`/`tokenizers` and a model file download in the
ingestion path to buy precision that the conservative cl100k over-estimate
already provides.
- **Character-based splitting** (no tokenizer at all): rejected — Persian
characters-per-token varies enough that a character budget cannot guarantee
the 512-token ceiling that actually matters.
- **Full Persian normalization** including digit unification (`۱۲۳`→`123`) and
punctuation mapping: rejected — it would improve lexical matching slightly
when a query uses the other digit form, at the cost of rendering Persian
citations with Western digits.
- **Storing raw and normalized text as separate payload fields**: rejected —
ADR-0001 fixes the payload field list, and doubling the stored text per point
is not justified when letter folding alone is visually lossless.
- **Row-chunking DOCX tables like spreadsheets**: rejected for v1 — the Farsi
`.doc` exports in this corpus use tables as page layout, with ordinary prose
inside cells, so treating every row as an atomic unit would shred paragraphs
mid-sentence. Revisit once real chunk output from the table-heavy documents
has been inspected.

109
docs/backlog.md Normal file
View File

@@ -0,0 +1,109 @@
# Backlog
Ideas and open questions not yet ready to be an ADR decision or a plan phase.
Each entry is short: what the idea is, which ADR/plan it would eventually
touch, and what's still unresolved. When an entry is picked up, turn it into
an ADR amendment (or a new ADR) and delete it from here — this file is not a
permanent record, `docs/adr/` is.
## Legacy `.doc` conversion
Relates to: [ADR-0018](adr/0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md).
`.doc` is rejected with `415`. ADR-0004 specified `soffice --headless
--convert-to docx`, which ADR-0018 rejected as a multi-second subprocess inside
an inline request. Gotenberg is the obvious candidate since it is already in
use elsewhere — **but verify before committing to it**: Gotenberg's LibreOffice
route is built for converting *to PDF*, and `.doc` → `.docx` output may not be
supported on that endpoint. If it is not, the options are a dedicated
LibreOffice sidecar or asking uploaders to re-save.
## Structural units ADR-0004 specifies but v1 does not emit
Relates to: [ADR-0004](adr/0004-docx-csv-chunking-strategy.md),
[ADR-0018](adr/0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md).
- **`qa_pair`**: one sample document alternates literal `سوال:`/`پاسخ:`
paragraphs. v1 chunks it as prose, so a chunk boundary can fall between a
question and its answer.
- **`image_caption`**: two sample documents embed images with no alt text.
ADR-0004 routes these through a vision API at ingest; v1 drops them silently.
Both need a decision on whether heuristic detection is worth the misfire risk —
the header-detection work showed that guessing structure is expensive when wrong.
## Tables whose header cannot be proven
Relates to: [ADR-0018](adr/0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md).
A table of short text over short text (`branch,city` with no numeric or long
column) is genuinely ambiguous, so v1 emits unlabeled `" | "` rows rather than
risk labeling every row from a data row. No file in the current corpus hits
this, but a future one will.
The honest fix is not a better heuristic — it is to stop guessing: let the
upload declare whether a sheet has a header, since the uploader knows. That is
a `POST /v1/files` contract change, so it belongs with plan 001 Phase 3 rather
than in the parser.
## Recalibrate chunk size against nomic's tokenizer
Relates to: [ADR-0018](adr/0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md).
**Revisit after retrieval quality is measurable** — deliberately deferred, not
forgotten.
ADR-0018 counts tokens with tiktoken `cl100k_base` and caps chunks at 512. But
512 is `nomic-embed-text-v2-moe`'s limit, measured in *nomic's* tokenizer, not
OpenAI's. Those are different units, and on Persian they differ by a lot.
Measured against a real production document (`bimeh_havades.docx`, 5,911 chars
of Farsi) via the Ollama server that already hosts the model:
| Sample | cl100k tokens | nomic tokens | ratio |
|---|---|---|---|
| 300 chars | 217 | 80 | 2.71 |
| 600 chars | 426 | 165 | 2.58 |
| 1,200 chars | 846 | 303 | 2.79 |
So **~2.7 cl100k tokens per nomic token** on Persian. The current
`chunk_size=400` is therefore about **148 nomic tokens — roughly 29% of the
512-token window**. Chunks land near 570 characters where ~1,500 would fit.
Two things this measurement also established:
- **Silent truncation is real, and now demonstrated.** Feeding 2,400 and 4,800
characters both returned `prompt_eval_count` of exactly 512, with no error
and no warning. This is what ADR-0004 meant by "silently truncated by the
model, not an error", confirmed on our own hardware.
- **Measuring nomic tokens needs no new dependency.** Ollama's `/api/embed`
returns `prompt_eval_count`, so the real count is obtainable from the
embedding call we already have to make. Note the value saturates at 512, so
it cannot measure anything longer than the window — calibration samples must
stay under it.
When picking this up, decide between: raising `chunk_size`/`max_chunk_tokens`
in cl100k terms using a calibration ratio (cheap, drifts if the corpus language
mix changes); counting with nomic's own tokenizer offline via HuggingFace
`tokenizers` and its `tokenizer.json` (exact, and lighter than ADR-0018
assumed — the tokenizer file only, not the 475M-param model weights); or
keeping small chunks because neighbor expansion recovers the context anyway.
Do not change this on the ratio alone. The reason to keep 400/60/512 for now is
that smaller chunks are not automatically worse for retrieval — measure
retrieval quality first, then decide.
Also note `nomic-embed-text:latest` (v1.5) is on the same Ollama server with a
2,048-token context, but it is the English-focused model; v2-moe is the
multilingual one and the reason ADR-0004 chose it for Farsi. Do not switch to
v1.5 just to get a bigger window.
## Get LLM usage/price from the OpenAI API
Relates to: [ADR-0009](adr/0009-postgres-sqlalchemy-alembic-schema.md)'s
`llm_calls`/`llm_pricing` tables.
Get token usage and price from the OpenAI API's response metadata, instead of
computing/tracking them ourselves. Need to check whether OpenAI actually
returns price, or only token counts — if only counts, we still need
`llm_pricing` for price and this only changes how `llm_calls` gets its
usage numbers.

View File

@@ -3,7 +3,7 @@
## Purpose ## Purpose
This plan turns the accepted architectural direction in the ADRs into the first This plan turns the accepted architectural direction in the ADRs into the first
working product slice: a tenant-scoped CSV upload is stored in MinIO, represented working product slice: a tenant-scoped DOCX/XLSX/CSV upload is stored in MinIO, represented
by durable Postgres records, parsed/chunked/embedded inline in the request by durable Postgres records, parsed/chunked/embedded inline in the request
(ADR-0017), and indexed as Qdrant points before the response returns. (ADR-0017), and indexed as Qdrant points before the response returns.
@@ -47,14 +47,19 @@ them.
### In scope ### In scope
- `POST /v1/files` for authenticated tenant-scoped **CSV** upload. - `POST /v1/files` for authenticated tenant-scoped **DOCX, XLSX, and CSV** upload.
`.doc` is rejected with `415` pending an out-of-process conversion service
(ADR-0018). An earlier revision of this plan scoped the slice to CSV only and
placed DOCX out of scope; the real corpus is DOCX and XLSX, so ADR-0018
corrects that.
- File validation, size limits, content hashing, and streaming upload to MinIO. - File validation, size limits, content hashing, and streaming upload to MinIO.
- Alembic-managed Postgres schema for the minimal tenant/auth, source file, - Alembic-managed Postgres schema for the minimal tenant/auth, source file,
ingestion job, and job event records needed by this slice. ingestion job, and job event records needed by this slice.
- Inline ingestion in `POST /v1/files`, with batched/bounded-concurrent - Inline ingestion in `POST /v1/files`, with batched/bounded-concurrent
embedding, thread-offloaded parsing, and enforced size/timeout/capacity embedding, thread-offloaded parsing, and enforced size/timeout/capacity
bounds. bounds.
- CSV parsing and deterministic chunk creation. - DOCX/XLSX/CSV parsing into structural units and deterministic chunk creation
(ADR-0018, implemented in `src/application/ingestion/`).
- Tenant-filtered Qdrant point upserts using deterministic point identifiers. - Tenant-filtered Qdrant point upserts using deterministic point identifiers.
- Job status/progress persistence and `GET /v1/files/{file_id}` status lookup. - Job status/progress persistence and `GET /v1/files/{file_id}` status lookup.
- Structured correlation logging at HTTP and ingestion-stage boundaries. - Structured correlation logging at HTTP and ingestion-stage boundaries.
@@ -63,7 +68,11 @@ them.
### Explicitly out of scope ### Explicitly out of scope
- XLSX, DOCX, and legacy DOC ingestion. - Legacy `.doc` ingestion — rejected with `415` until an out-of-process
conversion service exists (ADR-0018). DOCX and XLSX are **in** scope; they were
listed here before ADR-0018 corrected the scope line.
- `qa_pair` structural detection and embedded-image captioning (ADR-0004),
deferred by ADR-0018.
- The conversational LangGraph API and SSE streaming. - The conversational LangGraph API and SSE streaming.
- Final reranker selection, GPU deployment, or unresolved model licensing from - Final reranker selection, GPU deployment, or unresolved model licensing from
ADR-0005. ADR-0005.
@@ -116,9 +125,12 @@ them:
- Use `(tenant_id, domain, content_sha256)` to recognize identical uploads. - Use `(tenant_id, domain, content_sha256)` to recognize identical uploads.
- An identical active upload should return the existing source-file/job reference - An identical active upload should return the existing source-file/job reference
rather than create a duplicate ingestion. rather than create a duplicate ingestion.
- A changed upload creates a new ingestion job. Existing active Qdrant points are - A changed upload creates a new ingestion job. A failed re-ingestion never
replaced only after the new job completes successfully, so a failed re-ingestion removes a working index: the soft-delete sweep for a shortened file runs only
does not remove a working index. after every upsert has succeeded. Because ADR-0001's point ids are
deterministic, upserts overwrite in place, so an interrupted attempt can leave
a prefix updated — it cannot empty or partially delete the index, and a retry
converges. See ADR-0017, "Re-running an ingestion stays safe".
- Preserve the original filename in Postgres metadata. MinIO object keys remain - Preserve the original filename in Postgres metadata. MinIO object keys remain
internal ID-based paths. internal ID-based paths.
@@ -196,7 +208,7 @@ reads/writes and valid job transitions.
### Phase 3: MinIO upload and durable job creation ### Phase 3: MinIO upload and durable job creation
1. Implement API-key authentication and `AuthContext` tenant derivation. 1. Implement API-key authentication and `AuthContext` tenant derivation.
2. Implement `POST /v1/files` for CSV only, including streaming-size controls, 2. Implement `POST /v1/files` for DOCX, XLSX, and CSV, including streaming-size controls,
file-type validation, SHA-256 calculation, and a private MinIO upload using file-type validation, SHA-256 calculation, and a private MinIO upload using
an internal object key. an internal object key.
3. In one short Postgres transaction, persist `source_files` and create 3. In one short Postgres transaction, persist `source_files` and create
@@ -208,11 +220,11 @@ reads/writes and valid job transitions.
response that does not expose raw storage credentials or internal artifacts. response that does not expose raw storage credentials or internal artifacts.
6. Add cleanup/compensation handling for a MinIO upload that succeeds while the 6. Add cleanup/compensation handling for a MinIO upload that succeeds while the
database transaction fails. database transaction fails.
7. Add unit/API tests for trusted tenant derivation, CSV validation, idempotency, 7. Add unit/API tests for trusted tenant derivation, upload validation, idempotency,
the terminal `201 Created` response, and tenant-scoped status. Add MinIO adapter integration tests the terminal `201 Created` response, and tenant-scoped status. Add MinIO adapter integration tests
for server-derived private object paths and compensation behavior. for server-derived private object paths and compensation behavior.
**Exit criteria:** an authenticated CSV upload creates a private object and a **Exit criteria:** an authenticated upload creates a private object and a
`running` job row committed before any ingestion work; a tenant cannot retrieve `running` job row committed before any ingestion work; a tenant cannot retrieve
another tenant's file status. another tenant's file status.
@@ -238,21 +250,36 @@ code and a terminal job row.
### Phase 5: Ingestion execution and Qdrant Chunk/Point CRUD ### Phase 5: Ingestion execution and Qdrant Chunk/Point CRUD
> **Carried forward from Phase 4 — the `chunks` collection must create the
> `sparse` vector with `modifier="idf"`.** The BM25 adapter computes only
> term-frequency saturation client-side; IDF comes from Qdrant's
> collection-wide statistics. Omit the modifier and there is no error and no
> warning — sparse scoring silently loses its IDF term and lexical retrieval
> degrades. See ADR-0005, "Benchmark outcome".
>
> Collection creation must also use the pinned dimensions from ADR-0001:
> `dense_nomic` 768, `dense_openai` 3072.
1. Implement the ingestion service called by the route, using the 1. Implement the ingestion service called by the route, using the
application-lifetime database, MinIO, Qdrant, model, and logging clients. application-lifetime database, MinIO, Qdrant, model, and logging clients.
2. Validate the persisted records before fetching the MinIO object. 2. Validate the persisted records before fetching the MinIO object.
3. Append progress events, parse CSV, create deterministic chunks, embed them, 3. Append progress events, parse the document, create deterministic chunks, embed them,
and upsert tenant-scoped Qdrant points — without holding a Postgres session and upsert tenant-scoped Qdrant points — without holding a Postgres session
open across the work. open across the work.
4. In a second short transaction, mark the job `succeeded` with counters or 4. In a second short transaction, mark the job `succeeded` with counters or
`failed` with a safe error summary, then return the terminal response. `failed` with a safe error summary, then return the terminal response.
5. Make a retried upload safe: no duplicate logical chunks, no incorrect 5. Make a retried upload safe: no duplicate logical chunks, no incorrect
counters, and no transition from a terminal state back to `running`. counters, and no transition from a terminal state back to `running`.
6. Add unit tests for deterministic CSV chunks, point IDs, and terminal job 6. Add unit tests for deterministic chunks, point IDs, and terminal job
transitions. Add Testcontainers Qdrant and Postgres integration tests for transitions. Add Testcontainers Qdrant and Postgres integration tests for
tenant-filtered upserts, terminal state persistence, retrying an upload, and tenant-filtered upserts, terminal state persistence, retrying an upload, and
parser/Qdrant failure handling. parser/Qdrant failure handling.
The `chunks` collection itself is provisioned by a deployment step —
`uv run python -m src.cli.qdrant_bootstrap` — not by FastAPI startup, for the
same reason ADR-0009 keeps Alembic out of startup and ADR-0012 makes LangGraph's
`.setup()` a deployment step. See ADR-0001, "Collection provisioning".
**Exit criteria:** a successful upload returns `201` with a terminal status, and **Exit criteria:** a successful upload returns `201` with a terminal status, and
its points are retrievable only under the owning tenant's Qdrant filter. A forced its points are retrievable only under the owning tenant's Qdrant filter. A forced
failure mid-ingestion produces a `failed` job and the right HTTP status, and failure mid-ingestion produces a `failed` job and the right HTTP status, and
@@ -262,19 +289,30 @@ retrying the upload produces a correct final state without duplicate chunks.
1. Add an operator runbook covering local startup, migrations, MinIO bucket 1. Add an operator runbook covering local startup, migrations, MinIO bucket
setup, the run command, ingestion-bound tuning, the proxy/client timeout setup, the run command, ingestion-bound tuning, the proxy/client timeout
requirement, and how to retry a failed ingestion. requirement, and how to retry a failed ingestion. — `docs/runbook.md`.
2. Add a serialized Compose-based operational smoke test covering upload through 2. Add a serialized Compose-based operational smoke test covering upload through
indexed points against the running web process. Testcontainers remains the indexed points against the running web process. Testcontainers remains the
standard pytest mechanism for individual adapter integration tests. standard pytest mechanism for individual adapter integration tests. —
`scripts/smoke.sh` driving `tests/e2e/test_compose_smoke.py`, which skips
itself unless `SMOKE_BASE_URL` is set so `uv run pytest` never invokes
Compose.
3. Add end-to-end tests for duplicate upload, retrying a failed upload, tenant 3. Add end-to-end tests for duplicate upload, retrying a failed upload, tenant
isolation, capacity/timeout rejection, and failed parser/Qdrant behavior. isolation, capacity/timeout rejection, and failed parser/Qdrant behavior. —
`tests/e2e/test_ingestion_slice.py`, on Testcontainers, in the default suite.
4. Add health/readiness checks that distinguish process health from dependency 4. Add health/readiness checks that distinguish process health from dependency
readiness. readiness. — `/healthz` and `/readyz`; `/readyz` additionally requires the
`chunks` collection to exist, since a reachable but unbootstrapped Qdrant
would `502` on the first upload.
5. Update the README with local-start instructions and links to ADRs, this plan, 5. Update the README with local-start instructions and links to ADRs, this plan,
and the operations runbook. and the operations runbook.
Provisioning a tenant and its first API key turned out to be a prerequisite for
1 and 2 rather than a separate milestone: nothing over HTTP can create the first
tenant, so `src/cli/provision_tenant.py` was added alongside the other two
deployment-step commands.
**Exit criteria:** a new developer can start the stack, apply migrations, upload a **Exit criteria:** a new developer can start the stack, apply migrations, upload a
CSV, observe the job through completion, and understand how to investigate or a document, observe the job through completion, and understand how to investigate or
retry a failure. retry a failure.
## Definition of done for the vertical slice ## Definition of done for the vertical slice
@@ -283,7 +321,7 @@ The first slice is done when the following path works in local Compose and is
covered by automated tests: covered by automated tests:
```text ```text
POST /v1/files (authenticated CSV upload) POST /v1/files (authenticated DOCX/XLSX/CSV upload)
-> raw bytes stored privately in MinIO -> raw bytes stored privately in MinIO
-> source file and running job committed in Postgres, connection released -> source file and running job committed in Postgres, connection released
-> parse/chunk on threads, embed in bounded concurrent batches -> parse/chunk on threads, embed in bounded concurrent batches

View File

@@ -15,8 +15,9 @@ are; this document defines order, scope, and verification criteria.
## Prerequisite ## Prerequisite
Plan 001 must be complete through **Phase 5** before Phase 3 of this plan Plan 001 is complete through Phase 6, so this prerequisite is satisfied. It
starts. Specifically this plan depends on: the `chunks` collection and its required plan 001 through **Phase 5** before Phase 3 of this plan starts.
Specifically this plan depends on: the `chunks` collection and its
payload indexes actually existing, API-key authentication and `AuthContext` payload indexes actually existing, API-key authentication and `AuthContext`
tenant derivation, the application-lifetime Qdrant client from the FastAPI tenant derivation, the application-lifetime Qdrant client from the FastAPI
lifespan, and the request-lifetime `AsyncSession` wiring. Phases 1–2 below lifespan, and the request-lifetime `AsyncSession` wiring. Phases 1–2 below
@@ -71,7 +72,9 @@ project owner accepts them, and update the ADR rather than diverging silently.
(bulk soft delete of a file's points), from ADR-0008. (bulk soft delete of a file's points), from ADR-0008.
- Soft delete as the default for every delete path, with neighbor relinking. - Soft delete as the default for every delete path, with neighbor relinking.
- Optimistic concurrency on every mutating path via the `version` payload field. - Optimistic concurrency on every mutating path via the `version` payload field.
- Audit rows in Postgres for mutating operations. - Audit rows in Postgres for mutating operations: both ADR-0009 tables,
`api_request_logs` (one row per API call, written from the request middleware)
and `point_audit_events` with the real `api_request_log_id` foreign key.
- Automated tests for tenant isolation, pointer integrity, concurrency - Automated tests for tenant isolation, pointer integrity, concurrency
conflicts, and pagination. conflicts, and pagination.
@@ -113,40 +116,22 @@ project owner accepts them, and update the ADR rather than diverging silently.
9. Routers contain no Qdrant SDK calls and no filter construction. The Qdrant 9. Routers contain no Qdrant SDK calls and no filter construction. The Qdrant
client is injected from the lifespan (ADR-0012). client is injected from the lifespan (ADR-0012).
## Decisions needed before the affected phase ## Decisions resolved before implementation
### Re-embedding on content edit (blocks Phase 4) An earlier revision of this plan listed three open decisions here. All are now
settled, and one further question this plan deferred to a Phase 6 test has been
settled too. They are recorded in the ADRs — these lines are a pointer, not a
second source of truth.
`PUT /v1/points/{point_id}` can change `content`. The stored vectors then no | Question | Resolution | Recorded in |
longer match the text. Three options, in order of preference: |---|---|---|
| Re-embedding on content edit | Re-embed inline, reusing ingestion's ports and bounds and its `502`/`504` codes. The re-embed happens *before* the version-guarded write, so a stale edit still `409`s rather than re-embedding for nothing. | ADR-0002, "Re-embedding on content edit" |
| Fractional-key exhaustion | No renormalize endpoint in this slice. Log `points.order_id.gap_low` under a safety threshold; reject with `409` and a distinct error code if the gap would collapse onto a neighbor value. Recovery is a runbook operation. | ADR-0002, "`order_id` gap exhaustion" |
| Batch semantics | All-or-nothing, capped at 100 operations. Every operation's `version` precondition is validated before any is applied; one failure rejects the whole request and nothing reaches Qdrant. | ADR-0002, "`POST /points/batch` semantics" |
| Re-ingestion versus manual edits | The newly uploaded file wins. Surviving points are overwritten in place with an incremented `version`; points absent from the new version are flagged inactive, never removed; manually created points sit past the ingested `chunk_index` range and are swept by the same rule. Clobbered content is recorded in `point_audit_events` as `reingest_overwrite`. | ADR-0002, "Re-ingestion versus manual edits" |
1. **Re-embed inline** on content change, reusing plan 001's embedding ports and Phase 6's cross-slice end-to-end test therefore *verifies* the re-ingestion rule
bounds. Consistent, but puts embedder latency and `502`/`504` failure modes rather than forcing the decision.
on an admin edit path.
2. **Require caller-supplied vectors** when content changes, and reject the edit
otherwise. Simple and honest, but pushes model knowledge to the client.
3. **Mark the point stale** (a payload flag) and re-embed later. Needs
background work, which ADR-0017 currently rules out.
Default to (1) for parity with ingestion, with the same batch/semaphore bounds
and the same status codes. Record whichever is chosen in ADR-0002 before
implementing Phase 4 — this is a real behavioral contract, not an
implementation detail.
### Fractional-key exhaustion
ADR-0001 notes float keys eventually need renormalization. Decide now whether
this slice ships a renormalize path (an internal operation rewriting a file's
`order_id` values to `1000, 2000, 3000, ...`) or explicitly defers it with a
logged warning when the gap between neighbors falls under a threshold. Deferring
is acceptable; silently producing unrepresentable gaps is not.
### Batch semantics
`POST /v1/points/batch` must define, in the API schema and the tests: whether
operations are all-or-nothing, what happens when operation 3 of 5 fails a
version check, and the maximum operation count per request. Decide before
Phase 5; do not let the answer be "whatever Qdrant happened to do."
## Build order ## Build order
@@ -251,8 +236,8 @@ of mutations.
ordering behavior, isolated per test by unique collection or tenant keys. ordering behavior, isolated per test by unique collection or tenant keys.
2. An end-to-end test crossing plan 001 and this slice: ingest a CSV, list its 2. An end-to-end test crossing plan 001 and this slice: ingest a CSV, list its
points, reorder one, soft-delete another, re-upload the same file, and assert points, reorder one, soft-delete another, re-upload the same file, and assert
the manual edits interact with re-ingestion exactly as ADR-0001/0002 specify. the manual edits interact with re-ingestion exactly as ADR-0002's
If that interaction is not yet decided, this test is what forces the decision. "Re-ingestion versus manual edits" specifies.
3. Structured logging at the mutation boundary with stable event names 3. Structured logging at the mutation boundary with stable event names
(`points.updated`, `points.reordered`, `points.soft_deleted`) carrying (`points.updated`, `points.reordered`, `points.soft_deleted`) carrying
`request_id`, `tenant_id`, `file_id`, and the resulting version. `request_id`, `tenant_id`, `file_id`, and the resulting version.

303
docs/runbook.md Normal file
View File

@@ -0,0 +1,303 @@
# Operator runbook
How to start this service, configure its ingestion bounds, and investigate or
retry a failed upload. Architecture rationale lives in [`docs/adr/`](adr/); the
implementation milestone is
[plan 001](plans/001-ingestion-vertical-slice.md). This document covers
operating what those describe.
The service is a **single process with no background work**. `POST /v1/files`
parses, chunks, embeds, and indexes inline and returns a terminal result
(ADR-0017). There is no queue, no worker, and no automatic retry — the caller
owns the retry decision, which makes the request's duration a deployment
constraint. That fact drives most of this document.
## 1. Prerequisites and local startup
Docker, and [`uv`](https://docs.astral.sh/uv/) with Python 3.13.
```bash
cp .env.example .env # non-secret local defaults; .env is gitignored
uv sync
docker compose up -d --wait
```
`docker-compose.yml` runs Postgres (`127.0.0.1:5433`), MinIO
(`127.0.0.1:9100`, console `9101`), and Qdrant (`127.0.0.1:6343`). It is the
local development stack and says so in its header — it is not a production
deployment.
## 2. Deployment steps
Two schema steps run **before** the application, never at startup: FastAPI
performs no DDL, for Postgres (ADR-0009) or for Qdrant (ADR-0001, "Collection
provisioning"). Both commands and their reasoning are in the README's
[Provisioning the datastores](../README.md#provisioning-the-datastores)
section:
```bash
uv run alembic upgrade head # Postgres schema
uv run python -m src.cli.qdrant_bootstrap # the `chunks` collection
```
Both are idempotent. `qdrant_bootstrap` verifies an existing collection against
the pinned schema and **exits non-zero on a mismatch** rather than leaving a
silently degraded sparse index in place — the `sparse` vector's
`modifier="idf"` and the pinned dense dimensions (768 / 3072) fail silently if
wrong, which is why they are checked rather than assumed.
Run both again after every deploy that ships a migration or a collection-schema
change.
## 3. MinIO bucket
Under Compose the bucket already exists: the `app-minio` service's entrypoint
runs `mkdir -p /data/${MINIO_BUCKET:-chatbot-source-files}` before starting the
server, so first boot creates it. Nothing else needs to be done locally.
Outside Compose, create the bucket named by `MINIO_BUCKET` before the first
upload — the application never creates it. It must stay **private**; ADR-0013
keeps source bytes non-public and this slice ships no download API or presigned
URLs.
## 4. Provisioning a tenant, an API key, and its domains
Nothing over HTTP can bootstrap a tenant: every `/v1` route needs an API key,
and a key cannot exist before its tenant. So the first key is issued by an
operator command:
```bash
uv run python -m src.cli.provision_tenant \
--slug acme --domain fire --domain life --scopes files:write,domains:read
```
It prints `api_key=sk_...` **once**. Postgres stores only its SHA-256 hash
(ADR-0009), so a lost key is reissued by re-running the command, never
recovered. Structured logs carry only the non-secret `key_prefix` — a plaintext
key must never reach a log sink (ADR-0011).
Re-running with the same `--slug` reuses the tenant and any domains it already
has, and issues an **additional** key. Both keys stay valid; this adds a key, it
does not rotate one.
Scopes are the security boundary between uploading, reading chunks, and
managing the allowlist. Give an upload client `files:write` only. `domains:write`
lets its holder create new domains, which is exactly what the allowlist exists to
prevent an upload key from doing, and `points:read` lets its holder read the text
of every chunk of every file — so an upload-only key gets neither.
| Scope | Grants |
|---|---|
| `files:write` | Upload a document and read its ingestion status. |
| `points:read` | Read, list, count, and keyword-search this tenant's points, including `GET /v1/files/{file_id}/points`. |
| `points:write` | Create, edit, reorder, and soft-delete points (plan 002 Phases 3-5; no route uses it yet). |
| `domains:read` / `domains:write` | Inspect and manage the domain allowlist. |
| `admin` | Satisfies every scope check. |
The command's `--scopes` default issues all of the above except `admin`, which
suits a first operator key; narrow it explicitly for per-client keys.
### Domains after the first one
`POST /v1/files` rejects an unregistered or disabled `domain` with `400`
(`unknown_domain`) before anything is written. Ongoing domain management is the
`/v1/domains` API, under `domains:read` / `domains:write`:
```bash
curl -X POST http://localhost:8000/v1/domains \
-H "Authorization: Bearer $API_KEY" -H 'Content-Type: application/json' \
-d '{"domain": "fire", "display_name": "Fire insurance"}'
```
The `domain` key itself is immutable — it is denormalized into every Qdrant
point payload and into `source_files`, so renaming it is a migration, not an
edit (ADR-0009). Disabling a domain blocks new uploads; it does not delete
existing points.
## 5. Running the service
```bash
uv run fastapi dev src/main.py # local, reload
uv run uvicorn src.main:app --host 0.0.0.0 --port 8000 # deployed shape
```
Run more than one worker/replica only after reading §6: ingestion bounds are
**per process**, so `INGESTION_MAX_CONCURRENCY` multiplies by the number of
processes.
## 6. Ingestion bounds and tuning
Every bound is enforced server-side and maps to a status code. All are in
`.env.example`. Ingestion is CPU- and network-bound in the request, so these are
the numbers that decide whether the service degrades gracefully or falls over.
| Setting | Bounds | On breach | Size it against |
|---|---|---|---|
| `INGESTION_MAX_CONCURRENCY` | Ingestions in flight **per process** | `503` + `Retry-After` | Memory per in-flight upload (whole file plus its chunks and vectors are resident) and the embedder's capacity. Rejecting is deliberate: ADR-0017 refuses rather than queues. |
| `INGESTION_THREAD_POOL_SIZE` | Threads for blocking work (parse, chunk, hash, BM25, the sync `minio` SDK) | — (waits) | CPU cores. It exists to stop ingestion exhausting Starlette's own thread pool, so keep it below the total thread budget. |
| `INGESTION_TIMEOUT_SECONDS` | The whole work phase | `504`, job marked `failed` | The slowest legitimate document, plus headroom. See §7 — this must stay under every read timeout in front of it. |
| `INGESTION_MAX_UPLOAD_SIZE_MB` | Bytes accepted | `413` | Memory: the upload is read fully into the process before any work starts. |
| `INGESTION_MAX_CHUNKS_PER_FILE` | Chunks per file, checked before embedding | `413` | Embedder cost/time per chunk × `INGESTION_TIMEOUT_SECONDS`. This is the real defence against one pathological file eating a slot. |
| `INGESTION_EMBED_BATCH_SIZE` | Texts per embedder request | `502` on embedder failure | The provider's per-request limits. Batch before parallelizing. |
| `INGESTION_EMBED_CONCURRENCY` | Concurrent embed batches | `502` | Provider rate limits and the self-hosted embedder's throughput. Never unbounded. |
| `QDRANT_UPSERT_BATCH_SIZE` / `_CONCURRENCY` | Points per upsert and concurrent upserts | `502` (`index_error`) | Qdrant's ingest capacity; the batch size stays in ADR-0001's 64–256 band. |
Two settings that look like tuning knobs but are not:
- **`EMBEDDING_NOMIC_KEEP_ALIVE`** holds the self-hosted model resident. A cold
load of `nomic-embed-text-v2-moe` takes over 150 s — longer than any sane
`INGESTION_TIMEOUT_SECONDS` — so an idle period followed by an upload would
otherwise `504`. The lifespan also warms both dense embedders at startup for
the same reason.
- **The BM25 analyzer and weights** (`EMBEDDING_SPARSE_*`) are a measured
artifact ported from the `emet` evaluation lab, verified token-for-token
against it (ADR-0005). Re-benchmark; do not tune them in place.
## 7. The proxy and client read-timeout requirement
**Every read timeout in front of this service must exceed
`INGESTION_TIMEOUT_SECONDS`.** That includes the reverse proxy / ingress, any
load balancer, and the calling backend's own HTTP client.
If a proxy times out first, the client gets that proxy's error, the upload keeps
running in the process, and the caller learns nothing about the outcome from the
response. The job row still reaches a terminal status, so
`GET /v1/files/{file_id}` remains the way to find out what happened — but the
response contract is broken for that request. ADR-0017 names this the main cost
of inline ingestion.
A workable local ordering: client read timeout > proxy read timeout >
`INGESTION_TIMEOUT_SECONDS`.
## 8. Health and readiness
| Endpoint | Question it answers | Use for |
|---|---|---|
| `GET /healthz` | Is the process alive? | Liveness probes / restart policy. Never depends on Postgres, MinIO, or Qdrant. |
| `GET /readyz` | Can it actually serve? | Load-balancer admission and post-deploy gating. `200` with each dependency `true`, `503` if any is `false`. |
`/readyz` checks Postgres, MinIO, and Qdrant reachability **and** that the
`chunks` collection exists. A reachable-but-unbootstrapped Qdrant reports
`{"qdrant": false}` on purpose: uploads to it would fail with `502`, so it is
not ready, and this is how a skipped `qdrant_bootstrap` surfaces at deploy time
instead of on a user's first upload.
## 9. Investigating a failure
Logs are structured (`structlog`, JSON in production) with stable event names —
grep the event name, not prose (ADR-0011). Set `LOG_FILE_PATH` for a local
JSON file sink alongside the console renderer; leave it unset in production,
where stdout collection is preferred.
**Correlate by `request_id`.** Every request has one, echoed in the
`X-Request-Id` response header and included in every error envelope, and bound
into every log line emitted while handling that request. A client reporting a
failed upload should quote it. `tenant_id`, `file_id`, and `ingestion_job_id`
are the other join keys.
Events worth knowing:
| Event | Level | Means |
|---|---|---|
| `ingestion.job.started` | info | Txn A committed; work phase beginning. Carries `tenant_id`, `ingestion_job_id`, `file_id`, `domain`, `source_type`. |
| `ingestion.job.completed` | info | Terminal success, with `chunks_parsed`, `points_upserted`, `points_soft_deleted`. |
| `ingestion.job.failed` | warning | Terminal failure. **`error_code` says which stage**: `storage_upload_failed`, `parse_failed`, `chunk_limit_exceeded`, `embedding_failed`, `index_failed`, `timeout`. |
| `files.upload.duplicate` | info | Identical content already ingested; the existing file/job was returned and nothing was re-ingested. |
| `domain.rejected` | warning | Upload refused before any row was written; `reason` is `unregistered` or `disabled`. |
| `auth.failed` | warning | `reason` is `malformed_key`, `unknown_key`, `key_inactive`, `key_expired`, or `tenant_inactive`. Never contains key material. |
| `auth.succeeded` | info | Carries `tenant_id`, `api_key_id`, `actor_type`. |
| `lifespan.embedder.warm_failed` | warning | An embedder was unreachable at startup. Boot continues by design — `/readyz` and the first upload are where this bites. |
| `qdrant.bootstrap.schema_mismatch` | error | The existing collection diverges from the pinned schema. The bootstrap exits non-zero; do not start the app against it. |
| `api.unhandled_exception` | error | A bug: an exception with no mapping to the error envelope. Always worth a look. |
A `503` (`ingestion_at_capacity`) is rejected before a job row exists, so it
appears in the access log and metrics, not in `ingestion_jobs`.
### The durable record
Logs may roll; `ingestion_jobs` and `ingestion_job_events` do not. For one file:
```sql
SELECT id, status, error_code, error_message, points_created, points_soft_deleted,
created_at, updated_at
FROM ingestion_jobs
WHERE tenant_id = :tenant_id AND source_file_id = :file_id
ORDER BY created_at DESC;
SELECT stage, level, message, details, created_at
FROM ingestion_job_events
WHERE tenant_id = :tenant_id AND ingestion_job_id = :ingestion_job_id
ORDER BY created_at;
```
Recent failures across a tenant:
```sql
SELECT error_code, count(*), max(created_at)
FROM ingestion_jobs
WHERE tenant_id = :tenant_id AND status = 'failed' AND created_at > now() - interval '1 day'
GROUP BY error_code ORDER BY 2 DESC;
```
`GET /v1/files/{file_id}` reports the same terminal status over HTTP, scoped to
the owning tenant — a file belonging to another tenant returns `404`, not `403`.
## 10. Retrying a failed ingestion
**Re-upload the same bytes.** There is no retry endpoint and no automatic retry;
the client owns that decision (ADR-0017).
What that guarantees:
- Identical content with a **succeeded** job is recognized by
`(tenant_id, domain, content_sha256)` and returned as-is with `200` — no
re-ingestion, no duplicate points.
- Identical content whose last job **failed** starts a fresh job against the
same `source_files` row. A terminal job is never moved back to `running`.
- Point ids are deterministic from `file_id` + `chunk_index` (ADR-0001), so the
retry **overwrites in place** — it cannot duplicate chunks.
- A failed attempt never empties or partially deletes a working index: the
soft-delete sweep that retires a shortened file's leftover points runs only
after every upsert has succeeded. An interrupted attempt can leave a prefix
updated; a retry converges (ADR-0017, "Re-running an ingestion stays safe").
Fix the cause first — the `error_code` says where to look:
| `error_code` | Usual cause |
|---|---|
| `parse_failed` | The file is corrupt or is not really the type its extension claims. Retrying identical bytes will fail identically. |
| `chunk_limit_exceeded` | The file is genuinely too large for one inline ingestion. Split it, or raise `INGESTION_MAX_CHUNKS_PER_FILE` knowing what §6 says about the timeout. |
| `embedding_failed` | The embedder is down, rate-limiting, or unauthenticated. Fix it, then retry — this one usually succeeds unchanged. |
| `index_failed` | Qdrant is down, or the collection is missing (run `qdrant_bootstrap`). |
| `timeout` | The work exceeded `INGESTION_TIMEOUT_SECONDS`. Check whether the embedder was cold (see `lifespan.embedder.warm_failed` and `KEEP_ALIVE`) before raising the bound. |
| `storage_upload_failed` | MinIO is unreachable or the bucket is missing (§3). |
## 11. What to alert on
ADR-0017's own triggers for moving ingestion back off the request path. These
are the numbers that say the inline design has stopped fitting:
- **p95 ingestion duration** approaching `INGESTION_TIMEOUT_SECONDS`.
- **`503` and `504` rates** ceasing to be negligible.
- **Jobs stuck in `running` past the timeout** — every handled failure writes a
terminal status, so a non-zero count here means the process died mid-request:
```sql
SELECT count(*) FROM ingestion_jobs
WHERE status = 'running' AND created_at < now() - interval '5 minutes';
```
Also worth alerting: any `qdrant.bootstrap.schema_mismatch`, a sustained
`/readyz` `503`, and any `api.unhandled_exception`.
## 12. Verifying a deployment
```bash
./scripts/smoke.sh
```
Brings up Compose, runs both deployment steps, provisions a throwaway tenant,
starts the web process, and drives an upload through to indexed Qdrant points
against the **running process** — including asserting the structured log output
from §9. It is the only Compose-based test; everything else runs on
Testcontainers under `uv run pytest` (ADR-0016). Run it before a release.

View File

@@ -9,18 +9,21 @@ dependencies = [
"anyio>=4.11.0", "anyio>=4.11.0",
"asyncpg>=0.31.0", "asyncpg>=0.31.0",
"fastapi[standard]==0.141.1", "fastapi[standard]==0.141.1",
"httpx>=0.28.1",
"langgraph>=1.2.10", "langgraph>=1.2.10",
"minio>=7.2.20", "minio>=7.2.20",
"openpyxl>=3.1.5",
"pydantic-settings>=2.15.0", "pydantic-settings>=2.15.0",
"python-docx>=1.2.0",
"qdrant-client>=1.19.0", "qdrant-client>=1.19.0",
"sqlalchemy>=2.0.51", "sqlalchemy>=2.0.51",
"structlog>=26.1.0", "structlog>=26.1.0",
"tiktoken>=0.13.0",
] ]
[dependency-groups] [dependency-groups]
dev = [ dev = [
"asgi-lifespan>=2.1.0", "asgi-lifespan>=2.1.0",
"httpx>=0.28.1",
"pytest>=8.3.5", "pytest>=8.3.5",
"pytest-asyncio>=0.25.3", "pytest-asyncio>=0.25.3",
"pytest-cov>=6.0.0", "pytest-cov>=6.0.0",
@@ -34,6 +37,11 @@ dev = [
testpaths = ["tests"] testpaths = ["tests"]
asyncio_mode = "strict" asyncio_mode = "strict"
timeout = 10 timeout = 10
# Bound the test function only, not fixture setup. Testcontainers' container
# startup is charged to whichever test first pulls a session-scoped container
# fixture; on a cold Docker cache that is ~25s and would trip the 10s budget
# for every integration test, regardless of how fast the test itself is.
timeout_func_only = true
markers = [ markers = [
"unit: fast tests with no external services", "unit: fast tests with no external services",
"integration: tests against a real disposable service", "integration: tests against a real disposable service",

98
scripts/smoke.sh Executable file
View File

@@ -0,0 +1,98 @@
#!/usr/bin/env bash
# Serialized operational smoke test of the running web process (ADR-0016, plan
# 001 Phase 6).
#
# ./scripts/smoke.sh
#
# Brings up the Compose stack, runs both deployment steps for real, provisions
# a tenant, starts uvicorn, and drives `tests/e2e/test_compose_smoke.py`
# against it over a socket. This is the only Compose-based test: every other
# test uses Testcontainers and an in-process ASGI transport, which is exactly
# what makes this one worth having -- it is the only thing that exercises the
# deployment steps, the real logging configuration, and a real HTTP server.
#
# Not part of `uv run pytest`: the smoke test skips itself unless SMOKE_BASE_URL
# is set, so this script is the only way it runs. Run it before a release.
#
# Leaves the Compose stack running (it is the local dev stack); only the uvicorn
# process and the temporary log file are cleaned up.
set -euo pipefail
cd "$(dirname "${BASH_SOURCE[0]}")/.."
PORT="${SMOKE_PORT:-8021}"
SLUG="smoke-$(date +%s)"
DOMAIN="smoke"
LOG_FILE="$(mktemp -t smoke-app-log.XXXXXX.jsonl)"
CONSOLE_LOG="$(mktemp -t smoke-app-console.XXXXXX.log)"
APP_PID=""
cleanup() {
if [[ -n "${APP_PID}" ]] && kill -0 "${APP_PID}" 2>/dev/null; then
kill "${APP_PID}" 2>/dev/null || true
wait "${APP_PID}" 2>/dev/null || true
fi
rm -f "${LOG_FILE}" "${CONSOLE_LOG}"
}
trap cleanup EXIT
if [[ ! -f .env ]]; then
echo "no .env found; copy .env.example first (see docs/runbook.md)" >&2
exit 1
fi
echo "==> starting Postgres, MinIO, Qdrant"
docker compose up -d --wait
echo "==> applying deployment steps"
uv run alembic upgrade head
uv run python -m src.cli.qdrant_bootstrap
echo "==> provisioning tenant '${SLUG}'"
PROVISION_OUTPUT="$(uv run python -m src.cli.provision_tenant \
--slug "${SLUG}" --domain "${DOMAIN}" --scopes files:write 2>/dev/null)"
API_KEY="$(printf '%s\n' "${PROVISION_OUTPUT}" | sed -n 's/^api_key=//p')"
if [[ -z "${API_KEY}" ]]; then
echo "provisioning did not return an api_key" >&2
exit 1
fi
echo "==> starting the web process on port ${PORT}"
# JSON to a file sink, because the smoke test asserts the real ADR-0011 log
# output -- the one thing no in-process test can check.
LOG_JSON_FORMAT=true LOG_FILE_PATH="${LOG_FILE}" \
uv run python -m uvicorn src.main:app --host 127.0.0.1 --port "${PORT}" \
>"${CONSOLE_LOG}" 2>&1 &
APP_PID=$!
echo "==> waiting for /readyz"
for _ in $(seq 1 60); do
if curl -fsS "http://127.0.0.1:${PORT}/readyz" >/dev/null 2>&1; then
break
fi
if ! kill -0 "${APP_PID}" 2>/dev/null; then
echo "the web process exited before becoming ready:" >&2
tail -20 "${CONSOLE_LOG}" >&2
exit 1
fi
sleep 1
done
if ! curl -fsS "http://127.0.0.1:${PORT}/readyz" >/dev/null 2>&1; then
# Most often an unbootstrapped Qdrant or an unreachable embedder host; the
# runbook's health/readiness section covers reading this.
echo "the web process never became ready:" >&2
tail -20 "${CONSOLE_LOG}" >&2
exit 1
fi
echo "==> running the smoke test"
SMOKE_BASE_URL="http://127.0.0.1:${PORT}" \
SMOKE_API_KEY="${API_KEY}" \
SMOKE_DOMAIN="${DOMAIN}" \
SMOKE_LOG_PATH="${LOG_FILE}" \
SMOKE_QDRANT_URL="${QDRANT_URL:-http://127.0.0.1:6343}" \
SMOKE_QDRANT_COLLECTION="${QDRANT_COLLECTION:-chunks}" \
uv run python -m pytest tests/e2e/test_compose_smoke.py -q
echo "==> smoke test passed"

View File

@@ -0,0 +1,38 @@
"""Auth dependencies (ADR-0008): resolve `AuthContext` from a bearer token,
then gate routes on scope.
"""
from collections.abc import Awaitable, Callable
from typing import Annotated
from fastapi import Depends
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.application.auth.context import AuthContext
from src.application.auth.errors import InvalidApiKeyError, MissingScopeError
from src.application.auth.service import resolve_auth_context
from src.bootstrap.dependencies import get_sessionmaker
_bearer_scheme = HTTPBearer(auto_error=False)
async def get_auth_context(
credentials: Annotated[HTTPAuthorizationCredentials | None, Depends(_bearer_scheme)],
sessionmaker: Annotated[async_sessionmaker[AsyncSession], Depends(get_sessionmaker)],
) -> AuthContext:
if credentials is None:
raise InvalidApiKeyError("missing Authorization header")
return await resolve_auth_context(sessionmaker, credentials.credentials)
AuthContextDep = Annotated[AuthContext, Depends(get_auth_context)]
def require_scope(scope: str) -> Callable[[AuthContext], Awaitable[AuthContext]]:
async def _dependency(auth: AuthContextDep) -> AuthContext:
if not auth.has_scope(scope):
raise MissingScopeError(f"missing required scope '{scope}'")
return auth
return _dependency

138
src/api/errors.py Normal file
View File

@@ -0,0 +1,138 @@
"""Maps application exceptions to the ADR-0008 error envelope.
This is the single place that knows the exception-type -> status-code
mapping; application/infrastructure code never imports FastAPI or raises
`HTTPException` (ADR-0015).
"""
import structlog
from fastapi import FastAPI, Request, status
from fastapi.exceptions import RequestValidationError
from fastapi.responses import JSONResponse
from starlette.exceptions import HTTPException as StarletteHTTPException
from src.application.auth.errors import (
InvalidApiKeyError,
MissingScopeError,
TenantInactiveError,
)
from src.application.domains.errors import DomainAlreadyExistsError, UnknownDomainError
from src.application.files.errors import (
FileTooLargeError,
InvalidUploadError,
SourceFileNotFoundError,
)
from src.application.ingestion.errors import (
ChunkLimitExceededError,
DocumentParseError,
EmbedderError,
IngestionAtCapacityError,
IngestionTimeoutError,
PointIndexingError,
UnsupportedSourceTypeError,
)
from src.application.points.errors import PointVersionConflictError
from src.application.points.point import PointNotFoundError
logger = structlog.get_logger(__name__)
# A fixed backoff hint, not a computed retry budget: ADR-0017 rejects a
# request outright at capacity rather than queueing it, so there is no
# in-process estimate of when a slot will free up to report instead.
_CAPACITY_RETRY_AFTER_SECONDS = 1
# (exception type, status code, stable error code)
_MAPPING: tuple[tuple[type[Exception], int, str], ...] = (
(InvalidApiKeyError, status.HTTP_401_UNAUTHORIZED, "invalid_api_key"),
(TenantInactiveError, status.HTTP_401_UNAUTHORIZED, "tenant_not_found"),
(MissingScopeError, status.HTTP_403_FORBIDDEN, "missing_scope"),
(InvalidUploadError, status.HTTP_400_BAD_REQUEST, "validation_error"),
# 404, never 403: a cross-tenant point id must be indistinguishable from a
# nonexistent one, or the API becomes an existence oracle (ADR-0016).
(PointNotFoundError, status.HTTP_404_NOT_FOUND, "not_found"),
(SourceFileNotFoundError, status.HTTP_404_NOT_FOUND, "not_found"),
# Not "the version guard fired once" — that is retried. This is the service
# giving up after repeated re-plans, i.e. a genuinely contended point.
(PointVersionConflictError, status.HTTP_409_CONFLICT, "conflict"),
(UnknownDomainError, status.HTTP_400_BAD_REQUEST, "unknown_domain"),
(DomainAlreadyExistsError, status.HTTP_409_CONFLICT, "conflict"),
(DocumentParseError, status.HTTP_400_BAD_REQUEST, "validation_error"),
(UnsupportedSourceTypeError, status.HTTP_415_UNSUPPORTED_MEDIA_TYPE, "unsupported_media_type"),
(FileTooLargeError, status.HTTP_413_CONTENT_TOO_LARGE, "payload_too_large"),
(ChunkLimitExceededError, status.HTTP_413_CONTENT_TOO_LARGE, "payload_too_large"),
(EmbedderError, status.HTTP_502_BAD_GATEWAY, "embedder_error"),
(PointIndexingError, status.HTTP_502_BAD_GATEWAY, "index_error"),
(IngestionTimeoutError, status.HTTP_504_GATEWAY_TIMEOUT, "ingestion_timeout"),
)
def _request_id(request: Request) -> str | None:
return getattr(request.state, "request_id", None)
def _envelope(
code: str, message: str, request_id: str | None, details: dict[str, object] | None = None
) -> dict[str, object]:
return {
"error": {
"code": code,
"message": message,
"details": details or {},
"request_id": request_id,
}
}
def register_exception_handlers(app: FastAPI) -> None:
for exc_type, status_code, error_code in _MAPPING:
def _handler(
request: Request,
exc: Exception,
status_code: int = status_code,
error_code: str = error_code,
) -> JSONResponse:
return JSONResponse(
status_code=status_code,
content=_envelope(error_code, str(exc), _request_id(request)),
)
app.add_exception_handler(exc_type, _handler)
@app.exception_handler(IngestionAtCapacityError)
def _capacity_handler(request: Request, exc: IngestionAtCapacityError) -> JSONResponse:
return JSONResponse(
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
content=_envelope("ingestion_at_capacity", str(exc), _request_id(request)),
headers={"Retry-After": str(_CAPACITY_RETRY_AFTER_SECONDS)},
)
@app.exception_handler(RequestValidationError)
def _validation_handler(request: Request, exc: RequestValidationError) -> JSONResponse:
return JSONResponse(
status_code=status.HTTP_422_UNPROCESSABLE_CONTENT,
content=_envelope(
"validation_error",
"request validation failed",
_request_id(request),
details={"errors": exc.errors()},
),
)
@app.exception_handler(StarletteHTTPException)
def _http_exception_handler(request: Request, exc: StarletteHTTPException) -> JSONResponse:
code = "not_found" if exc.status_code == status.HTTP_404_NOT_FOUND else "http_error"
return JSONResponse(
status_code=exc.status_code,
content=_envelope(code, str(exc.detail), _request_id(request)),
)
@app.exception_handler(Exception)
def _unhandled_exception_handler(request: Request, exc: Exception) -> JSONResponse:
logger.exception("api.unhandled_exception", path=request.url.path)
return JSONResponse(
status_code=status.HTTP_500_INTERNAL_SERVER_ERROR,
content=_envelope(
"internal_error", "an unexpected error occurred", _request_id(request)
),
)

38
src/api/middleware.py Normal file
View File

@@ -0,0 +1,38 @@
"""Per-request correlation id (ADR-0008, ADR-0011).
Every request gets a `request_id`: reused from an incoming `X-Request-Id` if
the caller supplied one, otherwise generated. It is bound into structlog's
contextvars so every log line emitted while handling the request carries it,
stored on `request.state` for exception handlers, and echoed back in the
response header.
"""
import uuid
from collections.abc import Awaitable, Callable
from typing import override
import structlog
from starlette.middleware.base import BaseHTTPMiddleware
from starlette.requests import Request
from starlette.responses import Response
_HEADER = "X-Request-Id"
class RequestIdMiddleware(BaseHTTPMiddleware):
@override
async def dispatch(
self, request: Request, call_next: Callable[[Request], Awaitable[Response]]
) -> Response:
request_id = request.headers.get(_HEADER) or str(uuid.uuid4())
request.state.request_id = request_id
structlog.contextvars.clear_contextvars()
structlog.contextvars.bind_contextvars(request_id=request_id)
try:
response = await call_next(request)
finally:
structlog.contextvars.clear_contextvars()
response.headers[_HEADER] = request_id
return response

View File

@@ -1,3 +1,10 @@
from fastapi import APIRouter from fastapi import APIRouter
from src.api.routers.domains import router as domains_router
from src.api.routers.files import router as files_router
from src.api.routers.points import router as points_router
router = APIRouter() router = APIRouter()
router.include_router(domains_router)
router.include_router(files_router)
router.include_router(points_router)

110
src/api/routers/domains.py Normal file
View File

@@ -0,0 +1,110 @@
"""`/v1/domains` (ADR-0008, ADR-0009).
The management surface for a tenant's domain allowlist, used by the calling
backend rather than by an operator with a psql prompt.
Gated on `domains:read`/`domains:write`, deliberately **not** on `files:write`:
if an upload key could create domains, the allowlist would no longer prevent a
typo'd `domain` from creating a Qdrant partition, which is its only purpose.
"""
from typing import Annotated
from fastapi import APIRouter, Depends, status
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.api.dependencies.auth import require_scope
from src.api.schemas.domains import (
CreateDomainRequest,
DomainListResponse,
DomainResponse,
UpdateDomainRequest,
)
from src.application.auth.context import AuthContext
from src.application.domains import (
create_domain,
list_domains,
set_domain_status,
update_domain,
)
from src.bootstrap.dependencies import get_sessionmaker
router = APIRouter(prefix="/domains", tags=["domains"])
_RequireDomainsRead = Annotated[AuthContext, Depends(require_scope("domains:read"))]
_RequireDomainsWrite = Annotated[AuthContext, Depends(require_scope("domains:write"))]
_SessionmakerDep = Annotated[async_sessionmaker[AsyncSession], Depends(get_sessionmaker)]
@router.get("")
async def list_tenant_domains(
auth: _RequireDomainsRead,
sessionmaker: _SessionmakerDep,
include_disabled: bool = False,
) -> DomainListResponse:
results = await list_domains(
sessionmaker, tenant_id=auth.tenant_id, include_disabled=include_disabled
)
return DomainListResponse(domains=[DomainResponse.from_result(item) for item in results])
@router.post("", status_code=status.HTTP_201_CREATED)
async def create_tenant_domain(
request: CreateDomainRequest,
auth: _RequireDomainsWrite,
sessionmaker: _SessionmakerDep,
) -> DomainResponse:
result = await create_domain(
sessionmaker,
tenant_id=auth.tenant_id,
domain=request.domain,
display_name=request.display_name,
metadata=request.metadata,
)
return DomainResponse.from_result(result)
@router.patch("/{domain}")
async def update_tenant_domain(
domain: str,
request: UpdateDomainRequest,
auth: _RequireDomainsWrite,
sessionmaker: _SessionmakerDep,
) -> DomainResponse:
result = await update_domain(
sessionmaker,
tenant_id=auth.tenant_id,
domain=domain,
display_name=request.display_name,
)
return DomainResponse.from_result(result)
@router.delete("/{domain}")
async def disable_tenant_domain(
domain: str,
auth: _RequireDomainsWrite,
sessionmaker: _SessionmakerDep,
) -> DomainResponse:
"""Disable, not delete.
Blocks new uploads and drops the domain from pickers while leaving the
points already indexed under it intact and retrievable. Actually removing
them needs the tenant-erasure workflow plan 001 defers.
"""
result = await set_domain_status(
sessionmaker, tenant_id=auth.tenant_id, domain=domain, status="disabled"
)
return DomainResponse.from_result(result)
@router.post("/{domain}/enable")
async def enable_tenant_domain(
domain: str,
auth: _RequireDomainsWrite,
sessionmaker: _SessionmakerDep,
) -> DomainResponse:
result = await set_domain_status(
sessionmaker, tenant_id=auth.tenant_id, domain=domain, status="active"
)
return DomainResponse.from_result(result)

157
src/api/routers/files.py Normal file
View File

@@ -0,0 +1,157 @@
"""`POST /v1/files`, `GET /v1/files/{file_id}`, `DELETE /v1/files/{file_id}` (ADR-0008).
Routes adapt HTTP to `application/files` calls; they do not parse, hash,
touch MinIO/Qdrant, or otherwise carry ingestion business logic (ADR-0015).
"""
import uuid
from collections.abc import Sequence
from typing import Annotated
from anyio import CapacityLimiter, Semaphore
from fastapi import APIRouter, Depends, Form, HTTPException, Response, UploadFile, status
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.api.dependencies.auth import require_scope
from src.api.schemas.files import FileDeleteResponse, FileStatusResponse, FileUploadResponse
from src.api.schemas.points import DEFAULT_PAGE_SIZE, LimitQuery, PointListResponse
from src.application.auth.context import AuthContext
from src.application.files.deletion import delete_source_file
from src.application.files.status import get_file_status
from src.application.files.upload import upload_source_file
from src.application.points.queries import list_file_points
from src.application.ports.embedding import DenseEmbedder, SparseEmbedder
from src.application.ports.object_storage import ObjectStorage
from src.application.ports.point_repository import PointRepository
from src.application.ports.point_storage import PointStorage
from src.bootstrap.dependencies import (
get_dense_embedders,
get_ingestion_concurrency_limiter,
get_ingestion_limiter,
get_object_storage,
get_point_repository,
get_point_storage,
get_sessionmaker,
get_settings,
get_sparse_embedder,
)
from src.config import Settings
router = APIRouter(prefix="/files", tags=["files"])
_RequireFilesWrite = Annotated[AuthContext, Depends(require_scope("files:write"))]
_RequirePointsRead = Annotated[AuthContext, Depends(require_scope("points:read"))]
_RequirePointsWrite = Annotated[AuthContext, Depends(require_scope("points:write"))]
_SessionmakerDep = Annotated[async_sessionmaker[AsyncSession], Depends(get_sessionmaker)]
_ObjectStorageDep = Annotated[ObjectStorage, Depends(get_object_storage)]
_PointStorageDep = Annotated[PointStorage, Depends(get_point_storage)]
_PointRepositoryDep = Annotated[PointRepository, Depends(get_point_repository)]
_SettingsDep = Annotated[Settings, Depends(get_settings)]
_IngestionLimiterDep = Annotated[CapacityLimiter, Depends(get_ingestion_limiter)]
_ConcurrencyLimiterDep = Annotated[Semaphore, Depends(get_ingestion_concurrency_limiter)]
_DenseEmbeddersDep = Annotated[Sequence[DenseEmbedder], Depends(get_dense_embedders)]
_SparseEmbedderDep = Annotated[SparseEmbedder, Depends(get_sparse_embedder)]
@router.post("", status_code=status.HTTP_201_CREATED)
async def upload_file(
response: Response,
file: UploadFile,
domain: Annotated[str, Form()],
auth: _RequireFilesWrite,
sessionmaker: _SessionmakerDep,
storage: _ObjectStorageDep,
point_storage: _PointStorageDep,
settings: _SettingsDep,
limiter: _IngestionLimiterDep,
concurrency_limiter: _ConcurrencyLimiterDep,
dense_embedders: _DenseEmbeddersDep,
sparse_embedder: _SparseEmbedderDep,
) -> FileUploadResponse:
data = await file.read()
result = await upload_source_file(
sessionmaker=sessionmaker,
storage=storage,
point_storage=point_storage,
auth=auth,
domain=domain,
filename=file.filename or "",
data=data,
ingestion_settings=settings.ingestion,
chunking_settings=settings.chunking,
qdrant_settings=settings.qdrant,
thread_limiter=limiter,
concurrency_limiter=concurrency_limiter,
dense_embedders=dense_embedders,
sparse_embedder=sparse_embedder,
)
if not result.is_new_attempt:
response.status_code = status.HTTP_200_OK
return FileUploadResponse.from_result(result)
@router.get("/{file_id}")
async def get_file(
file_id: uuid.UUID,
auth: _RequireFilesWrite,
sessionmaker: _SessionmakerDep,
) -> FileStatusResponse:
result = await get_file_status(sessionmaker, tenant_id=auth.tenant_id, source_file_id=file_id)
if result is None:
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="file not found")
return FileStatusResponse.from_result(result)
@router.delete("/{file_id}")
async def delete_file(
file_id: uuid.UUID,
auth: _RequirePointsWrite,
sessionmaker: _SessionmakerDep,
repository: _PointRepositoryDep,
) -> FileDeleteResponse:
"""Soft-delete a file: every active point, then the `source_files` row.
Gated on `points:write` rather than `files:write` for the same reason as the
listing above — the data this destroys is points. Nothing is removed from
Qdrant (ADR-0002); the points are flagged inactive and the row is marked
`soft_deleted`, which is also what makes a later re-upload of the same bytes
ingest afresh instead of matching the duplicate path.
Deleting an already-deleted file is a success reporting `0` points.
"""
points_soft_deleted = await delete_source_file(
sessionmaker,
repository,
tenant_id=auth.tenant_id,
source_file_id=file_id,
actor=f"api_key:{auth.api_key_id}",
)
return FileDeleteResponse(
file_id=file_id, status="soft_deleted", points_soft_deleted=points_soft_deleted
)
@router.get("/{file_id}/points")
async def list_points_for_file(
file_id: uuid.UUID,
auth: _RequirePointsRead,
repository: _PointRepositoryDep,
limit: LimitQuery = DEFAULT_PAGE_SIZE,
cursor: str | None = None,
include_inactive: bool = False,
) -> PointListResponse:
"""The same listing as `GET /v1/points?file_id=...`, addressed by file.
Gated on `points:read`, not `files:write`: the resource being read is the
file's chunks, so the scope follows the data rather than the URL prefix. An
upload-only key must not become a way to read every chunk of every file.
"""
page = await list_file_points(
repository,
tenant_id=auth.tenant_id,
file_id=file_id,
limit=limit,
cursor=cursor,
include_inactive=include_inactive,
)
return PointListResponse.from_page(page)

View File

@@ -23,7 +23,9 @@ async def readyz(request: Request, response: Response) -> dict[str, bool]:
postgres_ready, minio_ready, qdrant_ready = await asyncio.gather( postgres_ready, minio_ready, qdrant_ready = await asyncio.gather(
ping_postgres(resources.db_engine, timeout), ping_postgres(resources.db_engine, timeout),
ping_minio(resources.minio_client, timeout), ping_minio(resources.minio_client, timeout),
ping_qdrant(resources.qdrant_client, timeout), ping_qdrant(
resources.qdrant_client, timeout, collection=resources.settings.qdrant.collection
),
) )
result = { result = {

162
src/api/routers/points.py Normal file
View File

@@ -0,0 +1,162 @@
"""`/v1/points` read and soft-delete paths (ADR-0002, ADR-0008).
Routes adapt HTTP to `application/points` calls. They build no Qdrant filters
and hold no CRUD semantics (ADR-0015), and they never read a tenant from the
request — `auth.tenant_id` is the only source, which is what makes ADR-0002's
isolation rule structural rather than a habit.
**Route order is load-bearing.** `/count` and `/search` are declared before
`/{point_id}`. FastAPI matches in declaration order, so with `/{point_id}` first
a request for `/v1/points/count` would try to parse `"count"` as a UUID and
fail with `422` instead of counting anything. The failure is loud but confusing,
and it comes back the moment someone reorders these for tidiness.
Gated on `points:read`, separately from `files:write`: a key that can upload
documents should not thereby be able to read every chunk of every file, and
plan 002's mutating paths will want `points:write` distinct again.
"""
import uuid
from typing import Annotated
from fastapi import APIRouter, Depends, Query
from src.api.dependencies.auth import require_scope
from src.api.schemas.points import (
DEFAULT_PAGE_SIZE,
LimitQuery,
PointCountResponse,
PointListResponse,
PointResponse,
PointSearchResponse,
)
from src.application.auth.context import AuthContext
from src.application.points.deletion import soft_delete_point
from src.application.points.queries import (
count_points,
get_point,
list_file_points,
search_points,
)
from src.application.ports.point_repository import PointRepository
from src.bootstrap.dependencies import get_point_repository
router = APIRouter(prefix="/points", tags=["points"])
_RequirePointsRead = Annotated[AuthContext, Depends(require_scope("points:read"))]
_RequirePointsWrite = Annotated[AuthContext, Depends(require_scope("points:write"))]
_PointRepositoryDep = Annotated[PointRepository, Depends(get_point_repository)]
@router.get("/count")
async def count_tenant_points(
auth: _RequirePointsRead,
repository: _PointRepositoryDep,
domain: str | None = None,
file_id: uuid.UUID | None = None,
include_inactive: bool = False,
) -> PointCountResponse:
count = await count_points(
repository,
tenant_id=auth.tenant_id,
domain=domain,
file_id=file_id,
include_inactive=include_inactive,
)
return PointCountResponse(count=count)
@router.get("/search")
async def search_tenant_points(
auth: _RequirePointsRead,
repository: _PointRepositoryDep,
q: Annotated[str, Query(min_length=1)],
limit: LimitQuery = DEFAULT_PAGE_SIZE,
cursor: str | None = None,
domain: str | None = None,
file_id: uuid.UUID | None = None,
include_inactive: bool = False,
) -> PointSearchResponse:
"""Keyword search over point content — **not** semantic retrieval.
Matches Qdrant's full-text payload index on `content`, combined with the
structured filters below. Results are unranked: the index filters rather
than scores, so there is no relevance order and no score to return. Callers
wanting ranked answers want the agent retrieval path (plan 003), not this.
"""
page = await search_points(
repository,
tenant_id=auth.tenant_id,
query=q,
limit=limit,
cursor=cursor,
domain=domain,
file_id=file_id,
include_inactive=include_inactive,
)
return PointSearchResponse.from_search(page, query=q)
@router.get("")
async def list_tenant_points(
auth: _RequirePointsRead,
repository: _PointRepositoryDep,
file_id: uuid.UUID,
limit: LimitQuery = DEFAULT_PAGE_SIZE,
cursor: str | None = None,
include_inactive: bool = False,
) -> PointListResponse:
"""A file's points in `order_id` order.
`file_id` is required rather than optional: the pagination cursor is an
`order_id` value, and `order_id` is only unique within one file. Listing
across files would silently drop or repeat rows at every page boundary.
"""
page = await list_file_points(
repository,
tenant_id=auth.tenant_id,
file_id=file_id,
limit=limit,
cursor=cursor,
include_inactive=include_inactive,
)
return PointListResponse.from_page(page)
@router.get("/{point_id}")
async def get_tenant_point(
point_id: uuid.UUID,
auth: _RequirePointsRead,
repository: _PointRepositoryDep,
with_vectors: bool = False,
) -> PointResponse:
point = await get_point(
repository, tenant_id=auth.tenant_id, point_id=point_id, with_vectors=with_vectors
)
return PointResponse.from_point(point)
@router.delete("/{point_id}")
async def delete_tenant_point(
point_id: uuid.UUID,
auth: _RequirePointsWrite,
repository: _PointRepositoryDep,
) -> PointResponse:
"""Soft-delete one point and relink its neighbours around the gap.
The point is never removed from Qdrant (ADR-0002): it is flagged
`is_active=false` with `deleted_at` set, and its old neighbours are pointed
at each other in the same batch, so context-window expansion never walks
into it.
Deleting an already-inactive point is a no-op success rather than a `404` —
the response is the point as it stands, so the resulting `version` and
`deleted_at` are visible either way.
"""
point = await soft_delete_point(
repository,
tenant_id=auth.tenant_id,
point_id=point_id,
actor=f"api_key:{auth.api_key_id}",
)
return PointResponse.from_point(point)

View File

@@ -0,0 +1,63 @@
"""Public request/response models for `/v1/domains` (ADR-0008, ADR-0009).
`tenant_id` appears in none of these: it comes from the authenticated key, and
accepting it from a body would break the isolation boundary (ADR-0002).
"""
import uuid
from datetime import datetime
from pydantic import BaseModel, Field, field_validator
from src.application.domains.models import DomainResult
# Lowercase alphanumerics plus - and _; the key is embedded in every Qdrant
# payload and filtered on as a keyword, so it stays boring on purpose.
_DOMAIN_PATTERN = r"^[a-z0-9][a-z0-9_-]*$"
class DomainResponse(BaseModel):
id: uuid.UUID
domain: str
display_name: str
status: str
metadata: dict[str, object]
created_at: datetime
updated_at: datetime
@classmethod
def from_result(cls, result: DomainResult) -> "DomainResponse":
return cls(
id=result.id,
domain=result.domain,
display_name=result.display_name,
status=result.status,
metadata=result.metadata,
created_at=result.created_at,
updated_at=result.updated_at,
)
class DomainListResponse(BaseModel):
domains: list[DomainResponse]
class CreateDomainRequest(BaseModel):
domain: str = Field(min_length=1, max_length=80, pattern=_DOMAIN_PATTERN)
display_name: str = Field(min_length=1, max_length=200)
metadata: dict[str, object] = Field(default_factory=dict)
@field_validator("domain")
@classmethod
def _normalize(cls, value: str) -> str:
return value.strip()
class UpdateDomainRequest(BaseModel):
"""`domain` is absent by design — the key is immutable.
It is denormalized into every point payload and into `source_files`, so
renaming it is a migration rather than an edit (ADR-0009).
"""
display_name: str = Field(min_length=1, max_length=200)

16
src/api/schemas/errors.py Normal file
View File

@@ -0,0 +1,16 @@
"""The ADR-0008 error envelope."""
from typing import Any
from pydantic import BaseModel
class ErrorDetail(BaseModel):
code: str
message: str
details: dict[str, Any] = {}
request_id: str | None = None
class ErrorResponse(BaseModel):
error: ErrorDetail

64
src/api/schemas/files.py Normal file
View File

@@ -0,0 +1,64 @@
"""Public request/response models for `/v1/files` (ADR-0008).
Separate from the SQLAlchemy ORM models and the `application/files` domain
dataclasses (ADR-0015): this is the shape callers see.
"""
import uuid
from pydantic import BaseModel
from src.application.files.models import UploadResult
from src.application.files.status import FileStatusResult
class FileUploadResponse(BaseModel):
file_id: uuid.UUID
ingestion_job_id: uuid.UUID
status: str
chunks_indexed: int
@classmethod
def from_result(cls, result: UploadResult) -> "FileUploadResponse":
return cls(
file_id=result.file_id,
ingestion_job_id=result.ingestion_job_id,
status=result.status,
chunks_indexed=result.chunks_indexed,
)
class FileDeleteResponse(BaseModel):
"""What `DELETE /v1/files/{file_id}` did.
`points_soft_deleted` is reported rather than left implicit because the
delete is a soft one: nothing is removed from Qdrant, and the count is the
only way a caller can tell "deactivated 40 points" from "the file was
already deleted" — both of which are successes.
"""
file_id: uuid.UUID
status: str
points_soft_deleted: int
class FileStatusResponse(BaseModel):
file_id: uuid.UUID
source_filename: str
domain: str
status: str
ingestion_job_id: uuid.UUID | None
ingestion_status: str | None
chunks_indexed: int
@classmethod
def from_result(cls, result: FileStatusResult) -> "FileStatusResponse":
return cls(
file_id=result.file_id,
source_filename=result.source_filename,
domain=result.domain,
status=result.status,
ingestion_job_id=result.ingestion_job_id,
ingestion_status=result.ingestion_status,
chunks_indexed=result.chunks_indexed,
)

165
src/api/schemas/points.py Normal file
View File

@@ -0,0 +1,165 @@
"""Public request/response models for `/v1/points` (ADR-0002, ADR-0008).
The shape callers see, kept separate from `application/points`' domain models
(ADR-0015). Two rules are encoded here rather than left to route code:
- **Vectors are opt-in.** `PointResponse` omits them unless the caller asked,
so a listing does not ship megabytes of floats nobody reads (ADR-0008).
- **Server-owned fields are not accepted on input.** The request models simply
do not declare `tenant_id`, `version`, or `chunk_index`, and forbid extra
keys, so a client that sends one gets `422` from Pydantic instead of having
it silently ignored — ADR-0002's isolation rule enforced at the boundary.
"""
import uuid
from datetime import datetime
from typing import Annotated
from fastapi import Query
from pydantic import BaseModel, ConfigDict, Field
from src.application.points.point import Point
from src.application.ports.point_repository import PointPage
# A page ceiling the caller cannot raise. Scroll pages are materialized in
# memory both here and in Qdrant, so an unbounded `limit` is a cheap way for one
# request to hurt every other tenant sharing the process. Declared once because
# two routers paginate points -- `/v1/points` and `/v1/files/{file_id}/points` --
# and a ceiling that differs between them is a ceiling in only one of them.
DEFAULT_PAGE_SIZE = 50
MAX_PAGE_SIZE = 200
LimitQuery = Annotated[int, Query(ge=1, le=MAX_PAGE_SIZE)]
class PointResponse(BaseModel):
point_id: uuid.UUID
domain: str
file_id: uuid.UUID
chunk_id: uuid.UUID
content: str
content_type: str
source_filename: str
source_type: str
order_id: float
chunk_index: int
previous_chunk_id: uuid.UUID | None
next_chunk_id: uuid.UUID | None
is_active: bool
deleted_at: datetime | None
created_at: datetime
updated_at: datetime
created_by: str
updated_by: str
version: int
content_hash: str
embedding_model_version: str
vectors: dict[str, object] | None = None
@classmethod
def from_point(cls, point: Point) -> "PointResponse":
# `tenant_id` is present on `Point` and deliberately absent here: the
# caller already knows which tenant it authenticated as, and echoing it
# back invites clients to start sending it.
return cls.model_validate(point.model_dump(exclude={"tenant_id"}))
class PointListResponse(BaseModel):
"""A page of points plus the cursor for the next one.
Cursor-based rather than `limit`/`offset`: an offset cursor silently skips
or repeats rows when a concurrent insert shifts positions, which is exactly
the pagination defect plan 002 requires a test for.
"""
points: list[PointResponse]
next_cursor: str | None = None
@classmethod
def from_page(cls, page: PointPage) -> "PointListResponse":
return cls(
points=[PointResponse.from_point(point) for point in page.points],
next_cursor=page.next_cursor,
)
class PointCountResponse(BaseModel):
count: int
class PointSearchResponse(PointListResponse):
"""Results of a **keyword** match, not of semantic retrieval.
Named and documented so it cannot be mistaken for ADR-0003's hybrid
retrieval: these points matched a full-text filter on `content`, they are
not ranked by relevance, and there is no score to report. Anything that
wants ranked results wants the agent retrieval path in plan 003.
"""
query: str
@classmethod
def from_search(cls, page: PointPage, *, query: str) -> "PointSearchResponse":
return cls(
query=query,
points=[PointResponse.from_point(point) for point in page.points],
next_cursor=page.next_cursor,
)
class PointCreateRequest(BaseModel):
"""Create one point. The server assigns identity, ordering, and provenance.
`after_point_id` positions the new point rather than a raw `order_id`: the
caller says where in the sequence it goes and the server computes the
fractional key and relinks neighbours, which a client-supplied `order_id`
could not do correctly (ADR-0002). `None` means "at the start of the file".
"""
model_config = ConfigDict(extra="forbid")
file_id: uuid.UUID
content: str = Field(min_length=1)
content_type: str = "paragraph"
after_point_id: uuid.UUID | None = None
class PointReplaceRequest(BaseModel):
"""Replace a point's content under a version guard.
`version` here is the *expected* version, not a value being written — the
optimistic-concurrency precondition. A mismatch is `409`.
"""
model_config = ConfigDict(extra="forbid")
content: str = Field(min_length=1)
content_type: str | None = None
version: int
class PointPayloadPatchRequest(BaseModel):
"""Payload-only update of caller-writable fields, under a version guard."""
model_config = ConfigDict(extra="forbid")
payload: dict[str, object]
version: int
class PointReorderRequest(BaseModel):
"""Move a point to sit immediately after `after_point_id`.
`None` moves it to the front of the file. Expressed as a neighbour rather
than an `order_id` for the same reason as `PointCreateRequest`.
"""
model_config = ConfigDict(extra="forbid")
after_point_id: uuid.UUID | None = None
version: int

View File

View File

@@ -0,0 +1,29 @@
"""API-key authentication and tenant resolution (ADR-0008).
`resolve_auth_context` is the entry point: it takes a bearer token and
returns a trusted `AuthContext`. Everything downstream of the FastAPI
boundary receives `tenant_id` only through that context — never from a
request body, query string, or object metadata.
"""
from src.application.auth.context import AuthContext
from src.application.auth.errors import (
AuthError,
InvalidApiKeyError,
MissingScopeError,
TenantInactiveError,
)
from src.application.auth.keys import generate_api_key, hash_secret, verify_secret
from src.application.auth.service import resolve_auth_context
__all__ = [
"AuthContext",
"AuthError",
"InvalidApiKeyError",
"MissingScopeError",
"TenantInactiveError",
"generate_api_key",
"hash_secret",
"resolve_auth_context",
"verify_secret",
]

View File

@@ -0,0 +1,16 @@
"""The trusted request-scoped auth/tenant context (ADR-0008)."""
import uuid
from dataclasses import dataclass
@dataclass(frozen=True)
class AuthContext:
tenant_id: uuid.UUID
tenant_slug: str
api_key_id: uuid.UUID
scopes: frozenset[str]
actor_type: str
def has_scope(self, scope: str) -> bool:
return scope in self.scopes or "admin" in self.scopes

View File

@@ -0,0 +1,22 @@
"""Auth failures (ADR-0008). No HTTP knowledge here — `src/api/errors.py` maps
these to status codes.
"""
class AuthError(Exception):
"""Base class for auth failures."""
class InvalidApiKeyError(AuthError):
"""The bearer token is missing, malformed, unknown, revoked, or expired.
Maps to `401`.
"""
class TenantInactiveError(AuthError):
"""The key's tenant is suspended or deleted. Maps to `401`."""
class MissingScopeError(AuthError):
"""The key is valid but lacks a scope the route requires. Maps to `403`."""

View File

@@ -0,0 +1,39 @@
"""API-key generation and hashing (ADR-0008, ADR-0009).
Keys are `sk_{prefix}_{secret}`. `prefix` is non-secret and indexed
(`api_keys.key_prefix`); `secret` is 256 bits of `secrets.token_urlsafe`
entropy, stored only as a SHA-256 hash. A random 256-bit secret does not
benefit from a slow password-hashing KDF the way a human-chosen password
does — the cost that defends against dictionary/brute-force guessing over a
low-entropy input has nothing to defend here, and would only tax every
request. Comparison is constant-time to avoid a hash-timing oracle.
"""
import hashlib
import hmac
import secrets
_PREFIX_LENGTH = 16
def generate_api_key() -> tuple[str, str, str]:
"""Return `(key_prefix, secret, full_key)` for a newly issued key."""
key_prefix = secrets.token_hex(_PREFIX_LENGTH // 2)
secret = secrets.token_urlsafe(32)
return key_prefix, secret, f"sk_{key_prefix}_{secret}"
def parse_api_key(full_key: str) -> tuple[str, str] | None:
"""Return `(key_prefix, secret)`, or `None` if the token is malformed."""
parts = full_key.split("_", 2)
if len(parts) != 3 or parts[0] != "sk" or not parts[1] or not parts[2]:
return None
return parts[1], parts[2]
def hash_secret(secret: str) -> str:
return hashlib.sha256(secret.encode("utf-8")).hexdigest()
def verify_secret(secret: str, key_hash: str) -> bool:
return hmac.compare_digest(hash_secret(secret), key_hash)

View File

@@ -0,0 +1,86 @@
"""Resolve a bearer token to a trusted `AuthContext` (ADR-0008).
This opens and releases its own session rather than borrowing a
request-scoped one, so auth resolution never pins a pool connection across
the rest of the request — including the ADR-0017 ingestion work phase, which
must run with no Postgres session held open at all.
"""
from datetime import UTC, datetime
import structlog
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.application.auth.context import AuthContext
from src.application.auth.errors import InvalidApiKeyError, TenantInactiveError
from src.application.auth.keys import parse_api_key, verify_secret
from src.infrastructure.postgres.repositories import api_keys as api_keys_repo
from src.infrastructure.postgres.repositories import tenants as tenants_repo
logger = structlog.get_logger(__name__)
async def resolve_auth_context(
sessionmaker: async_sessionmaker[AsyncSession], bearer_token: str
) -> AuthContext:
"""Resolve a bearer token, logging the outcome either way (ADR-0011).
This runs on every authenticated request, so `auth.failed` is the one
event most likely to matter first when diagnosing a client integration
issue -- and the reason string alone (never logged; it can echo back
attacker-supplied key material) is not enough to tell a malformed token
apart from a revoked one without this.
"""
parsed = parse_api_key(bearer_token)
if parsed is None:
logger.warning("auth.failed", reason="malformed_key")
raise InvalidApiKeyError("malformed API key")
key_prefix, secret = parsed
async with sessionmaker() as session:
api_key = await api_keys_repo.get_by_prefix(session, key_prefix)
if api_key is None or not verify_secret(secret, api_key.key_hash):
logger.warning("auth.failed", reason="unknown_key", key_prefix=key_prefix)
raise InvalidApiKeyError("unknown API key")
if api_key.status != "active":
logger.warning(
"auth.failed",
reason="key_inactive",
key_prefix=key_prefix,
api_key_id=str(api_key.id),
key_status=api_key.status,
)
raise InvalidApiKeyError(f"API key is {api_key.status}")
if api_key.expires_at is not None and api_key.expires_at <= datetime.now(UTC):
logger.warning(
"auth.failed",
reason="key_expired",
key_prefix=key_prefix,
api_key_id=str(api_key.id),
)
raise InvalidApiKeyError("API key has expired")
tenant = await tenants_repo.get_by_id(session, api_key.tenant_id)
if tenant is None or tenant.status != "active":
logger.warning(
"auth.failed",
reason="tenant_inactive",
key_prefix=key_prefix,
api_key_id=str(api_key.id),
tenant_id=str(api_key.tenant_id),
)
raise TenantInactiveError("tenant is not active")
logger.info(
"auth.succeeded",
tenant_id=str(tenant.id),
api_key_id=str(api_key.id),
actor_type=api_key.actor_type,
)
return AuthContext(
tenant_id=tenant.id,
tenant_slug=tenant.slug,
api_key_id=api_key.id,
scopes=frozenset(api_key.scopes),
actor_type=api_key.actor_type,
)

View File

@@ -0,0 +1,27 @@
"""Tenant-domain management and the upload-time allowlist check (ADR-0009)."""
from src.application.domains.errors import (
DomainAlreadyExistsError,
DomainsError,
UnknownDomainError,
)
from src.application.domains.models import DomainResult
from src.application.domains.service import (
create_domain,
ensure_domain_allowed,
list_domains,
set_domain_status,
update_domain,
)
__all__ = [
"DomainAlreadyExistsError",
"DomainResult",
"DomainsError",
"UnknownDomainError",
"create_domain",
"ensure_domain_allowed",
"list_domains",
"set_domain_status",
"update_domain",
]

View File

@@ -0,0 +1,21 @@
"""Domain-management failures (ADR-0009). No HTTP knowledge here —
`src/api/errors.py` maps these to status codes.
"""
class DomainsError(Exception):
"""Base class for tenant-domain failures."""
class UnknownDomainError(DomainsError):
"""The upload named a domain the tenant has not registered, or one that is
disabled. Maps to `400`.
Rejecting is the whole point: an unrecognized `domain` would otherwise
create a new Qdrant partition silently, and a file in a partition nothing
queries is invisible rather than failed (ADR-0009).
"""
class DomainAlreadyExistsError(DomainsError):
"""The tenant already has a domain with this key. Maps to `409`."""

View File

@@ -0,0 +1,16 @@
"""Transport-agnostic results for the domain-management service."""
import uuid
from dataclasses import dataclass
from datetime import datetime
@dataclass(frozen=True)
class DomainResult:
id: uuid.UUID
domain: str
display_name: str
status: str
metadata: dict[str, object]
created_at: datetime
updated_at: datetime

View File

@@ -0,0 +1,164 @@
"""Tenant-domain management (ADR-0009).
A tenant's domain set is per-tenant and varies in size — one may run 14
insurance lines, another 6 — so it is data, not an enum.
`ensure_domain_allowed` is the reason this package exists: it is the strict
allowlist check the upload path runs before anything is written. Everything
else here is the management surface the calling backend uses to populate that
allowlist, under its own `domains:write` scope so an upload key cannot create
partitions.
`tenant_id` is always a required parameter taken from `AuthContext`, never from
a request body (ADR-0002).
"""
import uuid
import structlog
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.application.domains.errors import DomainAlreadyExistsError, UnknownDomainError
from src.application.domains.models import DomainResult
from src.infrastructure.postgres.models.tenant_domain import TenantDomain
from src.infrastructure.postgres.repositories import tenant_domains as repo
logger = structlog.get_logger(__name__)
async def _flush_and_refresh(session: AsyncSession, tenant_domain: TenantDomain) -> None:
"""Materialize server-generated columns before the row leaves the session.
`updated_at` is `onupdate=func.now()`, so after an UPDATE its value lives in
the database, not in the instance. Reading it later would trigger a lazy
load outside any greenlet context (`MissingGreenlet`), so it is fetched here
while the session is still open.
"""
await session.flush()
await session.refresh(tenant_domain)
def _to_result(tenant_domain: TenantDomain) -> DomainResult:
return DomainResult(
id=tenant_domain.id,
domain=tenant_domain.domain,
display_name=tenant_domain.display_name,
status=tenant_domain.status,
metadata=tenant_domain.metadata_,
created_at=tenant_domain.created_at,
updated_at=tenant_domain.updated_at,
)
async def ensure_domain_allowed(
session: AsyncSession, *, tenant_id: uuid.UUID, domain: str
) -> None:
"""Raise `UnknownDomainError` unless the tenant has this domain active.
Takes a session rather than a sessionmaker: the upload path calls this
inside its existing txn A, so the check costs no extra connection and
cannot pass and then go stale before the row is written.
Logs the rejection here rather than at the call site: this runs before any
`ingestion_jobs` row exists, so `upload_source_file`'s job-level
`ingestion.job.failed` event (ADR-0011) never fires for it -- without a log
here, a rejected upload would leave no operational trace at all.
"""
tenant_domain = await repo.get(session, tenant_id=tenant_id, domain=domain)
if tenant_domain is None:
logger.warning(
"domain.rejected", tenant_id=str(tenant_id), domain=domain, reason="unregistered"
)
raise UnknownDomainError(
f"domain '{domain}' is not registered for this tenant; "
f"create it via POST /v1/domains before uploading to it"
)
if tenant_domain.status != "active":
logger.warning(
"domain.rejected", tenant_id=str(tenant_id), domain=domain, reason="disabled"
)
raise UnknownDomainError(f"domain '{domain}' is disabled for this tenant")
async def list_domains(
sessionmaker: async_sessionmaker[AsyncSession],
*,
tenant_id: uuid.UUID,
include_disabled: bool = False,
) -> list[DomainResult]:
async with sessionmaker() as session:
found = await repo.list_for_tenant(
session, tenant_id=tenant_id, include_disabled=include_disabled
)
return [_to_result(item) for item in found]
async def create_domain(
sessionmaker: async_sessionmaker[AsyncSession],
*,
tenant_id: uuid.UUID,
domain: str,
display_name: str,
metadata: dict[str, object] | None = None,
) -> DomainResult:
async with sessionmaker() as session:
if await repo.get(session, tenant_id=tenant_id, domain=domain) is not None:
raise DomainAlreadyExistsError(f"domain '{domain}' already exists for this tenant")
created = repo.create(
session,
tenant_id=tenant_id,
domain=domain,
display_name=display_name,
metadata=metadata,
)
await session.commit()
logger.info("domain.created", tenant_id=str(tenant_id), domain=domain)
return _to_result(created)
async def update_domain(
sessionmaker: async_sessionmaker[AsyncSession],
*,
tenant_id: uuid.UUID,
domain: str,
display_name: str,
) -> DomainResult:
"""Only the label is mutable — see `repo.update_display_name`."""
async with sessionmaker() as session:
found = await repo.get(session, tenant_id=tenant_id, domain=domain)
if found is None:
raise UnknownDomainError(f"domain '{domain}' is not registered for this tenant")
repo.update_display_name(found, display_name=display_name)
await _flush_and_refresh(session, found)
await session.commit()
result = _to_result(found)
logger.info("domain.updated", tenant_id=str(tenant_id), domain=domain)
return result
async def set_domain_status(
sessionmaker: async_sessionmaker[AsyncSession],
*,
tenant_id: uuid.UUID,
domain: str,
status: str,
) -> DomainResult:
"""Disable or re-enable a domain.
Disabling blocks new uploads and hides the domain from pickers. It does not
touch the points already indexed under it — removing those needs the
tenant-erasure workflow plan 001 defers.
"""
async with sessionmaker() as session:
found = await repo.get(session, tenant_id=tenant_id, domain=domain)
if found is None:
raise UnknownDomainError(f"domain '{domain}' is not registered for this tenant")
repo.set_status(found, status=status)
await _flush_and_refresh(session, found)
await session.commit()
result = _to_result(found)
logger.info("domain.status_changed", tenant_id=str(tenant_id), domain=domain, status=status)
return result

View File

@@ -0,0 +1,19 @@
"""Source-file upload and status use cases (ADR-0008, ADR-0009, ADR-0017)."""
from src.application.files.errors import FilesError, FileTooLargeError, InvalidUploadError
from src.application.files.models import UploadResult, ValidatedUpload
from src.application.files.status import FileStatusResult, get_file_status
from src.application.files.upload import upload_source_file
from src.application.files.validation import validate_and_hash_upload
__all__ = [
"FileStatusResult",
"FileTooLargeError",
"FilesError",
"InvalidUploadError",
"UploadResult",
"ValidatedUpload",
"get_file_status",
"upload_source_file",
"validate_and_hash_upload",
]

View File

@@ -0,0 +1,81 @@
"""`DELETE /v1/files/{file_id}` — retire a file and deactivate its points.
Two stores have to agree here, and the phase boundaries are the same ones
ingestion uses (ADR-0017): a short Postgres transaction to authorize, then the
Qdrant work with **no session held**, then a short transaction to record the
outcome. Holding a session across the sweep would pin a pool connection for the
length of a multi-page delete.
The order — points first, Postgres second — is deliberate. If the sweep dies
half way, the row stays `active` and a retried `DELETE` finishes the job, since
the sweep only ever looks at points that are still active. The reverse order
would leave a row marked deleted while its points are still live and still
retrievable by the agent, which is the failure that actually matters.
"""
import uuid
from datetime import UTC, datetime
from time import perf_counter
import structlog
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.application.files.errors import SourceFileNotFoundError
from src.application.points.deletion import soft_delete_file_points
from src.application.ports.point_repository import PointRepository
from src.infrastructure.postgres.repositories import source_files as source_files_repo
logger = structlog.get_logger(__name__)
async def delete_source_file(
sessionmaker: async_sessionmaker[AsyncSession],
repository: PointRepository,
*,
tenant_id: uuid.UUID,
source_file_id: uuid.UUID,
actor: str,
) -> int:
"""Soft-delete a file: every active point, then the `source_files` row.
Returns how many points the sweep deactivated. Raises
`SourceFileNotFoundError` (`404`) when the file is not this tenant's — the
check happens before anything is written, so a probe for another tenant's
file id cannot deactivate a single point.
Idempotent: a second call finds no active points and a row already marked
`soft_deleted`, and returns `0`.
"""
started = perf_counter()
async with sessionmaker() as session:
source_file = await source_files_repo.get_by_id(
session, tenant_id=tenant_id, source_file_id=source_file_id
)
if source_file is None:
raise SourceFileNotFoundError(f"file {source_file_id} not found")
points_soft_deleted = await soft_delete_file_points(
repository, tenant_id=tenant_id, file_id=source_file_id, actor=actor
)
async with sessionmaker() as session:
source_file = await source_files_repo.get_by_id(
session, tenant_id=tenant_id, source_file_id=source_file_id
)
if source_file is None:
raise SourceFileNotFoundError(f"file {source_file_id} not found")
source_files_repo.mark_soft_deleted(source_file, deleted_at=datetime.now(UTC))
await session.commit()
logger.info(
"files.soft_deleted",
tenant_id=str(tenant_id),
file_id=str(source_file_id),
points_soft_deleted=points_soft_deleted,
actor=actor,
# End to end, including both Postgres transactions. Comparing it with
# the sweep's own `duration_ms` on `points.file_soft_deleted` is what
# separates a slow Qdrant from a slow database.
duration_ms=round((perf_counter() - started) * 1000, 2),
)
return points_soft_deleted

View File

@@ -0,0 +1,25 @@
"""Upload-validation failures (ADR-0008). No HTTP knowledge here —
`src/api/errors.py` maps these to status codes.
"""
class FilesError(Exception):
"""Base class for file-upload failures."""
class InvalidUploadError(FilesError):
"""Missing domain, empty file, or content that doesn't match its
declared extension. Maps to `400`.
"""
class FileTooLargeError(FilesError):
"""The upload exceeds `INGESTION_MAX_UPLOAD_SIZE_MB`. Maps to `413`."""
class SourceFileNotFoundError(FilesError):
"""No such source file *within the requesting tenant*. Maps to `404`.
Same non-disclosure rule as points (ADR-0016): a cross-tenant file id and a
nonexistent one are indistinguishable to the caller, so this is never `403`.
"""

View File

@@ -0,0 +1,26 @@
"""Domain models for the upload use case (ADR-0008, ADR-0009)."""
import uuid
from dataclasses import dataclass
@dataclass(frozen=True)
class ValidatedUpload:
"""The result of extension/content validation, before any I/O."""
source_type: str
content_type: str
content_sha256: str
@dataclass(frozen=True)
class UploadResult:
"""What `upload_source_file` returns; the route maps this to `FileUploadResponse`."""
file_id: uuid.UUID
ingestion_job_id: uuid.UUID
status: str
chunks_indexed: int
is_new_attempt: bool
"""`False` when an identical active upload already succeeded and no new
ingestion attempt was made (route returns `200`, not `201`)."""

View File

@@ -0,0 +1,52 @@
"""`GET /v1/files/{file_id}` read model (ADR-0008, ADR-0009)."""
import uuid
from dataclasses import dataclass
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.infrastructure.postgres.repositories import ingestion_jobs as jobs_repo
from src.infrastructure.postgres.repositories import source_files as source_files_repo
@dataclass(frozen=True)
class FileStatusResult:
file_id: uuid.UUID
source_filename: str
domain: str
status: str
ingestion_job_id: uuid.UUID | None
ingestion_status: str | None
chunks_indexed: int
async def get_file_status(
sessionmaker: async_sessionmaker[AsyncSession],
*,
tenant_id: uuid.UUID,
source_file_id: uuid.UUID,
) -> FileStatusResult | None:
"""Returns `None` when the file doesn't exist under this tenant — the
route maps that to `404`, never `403` (ADR-0016: cross-tenant access
returns 404).
"""
async with sessionmaker() as session:
source_file = await source_files_repo.get_by_id(
session, tenant_id=tenant_id, source_file_id=source_file_id
)
if source_file is None:
return None
latest_job = await jobs_repo.get_latest_for_source_file(
session, tenant_id=tenant_id, source_file_id=source_file.id
)
return FileStatusResult(
file_id=source_file.id,
source_filename=source_file.source_filename,
domain=source_file.domain,
status=source_file.status,
ingestion_job_id=latest_job.id if latest_job else None,
ingestion_status=latest_job.status if latest_job else None,
chunks_indexed=latest_job.points_created if latest_job else 0,
)

View File

@@ -0,0 +1,11 @@
"""Object-storage key derivation (ADR-0013).
Object keys are internal identifiers, never the caller-supplied filename.
Pure and synchronous.
"""
import uuid
def source_file_object_key(tenant_id: uuid.UUID, source_file_id: uuid.UUID) -> str:
return f"tenants/{tenant_id}/source-files/{source_file_id}/original"

View File

@@ -0,0 +1,366 @@
"""`POST /v1/files` orchestration: the ADR-0017 three-phase upload.
This service owns two separate short-lived sessions/transactions rather than
one request-scoped session, because the request is two units of work
(ADR-0012, ADR-0017):
txn A (short): source_files [+ ingestion_jobs(status='running')], commit
no txn: store bytes in MinIO, parse/chunk (threads),
embed dense+sparse (bounded/batched)
txn B (short): ingestion_jobs -> succeeded/failed, append event, commit
No Postgres session is open during phase 2. A failure at any point between
txn A and txn B still leaves a durable, inspectable `failed` job — never a
job stuck in `running`. The whole request additionally holds one of
`INGESTION_MAX_CONCURRENCY` process-wide slots (`503` when exhausted) and
phase 2 is bounded by `INGESTION_TIMEOUT_SECONDS` (`504`) (ADR-0017, plan 001
Phase 4).
Phase 2 ends by upserting the embedded chunks as tenant-scoped Qdrant points
(`src/application/points/`), so a successful upload is searchable by the time
the `201` returns. The collection those points land in is provisioned by a
deployment step, not by this path — see `src/cli/qdrant_bootstrap.py`.
"""
import uuid
from collections.abc import Sequence
import structlog
from anyio import CapacityLimiter, Semaphore, fail_after, to_thread
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.application.auth.context import AuthContext
from src.application.domains import ensure_domain_allowed
from src.application.files.errors import InvalidUploadError
from src.application.files.models import UploadResult
from src.application.files.storage_keys import source_file_object_key
from src.application.files.validation import validate_and_hash_upload
from src.application.ingestion import (
ChunkTooLargeError,
DocumentParseError,
UnsupportedSourceTypeError,
parse_and_chunk_document,
)
from src.application.ingestion.bounds import acquire_ingestion_slot, enforce_chunk_limit
from src.application.ingestion.embedding import embed_chunks
from src.application.ingestion.errors import (
ChunkLimitExceededError,
EmbedderError,
IngestionTimeoutError,
PointIndexingError,
)
from src.application.points import index_chunks
from src.application.ports.embedding import DenseEmbedder, SparseEmbedder
from src.application.ports.object_storage import ObjectStorage
from src.application.ports.point_storage import PointStorage
from src.config import ChunkingSettings, IngestionSettings, QdrantSettings
from src.infrastructure.postgres.repositories import ingestion_jobs as jobs_repo
from src.infrastructure.postgres.repositories import source_files as source_files_repo
logger = structlog.get_logger(__name__)
async def _mark_job_failed(
sessionmaker: async_sessionmaker[AsyncSession],
*,
tenant_id: uuid.UUID,
ingestion_job_id: uuid.UUID,
error_code: str,
error_message: str,
) -> None:
"""Write the terminal `failed` job row and emit its log event together.
Every failure branch below calls this, so logging here once closes every
branch at once rather than duplicating a `logger.warning` at each call
site (CLAUDE.md, "prefer deep modules") -- previously only
`storage_upload_failed` and `timeout` did that ad hoc, and
`parse_failed`/`chunk_limit_exceeded`/`embedding_failed`/`index_failed`
logged nothing at all: visible in `ingestion_job_events` but invisible to
log-based alerting (ADR-0011).
"""
async with sessionmaker() as session:
job = await jobs_repo.mark_terminal(
session,
tenant_id=tenant_id,
ingestion_job_id=ingestion_job_id,
status="failed",
error_code=error_code,
error_message=error_message,
)
if job is not None:
jobs_repo.append_event(
session,
tenant_id=tenant_id,
ingestion_job_id=ingestion_job_id,
level="error",
stage="received",
message=error_message,
)
await session.commit()
logger.warning(
"ingestion.job.failed",
tenant_id=str(tenant_id),
ingestion_job_id=str(ingestion_job_id),
error_code=error_code,
error_message=error_message,
)
async def upload_source_file(
*,
sessionmaker: async_sessionmaker[AsyncSession],
storage: ObjectStorage,
point_storage: PointStorage,
auth: AuthContext,
domain: str,
filename: str,
data: bytes,
ingestion_settings: IngestionSettings,
chunking_settings: ChunkingSettings,
qdrant_settings: QdrantSettings,
thread_limiter: CapacityLimiter,
concurrency_limiter: Semaphore,
dense_embedders: Sequence[DenseEmbedder],
sparse_embedder: SparseEmbedder,
) -> UploadResult:
domain = domain.strip()
if not domain:
raise InvalidUploadError("domain is required")
validated = await to_thread.run_sync(
lambda: validate_and_hash_upload(
filename=filename, data=data, max_size_bytes=ingestion_settings.max_upload_size_bytes
),
limiter=thread_limiter,
)
async with acquire_ingestion_slot(concurrency_limiter):
async with sessionmaker() as session:
# Strict allowlist, checked inside txn A before anything is written
# (ADR-0009). An unregistered domain would otherwise create a new
# Qdrant partition silently, leaving the file invisible to
# retrieval rather than failing.
await ensure_domain_allowed(session, tenant_id=auth.tenant_id, domain=domain)
existing = await source_files_repo.find_active_by_content_hash(
session,
tenant_id=auth.tenant_id,
domain=domain,
content_sha256=validated.content_sha256,
)
if existing is not None:
latest_job = await jobs_repo.get_latest_for_source_file(
session, tenant_id=auth.tenant_id, source_file_id=existing.id
)
if latest_job is not None and latest_job.status == "succeeded":
logger.info(
"files.upload.duplicate",
tenant_id=str(auth.tenant_id),
file_id=str(existing.id),
)
return UploadResult(
file_id=existing.id,
ingestion_job_id=latest_job.id,
status=latest_job.status,
chunks_indexed=latest_job.points_created,
is_new_attempt=False,
)
source_file_id = existing.id
object_key = existing.storage_uri or source_file_object_key(
auth.tenant_id, source_file_id
)
else:
source_file_id = uuid.uuid4()
object_key = source_file_object_key(auth.tenant_id, source_file_id)
source_files_repo.create(
session,
source_file_id=source_file_id,
tenant_id=auth.tenant_id,
domain=domain,
source_filename=filename,
source_type=validated.source_type,
content_sha256=validated.content_sha256,
byte_size=len(data),
storage_uri=object_key,
created_by_api_key_id=auth.api_key_id,
)
# `ingestion_jobs.source_file_id` FKs to this row; flush so the
# insert below sees it, since the two mapped classes carry no
# ORM relationship for the unit of work to order by itself.
await session.flush()
job = jobs_repo.create_running(
session,
tenant_id=auth.tenant_id,
source_file_id=source_file_id,
requested_by_api_key_id=auth.api_key_id,
chunking_strategy=chunking_settings.strategy,
)
jobs_repo.append_event(
session,
tenant_id=auth.tenant_id,
ingestion_job_id=job.id,
level="info",
stage="received",
message="upload accepted, storing object",
)
await session.commit()
ingestion_job_id = job.id
logger.info(
"ingestion.job.started",
tenant_id=str(auth.tenant_id),
ingestion_job_id=str(ingestion_job_id),
file_id=str(source_file_id),
domain=domain,
source_type=validated.source_type,
)
# Phase 2: no Postgres session open across this work (ADR-0017),
# bounded end-to-end by INGESTION_TIMEOUT_SECONDS.
try:
with fail_after(ingestion_settings.timeout_seconds):
try:
await storage.put_object(
key=object_key, data=data, content_type=validated.content_type
)
except Exception as exc:
await _mark_job_failed(
sessionmaker,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
error_code="storage_upload_failed",
error_message=f"failed to store object: {exc}",
)
raise
try:
chunks = await parse_and_chunk_document(
data,
source_type=validated.source_type,
file_id=source_file_id,
settings=chunking_settings,
limiter=thread_limiter,
)
except (DocumentParseError, UnsupportedSourceTypeError, ChunkTooLargeError) as exc:
await _mark_job_failed(
sessionmaker,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
error_code="parse_failed",
error_message=str(exc),
)
raise
try:
enforce_chunk_limit(chunks, max_chunks=ingestion_settings.max_chunks_per_file)
except ChunkLimitExceededError as exc:
await _mark_job_failed(
sessionmaker,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
error_code="chunk_limit_exceeded",
error_message=str(exc),
)
raise
try:
embedded = await embed_chunks(
chunks,
dense_embedders=dense_embedders,
sparse_embedder=sparse_embedder,
settings=ingestion_settings,
thread_limiter=thread_limiter,
)
except EmbedderError as exc:
await _mark_job_failed(
sessionmaker,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
error_code="embedding_failed",
error_message=str(exc),
)
raise
try:
indexed = await index_chunks(
embedded,
storage=point_storage,
tenant_id=auth.tenant_id,
domain=domain,
file_id=source_file_id,
source_filename=filename,
source_type=validated.source_type,
actor=f"api_key:{auth.api_key_id}",
dense_embedders=dense_embedders,
sparse_embedder=sparse_embedder,
settings=qdrant_settings,
thread_limiter=thread_limiter,
)
except PointIndexingError as exc:
await _mark_job_failed(
sessionmaker,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
error_code="index_failed",
error_message=str(exc),
)
raise
except TimeoutError:
await _mark_job_failed(
sessionmaker,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
error_code="timeout",
error_message=f"ingestion exceeded {ingestion_settings.timeout_seconds}s",
)
raise IngestionTimeoutError(
f"ingestion exceeded {ingestion_settings.timeout_seconds}s"
) from None
async with sessionmaker() as session:
await jobs_repo.mark_terminal(
session,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
status="succeeded",
# An upsert with deterministic ids cannot tell an insert from
# an overwrite, so every written point is reported here and
# `points_updated` stays 0 rather than being guessed at.
points_created=indexed.points_upserted,
points_soft_deleted=indexed.points_soft_deleted,
)
jobs_repo.append_event(
session,
tenant_id=auth.tenant_id,
ingestion_job_id=ingestion_job_id,
level="info",
stage="completed",
message="chunks parsed, embedded, and indexed",
details={
"chunks_parsed": len(chunks),
"chunks_embedded": len(embedded),
"points_upserted": indexed.points_upserted,
"points_soft_deleted": indexed.points_soft_deleted,
},
)
await session.commit()
logger.info(
"ingestion.job.completed",
tenant_id=str(auth.tenant_id),
ingestion_job_id=str(ingestion_job_id),
file_id=str(source_file_id),
chunks_parsed=len(chunks),
points_upserted=indexed.points_upserted,
points_soft_deleted=indexed.points_soft_deleted,
)
return UploadResult(
file_id=source_file_id,
ingestion_job_id=ingestion_job_id,
status="succeeded",
chunks_indexed=indexed.points_upserted,
is_new_attempt=True,
)

View File

@@ -0,0 +1,55 @@
"""Upload extension/content-type/size validation (ADR-0008).
Pure and synchronous: no I/O. `content_sha256` computation lives here too —
hashing is blocking CPU work (ADR-0017), so the caller runs this whole
function through `anyio.to_thread.run_sync` with the ingestion
`CapacityLimiter`, the same rule applied to parsing/chunking.
"""
import hashlib
from src.application.files.errors import FileTooLargeError, InvalidUploadError
from src.application.files.models import ValidatedUpload
from src.application.ingestion.errors import UnsupportedSourceTypeError
_CONTENT_TYPES = {
"csv": "text/csv",
"xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
"docx": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
}
_OOXML_MAGIC = b"PK\x03\x04"
def _source_type_from_filename(filename: str) -> str:
suffix = filename.rsplit(".", 1)[-1].lower() if "." in filename else ""
if suffix == "doc":
raise UnsupportedSourceTypeError(
"legacy .doc is not ingestible until an out-of-process conversion "
"service exists (ADR-0018)"
)
if suffix not in _CONTENT_TYPES:
raise UnsupportedSourceTypeError(f"'.{suffix}' is not an ingestible file type")
return suffix
def validate_and_hash_upload(*, filename: str, data: bytes, max_size_bytes: int) -> ValidatedUpload:
source_type = _source_type_from_filename(filename)
if not data:
raise InvalidUploadError("uploaded file is empty")
if len(data) > max_size_bytes:
raise FileTooLargeError(
f"upload is {len(data)} bytes, over the {max_size_bytes}-byte limit"
)
is_ooxml = data[:4] == _OOXML_MAGIC
if source_type in ("docx", "xlsx") and not is_ooxml:
raise InvalidUploadError(f"content does not match the declared .{source_type} extension")
if source_type == "csv" and is_ooxml:
raise InvalidUploadError("content does not match the declared .csv extension")
return ValidatedUpload(
source_type=source_type,
content_type=_CONTENT_TYPES[source_type],
content_sha256=hashlib.sha256(data).hexdigest(),
)

View File

@@ -0,0 +1,51 @@
"""Document parsing and fixed-size chunking (ADR-0004, ADR-0018).
`parse_and_chunk_document` is the entry point callers outside this package
should use: it dispatches on source type and owns the
`anyio.to_thread.run_sync` + `CapacityLimiter` offload required by ADR-0017.
The individual parsers and `chunk_document` are pure, synchronous, and
exported mainly for their own unit tests — calling them directly from an
`async def` route or service is the defect ADR-0017 warns about.
"""
from src.application.ingestion.chunking import chunk_document, chunk_id_for, split_by_tokens
from src.application.ingestion.docx_parser import parse_docx
from src.application.ingestion.errors import (
ChunkLimitExceededError,
ChunkTooLargeError,
DocumentParseError,
IngestionError,
UnsupportedSourceTypeError,
)
from src.application.ingestion.models import (
Chunk,
ContentType,
ParsedDocument,
StructuralUnit,
)
from src.application.ingestion.normalization import normalize_persian_text
from src.application.ingestion.pipeline import parse_and_chunk_document
from src.application.ingestion.spreadsheet_parser import parse_csv, parse_xlsx
from src.application.ingestion.tokenizer import count_tokens, get_encoder
__all__ = [
"Chunk",
"ChunkLimitExceededError",
"ChunkTooLargeError",
"ContentType",
"DocumentParseError",
"IngestionError",
"ParsedDocument",
"StructuralUnit",
"UnsupportedSourceTypeError",
"chunk_document",
"chunk_id_for",
"count_tokens",
"get_encoder",
"normalize_persian_text",
"parse_and_chunk_document",
"parse_csv",
"parse_docx",
"parse_xlsx",
"split_by_tokens",
]

View File

@@ -0,0 +1,48 @@
"""Request bounds for inline ingestion (ADR-0017).
Three independent bounds, each mapping to its own status code: the chunk
ceiling (`413`, checked before embedding starts), process-wide concurrency
(`503` + `Retry-After`, rejected rather than queued), and the work-phase
deadline (`504`, and the caller must still write a terminal job status).
"""
from collections.abc import AsyncIterator, Sequence
from contextlib import asynccontextmanager
from anyio import Semaphore, WouldBlock
from src.application.ingestion.errors import ChunkLimitExceededError, IngestionAtCapacityError
from src.application.ingestion.models import Chunk
def enforce_chunk_limit(chunks: Sequence[Chunk], *, max_chunks: int) -> None:
"""Raise `ChunkLimitExceededError` if `chunks` exceeds `max_chunks`.
Call this immediately after parsing/chunking and before any embedding
call — the ceiling must be discovered up front, not mid-batch.
"""
if len(chunks) > max_chunks:
raise ChunkLimitExceededError(
f"document produced {len(chunks)} chunks, over the {max_chunks}-chunk limit"
)
@asynccontextmanager
async def acquire_ingestion_slot(limiter: Semaphore) -> AsyncIterator[None]:
"""Hold one of `INGESTION_MAX_CONCURRENCY` process-wide slots for the block.
`limiter` is an `anyio.Semaphore` created once in the lifespan. Rejects
immediately with `IngestionAtCapacityError` when the process is already at
capacity, rather than queueing the request behind an unbounded wait
(ADR-0017) — the semaphore's own async `acquire()` would do the latter.
"""
try:
limiter.acquire_nowait()
except WouldBlock:
raise IngestionAtCapacityError(
"ingestion is at capacity; retry after the configured backoff"
) from None
try:
yield
finally:
limiter.release()

View File

@@ -0,0 +1,129 @@
"""Fixed-size chunking with overlap (ADR-0018).
DOCX markdown is split on token windows; spreadsheet rows are already atomic
and bypass the splitter, falling through it only when a single row exceeds the
embedding model's sequence length.
"""
import uuid
from src.application.ingestion.errors import ChunkTooLargeError, DocumentParseError
from src.application.ingestion.models import Chunk, ContentType, ParsedDocument
from src.application.ingestion.tokenizer import count_tokens, get_encoder
from src.config import ChunkingSettings
# Fixed namespace so chunk ids stay stable across processes and releases.
CHUNK_ID_NAMESPACE = uuid.UUID("6f9619ff-8b86-d011-b42d-00c04fc964ff")
def chunk_id_for(file_id: uuid.UUID, chunk_index: int) -> uuid.UUID:
"""Return the deterministic point id for a chunk (ADR-0001).
Derived from `file_id` and the immutable ingestion ordinal, so re-ingesting
a file upserts its points instead of duplicating them.
"""
return uuid.uuid5(CHUNK_ID_NAMESPACE, f"{file_id}:{chunk_index}")
def split_by_tokens(text: str, *, chunk_size: int, overlap: int, encoding_name: str) -> list[str]:
"""Split text into overlapping windows of at most `chunk_size` tokens."""
if overlap >= chunk_size:
raise ValueError(f"overlap ({overlap}) must be smaller than chunk_size ({chunk_size})")
encoder = get_encoder(encoding_name)
tokens = encoder.encode(text)
if len(tokens) <= chunk_size:
return [text]
windows: list[str] = []
start = 0
while start < len(tokens):
end = min(start + chunk_size, len(tokens))
windows.append(encoder.decode(tokens[start:end]))
if end >= len(tokens):
break
start = end - overlap
return windows
def _split_oversized(text: str, settings: ChunkingSettings) -> list[str]:
"""Split a row only if it exceeds the cap; otherwise keep it atomic."""
if count_tokens(text, settings.encoding_name) <= settings.max_chunk_tokens:
return [text]
return split_by_tokens(
text,
chunk_size=settings.chunk_size,
overlap=settings.chunk_overlap,
encoding_name=settings.encoding_name,
)
def _content_units(
parsed: ParsedDocument, settings: ChunkingSettings
) -> list[tuple[str, ContentType]]:
"""Reduce a parsed document's structural units to ordered text pieces.
Prose is cut into token windows; a table row is atomic and survives whole
unless it alone exceeds the model's sequence length (ADR-0004).
"""
if not parsed.units:
raise DocumentParseError("Parsed document has no structural units")
pieces: list[tuple[str, ContentType]] = []
for unit in parsed.units:
if unit.content_type is ContentType.PARAGRAPH:
windows = split_by_tokens(
unit.text,
chunk_size=settings.chunk_size,
overlap=settings.chunk_overlap,
encoding_name=settings.encoding_name,
)
else:
windows = _split_oversized(unit.text, settings)
pieces.extend((window, unit.content_type) for window in windows)
return pieces
def chunk_document(
parsed: ParsedDocument,
*,
file_id: uuid.UUID,
settings: ChunkingSettings,
) -> list[Chunk]:
"""Turn a parsed document into ordered, neighbor-linked chunks."""
units = [
(text.strip(), content_type) for text, content_type in _content_units(parsed, settings)
]
# Drop blanks *before* assigning indices: an index gap would break the
# previous/next chain that retrieval-time expansion walks.
units = [(text, content_type) for text, content_type in units if text]
chunks: list[Chunk] = []
for index, (text, content_type) in enumerate(units):
token_count = count_tokens(text, settings.encoding_name)
if token_count > settings.max_chunk_tokens:
raise ChunkTooLargeError(
f"chunk {index} is {token_count} tokens, over the "
f"{settings.max_chunk_tokens}-token cap"
)
chunks.append(
Chunk(
chunk_id=chunk_id_for(file_id, index),
chunk_index=index,
order_id=float(index + 1),
content=text,
content_type=content_type,
token_count=token_count,
character_count=len(text),
)
)
for position, chunk in enumerate(chunks):
if position > 0:
chunk.previous_chunk_id = chunks[position - 1].chunk_id
if position < len(chunks) - 1:
chunk.next_chunk_id = chunks[position + 1].chunk_id
return chunks

View File

@@ -0,0 +1,165 @@
"""DOCX parsing into ordered structural units (ADR-0004, ADR-0018).
The body is walked in document order and decomposed into structural units
before any chunking runs: runs of flowing prose become `PARAGRAPH` units the
splitter cuts into token windows, and data-table rows become atomic
`TABLE_ROW` units.
The one classification this makes is between a *data* table and a table used
as page layout, and it is made structurally rather than by inspecting content:
a data cell fits inside a chunk by definition, so a table holding a cell that
alone exceeds `chunk_size`, or a cell containing nested tables, is a layout
container whose cells are prose.
"""
import io
from docx import Document
from docx.document import Document as DocxDocument
from docx.oxml.table import CT_Tbl
from docx.oxml.text.paragraph import CT_P
from docx.table import Table, _Cell
from docx.text.paragraph import Paragraph
from src.application.ingestion.errors import DocumentParseError
from src.application.ingestion.models import (
ContentType,
ParsedDocument,
StructuralUnit,
)
from src.application.ingestion.normalization import normalize_persian_text
from src.application.ingestion.tabular import clean_rows, render_rows
from src.application.ingestion.tokenizer import count_tokens
from src.config import ChunkingSettings
_NORMAL_STYLE = "Normal"
def heading_level_from_style(style_name: str) -> int | None:
"""Return the heading level of a paragraph style, or None for body text.
Word stores the styleId (`Heading1`), not the friendly name (`Heading 1`),
so both spellings must resolve. Only real Word styles count -- no heading
is ever inferred from the text itself (ADR-0018).
"""
normalized = style_name.strip().lower().replace(" ", "")
if not normalized.startswith("heading"):
return None
suffix = normalized.removeprefix("heading")
return int(suffix) if suffix.isdigit() else None
def _paragraph_style(paragraph: Paragraph) -> str:
style = paragraph.style.name if paragraph.style is not None else None
return style or _NORMAL_STYLE
def _is_layout_table(table: Table, settings: ChunkingSettings) -> bool:
"""Whether a table is page layout rather than data.
Structural, not a content heuristic: a data cell is small enough to be a
chunk, so a cell that alone overflows `chunk_size` -- or that nests another
table -- holds a document, not a field. In the sample corpus this separates
by two orders of magnitude (32-171 tokens for data tables against 54,007
for a cell containing a whole sub-document).
"""
for row in table.rows:
for cell in row.cells:
if cell.tables:
return True
if count_tokens(cell.text, settings.encoding_name) > settings.chunk_size:
return True
return False
def _cell_prose(cell: _Cell, settings: ChunkingSettings) -> list[StructuralUnit]:
"""Extract a layout cell's contents as units, recursing into nested tables."""
units: list[StructuralUnit] = []
for block in _iter_block_items(cell, settings):
units.append(block)
return units
def _table_units(table: Table, settings: ChunkingSettings) -> list[StructuralUnit]:
"""Convert a table to structural units."""
if _is_layout_table(table, settings):
units: list[StructuralUnit] = []
# Walk the physical `w:tc` elements rather than `row.cells`, which
# repeats a merged cell once per grid position it spans. Identity
# tracking is not an option here: lxml builds element proxies on
# demand, so `id()` is neither stable nor unique across them.
for row in table.rows:
for tc in row._tr.tc_lst:
units.extend(_cell_prose(_Cell(tc, table), settings))
return units
rows = clean_rows([[cell.text for cell in row.cells] for row in table.rows])
return [
StructuralUnit(text=text, content_type=ContentType.TABLE_ROW) for text in render_rows(rows)
]
def _iter_block_items(
container: DocxDocument | _Cell, settings: ChunkingSettings
) -> list[StructuralUnit]:
"""Walk a body or cell in document order, emitting structural units.
Consecutive paragraphs accumulate into one prose unit rather than becoming
one unit each: ADR-0004 treats "the whole remaining run of paragraphs" as a
single prose block, so the splitter sees flowing text instead of a series
of one-sentence fragments.
"""
element = container.element.body if isinstance(container, DocxDocument) else container._tc
units: list[StructuralUnit] = []
prose: list[str] = []
def flush() -> None:
if prose:
units.append(
StructuralUnit(text="\n\n".join(prose), content_type=ContentType.PARAGRAPH)
)
prose.clear()
for child in element:
if isinstance(child, CT_P):
paragraph = Paragraph(child, container)
# `Paragraph.text` includes hyperlink text, which a raw `w:r` walk
# silently drops.
text = normalize_persian_text(paragraph.text)
if not text:
continue
level = heading_level_from_style(_paragraph_style(paragraph))
prose.append(f"{'#' * level} {text}" if level else text)
elif isinstance(child, CT_Tbl):
table_units = _table_units(Table(child, container), settings)
# A layout table is prose; keep it in the surrounding prose block
# instead of fragmenting the document around it.
if table_units and all(
unit.content_type is ContentType.PARAGRAPH for unit in table_units
):
prose.extend(unit.text for unit in table_units)
else:
flush()
units.extend(table_units)
flush()
return units
def parse_docx(data: bytes, settings: ChunkingSettings) -> ParsedDocument:
"""Parse DOCX bytes into ordered structural units."""
try:
doc = Document(io.BytesIO(data))
except Exception as exc:
raise DocumentParseError(f"Could not open DOCX: {exc}") from exc
units = _iter_block_items(doc, settings)
if not units:
raise DocumentParseError("Document contains no text content")
return ParsedDocument(
units=units,
markdown="\n\n".join(unit.text for unit in units),
block_count=len(units),
)

View File

@@ -0,0 +1,112 @@
"""The one caller-facing entry point for embedding chunks (ADR-0001, ADR-0017).
`embed_chunks` is the only version of this step callers should reach for: it
owns batching, the `embed_concurrency` semaphore bounding in-flight dense
batches, and the `anyio.to_thread.run_sync` + `CapacityLimiter` offload for
the blocking BM25 pipeline. Composing these correctly at every call site is
exactly the obligation a deep module absorbs once (see CLAUDE.md's "prefer
deep modules").
Per-provider text shaping — task prefixes, `keep_alive`, request payload —
belongs to the adapters in `src/infrastructure/embedding/`, not here. This
module knows only that an embedder turns texts into vectors.
"""
import asyncio
from collections.abc import Sequence
from functools import partial
from anyio import CapacityLimiter, to_thread
from src.application.ingestion.errors import EmbedderError
from src.application.ingestion.models import Chunk, EmbeddedChunk
from src.application.ports.embedding import DenseEmbedder, SparseEmbedder
from src.config import IngestionSettings
def _batches(texts: Sequence[str], size: int) -> list[Sequence[str]]:
return [texts[i : i + size] for i in range(0, len(texts), size)]
async def _embed_dense_bounded(
embedder: DenseEmbedder,
batch: Sequence[str],
*,
semaphore: asyncio.Semaphore,
) -> list[list[float]]:
async with semaphore:
try:
return await embedder.embed_batch(batch)
except Exception as exc:
raise EmbedderError(f"{embedder.name} embedding batch failed: {exc}") from exc
async def _embed_dense_all(
embedder: DenseEmbedder,
texts: Sequence[str],
*,
batch_size: int,
semaphore: asyncio.Semaphore,
) -> list[list[float]]:
batches = _batches(texts, batch_size)
results = await asyncio.gather(
*(_embed_dense_bounded(embedder, batch, semaphore=semaphore) for batch in batches)
)
return [vector for batch_result in results for vector in batch_result]
def _embed_sparse_sync(embedder: SparseEmbedder, texts: Sequence[str]):
try:
return embedder.embed_batch(texts)
except Exception as exc:
raise EmbedderError(f"{embedder.name} embedding batch failed: {exc}") from exc
async def embed_chunks(
chunks: Sequence[Chunk],
*,
dense_embedders: Sequence[DenseEmbedder],
sparse_embedder: SparseEmbedder,
settings: IngestionSettings,
thread_limiter: CapacityLimiter,
) -> list[EmbeddedChunk]:
"""Embed every chunk into all dense vectors plus the sparse vector.
Dense embedders run concurrently with each other; each one's batches are
concurrent among themselves too, bounded by one `embed_concurrency`
semaphore shared across all dense embedders (ADR-0017: the limit exists
for both providers' rate limits and the self-hosted server's capacity —
not a per-provider budget). The sparse (BM25) pass is blocking and runs
once, off the event loop.
Raises `EmbedderError` (502) if any embedder call fails.
"""
if not chunks:
return []
texts = [chunk.content for chunk in chunks]
semaphore = asyncio.Semaphore(settings.embed_concurrency)
dense_task = asyncio.gather(
*(
_embed_dense_all(
embedder, texts, batch_size=settings.embed_batch_size, semaphore=semaphore
)
for embedder in dense_embedders
)
)
sparse_task = to_thread.run_sync(
partial(_embed_sparse_sync, sparse_embedder, texts), limiter=thread_limiter
)
dense_results, sparse_vectors = await asyncio.gather(dense_task, sparse_task)
dense_by_name = {
embedder.name: vectors
for embedder, vectors in zip(dense_embedders, dense_results, strict=True)
}
embedded: list[EmbeddedChunk] = []
for index, chunk in enumerate(chunks):
dense = {name: vectors[index] for name, vectors in dense_by_name.items()}
embedded.append(EmbeddedChunk(chunk=chunk, dense=dense, sparse=sparse_vectors[index]))
return embedded

View File

@@ -0,0 +1,73 @@
"""Errors raised by the parsing and chunking pipeline (ADR-0018).
These carry no HTTP knowledge — the API layer maps them to status codes
(ADR-0015: `application/` contains no FastAPI request objects).
"""
class IngestionError(Exception):
"""Base class for ingestion failures."""
class DocumentParseError(IngestionError):
"""A source file could not be decoded, opened, or yielded no text.
Maps to `400` per ADR-0017 ("unparseable file → 400").
"""
class UnsupportedSourceTypeError(IngestionError):
"""A source file's type is not ingestible in this version.
Maps to `415`. `.doc` lands here until an out-of-process conversion
service exists (ADR-0018).
"""
class ChunkLimitExceededError(IngestionError):
"""A document produced more chunks than `INGESTION_MAX_CHUNKS_PER_FILE`.
Maps to `413` per ADR-0017.
"""
class ChunkTooLargeError(IngestionError):
"""A chunk exceeded the embedding model's sequence length.
This is an internal invariant violation, not a user error: the splitter is
supposed to make it impossible. It exists because the failure it guards
against is silent — `nomic-embed-text-v2-moe` truncates over-long input
without raising (ADR-0004).
"""
class EmbedderError(IngestionError):
"""A dense or sparse embedder call failed (transport error, non-2xx, or
a malformed response).
Maps to `502` per ADR-0017.
"""
class PointIndexingError(IngestionError):
"""Upserting or soft-deleting Qdrant points failed.
Maps to `502` — like `EmbedderError`, this is an upstream dependency
failing, not a malformed request. Kept distinct from `EmbedderError` so the
job's `error_code` says which dependency broke.
"""
class IngestionAtCapacityError(IngestionError):
"""`INGESTION_MAX_CONCURRENCY` in-process ingestions are already running.
Maps to `503` with `Retry-After`, not a queued wait (ADR-0017).
"""
class IngestionTimeoutError(IngestionError):
"""The work phase (parse/embed/upsert) exceeded `INGESTION_TIMEOUT_SECONDS`.
Maps to `504`. The caller must still write a terminal `failed` job status
before this propagates (ADR-0017).
"""

View File

@@ -0,0 +1,91 @@
"""Domain models for parsing and chunking (ADR-0001, ADR-0004, ADR-0018)."""
import uuid
from enum import StrEnum
from pydantic import BaseModel, Field
class ContentType(StrEnum):
"""ADR-0004's finalized `content_type` value set.
v1 emits `PARAGRAPH` and `TABLE_ROW` only; `QA_PAIR` and `IMAGE_CAPTION`
are defined but not produced yet (ADR-0018).
"""
PARAGRAPH = "paragraph"
TABLE_ROW = "table_row"
QA_PAIR = "qa_pair"
IMAGE_CAPTION = "image_caption"
class StructuralUnit(BaseModel):
"""One structural unit of a document, in reading order (ADR-0004).
A document is decomposed into these *before* any chunking runs, because
the two kinds are chunked differently:
- `PARAGRAPH` is a run of flowing prose; the fixed-size splitter cuts it
into token windows.
- `TABLE_ROW` is already atomic; it becomes one chunk, and is split only
when a single row is too large for the embedding model.
"""
text: str
content_type: ContentType
class ParsedDocument(BaseModel):
"""The output of a parser, before chunking.
`units` is the content, in document order. `markdown` is the same content
rendered as one string, for eyeballing a parse; nothing chunks from it.
"""
units: list[StructuralUnit] = Field(default_factory=list)
markdown: str | None = None
block_count: int = 0
class Chunk(BaseModel):
"""One indexable unit of a document.
`chunk_id` is a deterministic UUIDv5 of `file_id` and `chunk_index`
(ADR-0001), so re-ingesting a file upserts its points rather than
duplicating them.
"""
chunk_id: uuid.UUID
chunk_index: int
order_id: float
content: str
content_type: ContentType
previous_chunk_id: uuid.UUID | None = None
next_chunk_id: uuid.UUID | None = None
token_count: int
character_count: int
class SparseVector(BaseModel):
"""A sparse (term-index -> weight) vector, Qdrant's `modifier="idf"` shape.
Kept free of the `qdrant_client` SDK (ADR-0015: ports carry no infra
imports) — `src/infrastructure/qdrant/` converts this to the SDK's own
`SparseVector` type at upsert time (Phase 5).
"""
indices: list[int]
values: list[float]
class EmbeddedChunk(BaseModel):
"""A chunk plus every vector it will be upserted with (ADR-0001).
`dense` is keyed by named-vector name (`dense_nomic`, `dense_openai`).
`late_interaction` is deliberately absent — not computed at ingest
(ADR-0017).
"""
chunk: Chunk
dense: dict[str, list[float]]
sparse: SparseVector

View File

@@ -0,0 +1,73 @@
"""Persian text normalization (ADR-0018).
Applied to every extracted text block before chunking, for all source formats.
The problem this solves is silent: Persian authored on mixed Arabic/Persian
keyboards contains both U+06A9 and U+0643 for "k", both U+06CC and U+064A for
"y". Those are distinct codepoints and therefore distinct tokens to every
embedding model, so the same word embeds two different ways depending on which
key the author pressed.
Letter folding only -- digits and punctuation are left as authored, because
chunk content is what citations render back to the reader and Western digits
inside Persian prose read as wrong.
"""
import re
import unicodedata
# Both tables are written as codepoints rather than character literals. Arabic
# letterforms are visually indistinguishable from one another (and alef from a
# Latin "l") in a monospace editor -- which is the very confusion this module
# exists to resolve -- and literals would render right-to-left, visually
# reordering the source line.
#
# `str.translate` accepts an ordinal->ordinal mapping directly, and an ordinal
# mapped to None is deleted.
_ARABIC_KAF = 0x0643
_ARABIC_YEH = 0x064A
_ALEF_MAKSURA = 0x0649
_ALEF_HAMZA_ABOVE = 0x0623
_ALEF_HAMZA_BELOW = 0x0625
_NOT_SIGN = 0x00AC
_PERSIAN_KEHEH = 0x06A9
_PERSIAN_YEH = 0x06CC
_ALEF = 0x0627
_SPACE = 0x0020
_LETTER_FOLDING: dict[int, int] = {
_ARABIC_KAF: _PERSIAN_KEHEH,
_ARABIC_YEH: _PERSIAN_YEH,
_ALEF_MAKSURA: _PERSIAN_YEH,
_ALEF_HAMZA_ABOVE: _ALEF,
_ALEF_HAMZA_BELOW: _ALEF,
# A soft-hyphen artifact from documents exported by older Word versions.
_NOT_SIGN: _SPACE,
}
_TATWEEL = 0x0640
_SUPERSCRIPT_ALEF = 0x0670
_HARAKAT = range(0x064B, 0x0660)
# Applied after NFKC, which can itself decompose presentation forms into a
# base letter plus a combining mark.
_MARK_REMOVAL: dict[int, int | None] = dict.fromkeys(_HARAKAT)
_MARK_REMOVAL[_TATWEEL] = None
_MARK_REMOVAL[_SUPERSCRIPT_ALEF] = None
_WHITESPACE = re.compile(r"\s+")
def normalize_persian_text(text: str) -> str:
"""Fold Arabic letterforms to Persian and collapse whitespace.
Call this per text block, **before** blocks are assembled into a document.
The whitespace collapse maps `\\n` to a space, so running it over assembled
markdown would flatten every heading and paragraph onto one line.
"""
text = text.translate(_LETTER_FOLDING)
text = unicodedata.normalize("NFKC", text)
text = text.translate(_MARK_REMOVAL)
return _WHITESPACE.sub(" ", text).strip()

View File

@@ -0,0 +1,65 @@
"""The one caller-facing entry point for parsing and chunking (ADR-0017).
`parse_docx`/`parse_csv`/`parse_xlsx`/`chunk_document` are blocking, pure
functions; calling any of them directly from an `async def` route or service
is the defect ADR-0017 names explicitly ("one large `python-docx` parse would
stall every concurrent request"). `parse_and_chunk_document` is the only
version of this pipeline callers should reach for: it owns source-type
dispatch and the `anyio.to_thread.run_sync` + `CapacityLimiter` offload, so
that obligation cannot be forgotten at a call site.
"""
import uuid
from functools import partial
from anyio import CapacityLimiter, to_thread
from src.application.ingestion.chunking import chunk_document
from src.application.ingestion.docx_parser import parse_docx
from src.application.ingestion.errors import UnsupportedSourceTypeError
from src.application.ingestion.models import Chunk, ParsedDocument
from src.application.ingestion.spreadsheet_parser import parse_csv, parse_xlsx
from src.config import ChunkingSettings
_PARSERS = {"csv", "xlsx", "docx"}
def _parse(data: bytes, source_type: str, settings: ChunkingSettings) -> ParsedDocument:
if source_type == "docx":
return parse_docx(data, settings)
if source_type == "xlsx":
return parse_xlsx(data)
if source_type == "csv":
return parse_csv(data)
raise UnsupportedSourceTypeError(f"'{source_type}' is not an ingestible source type")
def _parse_and_chunk(
data: bytes, source_type: str, file_id: uuid.UUID, settings: ChunkingSettings
) -> list[Chunk]:
parsed = _parse(data, source_type, settings)
return chunk_document(parsed, file_id=file_id, settings=settings)
async def parse_and_chunk_document(
data: bytes,
*,
source_type: str,
file_id: uuid.UUID,
settings: ChunkingSettings,
limiter: CapacityLimiter,
) -> list[Chunk]:
"""Parse and chunk a document off the event loop, bounded by `limiter`.
Raises `UnsupportedSourceTypeError` (415), `DocumentParseError` (400), or
`ChunkTooLargeError` — see `src/application/ingestion/errors.py`. Callers
map these to status codes; this module carries no HTTP knowledge
(ADR-0015). The `max_chunks_per_file` ceiling (413) is enforced by the
caller, not here — see Phase 4 of plan 001.
"""
if source_type not in _PARSERS:
raise UnsupportedSourceTypeError(f"'{source_type}' is not an ingestible source type")
return await to_thread.run_sync(
partial(_parse_and_chunk, data, source_type, file_id, settings),
limiter=limiter,
)

View File

@@ -0,0 +1,121 @@
"""CSV and XLSX parsing: row = chunk (ADR-0004, ADR-0018).
Both formats reduce to rows and hand them to the shared renderer in
`tabular`, so a Q&A sheet and a branch directory go through one code path with
no shape detection.
"""
import csv
import io
import openpyxl
from openpyxl.worksheet.worksheet import Worksheet
from src.application.ingestion.errors import DocumentParseError
from src.application.ingestion.models import ContentType, ParsedDocument, StructuralUnit
from src.application.ingestion.tabular import Row, clean_cell, clean_rows, render_rows
# Farsi exports from older Excel are frequently cp1256 (Windows Arabic).
_ENCODINGS = ("utf-8-sig", "utf-8", "cp1256")
_SNIFF_BYTES = 8192
def _decode(data: bytes) -> str:
for encoding in _ENCODINGS:
try:
return data.decode(encoding)
except UnicodeDecodeError:
continue
raise DocumentParseError(f"Could not decode file as any of: {', '.join(_ENCODINGS)}")
def _to_units(rendered: list[str]) -> list[StructuralUnit]:
return [StructuralUnit(text=text, content_type=ContentType.TABLE_ROW) for text in rendered]
def parse_csv(data: bytes) -> ParsedDocument:
"""Parse CSV bytes into one structural unit per row."""
text = _decode(data)
try:
dialect = csv.Sniffer().sniff(text[:_SNIFF_BYTES])
reader = csv.reader(io.StringIO(text), dialect)
except csv.Error:
# A single-column file has no delimiter to find; that is not an error.
reader = csv.reader(io.StringIO(text))
rows = clean_rows(reader)
if not rows:
raise DocumentParseError("File contains no rows")
units = _to_units(render_rows(rows))
if not units:
raise DocumentParseError("File contains no data rows")
return ParsedDocument(units=units, block_count=len(units))
def _forward_fill_merges(worksheet: Worksheet) -> dict[tuple[int, int], str]:
"""Map every cell of a merged range to the range's value.
openpyxl stores a merged range's value only in its top-left cell; the rest
read as None. Without this a branch row inherits nothing from the province
cell merged above it and silently loses that field (ADR-0004).
Iterate the range collection itself and read corners via `bounds`: `.ranges`
is a set subclass and `.min_row` and friends are descriptors, neither of
which resolves to an int for a type checker.
"""
filled: dict[tuple[int, int], str] = {}
for merged in worksheet.merged_cells:
min_col, min_row, max_col, max_row = merged.bounds
value = clean_cell(worksheet.cell(row=min_row, column=min_col).value)
if not value:
continue
for row in range(min_row, max_row + 1):
for column in range(min_col, max_col + 1):
filled[(row, column)] = value
return filled
def _sheet_rows(worksheet: Worksheet) -> list[Row]:
"""Read a worksheet into normalized text rows, merges resolved."""
merged = _forward_fill_merges(worksheet)
rows: list[Row] = []
for row_index, row in enumerate(worksheet.iter_rows(), start=1):
values = [
merged.get((row_index, column_index), clean_cell(cell.value))
for column_index, cell in enumerate(row, start=1)
]
if any(values):
rows.append(values)
return rows
def parse_xlsx(data: bytes) -> ParsedDocument:
"""Parse XLSX bytes into one structural unit per row, across all sheets."""
try:
workbook = openpyxl.load_workbook(io.BytesIO(data), data_only=True)
except Exception as exc:
raise DocumentParseError(f"Could not open XLSX: {exc}") from exc
units: list[StructuralUnit] = []
try:
for worksheet in workbook.worksheets:
rows = _sheet_rows(worksheet)
# The dead second sheet seen throughout the sample corpus.
if not rows:
continue
units.extend(_to_units(render_rows(rows)))
finally:
workbook.close()
if not units:
raise DocumentParseError("Workbook contains no data rows")
return ParsedDocument(units=units, block_count=len(units))

View File

@@ -0,0 +1,167 @@
"""Shared row-to-chunk rendering for every tabular source (ADR-0004, ADR-0018).
A table row is an atomic structural unit regardless of the container it
arrived in -- a docx table, an xlsx sheet, or a csv file -- so one renderer
serves all three and a Q&A sheet needs no special case against a branch
directory:
q: ... ردیف: 1
a: ... استان: اردبیل
شعبه: پارس آباد
The header rules exist because real tables are not uniform. Of the tables in
the sample corpus, some carry a header row, one is a bare list of values with
no header at all, and one is page decoration. The guiding constraint is
therefore: **never invent structure that is not provably there, and never
discard a row.** A wrongly-detected header turns every chunk into nonsense
(`80: 70`), which is worse than an unlabeled row.
"""
from collections.abc import Iterable, Sequence
from src.application.ingestion.normalization import normalize_persian_text
Row = Sequence[str]
# A header cell is a label, not a sentence.
MAX_HEADER_CELL_LENGTH = 80
# How much shorter a header cell must be than the column beneath it before
# length alone is taken as evidence of a header (the `q`/`a` case, where
# one-character labels sit above paragraph-long answers).
_HEADER_LENGTH_RATIO = 3.0
def clean_cell(value: object) -> str:
"""Normalize a cell to text; None and blank cells become the empty string.
`object` rather than a union: a spreadsheet cell holds whatever the
workbook stored -- str, int, float, bool, datetime, a formula error -- and
every one of them is handled the same way, by rendering it.
"""
if value is None:
return ""
return normalize_persian_text(str(value))
def _is_numeric(text: str) -> bool:
return bool(text) and text.replace(",", "").replace(".", "").replace("-", "").isdigit()
def _looks_like_title_row(row: Row) -> bool:
"""A merged title spanning the sheet resolves to one value, or repeats it."""
populated = [cell for cell in row if cell]
return len(populated) < 2 or len(set(populated)) == 1
def has_header(rows: Sequence[Row]) -> bool:
"""Whether the first row labels the columns beneath it.
Decided by comparing row 0 against the column below it, not by how row 0
looks on its own: a label row is *inconsistent* with its data (text above
numbers, or a short label above long prose), while a data row is
consistent with the rows that follow. This is the test `csv.Sniffer`
uses, and it is a property of the table rather than a pattern borrowed
from one document.
"""
if len(rows) < 2:
return False
candidate, data = rows[0], rows[1:]
if not all(cell for cell in candidate[: len(data[0])] if cell) and _looks_like_title_row(
candidate
):
return False
if any(len(cell) > MAX_HEADER_CELL_LENGTH for cell in candidate):
return False
for index, label in enumerate(candidate):
if not label:
continue
column = [row[index] for row in data if index < len(row) and row[index]]
if not column:
continue
# A text label above a numeric column.
if not _is_numeric(label) and all(_is_numeric(value) for value in column):
return True
# A short label above a column of much longer values.
mean_length = sum(len(value) for value in column) / len(column)
if mean_length > len(label) * _HEADER_LENGTH_RATIO:
return True
return False
def strip_title_rows(rows: Sequence[Row]) -> Sequence[Row]:
"""Drop leading merged-title and blank rows.
Structural, not a content guess: these rows are the artifact of a merged
range spanning the sheet width, so they carry one value across many cells.
Only *leading* rows are dropped, so no data row is ever lost.
"""
start = 0
while start < len(rows) and _looks_like_title_row(rows[start]):
start += 1
return rows[start:]
def _column_name(header: Row, index: int) -> str:
"""Return a header label, falling back positionally past the header width."""
if index < len(header) and header[index]:
return header[index]
return f"column_{index + 1}"
def _dedupe_horizontal_merge(row: Row) -> list[str]:
"""Collapse the repeats a horizontally merged cell produces.
Both python-docx and openpyxl report a merged cell once per grid column it
spans, so an unlabeled row would otherwise repeat the same value.
"""
collapsed: list[str] = []
for cell in row:
if cell and (not collapsed or collapsed[-1] != cell):
collapsed.append(cell)
return collapsed
def render_rows(rows: Sequence[Row]) -> list[str]:
"""Render table rows as text, one string per row.
With a provable header each row becomes `"{header}: {value}"` lines, which
makes it self-describing. Without one, cells are joined with `" | "` --
unlabeled, but never mislabeled.
"""
rows = strip_title_rows(rows)
if not rows:
return []
if not has_header(rows):
return [text for row in rows if (text := " | ".join(_dedupe_horizontal_merge(row)))]
header, data = rows[0], rows[1:]
# A header merged vertically across two rows resolves to the same text in
# the row below it; that duplicate is the header, not data.
while data and list(data[0]) == list(header):
data = data[1:]
rendered: list[str] = []
for row in data:
lines = [
f"{_column_name(header, index)}: {value}" for index, value in enumerate(row) if value
]
if lines:
rendered.append("\n".join(lines))
return rendered
def clean_rows(rows: Iterable[Iterable[object]]) -> list[Row]:
"""Normalize every cell and drop rows that are entirely empty."""
cleaned: list[Row] = []
for row in rows:
values = [clean_cell(cell) for cell in row]
if any(values):
cleaned.append(values)
return cleaned

View File

@@ -0,0 +1,33 @@
"""Token counting for chunk sizing (ADR-0018).
`cl100k_base` is a deliberate proxy for the embedding models' own tokenizers.
`text-embedding-3-large` has an 8191-token window and never binds;
`nomic-embed-text-v2-moe`'s 512-token sequence length is the only real
constraint. cl100k tokenizes Persian inefficiently while nomic's multilingual
tokenizer does not, so a cl100k count reliably over-estimates the nomic count --
safe in the conservative direction, without shipping a second tokenizer and its
model download into the ingestion path.
"""
from functools import lru_cache
import tiktoken
@lru_cache(maxsize=4)
def get_encoder(encoding_name: str) -> tiktoken.Encoding:
"""Return a cached tiktoken encoder.
Deliberately not a module-level constant: `tiktoken` fetches the BPE
vocabulary over the network the first time an encoding is used, and
ADR-0012 forbids external resource setup as an import-time side effect.
The lifespan warms this at startup so a process fails fast at boot rather
than inside the first ingestion request. Set `TIKTOKEN_CACHE_DIR` to a
pre-populated directory for offline deployments.
"""
return tiktoken.get_encoding(encoding_name)
def count_tokens(text: str, encoding_name: str) -> int:
"""Return the number of tokens `text` encodes to."""
return len(get_encoder(encoding_name).encode(text))

View File

@@ -0,0 +1,18 @@
"""Ingestion-generated Qdrant point CRUD (ADR-0001, ADR-0002).
`index_chunks` is the entry point callers outside this package should use: it
dispatches payload construction, batching, bounded-concurrency upserts, and the
post-success soft-delete sweep. `build_chunk_payload` and the batching helpers
stay internal, exported mainly for their own unit tests.
The `/v1/points` surface lives here too, in its own modules with their own
entry points: `queries.py` for the read paths and `deletion.py` for soft delete
with neighbour relinking. They share this package because they share ADR-0001's
payload schema, not because they share a caller — `index_chunks` writes a whole
file at once, while those serve one admin edit at a time.
"""
from src.application.points.indexing import IndexingResult, index_chunks
from src.application.points.models import ChunkPoint
__all__ = ["ChunkPoint", "IndexingResult", "index_chunks"]

View File

@@ -0,0 +1,253 @@
"""Soft delete for `/v1/points` and for a whole file's points (ADR-0002).
The caller-facing entry points are `soft_delete_point` and
`soft_delete_file_points`. Routers call these; `patches_for_removal` and the
planning helpers stay internal, because getting a delete right is exactly the
composition a caller should not have to reassemble: read the point, load its
neighbours, compute the patches still missing, send them in **one** batch,
verify they landed, and retry against fresh versions if they did not.
Why the retry exists. Qdrant has no multi-point transaction, so a batch whose
second operation loses a version race applies its first operation anyway — and
a filtered `set_payload` that matched nothing still reports success. That
combination means "did my write land?" is only answerable by reading back, and
a single-shot delete would be able to leave the deactivation applied and a
neighbour's pointer stale. Since `patches_for_removal` plans from current state
towards a fixed end state, simply re-planning emits precisely the patches that
did not land, so the loop converges instead of re-doing work. Only exhausting
the attempts raises `PointVersionConflictError` (`409`).
Deleting an already-inactive point falls out of the same machinery rather than
needing a special case: its neighbours were relinked by the first delete, so the
plan is empty and the call is a no-op success — not a `404`, and not a second
relink.
"""
import uuid
from collections.abc import Mapping, Sequence
from datetime import UTC, datetime
from time import perf_counter
import structlog
from src.application.points.errors import PointVersionConflictError
from src.application.points.point import Point, PointNotFoundError
from src.application.points.relinking import (
neighbour_ids,
patch_for_deactivation,
patches_for_removal,
)
from src.application.ports.point_repository import PayloadPatch, PointRepository
logger = structlog.get_logger(__name__)
# Three plan-apply rounds, then a final verifying plan. Each round only re-emits
# what a concurrent writer displaced, so a caller that legitimately needs more
# than this is contending on the same points continuously and deserves the
# `409` rather than an unbounded loop inside a request.
_MAX_ATTEMPTS = 3
# One sweep page. Matches ADR-0002's 100-operation batch cap, so a page of
# points is always expressible as a single `points/batch` request.
_SWEEP_BATCH_SIZE = 100
# A hard ceiling on sweep rounds, so a file being concurrently re-ingested while
# it is deleted cannot spin here for the life of the request.
_MAX_SWEEP_ROUNDS = 1_000
def _elapsed_ms(started: float) -> float:
"""Wall-clock milliseconds since `started` (ADR-0011's `duration_ms`).
Worth carrying on these events even though the relinking itself is O(1):
what a delete actually spends is Qdrant round trips, and the whole-file
sweep spends a number of them proportional to the file's length. Timing the
operation is the only way to tell a slow store from a contended one, which
`rounds` on the same event then disambiguates.
"""
return round((perf_counter() - started) * 1000, 2)
async def soft_delete_point(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
point_id: uuid.UUID,
actor: str,
) -> Point:
"""Deactivate one point and relink its neighbours around the gap.
Returns the point as it now stands. Raises `PointNotFoundError` (`404`) if
it is not this tenant's — the same non-disclosure rule the read paths
follow — or `PointVersionConflictError` (`409`) if concurrent writers keep
displacing the plan.
"""
started = perf_counter()
rounds = 0
point, patches = await _plan_removal(
repository, tenant_id=tenant_id, point_id=point_id, actor=actor
)
for _ in range(_MAX_ATTEMPTS):
if not patches:
break
await repository.apply_patches(tenant_id=tenant_id, patches=patches)
rounds += 1
# The next plan doubles as verification: anything that did not land is
# still missing from the end state and comes back as a patch.
point, patches = await _plan_removal(
repository, tenant_id=tenant_id, point_id=point_id, actor=actor
)
if patches:
logger.warning(
"points.soft_delete.conflict",
tenant_id=str(tenant_id),
point_id=str(point_id),
file_id=str(point.file_id),
unsettled_points=[str(patch.point_id) for patch in patches],
rounds=rounds,
duration_ms=_elapsed_ms(started),
)
raise PointVersionConflictError(
f"point {point_id} could not be soft-deleted under concurrent modification"
)
if rounds:
logger.info(
"points.soft_deleted",
tenant_id=str(tenant_id),
point_id=str(point_id),
file_id=str(point.file_id),
version=point.version,
actor=actor,
# `rounds` is 1 unless a concurrent writer forced a re-plan, so a
# rising value here is contention, not slow relinking.
rounds=rounds,
duration_ms=_elapsed_ms(started),
)
return point
async def soft_delete_file_points(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
file_id: uuid.UUID,
actor: str,
) -> int:
"""Deactivate every active point of one file, in batches.
No relinking: the whole file leaves the sequence at once, so no surviving
active point can be left pointing at a deactivated one, and the chain is
preserved intact for whoever reads the deleted file later.
Returns how many points were active when the sweep reached them. Each round
re-lists from the start rather than paging with a cursor — deactivated
points drop straight out of the default listing, so the listing itself is
the progress check, and a round that attempts the exact same ids as the one
before it made no progress and raises `PointVersionConflictError`.
"""
started = perf_counter()
swept: set[uuid.UUID] = set()
previous_attempt: frozenset[uuid.UUID] = frozenset()
rounds = 0
for _ in range(_MAX_SWEEP_ROUNDS):
page = await repository.list_by_file(
tenant_id=tenant_id, file_id=file_id, limit=_SWEEP_BATCH_SIZE
)
if not page.points:
if swept:
logger.info(
"points.file_soft_deleted",
tenant_id=str(tenant_id),
file_id=str(file_id),
points_soft_deleted=len(swept),
actor=actor,
# Two Qdrant round trips per round, so this is the delete
# path whose cost tracks the size of the file.
rounds=rounds,
duration_ms=_elapsed_ms(started),
)
return len(swept)
attempt = frozenset(point.point_id for point in page.points)
if attempt == previous_attempt:
logger.warning(
"points.file_soft_delete.conflict",
tenant_id=str(tenant_id),
file_id=str(file_id),
unsettled_points=[str(point_id) for point_id in sorted(attempt, key=str)],
rounds=rounds,
duration_ms=_elapsed_ms(started),
)
raise PointVersionConflictError(
f"file {file_id} could not be soft-deleted under concurrent modification"
)
previous_attempt = attempt
now = datetime.now(UTC)
await repository.apply_patches(
tenant_id=tenant_id,
patches=[patch_for_deactivation(point, actor=actor, now=now) for point in page.points],
)
swept |= attempt
rounds += 1
logger.warning(
"points.file_soft_delete.conflict",
tenant_id=str(tenant_id),
file_id=str(file_id),
reason="sweep_rounds_exhausted",
rounds=rounds,
duration_ms=_elapsed_ms(started),
)
raise PointVersionConflictError(f"file {file_id} still had active points after the sweep")
async def _plan_removal(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
point_id: uuid.UUID,
actor: str,
) -> tuple[Point, tuple[PayloadPatch, ...]]:
point = await repository.get(tenant_id=tenant_id, point_id=point_id)
if point is None:
raise PointNotFoundError(f"point {point_id} not found")
neighbours = await _load_neighbours(repository, tenant_id=tenant_id, point=point)
_warn_on_missing_neighbours(point, neighbours, tenant_id=tenant_id)
patches = patches_for_removal(point, neighbours, actor=actor, now=datetime.now(UTC))
return point, patches
async def _load_neighbours(
repository: PointRepository, *, tenant_id: uuid.UUID, point: Point
) -> dict[uuid.UUID, Point]:
wanted: Sequence[uuid.UUID] = neighbour_ids(point)
if not wanted:
return {}
found = await repository.get_many(tenant_id=tenant_id, point_ids=wanted)
return {neighbour.point_id: neighbour for neighbour in found}
def _warn_on_missing_neighbours(
point: Point, neighbours: Mapping[uuid.UUID, Point], *, tenant_id: uuid.UUID
) -> None:
"""A pointer naming a point that is not there means the chain is already broken.
Worth a log line rather than an exception: the delete can still complete the
part of the relink that does exist, and refusing would leave the caller with
a point it cannot remove through any endpoint.
"""
missing = [pointer for pointer in neighbour_ids(point) if pointer not in neighbours]
if missing:
logger.warning(
"points.relink.neighbour_missing",
tenant_id=str(tenant_id),
point_id=str(point.point_id),
file_id=str(point.file_id),
missing_neighbours=[str(pointer) for pointer in missing],
)

View File

@@ -0,0 +1,19 @@
"""Mutation failures for the `/v1/points` write paths (ADR-0002).
No HTTP knowledge here — `src/api/errors.py` owns the status mapping. Absence
lives on `PointNotFoundError` in `point.py`, next to the model whose read paths
raise it; this module is for the failures only a *write* can produce.
"""
class PointVersionConflictError(Exception):
"""A version-guarded write could not be landed against a moving target.
Raised when the service has re-read, recomputed, and re-applied its patches
the allowed number of times and the desired state still has not settled —
something else is writing the same points concurrently. Maps to `409`.
This is not "the guard fired once": a single stale guard is expected and is
retried, because Qdrant reports success for a filtered `set_payload` that
matched nothing. It means the retries were exhausted.
"""

View File

@@ -0,0 +1,203 @@
"""The one caller-facing entry point for indexing embedded chunks (ADR-0001, ADR-0017).
`index_chunks` is the only version of this step callers should reach for. It
owns the whole composition a correct upsert needs:
- building ADR-0001's payload for every chunk, with `tenant_id`/`domain` taken
from server-derived context;
- offloading that (and the per-chunk content hashing) to a thread, since it is
blocking CPU work (ADR-0017);
- batching at `QDRANT_UPSERT_BATCH_SIZE` inside ADR-0001's 64-256 band;
- bounding in-flight batches with an `asyncio.Semaphore` rather than an
unbounded `gather` (ADR-0017);
- running the soft-delete sweep for a shortened file **only after every batch
has succeeded**.
That last ordering is the point, not an implementation detail — see
`_deactivate_stale` below. `build_chunk_payload` and `_batches` stay internal;
pushing that composition onto every call site is exactly the obligation a deep
module absorbs once (CLAUDE.md, "prefer deep modules").
"""
import asyncio
import uuid
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import UTC, datetime
from anyio import CapacityLimiter, to_thread
from src.application.ingestion.errors import PointIndexingError
from src.application.ingestion.models import EmbeddedChunk
from src.application.points.models import ChunkPoint
from src.application.points.payload import build_chunk_payload
from src.application.ports.embedding import DenseEmbedder, SparseEmbedder
from src.application.ports.point_storage import PointStorage
from src.config import QdrantSettings
@dataclass(frozen=True)
class IndexingResult:
"""What one indexing pass wrote.
`points_upserted` counts points written, not points *created* — a
deterministic-id upsert cannot distinguish an insert from an overwrite, so
the ingestion job reports this as `points_created` and leaves
`points_updated` at zero rather than guessing.
"""
points_upserted: int
points_soft_deleted: int
def _embedding_model_version(
dense_embedders: Sequence[DenseEmbedder], sparse_embedder: SparseEmbedder
) -> str:
"""Compose the `embedding_model_version` payload value (ADR-0001).
Sorted so the string is stable regardless of the order the embedders were
wired in — an unstable value would make "which chunks need re-embedding?"
unanswerable, which is the field's only reason to exist.
"""
versions = sorted(
[embedder.model_version for embedder in dense_embedders] + [sparse_embedder.model_version]
)
return "+".join(versions)
def _batches(points: Sequence[ChunkPoint], size: int) -> list[Sequence[ChunkPoint]]:
return [points[i : i + size] for i in range(0, len(points), size)]
def _build_points(
embedded: Sequence[EmbeddedChunk],
*,
tenant_id: uuid.UUID,
domain: str,
file_id: uuid.UUID,
source_filename: str,
source_type: str,
actor: str,
embedding_model_version: str,
indexed_at: datetime,
) -> list[ChunkPoint]:
"""Blocking: hashes every chunk's content. Always called through a thread."""
return [
ChunkPoint(
point_id=item.chunk.chunk_id,
dense=item.dense,
sparse=item.sparse,
payload=build_chunk_payload(
item.chunk,
tenant_id=tenant_id,
domain=domain,
file_id=file_id,
source_filename=source_filename,
source_type=source_type,
actor=actor,
embedding_model_version=embedding_model_version,
indexed_at=indexed_at,
),
)
for item in embedded
]
async def _upsert_bounded(
storage: PointStorage, batch: Sequence[ChunkPoint], *, semaphore: asyncio.Semaphore
) -> None:
async with semaphore:
try:
await storage.upsert_points(batch)
except Exception as exc:
raise PointIndexingError(f"upserting {len(batch)} points failed: {exc}") from exc
async def _deactivate_stale(
storage: PointStorage,
*,
tenant_id: uuid.UUID,
file_id: uuid.UUID,
from_chunk_index: int,
actor: str,
deleted_at: datetime,
) -> int:
"""Soft-delete points left over from a longer previous version of this file.
Chunk indices are contiguous from 0, so "index >= the new chunk count" is
exactly the set of points the new version no longer produces.
This runs **only after every upsert has succeeded**, and that ordering is
what keeps a failed attempt from damaging a working index. ADR-0001's
deterministic point ids mean a re-ingestion overwrites in place, so literal
atomic replacement is not available; what *is* guaranteed is that a failed
attempt never removes content (it can only leave a prefix updated), and that
a retry converges to the correct state. See ADR-0017.
"""
try:
return await storage.deactivate_points_from_index(
tenant_id=tenant_id,
file_id=file_id,
from_chunk_index=from_chunk_index,
deleted_at=deleted_at,
updated_by=actor,
)
except Exception as exc:
raise PointIndexingError(f"soft-deleting stale points failed: {exc}") from exc
async def index_chunks(
embedded: Sequence[EmbeddedChunk],
*,
storage: PointStorage,
tenant_id: uuid.UUID,
domain: str,
file_id: uuid.UUID,
source_filename: str,
source_type: str,
actor: str,
dense_embedders: Sequence[DenseEmbedder],
sparse_embedder: SparseEmbedder,
settings: QdrantSettings,
thread_limiter: CapacityLimiter,
) -> IndexingResult:
"""Upsert every embedded chunk as a tenant-scoped point, then sweep leftovers.
Raises `PointIndexingError` (502) if any batch or the sweep fails.
"""
if not embedded:
return IndexingResult(points_upserted=0, points_soft_deleted=0)
indexed_at = datetime.now(UTC)
points = await to_thread.run_sync(
lambda: _build_points(
embedded,
tenant_id=tenant_id,
domain=domain,
file_id=file_id,
source_filename=source_filename,
source_type=source_type,
actor=actor,
embedding_model_version=_embedding_model_version(dense_embedders, sparse_embedder),
indexed_at=indexed_at,
),
limiter=thread_limiter,
)
semaphore = asyncio.Semaphore(settings.upsert_concurrency)
await asyncio.gather(
*(
_upsert_bounded(storage, batch, semaphore=semaphore)
for batch in _batches(points, settings.upsert_batch_size)
)
)
soft_deleted = await _deactivate_stale(
storage,
tenant_id=tenant_id,
file_id=file_id,
from_chunk_index=len(points),
actor=actor,
deleted_at=indexed_at,
)
return IndexingResult(points_upserted=len(points), points_soft_deleted=soft_deleted)

View File

@@ -0,0 +1,29 @@
"""Domain models for Qdrant points (ADR-0001).
Deliberately free of the `qdrant_client` SDK: `src/infrastructure/qdrant/`
converts these to `PointStruct`/`models.SparseVector` at upsert time
(ADR-0015 — application code and ports carry no infrastructure imports).
"""
import uuid
from pydantic import BaseModel, Field
from src.application.ingestion.models import SparseVector
class ChunkPoint(BaseModel):
"""One chunk, ready to upsert: its id, its named vectors, and its payload.
`point_id` is the chunk's deterministic UUIDv5 (`chunk_id_for`), so
re-ingesting a file overwrites its points rather than duplicating them
(ADR-0001).
`dense` is keyed by named-vector name (`dense_nomic`, `dense_openai`).
`late_interaction` is absent — ADR-0017 does not compute it at ingest.
"""
point_id: uuid.UUID
dense: dict[str, list[float]]
sparse: SparseVector
payload: dict[str, object] = Field(default_factory=dict)

View File

@@ -0,0 +1,70 @@
"""Builds ADR-0001's point payload from a chunk plus its ingestion context.
Internal to `src/application/points/` — callers use `index_chunks`, which owns
composing this with batching and the deactivation sweep. Exported for its own
unit tests, not as a surface to build payloads by hand.
"""
import uuid
from datetime import datetime
from hashlib import sha256
from src.application.ingestion.models import Chunk
def _optional_id(value: uuid.UUID | None) -> str | None:
return str(value) if value is not None else None
def build_chunk_payload(
chunk: Chunk,
*,
tenant_id: uuid.UUID,
domain: str,
file_id: uuid.UUID,
source_filename: str,
source_type: str,
actor: str,
embedding_model_version: str,
indexed_at: datetime,
) -> dict[str, object]:
"""Return ADR-0001's payload for one chunk.
`tenant_id` and `domain` are passed in from the server-derived `AuthContext`
and the validated request — never from anything the client could assert as
authority (ADR-0002's non-negotiable isolation rule).
UUIDs are serialized as strings because the `tenant_id`/`domain`/`file_id`/
`previous_chunk_id`/`next_chunk_id` payload indexes are *keyword* indexes;
a native UUID would not match a keyword filter.
**Known gap — `version` is always written as `1`.** ADR-0002 uses this field
for optimistic concurrency between ingestion and manual `/v1/points` edits,
which needs a read-check-write (one read per point). Ingestion is
authoritative for its own file today, so writing `1` is safe until
`/v1/points` exists; plan 002 owns closing this.
"""
timestamp = indexed_at.isoformat()
return {
"tenant_id": str(tenant_id),
"domain": domain,
"file_id": str(file_id),
"chunk_id": str(chunk.chunk_id),
"content": chunk.content,
"content_type": chunk.content_type.value,
"source_filename": source_filename,
"source_type": source_type,
"order_id": chunk.order_id,
"chunk_index": chunk.chunk_index,
"previous_chunk_id": _optional_id(chunk.previous_chunk_id),
"next_chunk_id": _optional_id(chunk.next_chunk_id),
"is_active": True,
"deleted_at": None,
"created_at": timestamp,
"updated_at": timestamp,
"created_by": actor,
"updated_by": actor,
"version": 1,
"content_hash": sha256(chunk.content.encode("utf-8")).hexdigest(),
"embedding_model_version": embedding_model_version,
}

View File

@@ -0,0 +1,126 @@
"""The caller-facing point model for `/v1/points` (ADR-0001, ADR-0002).
`ChunkPoint` in `models.py` is the *write* shape ingestion upserts: an id, its
named vectors, and an opaque payload dict. This module is the *read/edit* shape
the `/v1/points` surface works in, where the payload's individual fields matter
and the distinction between what a caller may write and what the server owns is
a security boundary rather than a convention.
That split is the reason this is a model and not a dict. ADR-0002's isolation
rule ("never accepted as client-supplied input") and its optimistic-concurrency
guard both fail open if a caller can smuggle `tenant_id` or `version` through a
payload update, so the writable field set is enumerated in one place here and
every write path validates against it.
"""
import uuid
from collections.abc import Mapping
from datetime import datetime
from typing import Self
from pydantic import BaseModel, ConfigDict, Field
# Fields the server derives and a caller may never set, patch, or override.
# `tenant_id` is authority, `version` is the concurrency guard, `chunk_index`
# derives the point id, and the rest are provenance the server timestamps.
SERVER_OWNED_FIELDS: frozenset[str] = frozenset(
{
"tenant_id",
"version",
"chunk_index",
"chunk_id",
"is_active",
"deleted_at",
"created_at",
"created_by",
"updated_at",
"updated_by",
"content_hash",
"embedding_model_version",
}
)
# Fields a caller may supply on create, replace, or payload patch. `order_id`
# is writable on create but moves only through `PATCH /v1/points/{id}/order`
# afterwards, because a bare `order_id` write would not relink neighbours.
CALLER_WRITABLE_FIELDS: frozenset[str] = frozenset(
{
"content",
"content_type",
"domain",
"source_filename",
"source_type",
"order_id",
"previous_chunk_id",
"next_chunk_id",
}
)
class Point(BaseModel):
"""One Qdrant point, read back with its ADR-0001 payload fields typed.
Vectors are deliberately absent: ADR-0008 returns them only when explicitly
requested, and every read path that does not ask for them should not pay to
deserialize them.
"""
model_config = ConfigDict(frozen=True)
point_id: uuid.UUID
tenant_id: uuid.UUID
domain: str
file_id: uuid.UUID
chunk_id: uuid.UUID
content: str
content_type: str
source_filename: str
source_type: str
order_id: float
chunk_index: int
previous_chunk_id: uuid.UUID | None = None
next_chunk_id: uuid.UUID | None = None
is_active: bool = True
deleted_at: datetime | None = None
created_at: datetime
updated_at: datetime
created_by: str
updated_by: str
version: int
content_hash: str
embedding_model_version: str
# Only populated when the caller explicitly asked for vectors.
vectors: dict[str, object] | None = Field(default=None)
@classmethod
def from_payload(
cls,
point_id: uuid.UUID,
payload: Mapping[str, object],
*,
vectors: Mapping[str, object] | None = None,
) -> Self:
"""Build a `Point` from a raw Qdrant payload dict.
Lives here rather than in the Qdrant adapter so the payload field names
are declared once, next to the model that mirrors them. The adapter
stays responsible for talking to the SDK, not for knowing ADR-0001's
schema twice.
"""
return cls.model_validate({**payload, "point_id": point_id, "vectors": vectors})
class PointNotFoundError(LookupError):
"""No such point *within the requesting tenant*.
Routes map this to `404`, never `403` — a caller must not be able to probe
for the existence of another tenant's point ids (ADR-0016). The error
deliberately carries no hint about which of the two cases occurred.
"""

View File

@@ -0,0 +1,115 @@
"""Read paths for `/v1/points` (ADR-0002, ADR-0008).
The caller-facing entry points for point reads. Routers call these; they never
touch the `PointRepository` directly, and never build a filter.
This module is thin on purpose but not empty, and the two things it does own are
exactly the ones a route would otherwise get wrong:
- **`tenant_id` always comes from the caller's `AuthContext`.** Every function
takes it as a required keyword and hands it to the repository. Nothing here
reads a tenant from a query string or body.
- **A keyword query is Persian-normalized before it reaches the index.**
Ingestion letter-folds chunk content (`normalize_persian_text`, ADR-0018), so
stored text contains Persian yeh/keheh. A query typed on an Arabic keyboard
carries U+064A/U+0643 and would match nothing at all — a silent empty result,
not an error. Folding the query the same way is what makes the two comparable.
Reads emit no log events. The request middleware already records every call, and
ADR-0011 reserves `INFO` for lifecycle events rather than per-read volume;
mutations get their own events when those paths land.
"""
import uuid
from src.application.ingestion.normalization import normalize_persian_text
from src.application.points.point import Point, PointNotFoundError
from src.application.ports.point_repository import PointPage, PointRepository
async def get_point(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
point_id: uuid.UUID,
with_vectors: bool = False,
) -> Point:
"""One point, or `PointNotFoundError` if it is not this tenant's.
Raises rather than returning `None` so a route cannot forget the check and
serve `200 null`. "Absent" and "another tenant's" are the same outcome by
design (ADR-0016: cross-tenant access is `404`, never `403`).
"""
point = await repository.get(tenant_id=tenant_id, point_id=point_id, with_vectors=with_vectors)
if point is None:
raise PointNotFoundError(f"point {point_id} not found")
return point
async def list_file_points(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
file_id: uuid.UUID,
limit: int,
cursor: str | None = None,
include_inactive: bool = False,
) -> PointPage:
"""One file's points in display (`order_id`) order.
An unknown or foreign `file_id` yields an empty page rather than an error:
the two are indistinguishable to the caller, which is the same
non-disclosure property `get_point` gets from raising.
"""
return await repository.list_by_file(
tenant_id=tenant_id,
file_id=file_id,
limit=limit,
cursor=cursor,
include_inactive=include_inactive,
)
async def count_points(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
domain: str | None = None,
file_id: uuid.UUID | None = None,
include_inactive: bool = False,
) -> int:
return await repository.count(
tenant_id=tenant_id,
domain=domain,
file_id=file_id,
include_inactive=include_inactive,
)
async def search_points(
repository: PointRepository,
*,
tenant_id: uuid.UUID,
query: str,
limit: int,
cursor: str | None = None,
domain: str | None = None,
file_id: uuid.UUID | None = None,
include_inactive: bool = False,
) -> PointPage:
"""Keyword match on `content`, within this tenant.
**Not semantic retrieval.** Qdrant's full-text index filters rather than
scores, so results carry no relevance ranking and their order is
unspecified. Ranked retrieval is ADR-0003's hybrid path in plan 003; this
function must not grow a semantic mode (ADR-0002).
"""
return await repository.keyword_search(
tenant_id=tenant_id,
query=normalize_persian_text(query),
limit=limit,
cursor=cursor,
domain=domain,
file_id=file_id,
include_inactive=include_inactive,
)

View File

@@ -0,0 +1,135 @@
"""Adjacency-pointer maintenance for a point leaving a file's sequence.
ADR-0001 keeps `previous_chunk_id`/`next_chunk_id` on every point so ADR-0003's
context-window expansion can walk a file in O(1) steps. ADR-0002 makes keeping
them correct an obligation of every operation that changes a point's position:
a partial relink is a defect, not a degraded-but-acceptable outcome.
The function below is the primitive that obligation reduces to. It is pure, and
it is written as **"what is still missing between the state I just read and the
state I want"** rather than "the patches a delete implies". That framing is what
makes the caller's retry loop correct: re-planning after a partial apply emits
exactly the patches that did not land, and re-planning after a completed delete
emits nothing at all. The three cases the plan calls out — a normal delete, a
second delete of an already-inactive point, and recovery from a half-applied
batch — are then one code path instead of three.
Note what is deliberately *not* patched: the departing point's own
`previous_chunk_id`/`next_chunk_id`. Nothing active points at it once its
neighbours are relinked, so those pointers are unreachable rather than stale,
and leaving them records where the point sat — which is what a later restore or
an audit reader would need. `src/application/points/deletion.py` relies on that
when it re-plans: the departing point's pointers are the only surviving record
of which two neighbours have to be joined.
"""
import uuid
from collections.abc import Mapping
from datetime import datetime
from src.application.points.point import Point
from src.application.ports.point_repository import PayloadPatch
def _optional_id(value: uuid.UUID | None) -> str | None:
return str(value) if value is not None else None
def _provenance(point: Point, *, actor: str, now: datetime) -> dict[str, object]:
"""The fields every mutation writes: who, when, and the next version.
Bumping `version` on a relinked *neighbour* is intentional. The neighbour's
payload really did change, so a concurrent editor holding the old version
must get a `409` rather than overwrite the pointer we just fixed.
"""
return {
"updated_at": now.isoformat(),
"updated_by": actor,
"version": point.version + 1,
}
def neighbour_ids(point: Point) -> tuple[uuid.UUID, ...]:
"""The ids `patches_for_removal` needs loaded, skipping the nulls."""
return tuple(
pointer for pointer in (point.previous_chunk_id, point.next_chunk_id) if pointer is not None
)
def patches_for_removal(
point: Point,
neighbours: Mapping[uuid.UUID, Point],
*,
actor: str,
now: datetime,
) -> tuple[PayloadPatch, ...]:
"""The patches still needed to remove `point` from its file's sequence.
Returns an empty tuple when the removal is already complete, which the
caller reads as both "converged" and "this was a no-op".
A neighbour absent from `neighbours` is skipped rather than patched blind:
its id came from the departing point's payload, so a missing one means the
chain was already broken, and inventing a patch for a point that is not
there would not fix it. The caller logs that case.
"""
patches: list[PayloadPatch] = []
if point.is_active:
patches.append(
PayloadPatch(
point_id=point.point_id,
payload={
"is_active": False,
"deleted_at": now.isoformat(),
**_provenance(point, actor=actor, now=now),
},
expected_version=point.version,
)
)
previous = neighbours.get(point.previous_chunk_id) if point.previous_chunk_id else None
if previous is not None and previous.next_chunk_id != point.next_chunk_id:
patches.append(
PayloadPatch(
point_id=previous.point_id,
payload={
"next_chunk_id": _optional_id(point.next_chunk_id),
**_provenance(previous, actor=actor, now=now),
},
expected_version=previous.version,
)
)
following = neighbours.get(point.next_chunk_id) if point.next_chunk_id else None
if following is not None and following.previous_chunk_id != point.previous_chunk_id:
patches.append(
PayloadPatch(
point_id=following.point_id,
payload={
"previous_chunk_id": _optional_id(point.previous_chunk_id),
**_provenance(following, actor=actor, now=now),
},
expected_version=following.version,
)
)
return tuple(patches)
def patch_for_deactivation(point: Point, *, actor: str, now: datetime) -> PayloadPatch:
"""Deactivate one point without touching any pointer.
Used by the whole-file sweep, where every point in the file leaves at once:
no active point survives to dangle, so there is no neighbour to relink and
the chain stays intact for a later reader of the deactivated file.
"""
return PayloadPatch(
point_id=point.point_id,
payload={
"is_active": False,
"deleted_at": now.isoformat(),
**_provenance(point, actor=actor, now=now),
},
expected_version=point.version,
)

View File

@@ -0,0 +1,7 @@
"""Narrow contracts for external side effects (ADR-0015).
Ports exist for external side effects/persistence that need a swappable or
fake-able boundary — not as a blanket wrapper around every database access.
`object_storage.py` is one: MinIO is a real external system with its own
failure modes, and ADR-0016 requires a hand-written fake for it in tests.
"""

View File

@@ -0,0 +1,63 @@
"""Embedding ports (ADR-0001, ADR-0017).
`src/infrastructure/embedding/` holds the production adapters; tests use
scripted fakes (ADR-0016). Application code depends on these Protocols, not
on `httpx`/provider SDKs directly.
"""
from collections.abc import Sequence
from typing import Protocol
from src.application.ingestion.models import SparseVector
class DenseEmbedder(Protocol):
"""One named dense vector's embedding client (`dense_nomic`/`dense_openai`).
`embed_batch` is a single batched network call — callers own concurrency
bounding (ADR-0017's `embed_concurrency` semaphore), not this Protocol.
"""
name: str
model_version: str
"""Identifies the model that produced these vectors (ADR-0001).
Written into every point's `embedding_model_version` payload field, which
exists so a future model swap can tell which chunks need re-embedding. The
embedder is what knows this, so it is reported here rather than
reconstructed from configuration at the call site.
"""
async def embed_batch(self, texts: Sequence[str]) -> list[list[float]]:
"""Return one vector per input text, same order. Raises `EmbedderError`
(see `src/application/ingestion/errors.py`) on transport/response
failure.
"""
...
class SparseEmbedder(Protocol):
"""The `sparse` (BM25) vector's embedding client.
Blocking/CPU-bound (ADR-0017): callers offload it via
`anyio.to_thread.run_sync` with the ingestion `CapacityLimiter`, not call
it directly from an `async def`.
"""
name: str
model_version: str
"""Identifies the analyzer/parameters that produced these vectors.
Same purpose as `DenseEmbedder.model_version`; for BM25 the "model" is the
analyzer choice (ADR-0005), which is equally a re-embedding trigger.
"""
def embed_batch(self, texts: Sequence[str], *, query: bool = False) -> list[SparseVector]:
"""Return one sparse vector per input text, same order.
`query=True` selects the query-side weighting, which omits document
length normalization. Ingestion always passes `False`; the flag exists
so retrieval (ADR-0003) encodes queries through this same port rather
than growing a second, silently divergent implementation.
"""
...

View File

@@ -0,0 +1,14 @@
"""The object-storage port (ADR-0013).
`src/infrastructure/minio/storage.py` is the production adapter; tests use a
hand-written fake (ADR-0016). Application code depends on this Protocol, not
on the `minio` SDK.
"""
from typing import Protocol
class ObjectStorage(Protocol):
async def put_object(self, *, key: str, data: bytes, content_type: str) -> None:
"""Store `data` privately under `key`. Overwrites an existing object."""
...

View File

@@ -0,0 +1,131 @@
"""The point read/edit port for `/v1/points` (ADR-0002, ADR-0015).
Separate from `PointStorage`, which stays exactly the two bulk operations
ingestion performs. Reads, single-point edits, reordering, and keyword search
have a different caller, a different failure vocabulary, and a different
tenant-filter obligation, so they get their own port rather than accreting onto
the ingestion one.
`tenant_id` is a required keyword argument on **every** method. That is not
style: ADR-0002's isolation rule has to hold on every code path that touches the
collection, and an optional tenant filter is one forgotten argument away from a
cross-tenant read. Making it required moves that from a review question to a
type error.
`src/infrastructure/qdrant/point_repository.py` is the production adapter;
`tests.fakes.FakePointRepository` is the test double.
"""
import uuid
from collections.abc import Sequence
from typing import Protocol
from pydantic import BaseModel
from src.application.points.point import Point
class PointPage(BaseModel):
"""One page of points plus the cursor that continues it.
`next_cursor` is opaque to callers and encoded by the adapter: ordered
scrolls and keyword searches paginate by different Qdrant mechanisms, and
neither is a plain integer offset. `None` means the listing is exhausted.
"""
points: tuple[Point, ...]
next_cursor: str | None = None
class PayloadPatch(BaseModel):
"""Set these payload fields on one point, optionally guarded by `version`.
When `expected_version` is set, the adapter attaches it to the operation's
filter, so a concurrent write that has already moved the version on means
this patch matches nothing rather than clobbering it. The guard is what
makes a lost update impossible; detecting that it fired is the service's
job (see `apply_patches`).
"""
point_id: uuid.UUID
payload: dict[str, object]
expected_version: int | None = None
class PointRepository(Protocol):
async def get(
self, *, tenant_id: uuid.UUID, point_id: uuid.UUID, with_vectors: bool = False
) -> Point | None:
"""One point, or `None` if it does not exist *under this tenant*.
The two cases are deliberately indistinguishable — the route maps both
to `404` so a caller cannot probe for another tenant's point ids.
"""
...
async def get_many(
self, *, tenant_id: uuid.UUID, point_ids: Sequence[uuid.UUID]
) -> tuple[Point, ...]:
"""The subset of `point_ids` that exists under this tenant.
Order is not guaranteed and missing ids are silently absent: callers are
neighbour-relinking and batch precondition checks, both of which match
on id rather than position.
"""
...
async def list_by_file(
self,
*,
tenant_id: uuid.UUID,
file_id: uuid.UUID,
limit: int,
cursor: str | None = None,
include_inactive: bool = False,
) -> PointPage:
"""One file's points in `order_id` order (ADR-0008's `scroll`).
Scoped to a single file because the cursor is an `order_id` value, and
`order_id` is only unique within a file.
"""
...
async def count(
self,
*,
tenant_id: uuid.UUID,
domain: str | None = None,
file_id: uuid.UUID | None = None,
include_inactive: bool = False,
) -> int: ...
async def keyword_search(
self,
*,
tenant_id: uuid.UUID,
query: str,
limit: int,
cursor: str | None = None,
domain: str | None = None,
file_id: uuid.UUID | None = None,
include_inactive: bool = False,
) -> PointPage:
"""Full-text payload match on `content`, plus structured filters.
Keyword matching, **not** semantic retrieval (ADR-0002). Qdrant's
full-text index is a filter, not a scorer, so results carry no relevance
ranking and their order is unspecified.
"""
...
async def apply_patches(self, *, tenant_id: uuid.UUID, patches: Sequence[PayloadPatch]) -> None:
"""Apply every patch in one Qdrant `points/batch` request.
Qdrant has no multi-point transaction, so this is not atomic and does
not pretend to be. ADR-0002's all-or-nothing rule is implemented one
layer up as validate-every-precondition-then-apply; the per-patch
`expected_version` guard here is what makes the residual window safe,
turning a lost update into a no-op the service can detect rather than a
silent clobber.
"""
...

View File

@@ -0,0 +1,45 @@
"""The point-storage port (ADR-0001, ADR-0015).
`src/infrastructure/qdrant/points.py` is the production adapter; tests use a
hand-written fake (ADR-0016). Application code depends on this Protocol, not on
the `qdrant_client` SDK.
Deliberately narrow: exactly the two operations ingestion performs. Reads,
single-point edits, reordering, and keyword search are plan 002's `/v1/points`
surface and belong on a port of their own rather than accreting here.
"""
import uuid
from collections.abc import Sequence
from datetime import datetime
from typing import Protocol
from src.application.points.models import ChunkPoint
class PointStorage(Protocol):
async def upsert_points(self, points: Sequence[ChunkPoint]) -> None:
"""Upsert one batch of points.
Callers own batching and concurrency bounding (ADR-0017's
`upsert_concurrency` semaphore), not this Protocol — the same division
`DenseEmbedder.embed_batch` uses.
"""
...
async def deactivate_points_from_index(
self,
*,
tenant_id: uuid.UUID,
file_id: uuid.UUID,
from_chunk_index: int,
deleted_at: datetime,
updated_by: str,
) -> int:
"""Soft-delete this file's points at or past `from_chunk_index`.
Sets `is_active=false`/`deleted_at` rather than removing the points
(ADR-0002: delete is soft by default). Tenant-filtered — a `file_id`
alone is never sufficient authority. Returns how many points matched.
"""
...

View File

@@ -0,0 +1,9 @@
"""Operator-run tenant provisioning (ADR-0008, ADR-0009)."""
from src.application.tenants.provisioning import (
DEFAULT_SCOPES,
ProvisionResult,
provision_tenant,
)
__all__ = ["DEFAULT_SCOPES", "ProvisionResult", "provision_tenant"]

View File

@@ -0,0 +1,129 @@
"""Provision a tenant, its first API key, and its domains (ADR-0008, ADR-0009).
Nothing in the HTTP surface can bootstrap a tenant: every `/v1` route needs a
key, and a key can only exist once a tenant does. That chicken-and-egg is why
this is an operator-run deployment step (`src/cli/provision_tenant.py`) rather
than an endpoint — the same reasoning that keeps `alembic upgrade head` and
`qdrant_bootstrap` off the request path.
This is the package's only caller-facing entry point. It owns the whole
composition — tenant reuse-or-create, key generation and hashing, domain
registration, and the single transaction the three share — so a caller cannot
get the order wrong or commit a key whose tenant never landed (CLAUDE.md,
"prefer deep modules"). The plaintext key is returned exactly once and is never
logged (ADR-0011 forbids plaintext keys in logs); only its non-secret
`key_prefix` appears in the event.
"""
import uuid
from dataclasses import dataclass
import structlog
from sqlalchemy.ext.asyncio import AsyncSession, async_sessionmaker
from src.application.auth.keys import generate_api_key, hash_secret
from src.infrastructure.postgres.repositories import api_keys as api_keys_repo
from src.infrastructure.postgres.repositories import tenant_domains as domains_repo
from src.infrastructure.postgres.repositories import tenants as tenants_repo
logger = structlog.get_logger(__name__)
DEFAULT_SCOPES = (
"files:write",
"domains:read",
"domains:write",
"points:read",
"points:write",
)
@dataclass(frozen=True)
class ProvisionResult:
tenant_id: uuid.UUID
tenant_slug: str
tenant_created: bool
api_key_id: uuid.UUID
api_key_prefix: str
api_key: str
"""The plaintext bearer token. Only ever returned here — never stored, never logged."""
domains_created: tuple[str, ...]
domains_existing: tuple[str, ...]
async def provision_tenant(
sessionmaker: async_sessionmaker[AsyncSession],
*,
slug: str,
name: str | None = None,
key_name: str = "bootstrap",
scopes: tuple[str, ...] = DEFAULT_SCOPES,
domains: tuple[str, ...] = (),
actor_type: str = "backend",
) -> ProvisionResult:
"""Create (or reuse) the tenant, issue a key, and register `domains`.
Re-running with the same `slug` reuses the tenant and its existing domains
rather than failing, so an operator can add a key to a live tenant with the
same command they used to create it. A *new* key is issued on every run —
keys are write-once by construction (only the hash is stored), so there is
nothing to return for an existing one.
"""
key_prefix, secret, full_key = generate_api_key()
async with sessionmaker() as session:
tenant = await tenants_repo.get_by_slug(session, slug)
tenant_created = tenant is None
if tenant is None:
tenant = tenants_repo.create(session, slug=slug, name=name or slug)
# `api_keys.tenant_id` and `tenant_domains.tenant_id` FK to this row
# and the mapped classes carry no ORM relationship for the unit of
# work to order by itself, so the insert has to land first.
await session.flush()
api_key = api_keys_repo.create(
session,
tenant_id=tenant.id,
name=key_name,
key_prefix=key_prefix,
key_hash=hash_secret(secret),
scopes=list(scopes),
actor_type=actor_type,
created_by="cli:provision_tenant",
)
created: list[str] = []
existing: list[str] = []
for domain in domains:
if await domains_repo.get(session, tenant_id=tenant.id, domain=domain) is not None:
existing.append(domain)
continue
domains_repo.create(session, tenant_id=tenant.id, domain=domain, display_name=domain)
created.append(domain)
await session.flush()
tenant_id, api_key_id, tenant_slug = tenant.id, api_key.id, tenant.slug
await session.commit()
if tenant_created:
logger.info("tenant.provisioned", tenant_id=str(tenant_id), tenant_slug=tenant_slug)
for domain in created:
logger.info("domain.created", tenant_id=str(tenant_id), domain=domain)
logger.info(
"api_key.provisioned",
tenant_id=str(tenant_id),
api_key_id=str(api_key_id),
key_prefix=key_prefix,
scopes=list(scopes),
actor_type=actor_type,
)
return ProvisionResult(
tenant_id=tenant_id,
tenant_slug=tenant_slug,
tenant_created=tenant_created,
api_key_id=api_key_id,
api_key_prefix=key_prefix,
api_key=full_key,
domains_created=tuple(created),
domains_existing=tuple(existing),
)

View File

@@ -1,11 +1,16 @@
from collections.abc import AsyncIterator from collections.abc import AsyncIterator, Sequence
from dataclasses import dataclass from dataclasses import dataclass
from anyio import CapacityLimiter, Semaphore
from fastapi import Request from fastapi import Request
from minio import Minio from minio import Minio
from qdrant_client import AsyncQdrantClient from qdrant_client import AsyncQdrantClient
from sqlalchemy.ext.asyncio import AsyncEngine, AsyncSession, async_sessionmaker from sqlalchemy.ext.asyncio import AsyncEngine, AsyncSession, async_sessionmaker
from src.application.ports.embedding import DenseEmbedder, SparseEmbedder
from src.application.ports.object_storage import ObjectStorage
from src.application.ports.point_repository import PointRepository
from src.application.ports.point_storage import PointStorage
from src.config import Settings from src.config import Settings
@@ -16,6 +21,13 @@ class AppResources:
db_sessionmaker: async_sessionmaker[AsyncSession] db_sessionmaker: async_sessionmaker[AsyncSession]
minio_client: Minio minio_client: Minio
qdrant_client: AsyncQdrantClient qdrant_client: AsyncQdrantClient
object_storage: ObjectStorage
point_storage: PointStorage
point_repository: PointRepository
ingestion_limiter: CapacityLimiter
dense_embedders: Sequence[DenseEmbedder]
sparse_embedder: SparseEmbedder
ingestion_concurrency_limiter: Semaphore
def _resources(request: Request) -> AppResources: def _resources(request: Request) -> AppResources:
@@ -34,6 +46,45 @@ def get_qdrant_client(request: Request) -> AsyncQdrantClient:
return _resources(request).qdrant_client return _resources(request).qdrant_client
def get_object_storage(request: Request) -> ObjectStorage:
return _resources(request).object_storage
def get_point_storage(request: Request) -> PointStorage:
return _resources(request).point_storage
def get_point_repository(request: Request) -> PointRepository:
return _resources(request).point_repository
def get_ingestion_limiter(request: Request) -> CapacityLimiter:
return _resources(request).ingestion_limiter
def get_dense_embedders(request: Request) -> Sequence[DenseEmbedder]:
return _resources(request).dense_embedders
def get_sparse_embedder(request: Request) -> SparseEmbedder:
return _resources(request).sparse_embedder
def get_ingestion_concurrency_limiter(request: Request) -> Semaphore:
return _resources(request).ingestion_concurrency_limiter
def get_sessionmaker(request: Request) -> async_sessionmaker[AsyncSession]:
"""The session *factory*, not a request-scoped session.
Application services that own more than one transaction in a single
request (ADR-0017's two-phase upload) need to open and close sessions
themselves rather than borrow one request-scoped session that would
otherwise stay open across the whole request.
"""
return _resources(request).db_sessionmaker
async def get_db_session(request: Request) -> AsyncIterator[AsyncSession]: async def get_db_session(request: Request) -> AsyncIterator[AsyncSession]:
sessionmaker = _resources(request).db_sessionmaker sessionmaker = _resources(request).db_sessionmaker
async with sessionmaker() as session: async with sessionmaker() as session:

View File

@@ -1,26 +1,75 @@
from collections.abc import AsyncIterator, Callable from collections.abc import AsyncIterator, Callable, Sequence
from contextlib import AbstractAsyncContextManager, asynccontextmanager from contextlib import AbstractAsyncContextManager, asynccontextmanager
import httpx
import structlog import structlog
from anyio import CapacityLimiter, Semaphore, to_thread
from fastapi import FastAPI from fastapi import FastAPI
from src.application.ingestion import get_encoder
from src.application.ports.embedding import DenseEmbedder
from src.bootstrap.dependencies import AppResources from src.bootstrap.dependencies import AppResources
from src.config import Settings from src.config import Settings
from src.infrastructure.embedding.bm25 import Bm25SparseEmbedder
from src.infrastructure.embedding.openai_compatible import (
OpenAICompatibleEmbedder,
is_ollama_base_url,
)
from src.infrastructure.minio.client import create_client as create_minio_client from src.infrastructure.minio.client import create_client as create_minio_client
from src.infrastructure.minio.storage import MinioObjectStorage
from src.infrastructure.observability.logging import configure_logging from src.infrastructure.observability.logging import configure_logging
from src.infrastructure.postgres.database import create_engine, create_sessionmaker from src.infrastructure.postgres.database import create_engine, create_sessionmaker
from src.infrastructure.qdrant.client import create_client as create_qdrant_client from src.infrastructure.qdrant.client import create_client as create_qdrant_client
from src.infrastructure.qdrant.point_repository import QdrantPointRepository
from src.infrastructure.qdrant.points import QdrantPointStorage
logger = structlog.get_logger(__name__) logger = structlog.get_logger(__name__)
def _auth_headers(api_key: str | None) -> dict[str, str]:
"""Bearer header, or none at all when no key is configured.
Sending an empty `Bearer ` is worse than sending nothing: some gateways
treat a malformed credential as an auth failure rather than as anonymous.
"""
return {"Authorization": f"Bearer {api_key}"} if api_key else {}
async def _warm_dense_embedders(embedders: Sequence[DenseEmbedder]) -> None:
"""Force each dense model to load before the first upload needs it.
Same rationale as the tiktoken warm-up above, but with the opposite
failure policy. A self-hosted embedder that has unloaded the model takes
minutes to serve its first request — longer than
`INGESTION_TIMEOUT_SECONDS` — so paying that once at boot keeps it off a
user's upload. Unlike the tokenizer this is best-effort: an embedder that
is merely *down* must not stop the process from booting and reporting its
own health, and `/readyz` is where that condition belongs.
"""
for embedder in embedders:
try:
await embedder.embed_batch(["warmup"])
logger.info("lifespan.embedder.warmed", embedder=embedder.name)
except Exception:
logger.warning("lifespan.embedder.warm_failed", embedder=embedder.name, exc_info=True)
def create_lifespan( def create_lifespan(
settings: Settings | None = None, settings: Settings | None = None,
) -> Callable[[FastAPI], AbstractAsyncContextManager[None, bool | None]]: ) -> Callable[[FastAPI], AbstractAsyncContextManager[None, bool | None]]:
@asynccontextmanager @asynccontextmanager
async def lifespan(app: FastAPI) -> AsyncIterator[None]: async def lifespan(app: FastAPI) -> AsyncIterator[None]:
resolved_settings = settings or Settings() resolved_settings = settings or Settings()
configure_logging(resolved_settings.logging) configure_logging(resolved_settings.logging, resolved_settings.app)
# tiktoken fetches its vocabulary over the network on first use, so warm
# it here: a missing vocabulary should fail the process at boot, not the
# first upload. Blocking, hence the thread.
await to_thread.run_sync(get_encoder, resolved_settings.chunking.encoding_name)
logger.info(
"lifespan.tokenizer.loaded",
encoding=resolved_settings.chunking.encoding_name,
)
db_engine = create_engine(resolved_settings.postgres) db_engine = create_engine(resolved_settings.postgres)
db_sessionmaker = create_sessionmaker(db_engine) db_sessionmaker = create_sessionmaker(db_engine)
@@ -30,14 +79,81 @@ def create_lifespan(
logger.info("lifespan.minio.client.created") logger.info("lifespan.minio.client.created")
qdrant_client = create_qdrant_client(resolved_settings.qdrant) qdrant_client = create_qdrant_client(resolved_settings.qdrant)
# No collection DDL here: `ensure_chunks_collection` is a deployment
# step (`python -m src.cli.qdrant_bootstrap`), for the same reason
# ADR-0009 keeps Alembic out of startup and ADR-0012 makes LangGraph's
# `.setup()` a deployment step.
point_storage = QdrantPointStorage(
qdrant_client, collection=resolved_settings.qdrant.collection
)
point_repository = QdrantPointRepository(
qdrant_client, collection=resolved_settings.qdrant.collection
)
logger.info("lifespan.qdrant.client.created") logger.info("lifespan.qdrant.client.created")
nomic_settings = resolved_settings.embedding.nomic
nomic_http_client = httpx.AsyncClient(
base_url=nomic_settings.base_url,
timeout=nomic_settings.timeout_seconds,
headers=_auth_headers(nomic_settings.api_key),
)
openai_settings = resolved_settings.embedding.openai
openai_http_client = httpx.AsyncClient(
base_url=openai_settings.base_url,
timeout=openai_settings.timeout_seconds,
headers=_auth_headers(openai_settings.api_key),
)
dense_embedders = (
OpenAICompatibleEmbedder(
nomic_http_client,
name="dense_nomic",
model=nomic_settings.model,
document_prefix=nomic_settings.document_prefix,
keep_alive=(
nomic_settings.keep_alive
if is_ollama_base_url(nomic_settings.base_url)
else None
),
),
OpenAICompatibleEmbedder(
openai_http_client,
name="dense_openai",
model=openai_settings.model,
dimensions=openai_settings.dimensions,
document_prefix=openai_settings.document_prefix,
),
)
sparse_embedder = Bm25SparseEmbedder(resolved_settings.embedding.sparse)
logger.info("lifespan.embedders.created")
await _warm_dense_embedders(dense_embedders)
# Bounds how many ingestions run in this process at once (ADR-0017);
# a distinct resource from ingestion_limiter, which bounds threads
# spent on blocking work within a single ingestion.
ingestion_concurrency_limiter = Semaphore(resolved_settings.ingestion.max_concurrency)
# Bounds threads spent on blocking ingestion work (parsing, chunking,
# hashing, the sync minio SDK) so it cannot exhaust Starlette's own
# thread pool (ADR-0017).
ingestion_limiter = CapacityLimiter(resolved_settings.ingestion.thread_pool_size)
object_storage = MinioObjectStorage(
minio_client, bucket=resolved_settings.minio.bucket, limiter=ingestion_limiter
)
app.state.resources = AppResources( app.state.resources = AppResources(
settings=resolved_settings, settings=resolved_settings,
db_engine=db_engine, db_engine=db_engine,
db_sessionmaker=db_sessionmaker, db_sessionmaker=db_sessionmaker,
minio_client=minio_client, minio_client=minio_client,
qdrant_client=qdrant_client, qdrant_client=qdrant_client,
object_storage=object_storage,
point_storage=point_storage,
point_repository=point_repository,
ingestion_limiter=ingestion_limiter,
dense_embedders=dense_embedders,
sparse_embedder=sparse_embedder,
ingestion_concurrency_limiter=ingestion_concurrency_limiter,
) )
try: try:
@@ -53,4 +169,14 @@ def create_lifespan(
except Exception: except Exception:
logger.exception("lifespan.qdrant.close.failed") logger.exception("lifespan.qdrant.close.failed")
try:
await nomic_http_client.aclose()
except Exception:
logger.exception("lifespan.embedding.nomic_client.close.failed")
try:
await openai_http_client.aclose()
except Exception:
logger.exception("lifespan.embedding.openai_client.close.failed")
return lifespan return lifespan

1
src/cli/__init__.py Normal file
View File

@@ -0,0 +1 @@
"""Operator entry points that run as deployment steps, not at app startup."""

View File

@@ -0,0 +1,94 @@
"""Create a tenant, issue its first API key, and register its domains.
uv run python -m src.cli.provision_tenant --slug acme --domain fire
A deployment step, like `alembic upgrade head` and `src.cli.qdrant_bootstrap`.
It exists because nothing over HTTP can bootstrap a tenant: every `/v1` route
requires an API key, and a key cannot exist before its tenant does.
The plaintext key is printed to **stdout once** and never stored or logged —
Postgres holds only its SHA-256 hash (ADR-0009), so a lost key is reissued by
re-running this command, not recovered. Structured logs go to stderr/the log
sink and carry only the non-secret `key_prefix` (ADR-0011).
Re-running with the same `--slug` reuses the tenant and any domains it already
has, and issues an additional key.
"""
import argparse
import asyncio
import sys
import structlog
from src.application.tenants import DEFAULT_SCOPES, provision_tenant
from src.config import Settings
from src.infrastructure.observability.logging import configure_logging
from src.infrastructure.postgres.database import create_engine, create_sessionmaker
logger = structlog.get_logger(__name__)
def _parse_args(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(
prog="python -m src.cli.provision_tenant",
description="Create a tenant, issue an API key, and register domains.",
)
parser.add_argument("--slug", required=True, help="URL-safe tenant identifier, e.g. 'acme'")
parser.add_argument("--name", default=None, help="Display name (defaults to --slug)")
parser.add_argument("--key-name", default="bootstrap", help="Label for the issued API key")
parser.add_argument(
"--scopes",
default=",".join(DEFAULT_SCOPES),
help=f"Comma-separated scopes for the key (default: {','.join(DEFAULT_SCOPES)})",
)
parser.add_argument(
"--domain",
action="append",
default=[],
dest="domains",
help="Domain to register; repeatable. Uploads reject an unregistered domain.",
)
return parser.parse_args(argv)
async def run(argv: list[str] | None = None, settings: Settings | None = None) -> int:
args = _parse_args(argv)
resolved = settings or Settings()
configure_logging(resolved.logging, resolved.app)
scopes = tuple(scope.strip() for scope in args.scopes.split(",") if scope.strip())
if not scopes:
logger.error("tenant.provision.failed", reason="no_scopes")
return 2
engine = create_engine(resolved.postgres)
try:
result = await provision_tenant(
create_sessionmaker(engine),
slug=args.slug,
name=args.name,
key_name=args.key_name,
scopes=scopes,
domains=tuple(args.domains),
)
finally:
await engine.dispose()
# stdout, not the logger: this is the one value the operator must copy, and
# it must never reach a log sink (ADR-0011).
print(f"tenant_id={result.tenant_id}")
print(f"tenant_slug={result.tenant_slug}")
print(f"api_key_id={result.api_key_id}")
print(f"domains={','.join(result.domains_created + result.domains_existing)}")
print(f"api_key={result.api_key}")
print("Store the api_key now -- only its hash is persisted and it cannot be shown again.")
return 0
def main() -> None:
sys.exit(asyncio.run(run()))
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,54 @@
"""Create the `chunks` collection — the Qdrant analogue of `alembic upgrade head`.
uv run python -m src.cli.qdrant_bootstrap
A deployment step, deliberately not part of the FastAPI lifespan: collection
creation is DDL, which ADR-0009 keeps out of application startup for Postgres
and ADR-0012 keeps out of it for LangGraph's `.setup()`. See
`src/infrastructure/qdrant/collection.py` for the full reasoning.
Idempotent and safe to re-run. Exits non-zero if an existing collection
diverges from the pinned schema, rather than leaving a silently degraded
sparse index behind.
"""
import asyncio
import sys
import structlog
from src.config import Settings
from src.infrastructure.observability.logging import configure_logging
from src.infrastructure.qdrant.client import create_client
from src.infrastructure.qdrant.collection import (
CollectionSchemaMismatchError,
ensure_chunks_collection,
)
logger = structlog.get_logger(__name__)
async def bootstrap(settings: Settings | None = None) -> int:
resolved = settings or Settings()
configure_logging(resolved.logging, resolved.app)
client = create_client(resolved.qdrant)
try:
created = await ensure_chunks_collection(client, collection=resolved.qdrant.collection)
except CollectionSchemaMismatchError as exc:
logger.error("qdrant.bootstrap.schema_mismatch", error=str(exc))
return 1
finally:
await client.close()
logger.info(
"qdrant.bootstrap.completed", collection=resolved.qdrant.collection, created=created
)
return 0
def main() -> None:
sys.exit(asyncio.run(bootstrap()))
if __name__ == "__main__":
main()

View File

@@ -1,9 +1,31 @@
from pydantic import Field from typing import Any
from pydantic import Field, model_validator
from pydantic_settings import BaseSettings, SettingsConfigDict from pydantic_settings import BaseSettings, SettingsConfigDict
# pydantic-settings does not cascade `env_file` from a parent `BaseSettings`
# to a nested one: each `BaseSettings` subclass only reads `.env` if its own
# `model_config` names it. `Settings` has no fields of its own -- only
# nested settings objects -- so every nested class below repeats
# `env_file=".env"` as its own default, or its fields would silently read
# only real process environment variables (fine under Docker Compose, broken
# for local `.env`-file development) while still *appearing* to work
# whenever a `.env.example` default happens to match the class default.
#
# That default alone isn't enough to let a caller point `Settings` at a
# *different* file (as the test suite does, to parse `.env.example`), since
# passing `_env_file=...` to `Settings(...)` only overrides `Settings`'s own
# `model_config` -- nested defaults still construct against their own
# hardcoded ".env". `Settings.__init__`/`EmbeddingSettings.__init__` below
# thread an explicit `_env_file` override through to every nested
# constructor so one override actually reaches the whole tree.
_UNSET: Any = object()
class PostgresSettings(BaseSettings): class PostgresSettings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="POSTGRES_", extra="ignore") model_config = SettingsConfigDict(
env_prefix="POSTGRES_", extra="ignore", env_file=".env", env_ignore_empty=True
)
host: str = "127.0.0.1" host: str = "127.0.0.1"
port: int = 5433 port: int = 5433
@@ -17,7 +39,9 @@ class PostgresSettings(BaseSettings):
class MinioSettings(BaseSettings): class MinioSettings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="MINIO_", extra="ignore") model_config = SettingsConfigDict(
env_prefix="MINIO_", extra="ignore", env_file=".env", env_ignore_empty=True
)
endpoint: str = "127.0.0.1:9100" endpoint: str = "127.0.0.1:9100"
access_key: str = "chatbot" access_key: str = "chatbot"
@@ -33,36 +57,203 @@ class IngestionSettings(BaseSettings):
timeouts, or callers give up on work that is still succeeding. timeouts, or callers give up on work that is still succeeding.
""" """
model_config = SettingsConfigDict(env_prefix="INGESTION_", extra="ignore") model_config = SettingsConfigDict(
env_prefix="INGESTION_", extra="ignore", env_file=".env", env_ignore_empty=True
)
max_concurrency: int = 4 max_concurrency: int = 4
thread_pool_size: int = 8 thread_pool_size: int = 8
timeout_seconds: float = 120.0 timeout_seconds: float = 120.0
max_upload_size_mb: int = 25
max_chunks_per_file: int = 5000 max_chunks_per_file: int = 5000
embed_batch_size: int = 128 embed_batch_size: int = 128
embed_concurrency: int = 4 embed_concurrency: int = 4
@property
def max_upload_size_bytes(self) -> int:
return self.max_upload_size_mb * 1024 * 1024
class ChunkingSettings(BaseSettings):
"""Parsing and chunking parameters (ADR-0018).
`max_chunk_tokens` is `nomic-embed-text-v2-moe`'s sequence length. Text past
it is silently truncated by the model rather than rejected, so the cap is
enforced here instead. `chunk_size` sits well under it to leave room for the
`search_document: ` task prefix and any heading text carried into a chunk.
"""
model_config = SettingsConfigDict(
env_prefix="CHUNKING_", extra="ignore", env_file=".env", env_ignore_empty=True
)
strategy: str = "fixed_size"
chunk_size: int = 400
chunk_overlap: int = 60
max_chunk_tokens: int = 512
encoding_name: str = "cl100k_base"
@model_validator(mode="after")
def _validate_sizes(self) -> "ChunkingSettings":
if self.chunk_overlap >= self.chunk_size:
raise ValueError(
f"chunk_overlap ({self.chunk_overlap}) must be smaller than "
f"chunk_size ({self.chunk_size}); otherwise splitting never advances"
)
if self.chunk_size > self.max_chunk_tokens:
raise ValueError(
f"chunk_size ({self.chunk_size}) must not exceed "
f"max_chunk_tokens ({self.max_chunk_tokens})"
)
return self
class QdrantSettings(BaseSettings): class QdrantSettings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="QDRANT_", extra="ignore") """Qdrant connection and bulk-upsert bounds (ADR-0001, ADR-0017).
`collection` names the single shared collection all tenants live in
(ADR-0001); it is deliberately configurable so tests can point at a
disposable one. The vector *dimensions* are not settings -- they are model
facts pinned in `src/infrastructure/qdrant/collection.py`, and changing one
is a re-embedding migration.
`upsert_batch_size` sits inside ADR-0001's 64-256 bulk-upload band, and
`upsert_concurrency` bounds in-flight batches so ingestion issues parallel
streams rather than an unbounded `gather` (ADR-0017).
"""
model_config = SettingsConfigDict(
env_prefix="QDRANT_", extra="ignore", env_file=".env", env_ignore_empty=True
)
url: str = "http://127.0.0.1:6343" url: str = "http://127.0.0.1:6343"
api_key: str | None = None api_key: str | None = None
collection: str = "chunks"
upsert_batch_size: int = 128
upsert_concurrency: int = 4
class NomicEmbeddingSettings(BaseSettings):
"""Self-hosted `nomic-embed-text-v2-moe` server (ADR-0001, 768-dim).
Talks the OpenAI-compatible `/embeddings` endpoint shape. The default
points at the Ollama host the `emet` benchmark used, whose OpenAI-compat
shim accepts any non-empty API key.
`document_prefix` is empty by design. ADR-0004 cites the model card's
`search_document: ` requirement, but emet's winning run used no prefix
(Ollama's template is a bare passthrough and injects none), and the
prefix is not cosmetic -- it moves the vector substantially. Turning it
on here obliges the query side to send `search_query: ` too (ADR-0003),
so it stays off until an emet run measures the pair together.
"""
model_config = SettingsConfigDict(
env_prefix="EMBEDDING_NOMIC_", extra="ignore", env_file=".env", env_ignore_empty=True
)
base_url: str = "http://192.168.10.10:11435/v1"
model: str = "nomic-embed-text-v2-moe"
api_key: str = "sk-not-set"
document_prefix: str = ""
keep_alive: str = "30m"
timeout_seconds: float = 30.0
class OpenaiEmbeddingSettings(BaseSettings):
"""OpenAI's hosted embedding API (ADR-0001's second dense signal).
`dimensions` is unset, so `text-embedding-3-large` returns its native
3072 dimensions -- the configuration emet benchmarked. Setting it would
truncate via Matryoshka and is a re-embedding migration, not a config
tweak.
"""
model_config = SettingsConfigDict(
env_prefix="EMBEDDING_OPENAI_", extra="ignore", env_file=".env", env_ignore_empty=True
)
base_url: str = "https://api.openai.com/v1"
model: str = "text-embedding-3-large"
api_key: str | None = None
dimensions: int | None = None
document_prefix: str = ""
timeout_seconds: float = 30.0
class SparseEmbeddingSettings(BaseSettings):
"""BM25 sparse-vector parameters (ADR-0001, ADR-0005).
`analyzer`, `k`, and `b` are the configuration emet benchmarked as
`bm25-fa-norm-stop`; changing them invalidates that result. `k`/`b`
saturation is applied client-side here, while IDF comes from Qdrant's
`modifier="idf"` sparse-vector config at query time.
`avg_len` is emet's placeholder average document length in analyzer
tokens, exposed as a setting so it can be recalibrated from real corpus
statistics without a code change.
"""
model_config = SettingsConfigDict(
env_prefix="EMBEDDING_SPARSE_", extra="ignore", env_file=".env", env_ignore_empty=True
)
analyzer: str = "fa_norm_stop"
k: float = 1.2
b: float = 0.75
avg_len: float = 256.0
class EmbeddingSettings(BaseSettings):
model_config = SettingsConfigDict(extra="ignore")
nomic: NomicEmbeddingSettings = Field(default_factory=NomicEmbeddingSettings)
openai: OpenaiEmbeddingSettings = Field(default_factory=OpenaiEmbeddingSettings)
sparse: SparseEmbeddingSettings = Field(default_factory=SparseEmbeddingSettings)
def __init__(self, _env_file: Any = _UNSET, **data: Any) -> None:
env_file = ".env" if _env_file is _UNSET else _env_file
data.setdefault("nomic", NomicEmbeddingSettings(_env_file=env_file))
data.setdefault("openai", OpenaiEmbeddingSettings(_env_file=env_file))
data.setdefault("sparse", SparseEmbeddingSettings(_env_file=env_file))
super().__init__(**data)
class AppLimitSettings(BaseSettings): class AppLimitSettings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="APP_", extra="ignore") model_config = SettingsConfigDict(
env_prefix="APP_", extra="ignore", env_file=".env", env_ignore_empty=True
)
env: str = "local" env: str = "local"
max_upload_size_mb: int = 25
readiness_check_timeout_seconds: float = 2.0 readiness_check_timeout_seconds: float = 2.0
# The deployed commit SHA or release tag (ADR-0011, "Bind process-level
# environment context"). Set by CI/CD at build/deploy time -- never
# computed at runtime by shelling out to git, which would fail in a
# container image with no .git directory.
service_version: str = "dev"
class LoggingSettings(BaseSettings): class LoggingSettings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="LOG_", extra="ignore") """Logging sinks (ADR-0011).
`json_format` controls stdout's renderer only. Production sets it `true`
so stdout is JSON for the container log collector; local development
leaves it `false` for a colored console renderer. `file_path`, when set,
is a second, independent handler that always renders JSON regardless of
`json_format` -- a developer can read a human console while still keeping
a machine-parseable file. Unset in production: stdout/stderr collection is
preferred there over a log file inside the container.
"""
model_config = SettingsConfigDict(
env_prefix="LOG_", extra="ignore", env_file=".env", env_ignore_empty=True
)
level: str = "INFO" level: str = "INFO"
json_format: bool = False json_format: bool = False
file_path: str | None = None
file_max_bytes: int = 10 * 1024 * 1024
file_backup_count: int = 5
class Settings(BaseSettings): class Settings(BaseSettings):
@@ -71,6 +262,20 @@ class Settings(BaseSettings):
postgres: PostgresSettings = Field(default_factory=PostgresSettings) postgres: PostgresSettings = Field(default_factory=PostgresSettings)
minio: MinioSettings = Field(default_factory=MinioSettings) minio: MinioSettings = Field(default_factory=MinioSettings)
ingestion: IngestionSettings = Field(default_factory=IngestionSettings) ingestion: IngestionSettings = Field(default_factory=IngestionSettings)
chunking: ChunkingSettings = Field(default_factory=ChunkingSettings)
qdrant: QdrantSettings = Field(default_factory=QdrantSettings) qdrant: QdrantSettings = Field(default_factory=QdrantSettings)
embedding: EmbeddingSettings = Field(default_factory=EmbeddingSettings)
app: AppLimitSettings = Field(default_factory=AppLimitSettings) app: AppLimitSettings = Field(default_factory=AppLimitSettings)
logging: LoggingSettings = Field(default_factory=LoggingSettings) logging: LoggingSettings = Field(default_factory=LoggingSettings)
def __init__(self, _env_file: Any = _UNSET, **data: Any) -> None:
env_file = ".env" if _env_file is _UNSET else _env_file
data.setdefault("postgres", PostgresSettings(_env_file=env_file))
data.setdefault("minio", MinioSettings(_env_file=env_file))
data.setdefault("ingestion", IngestionSettings(_env_file=env_file))
data.setdefault("chunking", ChunkingSettings(_env_file=env_file))
data.setdefault("qdrant", QdrantSettings(_env_file=env_file))
data.setdefault("embedding", EmbeddingSettings(_env_file=env_file))
data.setdefault("app", AppLimitSettings(_env_file=env_file))
data.setdefault("logging", LoggingSettings(_env_file=env_file))
super().__init__(_env_file=env_file, **data)

View File

View File

@@ -0,0 +1,150 @@
"""The `fa_norm_stop` BM25 analyzer (ADR-0001, ADR-0005).
Ported from the `emet` evaluation lab
(`src/chatbot_gh/adapters/sparse/analyzers.py`), which benchmarked four Farsi
analyzer variants on the real corpus and found `fa_norm_stop` the best
performer. This is a *measured* artifact: changing the normalization,
tokenization, or stopword list invalidates that result, so improvements belong
in a new emet benchmark run rather than in an edit here.
Deliberately independent of `src/application/ingestion/normalization.py`.
Those solve different problems: `normalize_persian_text` shapes chunk content
that gets cited back to the reader, so ADR-0018 has it preserve digits and
punctuation as authored. This module shapes index terms nobody ever sees, so
it folds digits and diacritics freely. Sharing one function between them would
let a display-motivated tweak silently perturb the benchmarked sparse index.
"""
import re
import unicodedata
# Several Arabic letterforms are visually indistinguishable from Latin ones in
# a monospace editor (alef from "l", heh from "o"), and literals render
# right-to-left, visually reordering the source line. `normalization.py` writes
# them as codepoints for that reason; this module follows the same convention.
_ZWNJ = 0x200C
_ARABIC_YEH = 0x064A
_ARABIC_KAF = 0x0643
_TEH_MARBUTA = 0x0629
_HAMZA_ON_WAW = 0x0624
_ALEF_HAMZA_BELOW = 0x0625
_ALEF_HAMZA_ABOVE = 0x0623
_PERSIAN_YEH = 0x06CC
_PERSIAN_KEHEH = 0x06A9
_HEH = 0x0647
_WAW = 0x0648
_ALEF = 0x0627
_SPACE = 0x0020
# Persian (U+06F0-U+06F9) and Arabic-Indic (U+0660-U+0669) digits both fold to
# ASCII, so the same number matches however it was authored.
_EASTERN_DIGITS = str.maketrans("۰۱۲۳۴۵۶۷۸۹٠١٢٣٤٥٦٧٨٩", "01234567890123456789")
# ZWNJ becomes a space (splitting compounds into separate terms) and the
# Arabic letterforms fold to their Persian equivalents. Every entry is a
# single codepoint mapping to a single codepoint over disjoint sources, so
# applying them in one pass is equivalent to emet's chained `str.replace`
# calls -- provided NFC runs first, since NFC is what composes the hamza
# forms this table then folds.
_FOLDING: dict[int, int] = {
_ZWNJ: _SPACE,
_ARABIC_YEH: _PERSIAN_YEH,
_ARABIC_KAF: _PERSIAN_KEHEH,
_TEH_MARBUTA: _HEH,
_HAMZA_ON_WAW: _WAW,
_ALEF_HAMZA_BELOW: _ALEF,
_ALEF_HAMZA_ABOVE: _ALEF,
}
_FOLDING.update(_EASTERN_DIGITS)
# Common Persian/Arabic stopwords (function words + FAQ noise), plus the
# English function words that appear in a mixed-script corpus. Kept small and
# explicit -- deliberately not a full hazm list. `_HEH_ALEF` is the plural
# suffix "ha"; written as a codepoint pair because both of its letters are
# Latin-confusable, which is exactly the case Ruff's RUF001 flags.
_HEH_ALEF = chr(_HEH) + chr(_ALEF)
_PERSIAN_STOPWORDS: frozenset[str] = frozenset(
{
"و",
"در",
"به",
"از",
"که",
"این",
"را",
"با",
"برای",
"آن",
"یک",
"است",
"شد",
"شده",
"می",
"های",
_HEH_ALEF,
"یا",
"تا",
"بر",
"اگر",
"هم",
"نیز",
"ولی",
"اما",
"چه",
"چون",
"روی",
"پس",
"پیش",
"هر",
"هیچ",
"بود",
"باشد",
"هست",
"نیست",
"کند",
"کرد",
"کردن",
"شود",
"the",
"a",
"an",
"of",
"to",
"and",
"in",
"on",
"for",
"is",
"are",
}
)
# Word characters minus underscore. Note this KEEPS digits: an insurance
# corpus is full of policy numbers, dates, and amounts, and those are exactly
# the tokens a lexical index should be able to match on.
_TOKEN_RE = re.compile(r"[^\W_]+", re.UNICODE)
FA_NORM_STOP = "fa_norm_stop"
def _normalize_fa(text: str) -> str:
return unicodedata.normalize("NFC", text).translate(_FOLDING)
def _tokenize_raw(text: str) -> list[str]:
return [m.group(0).lower() for m in _TOKEN_RE.finditer(text)]
def analyze(text: str, analyzer: str = FA_NORM_STOP) -> list[str]:
"""Tokenize `text` into sparse-index terms.
Only `fa_norm_stop` is implemented -- emet's other three variants
(`raw`, `fa_norm`, `fa_norm_stem`) lost the benchmark and exist there as
experiment arms, not as configurations this service should run.
"""
if analyzer != FA_NORM_STOP:
raise ValueError(f"Unknown analyzer '{analyzer}'")
tokens = _tokenize_raw(_normalize_fa(text))
return [token for token in tokens if token not in _PERSIAN_STOPWORDS]

View File

@@ -0,0 +1,93 @@
"""The `bm25-fa-norm-stop` sparse embedder (ADR-0001, ADR-0005).
Ported from the `emet` evaluation lab
(`src/chatbot_gh/adapters/sparse/bm25_embedder.py`), the configuration that
won its Farsi analyzer benchmark. Not Qdrant's hosted `Qdrant/bm25` FastEmbed
model, whose documented language support omits Farsi (ADR-0005).
**The BM25 work is split across two systems.** This adapter applies the
term-frequency saturation half client-side -- the `k` and `b` parameters,
including document-length normalization. IDF is *not* computed here: Qdrant
supplies it from collection-wide statistics when the sparse vector field is
created with `modifier="idf"`.
That split is load-bearing. A Qdrant collection created without
`modifier="idf"` will silently score these vectors as saturated term
frequencies with no IDF weighting at all -- no error, just materially worse
lexical retrieval. The collection bootstrap (plan 001 Phase 5) must set it.
"""
from collections import Counter
from collections.abc import Sequence
from hashlib import blake2b
from src.application.ingestion.models import SparseVector
from src.config import SparseEmbeddingSettings
from src.infrastructure.embedding.analyzers import analyze
# Qdrant sparse indices must be non-negative and fit a signed 32-bit int.
_INDEX_SPACE = 2**31 - 1
def _token_index(token: str) -> int:
"""Map a term to its sparse-vector index.
A hash rather than a vocabulary table, so the mapping needs no shared
state and stays identical across processes, restarts, and — critically —
between ingest-time and query-time encoding. `blake2b` rather than
`hash()`, which is PYTHONHASHSEED-salted and therefore differs per
process.
"""
digest = blake2b(token.encode("utf-8"), digest_size=8).digest()
return int.from_bytes(digest, "big") % _INDEX_SPACE
def text_to_sparse_vector(
text: str, *, settings: SparseEmbeddingSettings, query: bool = False
) -> SparseVector:
"""Encode one text as a BM25-saturated sparse vector (IDF applied by Qdrant).
Document and query sides differ in exactly one term: documents carry the
`b` length normalization, queries do not (standard BM25 practice — a
query's own length should not discount its terms).
"""
tokens = analyze(text, settings.analyzer)
if not tokens:
return SparseVector(indices=[], values=[])
frequencies = Counter(tokens)
doc_length = float(len(tokens))
k = settings.k
b = settings.b
indices: list[int] = []
values: list[float] = []
# Sorted so the emitted vector is deterministic for a given text, which
# keeps re-ingestion byte-stable and makes the output testable.
for token, freq in sorted(frequencies.items()):
if query:
weight = freq * (k + 1.0) / (freq + k)
else:
weight = freq * (k + 1.0) / (freq + k * (1.0 - b + b * doc_length / settings.avg_len))
indices.append(_token_index(token))
values.append(float(weight))
return SparseVector(indices=indices, values=values)
class Bm25SparseEmbedder:
"""A `SparseEmbedder` (see `src/application/ports/embedding.py`).
Pure CPU work with no network calls, so it is blocking: callers offload it
via `anyio.to_thread.run_sync` with the ingestion `CapacityLimiter`
(ADR-0017), never awaiting it directly on the event loop.
"""
name = "sparse"
def __init__(self, settings: SparseEmbeddingSettings) -> None:
self.model_version = f"bm25-{settings.analyzer}"
self._settings = settings
def embed_batch(self, texts: Sequence[str], *, query: bool = False) -> list[SparseVector]:
return [text_to_sparse_vector(text, settings=self._settings, query=query) for text in texts]

View File

@@ -0,0 +1,82 @@
"""Dense embedding adapter for OpenAI-compatible `/embeddings` endpoints.
Backs both `dense_nomic` (self-hosted `nomic-embed-text-v2-moe` behind
Ollama's OpenAI-compatible shim) and `dense_openai` (OpenAI's hosted API) —
both speak the same request/response shape, so one adapter serves both named
vectors with different config (ADR-0001).
Uses `httpx` directly rather than the `openai` SDK. The `emet` benchmark this
configuration comes from uses the SDK, but it is a synchronous batch tool;
ADR-0017 requires async, semaphore-bounded batches here, and the request shape
is small enough that the SDK earns nothing.
`httpx.AsyncClient` is application-lifetime (ADR-0012): built once in the
FastAPI lifespan and passed in, never constructed per call.
"""
from collections.abc import Sequence
from urllib.parse import urlparse
import httpx
# Ollama's default port, plus the alternate the benchmarked deployment uses.
_OLLAMA_PORTS = frozenset({11434, 11435})
def is_ollama_base_url(base_url: str) -> bool:
"""Whether `base_url` looks like an Ollama OpenAI-compatible endpoint.
Ollama unloads an idle model, and reloading `nomic-embed-text-v2-moe`
costs well over two minutes — longer than `INGESTION_TIMEOUT_SECONDS`, so
a cold upload would 504. `keep_alive` is how the model is kept resident,
and it is an Ollama extension, hence the sniffing.
"""
parsed = urlparse(base_url)
host = parsed.hostname or ""
return parsed.port in _OLLAMA_PORTS or "ollama" in host.lower()
class OpenAICompatibleEmbedder:
"""A `DenseEmbedder` (see `src/application/ports/embedding.py`) over one
OpenAI-compatible `/embeddings` endpoint.
"""
def __init__(
self,
client: httpx.AsyncClient,
*,
name: str,
model: str,
dimensions: int | None = None,
document_prefix: str = "",
keep_alive: str | None = None,
) -> None:
self.name = name
self.model_version = model
self._client = client
self._model = model
self._dimensions = dimensions
self._document_prefix = document_prefix
self._keep_alive = keep_alive
async def embed_batch(self, texts: Sequence[str]) -> list[list[float]]:
inputs = (
[f"{self._document_prefix}{text}" for text in texts]
if self._document_prefix
else list(texts)
)
payload: dict[str, object] = {"model": self._model, "input": inputs}
if self._dimensions is not None:
payload["dimensions"] = self._dimensions
if self._keep_alive is not None:
payload["keep_alive"] = self._keep_alive
response = await self._client.post("/embeddings", json=payload)
response.raise_for_status()
body = response.json()
# Sort by `index` rather than trusting response order: the contract
# guarantees the field, not the ordering, and a silently permuted
# batch would attach every vector to the wrong chunk.
data = sorted(body["data"], key=lambda item: item["index"])
return [item["embedding"] for item in data]

View File

@@ -0,0 +1,33 @@
"""MinIO adapter for the `ObjectStorage` port (ADR-0013, ADR-0017).
The `minio` SDK is synchronous, so every call runs through
`anyio.to_thread.run_sync` bounded by the ingestion `CapacityLimiter` — the
same rule ADR-0017 applies to parsing/chunking. Calling the SDK directly from
`async def` would block every concurrent request in the process.
"""
import io
from functools import partial
from anyio import CapacityLimiter, to_thread
from minio import Minio
class MinioObjectStorage:
def __init__(self, client: Minio, *, bucket: str, limiter: CapacityLimiter) -> None:
self._client = client
self._bucket = bucket
self._limiter = limiter
async def put_object(self, *, key: str, data: bytes, content_type: str) -> None:
await to_thread.run_sync(
partial(
self._client.put_object,
self._bucket,
key,
io.BytesIO(data),
length=len(data),
content_type=content_type,
),
limiter=self._limiter,
)

View File

@@ -1,14 +1,50 @@
"""Logging configuration: structlog + stdlib, dual local sinks (ADR-0011).
Console and an optional file are independent, simultaneous handlers on the
same logger, not a single renderer chosen by a flag -- the same structlog
event fans out to both. The console handler is always human-readable
(`ConsoleRenderer`); the file handler, when enabled via `LOG_FILE_PATH`,
always renders JSON regardless of `LOG_JSON_FORMAT`, so a saved log stays
machine-parseable even when the terminal next to it is not.
`LOG_JSON_FORMAT` controls *stdout's* renderer only: production sets it `true`
so the container log collector gets JSON; local development leaves it `false`
for the colored console. `LOG_FILE_PATH` is expected to be unset in
production -- stdout/stderr collection is preferred there over a log file
inside the container.
"""
import logging import logging
import logging.config import logging.config
import sys import sys
from collections.abc import Callable
import structlog import structlog
from src.config import LoggingSettings from src.config import AppLimitSettings, LoggingSettings
def configure_logging(settings: LoggingSettings) -> None: def _bind_environment(settings: AppLimitSettings) -> Callable[..., dict[str, object]]:
"""A static processor, not a contextvar: `env`/`service_version` don't
vary per request, and a contextvar bound before the first request would
be wiped by `RequestIdMiddleware`'s `clear_contextvars()` on that request.
Closing over `settings` at configure time makes every event carry them
instead, regardless of request context (ADR-0011).
"""
def processor(
logger: object, method_name: str, event_dict: dict[str, object]
) -> dict[str, object]:
event_dict["env"] = settings.env
event_dict["service_version"] = settings.service_version
return event_dict
return processor
def configure_logging(settings: LoggingSettings, app_settings: AppLimitSettings) -> None:
shared_processors = [ shared_processors = [
_bind_environment(app_settings),
structlog.contextvars.merge_contextvars, structlog.contextvars.merge_contextvars,
structlog.stdlib.add_log_level, structlog.stdlib.add_log_level,
structlog.stdlib.add_logger_name, structlog.stdlib.add_logger_name,
@@ -27,55 +63,84 @@ def configure_logging(settings: LoggingSettings) -> None:
cache_logger_on_first_use=True, cache_logger_on_first_use=True,
) )
renderer = ( console_renderer = (
structlog.processors.JSONRenderer() structlog.processors.JSONRenderer()
if settings.json_format if settings.json_format
else structlog.dev.ConsoleRenderer(colors=True) else structlog.dev.ConsoleRenderer(colors=True)
) )
formatters = {
"console": {
"()": structlog.stdlib.ProcessorFormatter,
"processors": [
structlog.stdlib.ProcessorFormatter.remove_processors_meta,
console_renderer,
],
"foreign_pre_chain": [
structlog.stdlib.ExtraAdder(),
*shared_processors,
],
},
}
handlers: dict[str, dict[str, object]] = {
"console": {
"class": "logging.StreamHandler",
"level": settings.level,
"formatter": "console",
"stream": sys.stdout,
},
}
root_handlers = ["console"]
if settings.file_path is not None:
# File handler always renders JSON, independent of the console
# renderer chosen above -- a saved log stays machine-parseable even
# when stdout is the colored, human-readable renderer.
formatters["file"] = {
"()": structlog.stdlib.ProcessorFormatter,
"processors": [
structlog.stdlib.ProcessorFormatter.remove_processors_meta,
structlog.processors.JSONRenderer(),
],
"foreign_pre_chain": [
structlog.stdlib.ExtraAdder(),
*shared_processors,
],
}
handlers["file"] = {
"class": "logging.handlers.RotatingFileHandler",
"level": settings.level,
"formatter": "file",
"filename": settings.file_path,
"maxBytes": settings.file_max_bytes,
"backupCount": settings.file_backup_count,
}
root_handlers.append("file")
logging.config.dictConfig( logging.config.dictConfig(
{ {
"version": 1, "version": 1,
"disable_existing_loggers": False, "disable_existing_loggers": False,
"formatters": { "formatters": formatters,
"default": { "handlers": handlers,
"()": structlog.stdlib.ProcessorFormatter,
"processors": [
structlog.stdlib.ProcessorFormatter.remove_processors_meta,
renderer,
],
"foreign_pre_chain": [
structlog.stdlib.ExtraAdder(),
*shared_processors,
],
},
},
"handlers": {
"console": {
"class": "logging.StreamHandler",
"level": settings.level,
"formatter": "default",
"stream": sys.stdout,
},
},
"loggers": { "loggers": {
"": { "": {
"handlers": ["console"], "handlers": root_handlers,
"level": settings.level, "level": settings.level,
"propagate": False, "propagate": False,
}, },
"uvicorn": { "uvicorn": {
"handlers": ["console"], "handlers": root_handlers,
"level": settings.level, "level": settings.level,
"propagate": False, "propagate": False,
}, },
"uvicorn.access": { "uvicorn.access": {
"handlers": ["console"], "handlers": root_handlers,
"level": settings.level, "level": settings.level,
"propagate": False, "propagate": False,
}, },
"sqlalchemy.engine": { "sqlalchemy.engine": {
"handlers": ["console"], "handlers": root_handlers,
"level": "WARNING", "level": "WARNING",
"propagate": False, "propagate": False,
}, },

View File

@@ -4,6 +4,7 @@ from src.infrastructure.postgres.models.ingestion_job import IngestionJob
from src.infrastructure.postgres.models.ingestion_job_event import IngestionJobEvent from src.infrastructure.postgres.models.ingestion_job_event import IngestionJobEvent
from src.infrastructure.postgres.models.source_file import SourceFile from src.infrastructure.postgres.models.source_file import SourceFile
from src.infrastructure.postgres.models.tenant import Tenant from src.infrastructure.postgres.models.tenant import Tenant
from src.infrastructure.postgres.models.tenant_domain import TenantDomain
__all__ = [ __all__ = [
"ApiKey", "ApiKey",
@@ -12,4 +13,5 @@ __all__ = [
"IngestionJobEvent", "IngestionJobEvent",
"SourceFile", "SourceFile",
"Tenant", "Tenant",
"TenantDomain",
] ]

View File

@@ -0,0 +1,52 @@
import uuid
from datetime import datetime
from sqlalchemy import CheckConstraint, DateTime, ForeignKey, String, UniqueConstraint, func
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.orm import Mapped, mapped_column
from src.infrastructure.postgres.models.base import Base
TENANT_DOMAIN_STATUSES = ("active", "disabled")
class TenantDomain(Base):
"""A domain a tenant is allowed to ingest into (ADR-0009).
Tenants do not share a domain list — one may run 14 insurance lines and
another 6 — so this is a per-tenant table rather than an enum or a global
lookup.
Its purpose is to stop an arbitrary caller-supplied `domain` from silently
creating a new Qdrant partition. `domain` is denormalized into every point's
payload and into `source_files`, and a typo like `fier` for `fire` produces
no error anywhere: the file indexes into a partition retrieval never queries,
so it is invisible rather than failed.
`domain` is the immutable key. Renaming it would mean rewriting every point
payload that carries it, which is a migration, not an edit — `display_name`
is the mutable human-facing label instead.
"""
__tablename__ = "tenant_domains"
__table_args__ = (
UniqueConstraint("tenant_id", "domain", name="uq_tenant_domains_tenant_id_domain"),
CheckConstraint(f"status IN {TENANT_DOMAIN_STATUSES}", name="ck_tenant_domains_status"),
)
id: Mapped[uuid.UUID] = mapped_column(primary_key=True)
tenant_id: Mapped[uuid.UUID] = mapped_column(
ForeignKey("tenants.id", ondelete="CASCADE"), index=True
)
domain: Mapped[str] = mapped_column(String(80))
display_name: Mapped[str] = mapped_column(String(200))
status: Mapped[str] = mapped_column(String(20), default="active", server_default="active")
metadata_: Mapped[dict[str, object]] = mapped_column(
"metadata", JSONB, default=dict, server_default="{}"
)
created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), server_default=func.now())
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), server_default=func.now(), onupdate=func.now()
)
disabled_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), default=None)

View File

@@ -0,0 +1,45 @@
"""API-key lookups and issuance (ADR-0008, ADR-0009).
Plain functions over an `AsyncSession` the caller owns. No function here
commits, rolls back, or closes the session (ADR-0012). Secret comparison
happens in `src/application/auth`, not here — this module only fetches rows
by their non-secret `key_prefix`.
"""
import uuid
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from src.infrastructure.postgres.models.api_key import ApiKey
async def get_by_prefix(session: AsyncSession, key_prefix: str) -> ApiKey | None:
result = await session.execute(select(ApiKey).where(ApiKey.key_prefix == key_prefix))
return result.scalar_one_or_none()
def create(
session: AsyncSession,
*,
tenant_id: uuid.UUID,
name: str,
key_prefix: str,
key_hash: str,
scopes: list[str],
actor_type: str = "backend",
created_by: str | None = None,
) -> ApiKey:
"""Persist an issued key. The caller hashes the secret; this never sees it."""
api_key = ApiKey(
id=uuid.uuid4(),
tenant_id=tenant_id,
name=name,
key_prefix=key_prefix,
key_hash=key_hash,
scopes=scopes,
actor_type=actor_type,
created_by=created_by,
)
session.add(api_key)
return api_key

View File

@@ -0,0 +1,119 @@
"""`ingestion_jobs`/`ingestion_job_events` persistence (ADR-0009, ADR-0017).
Plain functions over an `AsyncSession` the caller owns. No function here
commits, rolls back, or closes the session (ADR-0012) — the two-transaction
shape in `src/application/files/upload.py` depends on that. Every read is
tenant-scoped by a required `tenant_id` argument.
"""
import uuid
from datetime import UTC, datetime
from sqlalchemy import desc, select
from sqlalchemy.ext.asyncio import AsyncSession
from src.infrastructure.postgres.models.ingestion_job import IngestionJob
from src.infrastructure.postgres.models.ingestion_job_event import IngestionJobEvent
async def get_by_id(
session: AsyncSession, *, tenant_id: uuid.UUID, ingestion_job_id: uuid.UUID
) -> IngestionJob | None:
result = await session.execute(
select(IngestionJob).where(
IngestionJob.id == ingestion_job_id, IngestionJob.tenant_id == tenant_id
)
)
return result.scalar_one_or_none()
async def get_latest_for_source_file(
session: AsyncSession, *, tenant_id: uuid.UUID, source_file_id: uuid.UUID
) -> IngestionJob | None:
result = await session.execute(
select(IngestionJob)
.where(
IngestionJob.tenant_id == tenant_id,
IngestionJob.source_file_id == source_file_id,
)
.order_by(desc(IngestionJob.created_at))
.limit(1)
)
return result.scalar_one_or_none()
def create_running(
session: AsyncSession,
*,
tenant_id: uuid.UUID,
source_file_id: uuid.UUID,
requested_by_api_key_id: uuid.UUID | None,
chunking_strategy: str,
) -> IngestionJob:
job = IngestionJob(
id=uuid.uuid4(),
tenant_id=tenant_id,
source_file_id=source_file_id,
requested_by_api_key_id=requested_by_api_key_id,
status="running",
chunking_strategy=chunking_strategy,
started_at=datetime.now(UTC),
)
session.add(job)
return job
async def mark_terminal(
session: AsyncSession,
*,
tenant_id: uuid.UUID,
ingestion_job_id: uuid.UUID,
status: str,
points_created: int = 0,
points_updated: int = 0,
points_soft_deleted: int = 0,
points_skipped: int = 0,
error_code: str | None = None,
error_message: str | None = None,
) -> IngestionJob | None:
"""Move a job from `running` to a terminal status.
Never reads/writes a job whose current status is already terminal — a
terminal job must not transition back to `running` or to a different
terminal status (ADR-0017).
"""
job = await get_by_id(session, tenant_id=tenant_id, ingestion_job_id=ingestion_job_id)
if job is None or job.status != "running":
return None
job.status = status
job.completed_at = datetime.now(UTC)
job.points_created = points_created
job.points_updated = points_updated
job.points_soft_deleted = points_soft_deleted
job.points_skipped = points_skipped
job.error_code = error_code
job.error_message = error_message
return job
def append_event(
session: AsyncSession,
*,
tenant_id: uuid.UUID,
ingestion_job_id: uuid.UUID,
level: str,
stage: str,
message: str,
details: dict[str, object] | None = None,
) -> IngestionJobEvent:
event = IngestionJobEvent(
id=uuid.uuid4(),
tenant_id=tenant_id,
ingestion_job_id=ingestion_job_id,
level=level,
stage=stage,
message=message,
details=details or {},
)
session.add(event)
return event

View File

@@ -0,0 +1,86 @@
"""`source_files` persistence (ADR-0009).
Plain functions over an `AsyncSession` the caller owns. No function here
commits, rolls back, or closes the session (ADR-0012). Every read is
tenant-scoped by a required `tenant_id` argument, so a missing filter is a
signature error rather than a cross-tenant leak.
"""
import uuid
from datetime import datetime
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from src.infrastructure.postgres.models.source_file import SourceFile
async def get_by_id(
session: AsyncSession, *, tenant_id: uuid.UUID, source_file_id: uuid.UUID
) -> SourceFile | None:
result = await session.execute(
select(SourceFile).where(SourceFile.id == source_file_id, SourceFile.tenant_id == tenant_id)
)
return result.scalar_one_or_none()
async def find_active_by_content_hash(
session: AsyncSession, *, tenant_id: uuid.UUID, domain: str, content_sha256: str
) -> SourceFile | None:
result = await session.execute(
select(SourceFile).where(
SourceFile.tenant_id == tenant_id,
SourceFile.domain == domain,
SourceFile.content_sha256 == content_sha256,
SourceFile.status == "active",
)
)
return result.scalar_one_or_none()
def create(
session: AsyncSession,
*,
source_file_id: uuid.UUID,
tenant_id: uuid.UUID,
domain: str,
source_filename: str,
source_type: str,
content_sha256: str,
byte_size: int,
storage_uri: str,
created_by_api_key_id: uuid.UUID | None,
) -> SourceFile:
"""`source_file_id` is caller-generated: the upload service derives the
MinIO object key from it before this row exists, so the id has to be
chosen up front rather than assigned by the database.
"""
source_file = SourceFile(
id=source_file_id,
tenant_id=tenant_id,
domain=domain,
source_filename=source_filename,
source_type=source_type,
content_sha256=content_sha256,
byte_size=byte_size,
storage_uri=storage_uri,
created_by_api_key_id=created_by_api_key_id,
)
session.add(source_file)
return source_file
def mark_soft_deleted(source_file: SourceFile, *, deleted_at: datetime) -> None:
"""Retire a file: `status='soft_deleted'` plus `deleted_at` (ADR-0009).
Takes the already-loaded row rather than an id, because the caller fetched
it under its tenant filter and re-fetching here would be a second place
that could forget that filter.
Retiring the row matters beyond bookkeeping: `find_active_by_content_hash`
matches only `active` files, so a re-upload of the same bytes after a delete
creates a fresh file and re-ingests it, instead of taking the duplicate path
and returning a file whose points have all been deactivated.
"""
source_file.status = "soft_deleted"
source_file.deleted_at = deleted_at

View File

@@ -0,0 +1,73 @@
"""`tenant_domains` persistence (ADR-0009).
Plain functions over an `AsyncSession` the caller owns. No function here
commits, rolls back, or closes the session (ADR-0012). Every read and write is
tenant-scoped by a required `tenant_id` argument, so a missing filter is a
signature error rather than a cross-tenant leak.
"""
import uuid
from datetime import UTC, datetime
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from src.infrastructure.postgres.models.tenant_domain import TenantDomain
async def get(session: AsyncSession, *, tenant_id: uuid.UUID, domain: str) -> TenantDomain | None:
result = await session.execute(
select(TenantDomain).where(
TenantDomain.tenant_id == tenant_id, TenantDomain.domain == domain
)
)
return result.scalar_one_or_none()
async def list_for_tenant(
session: AsyncSession, *, tenant_id: uuid.UUID, include_disabled: bool = False
) -> list[TenantDomain]:
statement = select(TenantDomain).where(TenantDomain.tenant_id == tenant_id)
if not include_disabled:
statement = statement.where(TenantDomain.status == "active")
result = await session.execute(statement.order_by(TenantDomain.domain))
return list(result.scalars().all())
def create(
session: AsyncSession,
*,
tenant_id: uuid.UUID,
domain: str,
display_name: str,
metadata: dict[str, object] | None = None,
) -> TenantDomain:
tenant_domain = TenantDomain(
id=uuid.uuid4(),
tenant_id=tenant_id,
domain=domain,
display_name=display_name,
metadata_=metadata or {},
)
session.add(tenant_domain)
return tenant_domain
def update_display_name(tenant_domain: TenantDomain, *, display_name: str) -> TenantDomain:
"""`domain` itself is deliberately not updatable.
It is denormalized into every Qdrant point payload and into `source_files`,
so changing the key would mean rewriting all of them — a migration, not an
edit. The label is what callers actually want to change.
"""
tenant_domain.display_name = display_name
return tenant_domain
def set_status(tenant_domain: TenantDomain, *, status: str) -> TenantDomain:
"""Disable/re-enable a domain. Existing points are untouched either way —
disabling blocks new uploads, it is not a delete (ADR-0002).
"""
tenant_domain.status = status
tenant_domain.disabled_at = datetime.now(UTC) if status == "disabled" else None
return tenant_domain

View File

@@ -0,0 +1,27 @@
"""Tenant lookups and creation (ADR-0009).
Plain functions over an `AsyncSession` the caller owns. No function here
commits, rolls back, or closes the session (ADR-0012).
"""
import uuid
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from src.infrastructure.postgres.models.tenant import Tenant
async def get_by_id(session: AsyncSession, tenant_id: uuid.UUID) -> Tenant | None:
return await session.get(Tenant, tenant_id)
async def get_by_slug(session: AsyncSession, slug: str) -> Tenant | None:
result = await session.execute(select(Tenant).where(Tenant.slug == slug))
return result.scalar_one_or_none()
def create(session: AsyncSession, *, slug: str, name: str) -> Tenant:
tenant = Tenant(id=uuid.uuid4(), slug=slug, name=name)
session.add(tenant)
return tenant

View File

@@ -9,10 +9,22 @@ def create_client(settings: QdrantSettings) -> AsyncQdrantClient:
return AsyncQdrantClient(url=settings.url, api_key=settings.api_key) return AsyncQdrantClient(url=settings.url, api_key=settings.api_key)
async def ping(client: AsyncQdrantClient, timeout: float) -> bool: async def ping(client: AsyncQdrantClient, timeout: float, *, collection: str) -> bool:
"""Whether Qdrant is reachable **and** the `chunks` collection exists.
Reachability alone is not readiness here. The collection is created by a
deployment step (`python -m src.cli.qdrant_bootstrap`, see ADR-0001
"Collection provisioning"), so a process can boot against a healthy Qdrant
that has no collection at all. Without this check that misconfiguration
stays invisible until the first upload fails with a `502` — after the
request has already paid for the MinIO write and the embedding round trips.
This is the Qdrant analogue of an unapplied Alembic migration, and it
belongs in `/readyz` for the same reason: it is a dependency-readiness
condition, not a process-health one.
"""
try: try:
async with asyncio.timeout(timeout): async with asyncio.timeout(timeout):
await client.get_collections() return await client.collection_exists(collection)
except Exception: except Exception:
return False return False
return True

Some files were not shown because too many files have changed in this diff Show More