docs: add CLAUDE.md project guidance
This commit is contained in:
229
CLAUDE.md
Normal file
229
CLAUDE.md
Normal file
@@ -0,0 +1,229 @@
|
|||||||
|
# CLAUDE.md
|
||||||
|
|
||||||
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||||
|
|
||||||
|
## Project status
|
||||||
|
|
||||||
|
This repo is currently ADR-driven and mostly pre-implementation: `src/` contains
|
||||||
|
only an empty `main.py`/`config.py` scaffold and empty `api/routers`,
|
||||||
|
`api/dependencies`, `db`, and `schemas` directories. Architecture decisions live
|
||||||
|
in `docs/adr/` (17 ADRs plus the 0000 template; 0001–0004 are `Accepted`,
|
||||||
|
0014 is `Superseded by 0017`, and the rest — 0005–0013 and 0015–0017 — are
|
||||||
|
`Proposed`). Implementation plans live in `docs/plans/`:
|
||||||
|
`001-ingestion-vertical-slice.md` and
|
||||||
|
`002-point-crud-and-keyword-search.md`. **Read the relevant
|
||||||
|
ADR(s) before implementing anything** — the ADRs are the source of truth for
|
||||||
|
package layout, boundaries, and invariants; code should not silently diverge
|
||||||
|
from them. If an ADR needs to change to fit new work, update the ADR rather
|
||||||
|
than quietly contradicting it.
|
||||||
|
|
||||||
|
## Commands
|
||||||
|
|
||||||
|
Package manager is `uv`; Python 3.13.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
uv sync # install deps
|
||||||
|
uv run ruff check --fix <path> # lint (autofix)
|
||||||
|
uv run ruff format <path> # format
|
||||||
|
uv run ty check <path> # type check
|
||||||
|
uv run pytest # run tests
|
||||||
|
uv run pytest -m unit # fast, no external services (default dev loop)
|
||||||
|
uv run pytest -m "integration and postgres" # a single integration boundary
|
||||||
|
uv run pytest tests/path/to/test_file.py::test_name # single test
|
||||||
|
uv run alembic upgrade head # apply migrations (never create_all() at runtime)
|
||||||
|
uv run fastapi dev src/main.py # run the API locally (once main.py exists)
|
||||||
|
```
|
||||||
|
|
||||||
|
A PostToolUse hook (`.claude/hooks/python_quality.py`) already runs
|
||||||
|
`ruff check --fix`, `ruff format`, `ruff check`, and `ty check` on every
|
||||||
|
Python file you Write/Edit and reports remaining diagnostics back to you —
|
||||||
|
you don't need to run these manually after every edit, but do run the full
|
||||||
|
suite before considering a task done.
|
||||||
|
|
||||||
|
### Never silence a diagnostic
|
||||||
|
|
||||||
|
A `ty`/Ruff diagnostic is frequently a real bug — most often the checker is
|
||||||
|
right and the code is wrong. **Do not make a diagnostic go away by any means
|
||||||
|
other than correcting the code it points at.** Specifically, never do the
|
||||||
|
following to get a clean check:
|
||||||
|
|
||||||
|
- add `# ty: ignore`, `# type: ignore`, `# noqa`, or `# pragma: no cover`
|
||||||
|
- widen a type to `Any`, or wrap a value in `cast(...)`
|
||||||
|
- replace `x.attr` with `getattr(x, "attr")`, wrap the call in
|
||||||
|
`except Exception`, or otherwise restructure code so the checker can no
|
||||||
|
longer see the problem
|
||||||
|
- edit `[tool.ty.rules]`, `[[tool.ty.overrides]]`, `[tool.ruff.lint]`
|
||||||
|
`ignore`/`per-file-ignores`, or any `exclude` in `pyproject.toml`/`ty.toml`/
|
||||||
|
`ruff.toml` — checker configuration is the user's decision, not yours
|
||||||
|
- delete or skip a failing test, or narrow its assertions
|
||||||
|
|
||||||
|
If a dependency is missing, add it (`uv add`). If a third-party package genuinely
|
||||||
|
ships no types, say so and let the user decide on a scoped override. **If you
|
||||||
|
believe a diagnostic is a false positive, stop and tell the user — do not
|
||||||
|
suppress it yourself.** That claim is exactly the case where silencing is most
|
||||||
|
expensive, so it is theirs to confirm.
|
||||||
|
|
||||||
|
The hook audits every edit for these patterns and reports them, including on
|
||||||
|
config files. If it flags a line you added, either justify it explicitly in your
|
||||||
|
reply or revert it — never leave it unmentioned, and never describe a file as
|
||||||
|
clean when its diagnostics were suppressed rather than fixed.
|
||||||
|
|
||||||
|
Local Langfuse (observability) stack: `docker compose --env-file .env.langfuse -f docker-compose.langfuse.yml up -d`, UI at `http://localhost:3000` (see README.md).
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
This is a **modular monolith** (ADR-0015) backing a tenant-scoped RAG chatbot.
|
||||||
|
Today it is a **single process with no background work**: ingestion runs inline
|
||||||
|
in the `POST /v1/files` request and returns a terminal result (ADR-0017). There
|
||||||
|
is no message broker, no transactional outbox, no queue, and no worker process —
|
||||||
|
ADR-0014 describes that design and is deferred, not deleted.
|
||||||
|
|
||||||
|
### Package layout (target shape, created as needed)
|
||||||
|
|
||||||
|
```
|
||||||
|
src/
|
||||||
|
├── bootstrap/ # per-process composition: lifespan, dependency wiring
|
||||||
|
├── api/ # FastAPI HTTP adapter only (routers/, dependencies/, schemas/, router.py)
|
||||||
|
├── application/ # use-case services: files/, ingestion/, points/, retrieval/, threads/, ports/
|
||||||
|
├── agent/ # LangGraph graph.py, state.py, nodes/, prompts/, tools/, persistence.py
|
||||||
|
└── infrastructure/ # concrete adapters: postgres/, qdrant/, minio/, embedding/, langgraph/, observability/
|
||||||
|
```
|
||||||
|
|
||||||
|
`messaging/`, `workers/`, and `infrastructure/rabbitmq/` belong to the ADR-0014
|
||||||
|
target shape and are **not created** while ADR-0017 stands. The FastAPI route is
|
||||||
|
the only entry adapter; it calls `application/ingestion/` directly.
|
||||||
|
|
||||||
|
Dependency direction is one-way and enforced by convention, not tooling:
|
||||||
|
|
||||||
|
```
|
||||||
|
FastAPI routes / LangGraph nodes
|
||||||
|
-> application services -> application ports -> infrastructure adapters
|
||||||
|
```
|
||||||
|
|
||||||
|
Infrastructure never imports routers/graph nodes; LangGraph nodes never call route functions or construct SDK
|
||||||
|
clients directly; application services never import concrete
|
||||||
|
MinIO/Qdrant/SQLAlchemy client-construction code. Use ports only for
|
||||||
|
external side effects/persistence — not around pure local functions.
|
||||||
|
(ADR-0015)
|
||||||
|
|
||||||
|
### Resource lifetime rules (ADR-0012)
|
||||||
|
|
||||||
|
- Application-lifetime objects (SQLAlchemy engine/sessionmaker, Qdrant client,
|
||||||
|
LangGraph checkpointer/store, compiled graph, shared HTTP/model/embedding
|
||||||
|
clients, the ingestion `CapacityLimiter`) are created once in the FastAPI
|
||||||
|
lifespan and closed there. Never construct mutable network/DB clients at
|
||||||
|
import time.
|
||||||
|
- One `AsyncSession` per request/job unit of work — never shared globally.
|
||||||
|
- Routes/application services own transaction boundaries (explicit
|
||||||
|
`commit()`); repositories don't commit/rollback/close sessions they didn't
|
||||||
|
create.
|
||||||
|
- Lower layers receive dependencies as explicit parameters, not via imported
|
||||||
|
singletons — this is what makes FastAPI dependency overrides and test
|
||||||
|
fixtures work.
|
||||||
|
|
||||||
|
### Ingestion flow (the first vertical slice, ADR-0017 + plan 001)
|
||||||
|
|
||||||
|
Ingestion is **inline in the request**, in three phases — and the phase
|
||||||
|
boundaries are the point:
|
||||||
|
|
||||||
|
```
|
||||||
|
txn A (short): auth+tenant -> validate -> insert source_files,
|
||||||
|
ingestion_jobs(status=running) -> COMMIT, release connection
|
||||||
|
no txn: store bytes in MinIO -> parse+chunk (threads) ->
|
||||||
|
embed dense+sparse (async, batched, semaphore-bounded) ->
|
||||||
|
upsert Qdrant points (deterministic ids)
|
||||||
|
txn B (short): ingestion_jobs -> succeeded/failed + counters,
|
||||||
|
append ingestion_job_events -> COMMIT
|
||||||
|
201 Created { file_id, ingestion_job_id, status, chunks_indexed }
|
||||||
|
```
|
||||||
|
|
||||||
|
**Never hold a Postgres session/transaction open across the work phase** — it
|
||||||
|
pins a pool connection for the whole upload. `ingestion_jobs` is a record of an
|
||||||
|
attempt, not a queue.
|
||||||
|
|
||||||
|
Work placement is not optional:
|
||||||
|
- **Async + batched + `asyncio.Semaphore`**: dense embedders (network I/O; batch
|
||||||
|
before parallelizing, never an unbounded `gather`), Qdrant upserts.
|
||||||
|
- **`anyio.to_thread.run_sync` + `CapacityLimiter`**: `python-docx`/`csv`
|
||||||
|
parsing, chunking, hashing, the BM25 sparse pipeline, the sync `minio` SDK.
|
||||||
|
Calling these from `async def` directly is a defect — one big parse stalls
|
||||||
|
every concurrent request.
|
||||||
|
- **Not computed at ingest**: `late_interaction` (jina-colbert-v2, GPU).
|
||||||
|
Populating it is the trigger to move ingestion back off the request.
|
||||||
|
|
||||||
|
Bounds are enforced server-side and all have status codes: size/chunk ceiling
|
||||||
|
`413`, `INGESTION_MAX_CONCURRENCY` `503`, `INGESTION_TIMEOUT_SECONDS` `504`,
|
||||||
|
embedder failure `502`. A timeout must still write a terminal job status.
|
||||||
|
Retried uploads must be safe: deterministic point IDs, `(tenant_id, domain,
|
||||||
|
content_sha256)` idempotency, no terminal job returning to `running`.
|
||||||
|
|
||||||
|
### Retrieval / agent (ADR-0001, 0003, 0005, 0006, 0007)
|
||||||
|
|
||||||
|
- Single Qdrant collection `chunks`, shared across tenants, with named vectors
|
||||||
|
`dense_nomic` (nomic-embed-text-v2-moe), `dense_openai`, `sparse` (BM25,
|
||||||
|
Farsi-tuned), and `late_interaction` (jina-colbert-v2, rerank-only, on-disk).
|
||||||
|
Multitenancy via Qdrant's `is_tenant` payload index on `tenant_id`, `m: 0` +
|
||||||
|
`payload_m: 16` HNSW config — every query/prefetch carries a server-derived
|
||||||
|
`tenant_id`/`domain` filter, never client-supplied.
|
||||||
|
- Retrieval = 3 parallel prefetches (dense_nomic, dense_openai, sparse) → RRF
|
||||||
|
fusion → late-interaction rerank over the fused top-N only (bounded cost) →
|
||||||
|
context-window expansion via `previous_chunk_id`/`next_chunk_id` pointers.
|
||||||
|
- Conversational agent is an explicit LangGraph `StateGraph` (not a prebuilt
|
||||||
|
tool-calling agent): `load_memory -> triage -> {chitchat|out_of_scope|
|
||||||
|
handoff_request|account|knowledge->contextualize->retrieve->grade->
|
||||||
|
{clarify|generate->verify}} -> write_memory`. Escalation and retry-budget
|
||||||
|
policy are graph edges, not model discretion; `generate` answers only from
|
||||||
|
retrieved chunks with per-claim citations; `verify` checks groundedness.
|
||||||
|
- LangGraph persistence: single `AsyncPostgresSaver` backed by one
|
||||||
|
`psycopg_pool.AsyncConnectionPool`, graph compiled once at startup.
|
||||||
|
`.setup()` (DDL) runs as a deployment step, never at app startup. The
|
||||||
|
LangGraph `thread_id` **is** the conversation identifier this service
|
||||||
|
knows — no separate thread/session table; this service stores no
|
||||||
|
conversation metadata beyond the checkpoint sequence.
|
||||||
|
|
||||||
|
### Postgres conventions (ADR-0009)
|
||||||
|
|
||||||
|
UUID primary keys (app-generated), `timestamptz` for all timestamps,
|
||||||
|
`Numeric(18, 8)` for money (never floats), `JSONB` for flexible metadata but
|
||||||
|
typed/indexed columns for query-critical fields, string status columns with
|
||||||
|
`CHECK` constraints (not native Postgres enums — they're painful to migrate).
|
||||||
|
`metadata` column maps to a `metadata_` attribute (reserved name on
|
||||||
|
Declarative models). All DDL goes through Alembic; FastAPI never creates or
|
||||||
|
alters tables at startup.
|
||||||
|
|
||||||
|
### Observability (ADR-0010, 0011)
|
||||||
|
|
||||||
|
Langfuse is the observability/evaluation plane (traces, prompt versions,
|
||||||
|
datasets/experiments, user feedback) — it is not the transactional database.
|
||||||
|
Postgres remains system of record for tenants, API keys, audit, jobs,
|
||||||
|
`graph_runs`, `llm_calls`/`llm_pricing`. Correlate the two via `request_id`,
|
||||||
|
`tenant_id`, `thread_id`, `run_id`. Use `structlog` with stable event names
|
||||||
|
and structured fields (`logger.info("graph.run.completed", ...)`), not
|
||||||
|
interpolated prose; JSON logs by default in production.
|
||||||
|
|
||||||
|
## Testing (ADR-0016)
|
||||||
|
|
||||||
|
- pytest, `pytest-asyncio` strict mode, `httpx.AsyncClient` + `ASGITransport`
|
||||||
|
+ `LifespanManager` for API tests (not `TestClient`) so lifespan/ADR-0012
|
||||||
|
wiring is actually exercised.
|
||||||
|
- One primary marker per test: `unit`, `integration`, or `e2e`; integration
|
||||||
|
tests also carry `postgres`/`minio`/`qdrant`; `live_provider` for
|
||||||
|
opt-in credential-gated external calls (never in routine runs); `slow` only
|
||||||
|
when materially over the normal integration budget.
|
||||||
|
- Test naming: `test_<unit>_<scenario>_<outcome>`, Arrange–Act–Assert.
|
||||||
|
- Layout mirrors architecture: `tests/unit/{application,agent}`,
|
||||||
|
`tests/integration/{postgres,minio,qdrant}`, `tests/e2e/`.
|
||||||
|
- Integration tests use **Testcontainers** (never a developer's local
|
||||||
|
services or Langfuse-owned storage/credentials) — this is the standard
|
||||||
|
automated mechanism, not Docker Compose. Isolate data per test via unique
|
||||||
|
keys/queue/collection names; parallel integration execution is disabled
|
||||||
|
until fixture isolation is proven safe.
|
||||||
|
- Pytest never calls a live/paid model provider in routine runs — that's
|
||||||
|
Langfuse's job (dataset experiments), not pytest's. Migrations are always
|
||||||
|
tested through real Alembic upgrade, never `Base.metadata.create_all()`.
|
||||||
|
- Invariants worth testing explicitly wherever they apply: tenant identity
|
||||||
|
only from server-side auth context (never request-suppliable); cross-tenant
|
||||||
|
access returns 404, not 403; the file + job row commit atomically before the
|
||||||
|
work phase; no session is held across parse/embed/upsert; a timeout or failure
|
||||||
|
always writes a terminal job status; retrying an upload must not duplicate
|
||||||
|
chunks or regress terminal job state.
|
||||||
Reference in New Issue
Block a user