docs(architecture): define application resource ownership

Why:
- Establish explicit ownership and lifetime rules for application, request/job, and operation-scoped resources.

Changes:
- Define FastAPI lifespan ownership for engines, clients, graphs, and pools.
- Define dependency-managed sessions and request context.
- Require explicit transaction boundaries and dependency passing.
- Prohibit shared global SQLAlchemy sessions and import-time network clients.

Impact:
- Application resources are created and closed by their process owner.
- Request and job resources must not be shared across concurrent units of work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-10 11:50:04 +03:30
parent e9caeaa4d8
commit 0ca698acfa
3 changed files with 734 additions and 0 deletions

View File

@@ -0,0 +1,270 @@
# 0012. Application resource lifetime and dependency ownership
## Status
Proposed
## Context
ADR-0008 defines the FastAPI REST boundary and says the app lifespan owns
long-lived resources. ADR-0009 chooses SQLAlchemy 2.x, Alembic migrations, and
async sessions for the Postgres layer. We need a project-wide rule for how those
resources, and future resources with similar lifetimes, are created, shared, and
closed.
The service will need objects with different lifetimes:
- application-lifetime objects such as SQLAlchemy engines, session factories,
Qdrant clients, LangGraph checkpointers/stores, compiled graphs, HTTP clients,
observability clients, model clients, Redis clients, or other connection pools;
- request/job-lifetime objects such as SQLAlchemy sessions, transactions,
authentication context, tenant context, request ids, and unit-of-work state;
- operation-lifetime objects such as temporary files, one-off streams, locks, or
short-lived connections checked out for a single block.
Using module-level mutable resource instances for everything would make startup,
shutdown, testing, concurrency, and transaction boundaries hard to reason about.
The most dangerous example is a global SQLAlchemy `Session`: sessions track
identity state and transaction state, and are not safe to share across concurrent
requests.
At the same time, creating expensive clients or connection pools for every
request would waste resources and lose pooling benefits. We need explicit
ownership: the component that creates a resource is responsible for closing it.
## Decision
### Use FastAPI lifespan for application-lifetime resources
Create long-lived shared resources in the FastAPI lifespan function and close
them after the lifespan `yield`.
Application-lifetime resources include:
- SQLAlchemy async `Engine`;
- SQLAlchemy `async_sessionmaker` bound to that engine;
- Qdrant client;
- LangGraph checkpointer/store and compiled graph;
- shared `httpx.AsyncClient` instances;
- Redis or cache clients;
- Langfuse/observability clients;
- embedding, reranking, LLM, or other expensive model clients.
The lifespan function is the owner of these resources. It is responsible for
calling cleanup methods such as `dispose()`, `aclose()`, `close()`, or equivalent
shutdown hooks.
Indicative shape:
```python
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
import httpx
from fastapi import FastAPI
@asynccontextmanager
async def lifespan(app: FastAPI) -> AsyncIterator[None]:
app.state.database = Database(settings.database_url)
app.state.http_client = httpx.AsyncClient()
try:
yield
finally:
await app.state.http_client.aclose()
await app.state.database.close()
app = FastAPI(lifespan=lifespan)
```
Avoid opening external resource clients as import-time side effects. Module-level
constants, settings objects, pure functions, type aliases, and stateless helpers
are acceptable. Mutable network/database clients should be constructed by the
application owner, not by importing a module.
### Treat SQLAlchemy engine and session as different lifetimes
A SQLAlchemy `Engine` is an application-lifetime pool manager, not a single
request's transaction. It may keep database connections open in a pool and reuse
them across requests.
A SQLAlchemy `AsyncSession` is a request/job-lifetime unit of work. It tracks ORM
identity state and transaction state and may check out a database connection from
the engine when needed.
Therefore:
- create one engine/session factory per FastAPI process during lifespan;
- create one `AsyncSession` per request or background-job unit of work;
- never share one `AsyncSession` globally across concurrent requests;
- never run DDL such as `create_all()` at FastAPI startup; use Alembic as decided
in ADR-0009.
### Use dependencies for request-lifetime resources
Use FastAPI dependencies, usually `yield` dependencies, for objects that should
exist for one request and then be cleaned up.
Indicative database session dependency:
```python
from collections.abc import AsyncIterator
from fastapi import Request
from sqlalchemy.ext.asyncio import AsyncSession
async def get_db_session(request: Request) -> AsyncIterator[AsyncSession]:
database: Database = request.app.state.database
async with database.session_factory() as session:
try:
yield session
except Exception:
await session.rollback()
raise
```
The dependency owns the session lifetime. Route handlers, services, and
repositories must not close a session they did not create.
Authentication and tenant context follow the same rule: resolve them through
request dependencies, then pass the trusted context into code that needs it.
### Keep transaction boundaries explicit
Routes or application-service functions own transaction boundaries. Repository
and CRUD functions receive a session and perform database work, but they should
not secretly commit, rollback, or close the session unless their contract
explicitly says they own a complete unit of work.
Preferred shape:
```python
async def create_user(session: AsyncSession, data: CreateUserRequest) -> User:
user = User(email=data.email)
session.add(user)
return user
@router.post("/users")
async def create_user_endpoint(
body: CreateUserRequest,
session: SessionDep,
) -> UserResponse:
user = await create_user(session, body)
await session.commit()
return UserResponse.model_validate(user)
```
Avoid hidden global access:
```python
# Do not do this.
session = SessionLocal()
async def create_user(data: CreateUserRequest) -> User:
session.add(User(email=data.email))
await session.commit()
```
This keeps unit-of-work ownership visible, makes multi-step transactions easier,
and lets tests pass in their own sessions.
### Use context managers for operation-lifetime resources
Objects needed only inside one function or block should use `with` or
`async with` rather than app state or module globals.
Examples include file handles, temporary streams, one-off SQLAlchemy connections,
locks, and response streams.
```python
async with http_client.stream("GET", url) as response:
...
```
or:
```python
async with engine.connect() as connection:
...
```
The context manager owns and releases the resource at block exit.
### Pass dependencies explicitly through service boundaries
Routers adapt HTTP requests into typed application calls. Services and
repositories should receive the resources they use as explicit parameters:
```python
async def get_point(
session: AsyncSession,
qdrant: QdrantClient,
auth: AuthContext,
point_id: str,
) -> PointResponse:
...
```
Do not have lower layers import mutable resource singletons. Explicit parameters
make ownership clear and allow FastAPI dependency overrides, test fixtures, and
background workers to supply equivalent resources.
### Account for worker process multiplicity
FastAPI lifespan runs once per worker process. If the service runs with multiple
Uvicorn/Gunicorn workers, each worker has its own app instance, engine, and
connection pool. Database pool sizes and external client limits must be chosen
with worker count in mind.
## Consequences
### Positive
- Resource ownership and cleanup become explicit: lifespan closes app resources,
dependencies close request resources, and context managers close block-scoped
resources.
- SQLAlchemy sessions are not shared across concurrent requests, avoiding unsafe
identity-map and transaction-state reuse.
- Connection pooling is preserved because expensive clients and engines are
created once per process instead of once per request.
- Transaction boundaries are easier to reason about because routes or
application services decide when to commit or rollback.
- Services and repositories are easier to test because sessions, clients, and
auth contexts can be injected directly.
- The same policy can be reused for future Redis, Qdrant, Langfuse, HTTP, model,
and worker resources.
### Negative
- This adds boilerplate: lifespan setup, app state typing/conventions,
dependencies, and explicit parameters must be maintained.
- Developers must distinguish engine/session, application/request, and
owner/borrower lifetimes instead of importing a convenient global.
- App state access needs discipline and typing helpers so resources do not become
an unstructured service locator.
- Multi-worker deployments require explicit pool sizing because each process owns
its own pools and clients.
## Alternatives Considered
- **Module-level resource singletons**: rejected as the default for mutable
external resources. Import-time construction makes startup order, tests,
reloads, and shutdown cleanup harder to control. Module-level constants and
stateless helpers remain acceptable.
- **One global SQLAlchemy session**: rejected. A session is a unit of work with
identity and transaction state, not an application-wide pool. Sharing it across
requests is unsafe and makes rollback/error handling ambiguous.
- **Create every client/session inside each CRUD function**: rejected. It hides
ownership and transaction boundaries, prevents multi-step units of work, and
wastes pooling benefits for expensive clients.
- **Automatically commit every successful request in the session dependency**:
deferred as a default. It reduces route boilerplate, but it can make write
boundaries too implicit and does not fit every read/write or multi-transaction
flow. Routes or application services should own commits unless a future ADR
adopts a unit-of-work abstraction.
- **Use a generic global service locator**: rejected for now. It can reduce
parameter lists, but it obscures dependencies and ownership. FastAPI
dependencies plus explicit function parameters are clearer at this stage.

View File

@@ -0,0 +1,182 @@
# 0013. S3-compatible object storage with MinIO
## Status
Proposed
## Context
ADR-0008 defines source-file upload endpoints and returns a job-shaped ingestion
response. ADR-0009 models `source_files` and leaves room for a `storage_uri`
where original files are retained. The service now needs a clear place to store
uploaded file bytes and any derived ingestion artifacts without putting large
binary/text payloads in Postgres or broker messages.
The storage decision has to satisfy several constraints:
- Uploaded files may be larger than is appropriate for Postgres rows.
- Ingestion workers need to fetch the exact file bytes after the HTTP request has
completed.
- Postgres should remain the source of truth for tenant metadata, file metadata,
ingestion job state, audit records, and retention decisions.
- Broker messages should carry durable identifiers, not raw file content.
- Local/self-hosted deployment should stay practical and avoid coupling this app
to Langfuse's internal object-storage services.
## Decision
Use **MinIO** as the application's S3-compatible object store for uploaded source
files and derived ingestion blobs.
MinIO stores bytes. Postgres stores metadata and authoritative state.
### Store source-file bytes in MinIO
When `POST /v1/files` receives an upload, the service stores the raw file in an
application-owned MinIO bucket before or within the same application flow that
creates the `source_files` and `ingestion_jobs` rows.
`source_files` keeps the durable pointer and metadata, including fields already
sketched in ADR-0009:
- `tenant_id`;
- `domain`;
- `source_filename`;
- `source_type`;
- `content_sha256`;
- `byte_size`;
- `storage_uri` or equivalent bucket/key fields;
- lifecycle status and audit timestamps.
Object keys are internal identifiers, not trusted user filenames. Prefer a shape
like:
```text
tenants/{tenant_id}/source-files/{file_id}/original
```
or, when content-addressing is useful:
```text
tenants/{tenant_id}/source-files/{file_id}/{content_sha256}
```
The original filename is stored as metadata in Postgres, not used as the object
key authority.
### Keep MinIO separate from authoritative state
MinIO is not the source of truth for:
- tenant identity or authorization;
- file ownership;
- ingestion job status;
- ingestion progress events;
- audit logs;
- LangGraph checkpoints or Store memory;
- Qdrant point state;
- broker delivery state.
Workers always load `source_files` and `ingestion_jobs` from Postgres by id, then
use the stored object pointer to fetch bytes from MinIO.
### Use MinIO for retained blobs, not queue payloads
Broker messages reference ids such as `file_id` and `ingestion_job_id`; they do
not contain raw file bytes, extracted text, chunks, embeddings, or large parser
outputs.
MinIO may also store derived blobs when retention policy allows it, for example:
- converted `.doc` -> `.docx` outputs from ADR-0004;
- extracted text snapshots used for debugging failed ingestion;
- parser diagnostics or quarantined uploads;
- exported artifacts generated by maintenance jobs.
Those derived objects still need Postgres metadata if they are user-visible,
auditable, or subject to retention/erasure policy.
### Use application-owned buckets and credentials
The app must not casually reuse Langfuse's internal MinIO bucket, credentials, or
lifecycle. Langfuse object storage belongs to the Langfuse stack from ADR-0010;
this service needs its own application bucket(s) and credentials.
A local deployment may run MinIO in the same Docker environment, but with a
separate bucket such as:
```text
chatbot-source-files
```
Production may use MinIO or another S3-compatible service behind the same object
storage interface, provided it preserves tenant isolation, encryption, backup,
and retention requirements.
### Manage the client as an application-lifetime resource
The object-storage client is an application-lifetime external client under
ADR-0012. It is created by the FastAPI lifespan or worker process startup and
closed on shutdown when the client library requires explicit cleanup.
Lower-level ingestion services receive an object-storage client or interface
explicitly. They do not import a module-level mutable MinIO singleton.
For async code, object-storage operations must not block the event loop. Use an
async-capable S3 client, or isolate synchronous SDK calls in the appropriate
threadpool boundary.
### Enforce privacy, retention, and integrity rules
- Compute and store `content_sha256` for uploaded bytes.
- Enforce allowed content types/extensions and maximum file size before retaining
files, as required by ADR-0008.
- Prefer private buckets; generate short-lived presigned URLs only for explicit
internal workflows that need them.
- Apply server-side encryption and backups according to deployment policy.
- Deletion and tenant-erasure workflows must delete both Postgres metadata and
MinIO objects, while preserving audit records according to tenant/legal
retention requirements.
## Consequences
### Positive
- Large uploaded files and derived blobs stay out of Postgres rows and broker
payloads.
- Ingestion can run in separate workers after the HTTP request finishes because
workers can fetch the retained object by durable id.
- The service keeps S3-compatible portability while using MinIO for local and
self-hosted deployments.
- `source_files.storage_uri` from ADR-0009 now has a concrete storage backend.
- Separating app MinIO from Langfuse MinIO avoids accidental coupling to
Langfuse's internal lifecycle and credentials.
### Negative
- Adds another external dependency to configure, secure, back up, and monitor.
- Retention and tenant erasure now have to coordinate Postgres rows, Qdrant
points, and MinIO objects.
- Upload flows need cleanup handling for partial failures, such as object upload
succeeding but the Postgres transaction failing.
- Async FastAPI handlers need care if the chosen S3/MinIO client is synchronous.
- Object keys and bucket policies become part of the security boundary; mistakes
can expose cross-tenant files.
## Alternatives Considered
- **Store uploaded files directly in Postgres**: rejected. It simplifies
transactional consistency, but bloats the relational database, complicates
backups, and mixes large binary/blob storage with metadata and audit queries.
- **Store only extracted chunks in Qdrant and discard originals immediately**:
rejected as the default. It makes failed-ingestion debugging, reprocessing,
parser upgrades, and retention review harder. Tenants may still configure
shorter retention later.
- **Use the broker message as file transport**: rejected. Brokers should carry
small work identifiers and metadata; raw files, extracted text, chunks, and
embeddings belong in object storage or the appropriate durable store.
- **Reuse Langfuse's internal MinIO service/bucket**: rejected. Langfuse object
storage is part of the observability stack, not this application's file store.
Sharing it would couple credentials, retention, backups, and operational
lifecycle across unrelated systems.
- **Bind directly to a cloud-only S3 provider**: rejected for now. The project is
self-hosted/local-development oriented, and MinIO gives an S3-compatible API
while keeping deployment portable.

View File

@@ -0,0 +1,282 @@
# 0014. Durable job dispatch with NATS JetStream
## Status
Proposed
## Context
ADR-0008 makes file ingestion job-shaped: `POST /v1/files` returns `202 Accepted`
with `file_id`, `ingestion_job_id`, and a queued/running status. ADR-0009 makes
Postgres the durable source of truth for `source_files`, `ingestion_jobs`, and
`ingestion_job_events`. ADR-0013 stores uploaded file bytes in MinIO so ingestion
workers can fetch them after the request completes.
What is still missing is the delivery mechanism that tells workers which durable
jobs are ready to process.
The broker decision has to preserve existing boundaries:
- Postgres remains authoritative for job status, progress, audit, tenancy, and
retention-sensitive state.
- MinIO stores file bytes and derived blobs, not queue state.
- Broker messages carry ids and correlation metadata, not raw files or chunks.
- Chat runs are not queued by this service: ADR-0007/0008 allow one in-flight run
per `thread_id` and return `409 Conflict` for concurrent runs.
- Application-lifetime clients are created and closed by lifespan/startup owners
under ADR-0012, not import-time globals.
## Decision
Use **NATS JetStream** as the durable pub/sub and job-dispatch broker for
asynchronous ingestion and maintenance work.
NATS core pub/sub alone is not enough for ingestion dispatch because ingestion
jobs need durable delivery, acknowledgement, redelivery, and worker restart
recovery. JetStream provides those broker-level mechanics while Postgres remains
the source of truth for application state.
### Start with ingestion jobs
The first broker-backed workflow is source-file ingestion:
```text
POST /v1/files
-> authenticate and resolve tenant
-> store source bytes in MinIO (ADR-0013)
-> create/update source_files row
-> create ingestion_jobs row with status='queued'
-> commit the transaction
-> publish a JetStream message containing ingestion_job_id
-> return 202 Accepted
```
A dedicated ingestion worker process consumes the message:
```text
JetStream message
-> load ingestion_jobs/source_files from Postgres
-> bind logging context from durable ids
-> fetch source object from MinIO
-> parse, chunk, embed, and mutate Qdrant points
-> append ingestion_job_events
-> mark ingestion_jobs succeeded/failed/cancelled
-> ack or negatively acknowledge the message
```
### Worker-owned ingestion chunk CRUD
For file ingestion and re-ingestion, generated chunk/point CRUD is performed by
the ingestion worker, not by the FastAPI publisher/request path.
The publisher only owns the lightweight durable handoff:
- authenticate the caller and derive tenant context;
- store uploaded bytes in MinIO;
- create or update `source_files`;
- create `ingestion_jobs(status='queued')`;
- commit the Postgres transaction;
- publish the JetStream message containing durable ids.
The ingestion worker owns the expensive and retryable side effects:
- parsing source files;
- chunking;
- embedding and reranking-vector generation when applicable;
- creating, updating, reordering, and soft-deleting generated Qdrant points;
- appending `ingestion_job_events`;
- updating ingestion job status, counters, and failure summaries.
Shared chunk/point mutation logic should live in a service/repository layer that
can be reused by both direct API routes and ingestion workers. Direct
`/v1/points` CRUD remains a synchronous API responsibility unless the operation
is explicitly bulk or job-shaped. Ingestion-generated mutations are asynchronous
worker responsibility because they must be safe to retry and resume by
`ingestion_job_id`.
### Keep messages small and versioned
Broker payloads carry identifiers and correlation metadata only. A typical
message is:
```json
{
"type": "ingestion.job.created",
"version": 1,
"tenant_id": "...",
"file_id": "...",
"ingestion_job_id": "...",
"request_id": "...",
"api_key_id": "..."
}
```
Do not put raw file bytes, extracted text, chunks, embeddings, raw prompts, model
outputs, or secrets in JetStream messages.
Workers must treat message metadata as routing/correlation input, not as the only
authority. Before doing tenant-scoped work, workers reload the durable job and
file rows from Postgres and verify the ids are consistent.
### Use explicit subjects, streams, and durable consumers
Use stable, dot-separated subjects. Initial subjects:
```text
ingestion.job.created
ingestion.job.retry_requested
maintenance.retention.requested
maintenance.erasure.requested
```
Create an application-owned JetStream stream for durable work messages, for
example:
```text
CHATBOT_JOBS
subjects: ingestion.job.*, maintenance.*
```
Workers use durable consumers or queue groups so multiple worker processes can
share work without processing every message independently.
Ack messages only after the worker has persisted the resulting job state and
progress events to Postgres. Redelivery must be safe: ingestion processors should
be idempotent by `ingestion_job_id` and Qdrant point ids/upsert semantics.
### Run subscribers as worker processes, not FastAPI side effects
FastAPI may create a JetStream publisher/client during application lifespan so
routes can publish job notifications after durable state is committed.
Subscribers should run in separate worker process entrypoints. Do not start a
subscriber inside every FastAPI web worker by default: with multiple
Uvicorn/Gunicorn workers, each process has its own lifespan and could create
surprising duplicate consumers or resource pressure.
Worker processes create their own application-lifetime resources at startup:
- SQLAlchemy engine/sessionmaker;
- MinIO/S3 object-storage client;
- Qdrant client;
- NATS/JetStream client;
- observability/logging clients;
- ingestion model clients.
Each message gets its own request/job-lifetime SQLAlchemy `AsyncSession` and
explicit transaction boundaries, following ADR-0012.
### Handle commit/publish consistency deliberately
The initial implementation may publish the JetStream message after the Postgres
transaction commits. This avoids workers observing uncommitted jobs, but it means
a process crash between commit and publish can leave a job in `queued` state with
no broker message.
To handle that failure mode, workers or a small scheduler should periodically
scan Postgres for queued/stale ingestion jobs and publish or process them again.
This keeps Postgres authoritative and makes JetStream a wakeup/delivery
mechanism rather than the only source of pending work.
If stricter guarantees become necessary, introduce a transactional outbox table
in Postgres via a future ADR/migration:
```text
same DB transaction:
insert/update ingestion_jobs
insert outbox_events
publisher process:
read unpublished outbox_events
publish to JetStream
mark outbox event published
```
Do not add DDL at FastAPI startup; all outbox schema changes must go through
Alembic as required by ADR-0009.
### Configure retries, dead letters, and observability
JetStream consumers should configure explicit acknowledgement, redelivery, and
maximum delivery behavior. Poison messages or repeatedly failing jobs should end
as durable Postgres failures with an `ingestion_job_events` error entry, not an
infinite invisible retry loop.
At worker ingress, bind structured logging context from durable identifiers:
- `tenant_id`;
- `request_id`;
- `api_key_id`;
- `file_id`;
- `ingestion_job_id`;
- JetStream stream/subject/sequence when available.
Use stable event names such as:
- `ingestion.job.dispatched`;
- `ingestion.job.received`;
- `ingestion.job.completed`;
- `ingestion.job.failed`;
- `broker.publish.failed`;
- `broker.message.redelivered`.
### Do not queue chat runs through JetStream yet
ADR-0007 and ADR-0008 intentionally reject queueing concurrent chat runs per
thread: one run may be in flight, and concurrent runs return `409 Conflict`.
This ADR does not change that decision.
A future ADR may introduce broker-backed work for chat-adjacent tasks such as
summarization, retention, erasure, or offline evaluation, but synchronous chat run
execution remains owned by the FastAPI/LangGraph boundary unless superseded.
## Consequences
### Positive
- File ingestion can move out of HTTP request handling into scalable worker
processes without changing the REST contract from ADR-0008.
- JetStream gives durable delivery, ack/redelivery, and worker-group semantics
without adopting Kafka-level operational complexity.
- Postgres remains the auditable source of truth for jobs, progress, tenant
ownership, and failure state.
- MinIO, Postgres, Qdrant, and the broker each have distinct responsibilities:
bytes, metadata/state, vector search, and work delivery.
- Separate worker processes avoid surprising subscriptions in every FastAPI web
worker and make GPU/model-heavy ingestion resources easier to size.
### Negative
- Adds another external service to run, secure, monitor, and back up where stream
durability is required.
- Developers must reason about redelivery and idempotency; handlers may run more
than once for the same `ingestion_job_id`.
- Publishing after commit has a crash window unless a stale-job scanner or future
transactional outbox is implemented.
- NATS/JetStream stream, consumer, retention, and dead-letter configuration must
be managed explicitly per environment.
- The app now needs worker deployment/runbook conventions in addition to the
FastAPI web process.
## Alternatives Considered
- **NATS core pub/sub only**: rejected. Core pub/sub is useful for ephemeral
events, but ingestion dispatch needs durable delivery and acknowledgement.
JetStream is the NATS feature that provides those semantics.
- **Redis Streams**: rejected for this decision. Redis Streams would work for an
MVP and has simple local deployment, but NATS JetStream is a cleaner dedicated
broker for durable pub/sub and worker coordination without also mixing cache
responsibilities into the same service.
- **RabbitMQ**: rejected for now. It is mature for work queues and routing, but
the project does not currently need its exchange/routing complexity, and NATS
JetStream keeps the broker footprint lighter.
- **Kafka**: rejected. Kafka is an excellent durable event log at high scale, but
its operational overhead is unnecessary for this service's initial ingestion
and maintenance jobs.
- **FastAPI `BackgroundTasks` or in-process asyncio tasks**: rejected as the
durable design. They are simple, but work can be lost on process restart and
they do not coordinate multiple workers cleanly.
- **Database polling only**: rejected as the main dispatch mechanism. Polling
Postgres for queued jobs is a useful repair path, but relying on polling alone
adds latency and unnecessary database load once a broker is available.
- **Use the broker as the source of truth for job state**: rejected. Broker state
is delivery state. Application job state, progress, audit, and tenant ownership
remain in Postgres.