docs(data): define Postgres schema and migration conventions

Why:
- Establish the application-owned relational source of truth for tenants, authentication, ingestion, audit, graph runs, LLM usage, and feedback.

Changes:
- Define SQLAlchemy 2.x and Alembic conventions.
- Specify tenant, API-key, ingestion, audit, graph-run, pricing, usage, and feedback tables.
- Document indexing, retention, privacy, and multitenancy rules.

Impact:
- Postgres schema changes must be implemented through Alembic migrations.
- FastAPI must not run DDL during startup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-10 11:45:19 +03:30
parent 50f6c33c0c
commit ed96f20ebf

View File

@@ -0,0 +1,444 @@
# 0009. Postgres schema with SQLAlchemy 2 and Alembic
## Status
Proposed
## Context
ADR-0008 defines the FastAPI REST boundary: API-key authentication, tenant
resolution from Postgres, source-file ingestion, point management, and chat
thread runs. ADR-0007 defines LangGraph checkpoint persistence and explicitly
avoids making this service the source of truth for chat sessions. We now need
the application-owned Postgres schema that supports those decisions.
The schema has to serve several purposes at once:
- **Multitenancy**: the chatbot serves multiple tenants; every API key belongs
to one tenant and every request resolves to that tenant before touching
Qdrant or LangGraph.
- **Authentication**: each tenant needs API keys. In practice this should mean
*one or more* API keys per tenant, so keys can be rotated, scoped, and
revoked independently.
- **Ingestion records**: when a caller uploads a source file (`.csv`, `.xlsx`,
`.docx`, legacy `.doc` via conversion), the service records the file and the
ingestion job that parsed, chunked, embedded, and wrote Qdrant points.
- **Point/API audit**: point CRUD and batch operations mutate Qdrant, which is
not an audit-log database. The service needs its own record of who changed
what and when.
- **LLM usage and cost**: every model call inside the graph (triage,
contextualize, grade, generate, verify, summary/memory extraction) should be
attributable to a tenant, thread, run, graph node, model, token counts, and
price. For debugging/evals, the system also needs a controlled way to keep
the model input and output.
The project uses SQLAlchemy 2.x and Alembic. The schema should therefore be
specified in SQLAlchemy 2 style (`DeclarativeBase`, `Mapped[...]`,
`mapped_column(...)`) and migrated only through Alembic — never `create_all()`
at FastAPI startup.
## Decision
### SQLAlchemy and migration conventions
Use SQLAlchemy 2.x ORM models with typed mappings:
```python
class Base(DeclarativeBase):
pass
class Tenant(Base):
__tablename__ = "tenants"
id: Mapped[uuid.UUID] = mapped_column(primary_key=True)
slug: Mapped[str] = mapped_column(String(80), unique=True, index=True)
```
Conventions:
- Use SQLAlchemy async sessions in FastAPI (`AsyncSession`) and one session per
request/job unit of work.
- Use Alembic for all DDL. FastAPI startup opens connections and checks
readiness; it does not create or alter tables.
- Prefer UUID primary keys generated by the application. Avoid integer IDs that
leak tenant size and make distributed workers harder to compose.
- Use `timestamptz`/`DateTime(timezone=True)` for all timestamps.
- Use `Numeric(18, 8)` or finer for monetary/cost fields; never floats for
money.
- Store flexible metadata in `JSONB`, but keep relational identifiers and
query-critical fields as typed columns with indexes. If the database column
is named `metadata`, map it with a safe SQLAlchemy attribute such as
`metadata_ = mapped_column("metadata", JSONB, ...)` because `metadata` is
reserved on Declarative models.
- Use string status columns with SQLAlchemy/Pydantic enums and database
`CHECK` constraints rather than PostgreSQL native enums. Status sets change
often during early product work, and native enum migrations are painful.
- Every tenant-owned table has `tenant_id` and an index beginning with
`tenant_id`. Foreign keys include `ondelete` behaviour deliberately, not by
accident.
Postgres uses a shared schema with tenant foreign keys, not database-per-tenant
or schema-per-tenant. This matches ADR-0001's shared Qdrant collection and
keeps tenant count scalable.
### Core tenant and API-key tables
#### `tenants`
One row per customer/tenant.
| Column | Notes |
|---|---|
| `id` | UUID PK. Used as the canonical `tenant_id` injected into Qdrant filters and LangGraph config. |
| `slug` | Stable short name, unique, human-readable. |
| `name` | Display name. |
| `status` | `active` \| `suspended` \| `deleted`. Suspended tenants authenticate to a clear error but cannot run work. |
| `settings` | JSONB for tenant-level feature flags/limits (max upload size, enabled file types, allowed domains, etc.). |
| `created_at`, `updated_at`, `deleted_at` | Audit/soft-delete timestamps. |
#### `tenant_domains`
Optional but recommended. Validates the `domain` values used throughout Qdrant
payloads (`car`, `fire`, etc.) per tenant.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `tenant_id` | FK to `tenants.id`. |
| `domain` | Tenant-local domain key. Unique with `tenant_id`. |
| `display_name` | Human-readable label. |
| `status` | `active` \| `disabled`. |
| `metadata` | JSONB for domain-specific ingestion/retrieval settings. |
This prevents arbitrary caller-supplied domains from silently creating new
partitions in Qdrant.
#### `api_keys`
One tenant can have multiple active keys for rotation and scoped access.
| Column | Notes |
|---|---|
| `id` | UUID PK, also usable as the API key lookup id/prefix. |
| `tenant_id` | FK to `tenants.id`. |
| `name` | Human label, e.g. `main-backend-prod`. |
| `key_prefix` | Short non-secret prefix shown in logs/admin UI, unique. |
| `key_hash` | Hash of the secret key material. Plaintext API keys are never stored. |
| `scopes` | JSONB or text array: `threads:run`, `points:read`, `points:write`, `files:write`, `memory:read`, `admin`. |
| `actor_type` | `backend` \| `admin` \| `worker`. |
| `status` | `active` \| `revoked` \| `expired`. |
| `expires_at`, `revoked_at`, `last_used_at` | Lifecycle timestamps. |
| `created_by`, `created_at`, `updated_at` | Audit fields. |
Authentication dependency in ADR-0008 queries by `key_prefix`/id, verifies
`key_hash` with constant-time comparison, checks tenant/key status and scopes,
and returns `AuthContext`.
### Request and mutation audit tables
#### `api_request_logs`
Append-only request log for authenticated `/v1` calls. This is not a
replacement for structured application logs; it is the durable queryable audit
record.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `tenant_id` | Denormalized from API key for fast tenant queries. |
| `api_key_id` | FK to `api_keys.id`, nullable only for failed auth where key is unknown. |
| `request_id` | Correlation id, unique. |
| `method`, `path_template`, `status_code` | Route identity and result. |
| `scopes_required` | JSONB/text array. |
| `external_user_id` | User id supplied by the main backend, when present. |
| `thread_id` | Present for thread/run routes. No FK — ADR-0007 says this service owns no thread table. |
| `source_ip_hash`, `user_agent` | Optional operational metadata; avoid storing raw IP unless required. |
| `request_summary`, `response_summary` | JSONB summaries, not raw bodies by default. |
| `error_code` | Stable error code from ADR-0008, nullable. |
| `duration_ms`, `created_at` | Timing. |
This table records that an API call happened. Tables below record domain-level
side effects.
#### `point_audit_events`
Append-only audit of `/v1/points` and `/v1/files/{file_id}/points` mutations.
Qdrant remains the storage/search engine; this table records mutation intent
and result.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `tenant_id` | FK to `tenants.id`. |
| `api_request_log_id` | FK to `api_request_logs.id`. |
| `api_key_id` | FK to `api_keys.id`. |
| `operation` | `create` \| `update` \| `payload_patch` \| `soft_delete` \| `hard_delete` \| `reorder` \| `batch`. |
| `point_id` | Qdrant point id, nullable for batch/file-wide operations. |
| `file_id` | Source file id when relevant. |
| `domain` | Qdrant payload domain. |
| `before_version`, `after_version` | Optimistic-concurrency versions when available. |
| `changed_fields` | JSONB list/summary; no large content or vectors. |
| `qdrant_operation_id`, `qdrant_status` | Result returned by Qdrant, if available. |
| `created_at` | Event time. |
### File and ingestion tables
#### `source_files`
One logical source document uploaded by a tenant. Re-ingestion of the same file
creates new jobs against the same or replacement `source_files` row depending
on the `content_hash` policy.
| Column | Notes |
|---|---|
| `id` | UUID PK; this is the `file_id` copied into Qdrant point payloads. |
| `tenant_id` | FK to `tenants.id`. |
| `domain` | Tenant domain, validated by `tenant_domains` where enabled. |
| `source_filename` | Original filename. |
| `source_type` | `csv` \| `xlsx` \| `docx` \| `doc`. |
| `content_sha256` | Hash of the uploaded file bytes for idempotency/change detection. |
| `byte_size` | Upload size. |
| `storage_uri` | Where the original file is stored, if retained. Nullable if not retaining originals. |
| `status` | `active` \| `superseded` \| `soft_deleted` \| `purged`. |
| `created_by_api_key_id`, `created_at`, `updated_at`, `deleted_at` | Audit fields. |
#### `ingestion_jobs`
One attempt to parse/chunk/embed/upsert a source file. This table is required
even if the first implementation processes inline, because ADR-0008's file
upload contract is job-shaped.
| Column | Notes |
|---|---|
| `id` | UUID PK; returned as `ingestion_job_id`. |
| `tenant_id` | FK to `tenants.id`. |
| `source_file_id` | FK to `source_files.id`. |
| `api_request_log_id` | FK to the upload request log. |
| `requested_by_api_key_id` | FK to `api_keys.id`. |
| `status` | `queued` \| `running` \| `succeeded` \| `failed` \| `cancelled`. |
| `chunking_strategy` | `semantic` \| `fixed_size`; matches ADR-0004. |
| `embedding_model_versions` | JSONB map of vector name → model/version. |
| `started_at`, `completed_at` | Lifecycle timestamps. |
| `points_created`, `points_updated`, `points_soft_deleted`, `points_skipped` | Result counters. |
| `error_code`, `error_message` | Failure summary. |
| `metadata` | JSONB for parser/chunker options. |
| `created_at`, `updated_at` | Audit timestamps. |
#### `ingestion_job_events`
Append-only progress/error stream for a job.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `tenant_id` | FK to `tenants.id`. |
| `ingestion_job_id` | FK to `ingestion_jobs.id`. |
| `level` | `info` \| `warning` \| `error`. |
| `stage` | `received` \| `parsed` \| `chunked` \| `embedded` \| `upserted` \| `completed`. |
| `message` | Short human-readable event. |
| `details` | JSONB structured details. |
| `created_at` | Event time. |
### Graph run, LLM usage, and feedback tables
#### `graph_runs`
One row per `POST /v1/threads/{thread_id}/runs`. This is not a session/thread
table: it records one execution for audit, feedback, usage aggregation, and
cost reporting.
| Column | Notes |
|---|---|
| `id` | UUID PK; this is the `run_id` returned by the run endpoint. |
| `tenant_id` | FK to `tenants.id`. |
| `api_request_log_id` | FK to `api_request_logs.id`. |
| `api_key_id` | FK to `api_keys.id`. |
| `thread_id` | LangGraph thread id from the path. No FK. |
| `external_user_id` | User id supplied by the main backend. |
| `status` | `running` \| `answered` \| `clarifying` \| `escalate` \| `failed`. |
| `escalation_reason` | ADR-0006 reason when `status='escalate'`. |
| `input_message_hash` | Hash for idempotency/debug correlation without storing raw text here. |
| `output_message_hash` | Hash of final answer/clarifying/escalation message. |
| `llm_input_tokens`, `llm_output_tokens`, `llm_total_cost` | Denormalized totals from `llm_calls`. |
| `started_at`, `completed_at`, `duration_ms` | Timing. |
| `metadata` | JSONB for graph version, prompt version, retrieved chunk IDs, etc. |
`graph_runs` solves the practical problem left by ADR-0007's feedback endpoint:
feedback needs a stable `run_id`, but this service still does not need a table
that represents chat sessions.
#### `llm_pricing`
Versioned model pricing table so historical cost calculations remain
explainable when model prices change.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `provider` | `anthropic` \| `openai` \| other. |
| `model` | Provider model id. |
| `currency` | Usually `USD`. |
| `input_price_per_1m_tokens`, `output_price_per_1m_tokens` | Numeric. |
| `effective_from`, `effective_to` | Time-bounded price validity. |
| `created_at` | Audit timestamp. |
#### `llm_calls`
One row per provider model call made inside the graph or ingestion pipeline.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `tenant_id` | FK to `tenants.id`. |
| `graph_run_id` | FK to `graph_runs.id`, nullable for ingestion-time LLM calls such as image extraction. |
| `ingestion_job_id` | FK to `ingestion_jobs.id`, nullable for chat-time calls. |
| `api_request_log_id` | FK to the originating request when available. |
| `thread_id`, `external_user_id` | Denormalized for query convenience; nullable outside chat. |
| `node_name` | `triage`, `contextualize`, `grade`, `generate`, `verify`, `summarize`, `memory_extract`, `image_extract`, etc. |
| `provider`, `model`, `model_version` | Provider identity. |
| `pricing_id` | FK to `llm_pricing.id`, nullable if price was configured externally. |
| `input_tokens`, `output_tokens`, `total_tokens` | Provider usage numbers. |
| `input_cost`, `output_cost`, `total_cost`, `currency` | Cost at call time. |
| `latency_ms` | Provider round-trip. |
| `status` | `succeeded` \| `failed` \| `cancelled`. |
| `error_code`, `error_message` | Failure summary. |
| `prompt_version`, `schema_version` | Version of prompt/structured-output schema used. |
| `input_hash`, `output_hash` | Hashes of stored/redacted payloads. |
| `created_at` | Call start time. |
Costs are computed and stored at call time from `llm_pricing` (or explicit
runtime pricing config), not recomputed later from a mutable current price.
#### `llm_call_payloads`
Stores the actual model input/output only when allowed by tenant policy. This
is deliberately separate from `llm_calls` so usage/billing queries never touch
large or sensitive payloads.
| Column | Notes |
|---|---|
| `llm_call_id` | PK/FK to `llm_calls.id`. |
| `tenant_id` | FK to `tenants.id`, repeated for partition/index convenience. |
| `input_redacted` | JSONB/text redacted prompt/messages/tool input. |
| `output_redacted` | JSONB/text redacted model output/tool call result. |
| `input_encrypted`, `output_encrypted` | Optional encrypted raw payload bytes/text if raw retention is enabled. |
| `redaction_version` | Which redaction policy produced the redacted fields. |
| `retention_until` | When payloads must be deleted, independent of usage rows. |
| `created_at` | Timestamp. |
Default policy: store token counts/costs for every call, store **redacted**
input/output for debugging/evals, and store raw encrypted payloads only for
tenants that explicitly enable it. Insurance chat can contain PII and sensitive
claim/coverage information; raw prompt logging cannot be an accidental default.
#### `run_feedback`
Feedback from `POST /v1/threads/{thread_id}/runs/{run_id}/feedback`.
| Column | Notes |
|---|---|
| `id` | UUID PK. |
| `tenant_id` | FK to `tenants.id`. |
| `graph_run_id` | FK to `graph_runs.id`. |
| `external_user_id` | User id from main backend, if present. |
| `rating` | `thumbs_up` \| `thumbs_down` \| numeric score. |
| `reason_codes` | JSONB/text array. |
| `comment` | Optional free text. |
| `created_at` | Timestamp. |
### What is intentionally not modeled
- **No `threads`/`sessions` table.** ADR-0007 remains in force: the main
backend owns session records and LangGraph owns thread checkpoints. This
schema records runs and usage, not conversation ownership.
- **No Postgres copy of Qdrant point content/vectors.** Qdrant remains the
source of truth for point payloads/vectors. Postgres stores source-file,
ingestion, and audit records.
- **No plaintext API keys.** Only hashes and non-secret prefixes.
- **No automatic raw prompt retention.** Raw LLM input/output is opt-in,
encrypted, and retention-limited.
### Indexing and retention
Required indexes:
- `api_keys(key_prefix)` unique; `api_keys(tenant_id, status)`.
- `tenant_domains(tenant_id, domain)` unique.
- `api_request_logs(tenant_id, created_at desc)`, `api_request_logs(request_id)` unique.
- `source_files(tenant_id, domain, created_at desc)`, `source_files(tenant_id, content_sha256)`.
- `ingestion_jobs(tenant_id, status, created_at desc)`, `ingestion_jobs(source_file_id, created_at desc)`.
- `point_audit_events(tenant_id, point_id, created_at desc)`, `point_audit_events(tenant_id, file_id, created_at desc)`.
- `graph_runs(tenant_id, thread_id, started_at desc)`, `graph_runs(tenant_id, external_user_id, started_at desc)`.
- `llm_calls(tenant_id, created_at desc)`, `llm_calls(graph_run_id)`, `llm_calls(ingestion_job_id)`.
- `run_feedback(tenant_id, graph_run_id)`.
Retention:
- Usage/cost rows (`llm_calls`) live longer than payload rows.
- `llm_call_payloads` has the shortest retention and is purged by
`retention_until`.
- API logs and point audit events follow tenant contract/legal retention.
- Deleted tenants are soft-deleted first; hard purge removes API keys, Store
namespaces, checkpointer threads, payload logs, and Qdrant points according
to a separate erasure runbook.
## Consequences
### Positive
- Tenant/API-key authentication has a clear relational source of truth, and
FastAPI dependencies can resolve `AuthContext` with one indexed lookup.
- Ingestion becomes observable and supportable: users can see whether a file
is queued, running, failed, or succeeded, and developers can inspect stage
events without scraping logs.
- Qdrant mutations become auditable even though Qdrant remains the actual
vector/payload store.
- LLM usage is attributable by tenant, thread, run, graph node, model, and
ingestion job, enabling cost reports and per-node optimization.
- Separating `llm_calls` from `llm_call_payloads` keeps billing/analytics fast
and makes sensitive prompt retention a deliberate policy choice.
- `graph_runs` gives the feedback endpoint a stable target while preserving
ADR-0007's decision not to own chat sessions.
### Negative
- This is a larger schema than the minimum needed to answer chat requests.
Implementing all tables up front adds migration and repository code before
the first end-to-end demo.
- There is partial duplication between structured logs and `api_request_logs`;
the former is operational, the latter is durable audit. Both must use the
same `request_id` or they become hard to correlate.
- Storing redacted LLM inputs/outputs still carries privacy risk: redaction can
miss sensitive details, especially in insurance text. Raw encrypted payloads
raise the risk further and need strict access controls.
- Cost calculation depends on pricing data being kept current. If pricing is
wrong at call time, historical costs are wrong unless corrected explicitly.
- Shared-schema tenancy relies on every query and foreign key carrying
`tenant_id`; a missed filter is a data leak. RLS could add defense in depth
later, but it is not part of the initial decision.
## Alternatives Considered
- **Exactly one API key per tenant**: rejected. It makes rotation and scope
separation painful. The requirement is that each tenant can authenticate;
allowing multiple keys per tenant is the safer implementation.
- **Database/schema per tenant**: rejected. It adds migration and operational
overhead per tenant and diverges from ADR-0001's shared Qdrant multitenancy
model. A shared schema with `tenant_id` indexes is simpler and scales better
for this stage.
- **PostgreSQL Row Level Security from day one**: deferred. RLS is useful
defense in depth, but it adds session-variable plumbing and migration/test
complexity. The initial boundary is FastAPI dependency resolution plus
explicit tenant filters and indexes; revisit RLS when the schema stabilizes.
- **Use SQLModel instead of SQLAlchemy ORM**: rejected because the project
explicitly wants SQLAlchemy 2 and Alembic table design. Pydantic request/
response models remain separate from ORM models.
- **Store all HTTP request/response bodies in `api_request_logs`**: rejected.
It would duplicate large payloads, accidentally retain files/prompts, and
raise privacy risk. Store summaries in `api_request_logs`; store controlled
LLM payloads in `llm_call_payloads`; store original files only via
`source_files.storage_uri` if retention policy allows it.
- **Store full Qdrant payloads/vectors in Postgres for audit**: rejected. It
doubles storage and creates two sources of truth. Audit records store change
summaries, ids, versions, and operation outcomes.
- **Compute LLM cost later from token counts**: rejected. Pricing changes over
time. Store the price used and the computed cost with each call so invoices
and reports are reproducible.