diff --git a/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md b/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md new file mode 100644 index 0000000..4df0295 --- /dev/null +++ b/docs/adr/0009-postgres-sqlalchemy-alembic-schema.md @@ -0,0 +1,444 @@ +# 0009. Postgres schema with SQLAlchemy 2 and Alembic + +## Status + +Proposed + +## Context + +ADR-0008 defines the FastAPI REST boundary: API-key authentication, tenant +resolution from Postgres, source-file ingestion, point management, and chat +thread runs. ADR-0007 defines LangGraph checkpoint persistence and explicitly +avoids making this service the source of truth for chat sessions. We now need +the application-owned Postgres schema that supports those decisions. + +The schema has to serve several purposes at once: + +- **Multitenancy**: the chatbot serves multiple tenants; every API key belongs + to one tenant and every request resolves to that tenant before touching + Qdrant or LangGraph. +- **Authentication**: each tenant needs API keys. In practice this should mean + *one or more* API keys per tenant, so keys can be rotated, scoped, and + revoked independently. +- **Ingestion records**: when a caller uploads a source file (`.csv`, `.xlsx`, + `.docx`, legacy `.doc` via conversion), the service records the file and the + ingestion job that parsed, chunked, embedded, and wrote Qdrant points. +- **Point/API audit**: point CRUD and batch operations mutate Qdrant, which is + not an audit-log database. The service needs its own record of who changed + what and when. +- **LLM usage and cost**: every model call inside the graph (triage, + contextualize, grade, generate, verify, summary/memory extraction) should be + attributable to a tenant, thread, run, graph node, model, token counts, and + price. For debugging/evals, the system also needs a controlled way to keep + the model input and output. + +The project uses SQLAlchemy 2.x and Alembic. The schema should therefore be +specified in SQLAlchemy 2 style (`DeclarativeBase`, `Mapped[...]`, +`mapped_column(...)`) and migrated only through Alembic — never `create_all()` +at FastAPI startup. + +## Decision + +### SQLAlchemy and migration conventions + +Use SQLAlchemy 2.x ORM models with typed mappings: + +```python +class Base(DeclarativeBase): + pass + +class Tenant(Base): + __tablename__ = "tenants" + + id: Mapped[uuid.UUID] = mapped_column(primary_key=True) + slug: Mapped[str] = mapped_column(String(80), unique=True, index=True) +``` + +Conventions: + +- Use SQLAlchemy async sessions in FastAPI (`AsyncSession`) and one session per + request/job unit of work. +- Use Alembic for all DDL. FastAPI startup opens connections and checks + readiness; it does not create or alter tables. +- Prefer UUID primary keys generated by the application. Avoid integer IDs that + leak tenant size and make distributed workers harder to compose. +- Use `timestamptz`/`DateTime(timezone=True)` for all timestamps. +- Use `Numeric(18, 8)` or finer for monetary/cost fields; never floats for + money. +- Store flexible metadata in `JSONB`, but keep relational identifiers and + query-critical fields as typed columns with indexes. If the database column + is named `metadata`, map it with a safe SQLAlchemy attribute such as + `metadata_ = mapped_column("metadata", JSONB, ...)` because `metadata` is + reserved on Declarative models. +- Use string status columns with SQLAlchemy/Pydantic enums and database + `CHECK` constraints rather than PostgreSQL native enums. Status sets change + often during early product work, and native enum migrations are painful. +- Every tenant-owned table has `tenant_id` and an index beginning with + `tenant_id`. Foreign keys include `ondelete` behaviour deliberately, not by + accident. + +Postgres uses a shared schema with tenant foreign keys, not database-per-tenant +or schema-per-tenant. This matches ADR-0001's shared Qdrant collection and +keeps tenant count scalable. + +### Core tenant and API-key tables + +#### `tenants` + +One row per customer/tenant. + +| Column | Notes | +|---|---| +| `id` | UUID PK. Used as the canonical `tenant_id` injected into Qdrant filters and LangGraph config. | +| `slug` | Stable short name, unique, human-readable. | +| `name` | Display name. | +| `status` | `active` \| `suspended` \| `deleted`. Suspended tenants authenticate to a clear error but cannot run work. | +| `settings` | JSONB for tenant-level feature flags/limits (max upload size, enabled file types, allowed domains, etc.). | +| `created_at`, `updated_at`, `deleted_at` | Audit/soft-delete timestamps. | + +#### `tenant_domains` + +Optional but recommended. Validates the `domain` values used throughout Qdrant +payloads (`car`, `fire`, etc.) per tenant. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `tenant_id` | FK to `tenants.id`. | +| `domain` | Tenant-local domain key. Unique with `tenant_id`. | +| `display_name` | Human-readable label. | +| `status` | `active` \| `disabled`. | +| `metadata` | JSONB for domain-specific ingestion/retrieval settings. | + +This prevents arbitrary caller-supplied domains from silently creating new +partitions in Qdrant. + +#### `api_keys` + +One tenant can have multiple active keys for rotation and scoped access. + +| Column | Notes | +|---|---| +| `id` | UUID PK, also usable as the API key lookup id/prefix. | +| `tenant_id` | FK to `tenants.id`. | +| `name` | Human label, e.g. `main-backend-prod`. | +| `key_prefix` | Short non-secret prefix shown in logs/admin UI, unique. | +| `key_hash` | Hash of the secret key material. Plaintext API keys are never stored. | +| `scopes` | JSONB or text array: `threads:run`, `points:read`, `points:write`, `files:write`, `memory:read`, `admin`. | +| `actor_type` | `backend` \| `admin` \| `worker`. | +| `status` | `active` \| `revoked` \| `expired`. | +| `expires_at`, `revoked_at`, `last_used_at` | Lifecycle timestamps. | +| `created_by`, `created_at`, `updated_at` | Audit fields. | + +Authentication dependency in ADR-0008 queries by `key_prefix`/id, verifies +`key_hash` with constant-time comparison, checks tenant/key status and scopes, +and returns `AuthContext`. + +### Request and mutation audit tables + +#### `api_request_logs` + +Append-only request log for authenticated `/v1` calls. This is not a +replacement for structured application logs; it is the durable queryable audit +record. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `tenant_id` | Denormalized from API key for fast tenant queries. | +| `api_key_id` | FK to `api_keys.id`, nullable only for failed auth where key is unknown. | +| `request_id` | Correlation id, unique. | +| `method`, `path_template`, `status_code` | Route identity and result. | +| `scopes_required` | JSONB/text array. | +| `external_user_id` | User id supplied by the main backend, when present. | +| `thread_id` | Present for thread/run routes. No FK — ADR-0007 says this service owns no thread table. | +| `source_ip_hash`, `user_agent` | Optional operational metadata; avoid storing raw IP unless required. | +| `request_summary`, `response_summary` | JSONB summaries, not raw bodies by default. | +| `error_code` | Stable error code from ADR-0008, nullable. | +| `duration_ms`, `created_at` | Timing. | + +This table records that an API call happened. Tables below record domain-level +side effects. + +#### `point_audit_events` + +Append-only audit of `/v1/points` and `/v1/files/{file_id}/points` mutations. +Qdrant remains the storage/search engine; this table records mutation intent +and result. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `tenant_id` | FK to `tenants.id`. | +| `api_request_log_id` | FK to `api_request_logs.id`. | +| `api_key_id` | FK to `api_keys.id`. | +| `operation` | `create` \| `update` \| `payload_patch` \| `soft_delete` \| `hard_delete` \| `reorder` \| `batch`. | +| `point_id` | Qdrant point id, nullable for batch/file-wide operations. | +| `file_id` | Source file id when relevant. | +| `domain` | Qdrant payload domain. | +| `before_version`, `after_version` | Optimistic-concurrency versions when available. | +| `changed_fields` | JSONB list/summary; no large content or vectors. | +| `qdrant_operation_id`, `qdrant_status` | Result returned by Qdrant, if available. | +| `created_at` | Event time. | + +### File and ingestion tables + +#### `source_files` + +One logical source document uploaded by a tenant. Re-ingestion of the same file +creates new jobs against the same or replacement `source_files` row depending +on the `content_hash` policy. + +| Column | Notes | +|---|---| +| `id` | UUID PK; this is the `file_id` copied into Qdrant point payloads. | +| `tenant_id` | FK to `tenants.id`. | +| `domain` | Tenant domain, validated by `tenant_domains` where enabled. | +| `source_filename` | Original filename. | +| `source_type` | `csv` \| `xlsx` \| `docx` \| `doc`. | +| `content_sha256` | Hash of the uploaded file bytes for idempotency/change detection. | +| `byte_size` | Upload size. | +| `storage_uri` | Where the original file is stored, if retained. Nullable if not retaining originals. | +| `status` | `active` \| `superseded` \| `soft_deleted` \| `purged`. | +| `created_by_api_key_id`, `created_at`, `updated_at`, `deleted_at` | Audit fields. | + +#### `ingestion_jobs` + +One attempt to parse/chunk/embed/upsert a source file. This table is required +even if the first implementation processes inline, because ADR-0008's file +upload contract is job-shaped. + +| Column | Notes | +|---|---| +| `id` | UUID PK; returned as `ingestion_job_id`. | +| `tenant_id` | FK to `tenants.id`. | +| `source_file_id` | FK to `source_files.id`. | +| `api_request_log_id` | FK to the upload request log. | +| `requested_by_api_key_id` | FK to `api_keys.id`. | +| `status` | `queued` \| `running` \| `succeeded` \| `failed` \| `cancelled`. | +| `chunking_strategy` | `semantic` \| `fixed_size`; matches ADR-0004. | +| `embedding_model_versions` | JSONB map of vector name → model/version. | +| `started_at`, `completed_at` | Lifecycle timestamps. | +| `points_created`, `points_updated`, `points_soft_deleted`, `points_skipped` | Result counters. | +| `error_code`, `error_message` | Failure summary. | +| `metadata` | JSONB for parser/chunker options. | +| `created_at`, `updated_at` | Audit timestamps. | + +#### `ingestion_job_events` + +Append-only progress/error stream for a job. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `tenant_id` | FK to `tenants.id`. | +| `ingestion_job_id` | FK to `ingestion_jobs.id`. | +| `level` | `info` \| `warning` \| `error`. | +| `stage` | `received` \| `parsed` \| `chunked` \| `embedded` \| `upserted` \| `completed`. | +| `message` | Short human-readable event. | +| `details` | JSONB structured details. | +| `created_at` | Event time. | + +### Graph run, LLM usage, and feedback tables + +#### `graph_runs` + +One row per `POST /v1/threads/{thread_id}/runs`. This is not a session/thread +table: it records one execution for audit, feedback, usage aggregation, and +cost reporting. + +| Column | Notes | +|---|---| +| `id` | UUID PK; this is the `run_id` returned by the run endpoint. | +| `tenant_id` | FK to `tenants.id`. | +| `api_request_log_id` | FK to `api_request_logs.id`. | +| `api_key_id` | FK to `api_keys.id`. | +| `thread_id` | LangGraph thread id from the path. No FK. | +| `external_user_id` | User id supplied by the main backend. | +| `status` | `running` \| `answered` \| `clarifying` \| `escalate` \| `failed`. | +| `escalation_reason` | ADR-0006 reason when `status='escalate'`. | +| `input_message_hash` | Hash for idempotency/debug correlation without storing raw text here. | +| `output_message_hash` | Hash of final answer/clarifying/escalation message. | +| `llm_input_tokens`, `llm_output_tokens`, `llm_total_cost` | Denormalized totals from `llm_calls`. | +| `started_at`, `completed_at`, `duration_ms` | Timing. | +| `metadata` | JSONB for graph version, prompt version, retrieved chunk IDs, etc. | + +`graph_runs` solves the practical problem left by ADR-0007's feedback endpoint: +feedback needs a stable `run_id`, but this service still does not need a table +that represents chat sessions. + +#### `llm_pricing` + +Versioned model pricing table so historical cost calculations remain +explainable when model prices change. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `provider` | `anthropic` \| `openai` \| other. | +| `model` | Provider model id. | +| `currency` | Usually `USD`. | +| `input_price_per_1m_tokens`, `output_price_per_1m_tokens` | Numeric. | +| `effective_from`, `effective_to` | Time-bounded price validity. | +| `created_at` | Audit timestamp. | + +#### `llm_calls` + +One row per provider model call made inside the graph or ingestion pipeline. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `tenant_id` | FK to `tenants.id`. | +| `graph_run_id` | FK to `graph_runs.id`, nullable for ingestion-time LLM calls such as image extraction. | +| `ingestion_job_id` | FK to `ingestion_jobs.id`, nullable for chat-time calls. | +| `api_request_log_id` | FK to the originating request when available. | +| `thread_id`, `external_user_id` | Denormalized for query convenience; nullable outside chat. | +| `node_name` | `triage`, `contextualize`, `grade`, `generate`, `verify`, `summarize`, `memory_extract`, `image_extract`, etc. | +| `provider`, `model`, `model_version` | Provider identity. | +| `pricing_id` | FK to `llm_pricing.id`, nullable if price was configured externally. | +| `input_tokens`, `output_tokens`, `total_tokens` | Provider usage numbers. | +| `input_cost`, `output_cost`, `total_cost`, `currency` | Cost at call time. | +| `latency_ms` | Provider round-trip. | +| `status` | `succeeded` \| `failed` \| `cancelled`. | +| `error_code`, `error_message` | Failure summary. | +| `prompt_version`, `schema_version` | Version of prompt/structured-output schema used. | +| `input_hash`, `output_hash` | Hashes of stored/redacted payloads. | +| `created_at` | Call start time. | + +Costs are computed and stored at call time from `llm_pricing` (or explicit +runtime pricing config), not recomputed later from a mutable current price. + +#### `llm_call_payloads` + +Stores the actual model input/output only when allowed by tenant policy. This +is deliberately separate from `llm_calls` so usage/billing queries never touch +large or sensitive payloads. + +| Column | Notes | +|---|---| +| `llm_call_id` | PK/FK to `llm_calls.id`. | +| `tenant_id` | FK to `tenants.id`, repeated for partition/index convenience. | +| `input_redacted` | JSONB/text redacted prompt/messages/tool input. | +| `output_redacted` | JSONB/text redacted model output/tool call result. | +| `input_encrypted`, `output_encrypted` | Optional encrypted raw payload bytes/text if raw retention is enabled. | +| `redaction_version` | Which redaction policy produced the redacted fields. | +| `retention_until` | When payloads must be deleted, independent of usage rows. | +| `created_at` | Timestamp. | + +Default policy: store token counts/costs for every call, store **redacted** +input/output for debugging/evals, and store raw encrypted payloads only for +tenants that explicitly enable it. Insurance chat can contain PII and sensitive +claim/coverage information; raw prompt logging cannot be an accidental default. + +#### `run_feedback` + +Feedback from `POST /v1/threads/{thread_id}/runs/{run_id}/feedback`. + +| Column | Notes | +|---|---| +| `id` | UUID PK. | +| `tenant_id` | FK to `tenants.id`. | +| `graph_run_id` | FK to `graph_runs.id`. | +| `external_user_id` | User id from main backend, if present. | +| `rating` | `thumbs_up` \| `thumbs_down` \| numeric score. | +| `reason_codes` | JSONB/text array. | +| `comment` | Optional free text. | +| `created_at` | Timestamp. | + +### What is intentionally not modeled + +- **No `threads`/`sessions` table.** ADR-0007 remains in force: the main + backend owns session records and LangGraph owns thread checkpoints. This + schema records runs and usage, not conversation ownership. +- **No Postgres copy of Qdrant point content/vectors.** Qdrant remains the + source of truth for point payloads/vectors. Postgres stores source-file, + ingestion, and audit records. +- **No plaintext API keys.** Only hashes and non-secret prefixes. +- **No automatic raw prompt retention.** Raw LLM input/output is opt-in, + encrypted, and retention-limited. + +### Indexing and retention + +Required indexes: + +- `api_keys(key_prefix)` unique; `api_keys(tenant_id, status)`. +- `tenant_domains(tenant_id, domain)` unique. +- `api_request_logs(tenant_id, created_at desc)`, `api_request_logs(request_id)` unique. +- `source_files(tenant_id, domain, created_at desc)`, `source_files(tenant_id, content_sha256)`. +- `ingestion_jobs(tenant_id, status, created_at desc)`, `ingestion_jobs(source_file_id, created_at desc)`. +- `point_audit_events(tenant_id, point_id, created_at desc)`, `point_audit_events(tenant_id, file_id, created_at desc)`. +- `graph_runs(tenant_id, thread_id, started_at desc)`, `graph_runs(tenant_id, external_user_id, started_at desc)`. +- `llm_calls(tenant_id, created_at desc)`, `llm_calls(graph_run_id)`, `llm_calls(ingestion_job_id)`. +- `run_feedback(tenant_id, graph_run_id)`. + +Retention: + +- Usage/cost rows (`llm_calls`) live longer than payload rows. +- `llm_call_payloads` has the shortest retention and is purged by + `retention_until`. +- API logs and point audit events follow tenant contract/legal retention. +- Deleted tenants are soft-deleted first; hard purge removes API keys, Store + namespaces, checkpointer threads, payload logs, and Qdrant points according + to a separate erasure runbook. + +## Consequences + +### Positive +- Tenant/API-key authentication has a clear relational source of truth, and + FastAPI dependencies can resolve `AuthContext` with one indexed lookup. +- Ingestion becomes observable and supportable: users can see whether a file + is queued, running, failed, or succeeded, and developers can inspect stage + events without scraping logs. +- Qdrant mutations become auditable even though Qdrant remains the actual + vector/payload store. +- LLM usage is attributable by tenant, thread, run, graph node, model, and + ingestion job, enabling cost reports and per-node optimization. +- Separating `llm_calls` from `llm_call_payloads` keeps billing/analytics fast + and makes sensitive prompt retention a deliberate policy choice. +- `graph_runs` gives the feedback endpoint a stable target while preserving + ADR-0007's decision not to own chat sessions. + +### Negative +- This is a larger schema than the minimum needed to answer chat requests. + Implementing all tables up front adds migration and repository code before + the first end-to-end demo. +- There is partial duplication between structured logs and `api_request_logs`; + the former is operational, the latter is durable audit. Both must use the + same `request_id` or they become hard to correlate. +- Storing redacted LLM inputs/outputs still carries privacy risk: redaction can + miss sensitive details, especially in insurance text. Raw encrypted payloads + raise the risk further and need strict access controls. +- Cost calculation depends on pricing data being kept current. If pricing is + wrong at call time, historical costs are wrong unless corrected explicitly. +- Shared-schema tenancy relies on every query and foreign key carrying + `tenant_id`; a missed filter is a data leak. RLS could add defense in depth + later, but it is not part of the initial decision. + +## Alternatives Considered + +- **Exactly one API key per tenant**: rejected. It makes rotation and scope + separation painful. The requirement is that each tenant can authenticate; + allowing multiple keys per tenant is the safer implementation. +- **Database/schema per tenant**: rejected. It adds migration and operational + overhead per tenant and diverges from ADR-0001's shared Qdrant multitenancy + model. A shared schema with `tenant_id` indexes is simpler and scales better + for this stage. +- **PostgreSQL Row Level Security from day one**: deferred. RLS is useful + defense in depth, but it adds session-variable plumbing and migration/test + complexity. The initial boundary is FastAPI dependency resolution plus + explicit tenant filters and indexes; revisit RLS when the schema stabilizes. +- **Use SQLModel instead of SQLAlchemy ORM**: rejected because the project + explicitly wants SQLAlchemy 2 and Alembic table design. Pydantic request/ + response models remain separate from ORM models. +- **Store all HTTP request/response bodies in `api_request_logs`**: rejected. + It would duplicate large payloads, accidentally retain files/prompts, and + raise privacy risk. Store summaries in `api_request_logs`; store controlled + LLM payloads in `llm_call_payloads`; store original files only via + `source_files.storage_uri` if retention policy allows it. +- **Store full Qdrant payloads/vectors in Postgres for audit**: rejected. It + doubles storage and creates two sources of truth. Audit records store change + summaries, ids, versions, and operation outcomes. +- **Compute LLM cost later from token counts**: rejected. Pricing changes over + time. Store the price used and the computed cost with each call so invoices + and reports are reproducible.