initial commit
This commit is contained in:
129
docs/adr/0002-chunk-crud-and-search-api.md
Normal file
129
docs/adr/0002-chunk-crud-and-search-api.md
Normal file
@@ -0,0 +1,129 @@
|
||||
# 0002. Chunk CRUD and payload/keyword search API
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Beyond bulk ingestion ([0001](0001-ingestion-pipeline-and-collection-schema.md)),
|
||||
we need FastAPI endpoints for direct, fine-grained management of individual
|
||||
chunks: create a single chunk, update a chunk's content/vectors/payload,
|
||||
delete chunks (by ID or in bulk by filter), and list/search chunks by keyword
|
||||
and payload metadata (tenant, domain, file, etc.) — independent of the
|
||||
semantic hybrid retrieval used by the AI agent (ADR-0003).
|
||||
|
||||
This pipeline operates on the same collection and payload schema defined in
|
||||
ADR-0001, and Qdrant's Points API already provides the primitives needed —
|
||||
there's no need for a custom data layer on top.
|
||||
|
||||
## Decision
|
||||
|
||||
Expose a thin FastAPI layer directly over Qdrant's Points API, operating on
|
||||
the `chunks` collection and payload schema from ADR-0001:
|
||||
|
||||
| Operation | FastAPI endpoint (indicative) | Qdrant primitive |
|
||||
|---|---|---|
|
||||
| Create a chunk | `POST /chunks` | upsert (single point) |
|
||||
| Update a chunk's vectors/content | `PUT /chunks/{chunk_id}` | upsert (`update_only` mode) |
|
||||
| Partially update payload | `PATCH /chunks/{chunk_id}/payload` | `set_payload` / `overwrite_payload` |
|
||||
| Delete one chunk | `DELETE /chunks/{chunk_id}` | delete by ID |
|
||||
| Delete many chunks | `DELETE /chunks?file_id=...` | delete by filter |
|
||||
| List/paginate chunks, in order | `GET /chunks?file_id=...` | `scroll` with filter + pagination, `order_by: order_id` |
|
||||
| Reorder/insert a chunk | `PATCH /chunks/{chunk_id}/order` | `set_payload` on `order_id` only |
|
||||
| Count chunks | `GET /chunks/count` | `count` |
|
||||
| Keyword search | `GET /chunks/search?q=...` | full-text payload index on `content` (match, not semantic) |
|
||||
| Bulk multi-op edits | `POST /chunks/batch` | Qdrant `points/batch` |
|
||||
|
||||
### Ordering and reordering
|
||||
|
||||
Chunks belonging to a `file_id` are listed via `scroll` ordered by the
|
||||
mutable `order_id` payload field (ADR-0001) — this is what lets the frontend
|
||||
display and let a backend user modify chunk order. Because `order_id` is a
|
||||
fractional float key rather than a sequential integer position, moving or
|
||||
inserting a chunk only requires assigning it a new value between its two new
|
||||
neighbors — no renumbering of siblings, and no effect on the chunk's stable
|
||||
point ID (which is derived from the immutable `chunk_index`, not `order_id`).
|
||||
|
||||
`PATCH /chunks/{chunk_id}/order` does more than set one field, though: since
|
||||
ADR-0001 also maintains `previous_chunk_id`/`next_chunk_id` pointers for O(1)
|
||||
adjacency lookups, a single reorder touches **up to four chunks** in one
|
||||
`points/batch` call:
|
||||
1. The moved chunk: new `order_id`, new `previous_chunk_id`/`next_chunk_id`.
|
||||
2. Its old neighbors: re-link their `previous_chunk_id`/`next_chunk_id` to
|
||||
skip over the moved chunk.
|
||||
3. Its new neighbors: re-link them to point at the moved chunk.
|
||||
|
||||
The same fix-up applies to insert (new chunk) and delete (soft-deleted chunk
|
||||
— see below) operations: any operation that changes a chunk's position or
|
||||
removes it from the sequence must atomically update the neighbors' pointers
|
||||
in the same `points/batch` request, or `previous_chunk_id`/`next_chunk_id`
|
||||
go stale.
|
||||
|
||||
### Keyword search is not semantic search
|
||||
|
||||
`GET /chunks/search` matches against the full-text payload index on
|
||||
`content` (and structured filters on `tenant_id`, `domain`, `file_id`,
|
||||
etc.) — it is filter/match-based keyword search, not embedding-based
|
||||
retrieval. This is deliberately distinct from the hybrid dense+sparse
|
||||
retrieval used by the AI agent in ADR-0003; the two "search" concepts serve
|
||||
different callers (a human/admin managing chunks vs. an agent retrieving
|
||||
context) and should not be conflated in the API or in future discussion.
|
||||
|
||||
### Delete is soft by default
|
||||
|
||||
`DELETE /chunks/{chunk_id}` and `DELETE /chunks?file_id=...` set
|
||||
`is_active: false` and `deleted_at` (via `set_payload`) rather than removing
|
||||
the point from Qdrant — consistent with ADR-0001's soft-delete fields. This
|
||||
keeps deleted chunks available for audit and lets `GET /chunks` and
|
||||
`GET /chunks/search` filter them out by default (`is_active: true` implied
|
||||
unless the caller explicitly asks to include inactive chunks). A hard delete
|
||||
(actual point removal, e.g. `points/delete`) is available separately for
|
||||
compliance-driven purges, not as the default CRUD behavior. Either way, the
|
||||
deleted chunk's `previous_chunk_id`/`next_chunk_id` neighbors are relinked to
|
||||
point at each other in the same `points/batch` call, so context-window
|
||||
expansion (ADR-0003) never walks into a deactivated or removed chunk.
|
||||
|
||||
### Tenant isolation
|
||||
|
||||
Every endpoint (read and write) has its `tenant_id` filter injected
|
||||
server-side from the authenticated request context — never accepted as
|
||||
client-supplied input in the request body or query string. This is
|
||||
non-negotiable: it's the same isolation boundary ADR-0001 relies on for
|
||||
multitenancy, and it must hold for every code path that touches the
|
||||
collection, not just ingestion.
|
||||
|
||||
### Concurrency
|
||||
|
||||
Updates use the `version` payload field (from ADR-0001) together with
|
||||
Qdrant's `update_filter`, giving an optimistic-concurrency-style guard
|
||||
against races between a concurrent ingestion re-run (ADR-0001) and a manual
|
||||
edit through this API.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
- Reuses the exact collection/schema from ADR-0001 — no parallel data model
|
||||
to keep in sync.
|
||||
- Thin mapping to Qdrant primitives keeps the API predictable and easy to
|
||||
extend as new filter/sort needs arise.
|
||||
- Clear separation from ADR-0003's semantic search avoids API consumers
|
||||
confusing "find chunks matching this filter" with "retrieve chunks relevant
|
||||
to this query."
|
||||
|
||||
### Negative
|
||||
- Keyword search here is limited to what Qdrant's full-text payload index
|
||||
supports — no stemming/synonym handling beyond what that index offers.
|
||||
- Optimistic concurrency via `version` requires every writer (ingestion and
|
||||
this API) to consistently read-check-write; a writer that skips this can
|
||||
silently clobber concurrent edits.
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
- **Separate keyword-search infrastructure (e.g. Elasticsearch)**: rejected —
|
||||
Qdrant's native full-text payload index already covers filter-style keyword
|
||||
search on `content`, and adding a second search system would duplicate data
|
||||
and infrastructure for no clear benefit at current scale.
|
||||
- **Client-supplied `tenant_id` in request body**: rejected — trusting
|
||||
client input for the isolation boundary is a direct multitenancy security
|
||||
risk; it must come from server-side auth context.
|
||||
Reference in New Issue
Block a user