Why: - Design the LangGraph-based insurance chatbot: retrieval-grounded answering, grounded clarification, and handoff to a human assistant owned by the existing main backend. Changes: - ADR-0006: graph nodes/edges, escalation triggers, and the backend-owned handoff signal contract. - ADR-0007: Postgres checkpointer, thread model (thread_id as the shared identifier), history trimming strategy, and cross-thread user memory via the Store. Impact: - Design-only; no code changes yet.
14 KiB
0007. Agent persistence, thread model, history trimming, and cross-thread memory
Status
Proposed
Context
ADR-0006 defines the conversational graph but deliberately leaves its persistence to this ADR. Four separate concerns need deciding, and they are easy to conflate:
- Durable conversation state. Each turn must see the previous turns, and
that history must survive a process restart and be visible to every
replica. (Note that ADR-0006 uses no
interrupt(), so persistence is needed for memory, not for resuming paused runs.) - Conversation identity. The main backend calls this service with the identifier of one chat session; something must map that to a LangGraph thread, and it must be clear which system owns the session records and which vocabulary this service's API speaks.
- History size. The frontend displays at most 20 messages. It is not obvious whether that number should also govern what the model sees, what Postgres stores, or neither. Getting this wrong either blows up context cost or silently amputates conversations.
user_id. The backend supplies auser_idthat is stable across all of a user's threads. It currently has no defined role.
The retrieval path already carries its own tenancy boundary
(0002): tenant_id is injected
server-side from auth and never accepted from an end client. Anything decided
here must preserve that.
Decision
Checkpointer: AsyncPostgresSaver, one pool, compiled once
Postgres via langgraph-checkpoint-postgres (AsyncPostgresSaver), because
the API is async FastAPI. Concretely:
- A single
psycopg_pool.AsyncConnectionPool(autocommit=True,row_factory=dict_row) opened in the FastAPI lifespan, shared by the checkpointer and the Store. - The graph is compiled once at startup and reused; it is not rebuilt per request.
.setup()runs as a deployment/migration step, not at application startup — it issues DDL, and every replica racing to create tables on boot is a deployment hazard.
InMemorySaver remains acceptable in tests only.
Thread model: the thread is the only conversation record this service keeps
The identifier the main backend uses for a chat session is the LangGraph
thread_id. No indirection, no mapping table.
This service's API therefore speaks thread_id — in paths, request bodies,
and responses alike — rather than introducing a second noun for the same
entity. The backend passes its session identifier as the thread_id and that
is the whole of the mapping. Internal vocabulary matching the runtime keeps
the resource model honest: a thread here really is a LangGraph thread, with
LangGraph's semantics (a durable checkpoint sequence), not a product concept
that merely resembles one.
The main backend already owns users, sessions, and the human-assistant transfer (ADR-0006). Duplicating that as a thread-metadata table here would be a second set of records to keep in sync for no benefit, so this service stores no conversation metadata of its own — no owner column, no title, no status. Listing a user's conversations, showing titles, and knowing whether a conversation was transferred to a human are all the backend's concerns. The checkpoint sequence is the only record this service keeps.
The consequence is a trust boundary: thread_id, user_id, and
tenant_id arrive as trusted inputs on every call, and this service performs
no ownership check on them. That is only safe if the service is reachable
solely by the main backend — private network or service-to-service auth, never
exposed to browsers. Without that, a thread_id becomes a bearer token for
someone else's conversation. This is the single most important operational
constraint in this ADR.
History: persist everything, trim only at the model call
The three "history sizes" are decoupled, and each gets its own answer:
| Layer | Size | Rationale |
|---|---|---|
| Postgres checkpoints | Unbounded (subject to retention) | Audit, handoff context, and eval datasets all need the full transcript. Insurance conversations are records. |
| Model context | Trimmed per call | Cost and latency, and Persian tokenizes expensively. |
| Frontend | 20 messages | A display/pagination concern only. It must not drive either of the above. |
Trimming happens in a pre-model hook — trim_messages on the way into
generate, keeping the system prompt plus the most recent turns under a token
budget, and always starting the trimmed window on a HumanMessage with no
orphaned tool call/result pairs. It is applied to the messages passed to the
model; it does not mutate persisted state. Nothing is deleted.
Two additions on top of plain trimming:
- Rolling summary. Once a thread exceeds a threshold (~20 turns), a
summarization step folds the dropped older turns into a
summaryfield in state, which is prepended to the trimmed window. This is what stops a long conversation from losing its own beginning — pure trimming does. It also supplieshandoff_context.conversation_summaryin ADR-0006's escalation payload for free. - Retrieved chunks are not messages. Chunk text from ADR-0003 lands in a
dedicated
retrieved_chunksstate field scoped to the current turn, not in the message list. Persisting full reranked chunk text as tool messages on every turn is what actually blows up a RAG thread's state size; only thechunk_idcitations persist in the transcript.
The frontend's 20 messages are served by a paginated read endpoint over checkpoint state, not by capping what is stored.
user_id: long-term cross-thread memory, handoff context, and operations
user_id is passed per request by the main backend and does three jobs:
- Cross-thread memory (the main one). An
AsyncPostgresStore, same pool, namespaced(tenant_id, "user", user_id, "profile")and(… , "facts"). ADR-0006'sload_memoryreads it at the top of every turn;write_memorywrites after a successful answer. What belongs there: language and formality preference, which products the user holds or has asked about, household context they volunteered, recurring topics. What does not: national ID, payment details, health information, or anything derived from a single turn that shouldn't outlive it. Memory writes are extraction-based and bounded, not "append every message." - Handoff context — the escalation payload in ADR-0006 carries the profile so the human assistant starts informed rather than from zero.
- Operations — per-user rate limiting and cohort segmentation in analytics and evals.
tenant_id leads every Store namespace, for the same reason ADR-0002 gives:
a shared store partitioned only by user_id leaks across tenants the moment
user IDs are tenant-scoped rather than globally unique.
Cross-thread memory is the only reason this service needs user_id at all;
if long-term memory is dropped, thread_id and tenant_id would suffice.
FastAPI surface
Resources are named after the LangGraph concepts they actually are, matching
LangGraph Platform's own /threads/{thread_id}/runs shape. One user message is
one graph execution, so it is a run, not a bespoke noun.
| Endpoint | Purpose |
|---|---|
POST /v1/threads/{thread_id}/runs |
Execute one graph run for the user's message. Body carries the message; user_id, tenant_id, and locale come from the service-to-service context. SSE stream (stream_mode="messages" for tokens, "custom" for retrieval progress). Terminal status: answered | clarifying | escalate. |
GET /v1/threads/{thread_id}/messages?limit=20&before=… |
Paginated transcript from checkpoint state, for the frontend's 20-message window. |
DELETE /v1/threads/{thread_id} |
Purge a thread's checkpoints (user-initiated "clear chat", or erasure). |
POST /v1/threads/{thread_id}/runs/{run_id}/feedback |
Thumbs up/down + optional text, for the eval dataset. |
GET /v1/users/{user_id}/memory / DELETE … |
Inspect and reset what the Store remembers about a user. Not thread-scoped — memory deliberately outlives any one thread. |
The thread is not created by an explicit endpoint: the first POST …/runs
against an unseen thread_id creates the checkpoint sequence implicitly, which
is what the checkpointer does anyway. Adding a POST /v1/threads would be a
row this service has already decided not to store.
There are no handoff endpoints. Escalation is a terminal status on the run
response (ADR-0006); the main backend acts on it and then stops issuing runs
against that thread_id.
One in-flight run per thread. Concurrent runs on the same thread_id race
on checkpoint writes. The run endpoint takes a Postgres advisory lock keyed on
thread_id and returns 409 Conflict rather than queueing — a user
double-sending is a UI problem, not something to serialize silently.
user_id, tenant_id, and locale are passed per invocation in
config["configurable"], not stored in graph state, so a change in the
caller's context is never overridden by a stale checkpoint.
Retention
Checkpoints are not kept forever. A scheduled job deletes threads whose most recent checkpoint is older than the retention window (to be set with the business — insurance record-keeping obligations likely dominate this choice), and the erasure endpoints above provide the per-user path that removes both the thread checkpoints and the Store namespaces.
Consequences
Positive
- One Postgres holds checkpoints and long-term memory — one backup, one connection pool, one transactional story — and no conversation records are duplicated between this service and the main backend.
- The API names the runtime's own concepts (
threads,runs) rather than inventing parallel ones, so there is no translation layer between the URL, the checkpointer, and the LangGraph SDK — and a future move onto LangGraph Platform would not be a rename. - Decoupling the three history sizes means the frontend's 20-message window can change without touching model cost or the durability guarantee.
- Keeping chunk text out of the message list is what keeps a long RAG thread's checkpoint size flat, and it is far cheaper to get right now than to retrofit.
- Checkpoint history gives support a way to replay exactly what the agent saw when a user disputes an answer.
- With no
interrupt()in the design, the checkpointer serves one purpose only (conversation memory), which keeps the persistence story small.
Negative
- Authorization is entirely delegated. This service will happily read or
append to any
thread_idit is given. If it is ever exposed beyond the main backend, that is a direct data-leak path. - The API and the main backend use different words for the same identifier
(
thread_idhere, session elsewhere). The mapping is trivial and stated once, but it is one more thing an integrator has to know. - Unbounded checkpoint growth is a real storage cost, and it is deferred to a retention policy that is not yet decided.
- The rolling summary is lossy by construction — a fact stated 40 turns ago and dropped from the summary is gone from the model's view even though it is still in Postgres.
- Advisory-lock-and-409 pushes retry handling onto the caller.
- Long-term memory is a new correctness surface: a wrong fact written to the Store silently contaminates every future session until something removes it. Hence the inspect/reset endpoints, which the backend should surface to users.
- Nothing here records that a conversation was transferred to a human, so this service cannot distinguish an abandoned thread from a transferred one in its own analytics; that join has to happen on the backend's data.
Alternatives Considered
InMemorySaver/SqliteSaver: rejected — history must survive restarts and be shared across replicas.- An application-owned thread-metadata table in this service (user, title, status, timestamps): rejected now that the main backend owns sessions and the human transfer. It would be a second source of truth to keep in sync. It becomes worth revisiting only if this service ever needs to answer ownership or listing queries itself.
- Naming the API resource
sessionrather thanthread(e.g./v1/sessions/{session_id}/turns, mirroring the main backend's vocabulary): rejected. This is an internal service-to-service API, not a public one, so there is no audience it needs to shield from the runtime's concepts — and the entity genuinely is a LangGraph thread with a checkpoint sequence, not a session record, since this service stores no session record at all. "Thread" is also ordinary conversation vocabulary (email, chat), so it carries little framework lock-in as a name. The main backend keeps saying "session"; the translation happens at its call site, in one place. - A separate
messagestable as the source of truth, with LangGraph state rebuilt per request: rejected — two transcripts to keep in sync, and it discards the checkpoint history that makes replay and audit possible. - Hard-capping stored history at 20 messages to match the frontend: rejected — it conflates a display window with a retention policy, destroys audit and eval data, and would hand a human assistant a truncated transcript at exactly the moment they need the whole thing.
RemoveMessage/REMOVE_ALL_MESSAGESto prune persisted state: rejected as an automatic policy — it is destructive trimming where non-destructive trimming costs nothing extra. It remains the right tool behind the explicitDELETE /v1/threads/{thread_id}action.- Summarization instead of trimming: rejected as a replacement, adopted as a supplement. Summarizing every turn is an extra model call for short conversations that don't need one; trimming is free and handles the common case.
- Redis for checkpoints: rejected — the Store and the checkpoints want the same transactional store, and durability matters more here than checkpoint write latency.
- A single long-lived thread per user instead of one per conversation: rejected — it would make one escalation apply to all of a user's future conversations, and it removes any natural retention or "clear chat" boundary. Cross-thread continuity is what the Store is for.