Files
Research/chunking/docs/adr/0012-observability-auth.md

3.3 KiB

ADR 0012: Observability and Authentication

Status

Accepted

Context

The system requires:

  • Full observability for retrieval and generation pipeline (tracing, metrics, debugging)
  • Authentication for API access control
  • Frontend handled separately — API backend only

Decision

Observability: Langfuse

Langfuse provides comprehensive LLM observability:

  • Trace every request end-to-end
  • Log retrieval steps (embedding, search, reranking)
  • Log generation steps (prompt, context, completion)
  • Track costs and latency
  • Debug failed requests
  • Evaluate quality with scores

Integration:

from langfuse import Langfuse
from langfuse.decorators import observe

langfuse = Langfuse(
    public_key="...",
    secret_key="...",
    host="http://localhost:3000"  # Self-hosted
)

@observe
def query_pipeline(question: str):
    # Traces entire pipeline automatically
    ...

Deployment: Self-hosted Langfuse via Docker (matches Qdrant deployment pattern).

Tracked Metrics

Metric Level Description
Query latency Request End-to-end response time
Embedding latency Step Time to embed query
Retrieval latency Step Vector search time
Reranking latency Step Rerank processing time
Generation latency Step LLM inference time
Token usage Request Input + output tokens
Model used Request Which embedding/reranker/LLM
Retrieval count Request Chunks retrieved before/after reranking
User feedback Request Thumbs up/down on answers

Authentication: API Key

Simple API key authentication for the ingestion and query endpoints:

from fastapi import Security, HTTPException
from fastapi.security import APIKeyHeader

api_key_header = APIKeyHeader(name="X-API-Key")

async def verify_api_key(api_key: str = Security(api_key_header)):
    if api_key not in VALID_API_KEYS:
        raise HTTPException(status_code=401, detail="Invalid API key")
    return api_key

API Key Management:

  • Store hashed keys in database
  • Generate keys via admin endpoint
  • Revoke keys when needed
  • Rate limit per key (optional)

API Structure

POST /ingest          - Upload documents (API key required)
GET  /status/{job_id} - Check ingestion status (API key required)
POST /query           - Ask question (API key required)
GET  /health          - Health check (no auth)

Consequences

Positive

  • Full traceability — every request traced from query to answer
  • Debugging — identify bottlenecks and failures quickly
  • Cost tracking — monitor token usage and model costs
  • Simple auth — API keys are easy to manage and rotate
  • Self-hosted — Langfuse runs on-prem, no external dependency

Negative

  • Storage overhead — traces consume storage
  • Performance impact — tracing adds minimal latency
  • Key management — need secure key storage and rotation process

Neutral

  • Need to set trace retention policy (e.g., 30 days)
  • May need to sample traces under high load

Implementation Notes

  • Deploy Langfuse with Docker Compose alongside Qdrant
  • Use Langfuse's Python SDK decorators for automatic tracing
  • Integrate user feedback collection (thumbs up/down)
  • Set up dashboards for monitoring key metrics
  • Implement API key rotation schedule