Files
Research/chunking/docs/adr/0008-metadata-schema.md

2.6 KiB

ADR 0008: Chunk Metadata Schema

Status

Accepted

Context

Each chunk in Qdrant needs metadata for filtering, citation, and debugging. The schema must balance flexibility with consistency, and handle both document-based (PDF, Word) and tabular (CSV, Excel) sources.

Decision

Core Metadata Schema

Field Type Required Description
domain string Yes Insurance domain: car, health, fire, life, travel, etc.
document_type string Yes Document category: policy, faq, claim, terms, guidance, etc.
language string Yes Language code: fa (Persian), en (English)
source_file string Yes Original filename with extension
chunk_index int Yes Zero-based position in source document
content_hash string Yes SHA-256 hash of chunk content
created_at datetime Yes Ingestion timestamp

Tabular Data Extensions

For CSV and Excel sources, additional fields:

Field Type Description
row_number int Row index in original file
column_names list[str] Column headers for context
original_format string Source format: csv, xlsx, xls

Example payload for a CSV row:

{
  "domain": "car",
  "document_type": "faq",
  "language": "fa",
  "source_file": "car_insurance_faq.csv",
  "chunk_index": 0,
  "content_hash": "a1b2c3...",
  "created_at": "2025-01-15T10:30:00Z",
  "row_number": 1,
  "column_names": ["question", "answer", "category"],
  "original_format": "csv"
}

Document Metadata (Optional, Future)

Reserved fields for future use:

  • policy_version: string
  • effective_date: date
  • department: string
  • region: string

Consequences

Positive

  • Consistent filtering — all chunks have core fields for retrieval filtering
  • Citation support — source file + chunk index enables precise citations
  • Deduplication — content hash identifies duplicate content
  • Tabular context — column names preserved for row-based chunks

Negative

  • Manual tagging — domain and document_type require human input or heuristics at ingestion
  • Storage overhead — metadata stored with each chunk

Neutral

  • Schema is minimal but extensible
  • Column names may vary across files — need normalization strategy

Implementation Notes

  • Define metadata validator at ingestion time
  • Auto-detect language using langdetect or similar
  • Derive domain from file path or manual mapping
  • Compute content_hash before insertion
  • Store created_at in ISO 8601 format for consistency