2.6 KiB
2.6 KiB
ADR 0008: Chunk Metadata Schema
Status
Accepted
Context
Each chunk in Qdrant needs metadata for filtering, citation, and debugging. The schema must balance flexibility with consistency, and handle both document-based (PDF, Word) and tabular (CSV, Excel) sources.
Decision
Core Metadata Schema
| Field | Type | Required | Description |
|---|---|---|---|
domain |
string | Yes | Insurance domain: car, health, fire, life, travel, etc. |
document_type |
string | Yes | Document category: policy, faq, claim, terms, guidance, etc. |
language |
string | Yes | Language code: fa (Persian), en (English) |
source_file |
string | Yes | Original filename with extension |
chunk_index |
int | Yes | Zero-based position in source document |
content_hash |
string | Yes | SHA-256 hash of chunk content |
created_at |
datetime | Yes | Ingestion timestamp |
Tabular Data Extensions
For CSV and Excel sources, additional fields:
| Field | Type | Description |
|---|---|---|
row_number |
int | Row index in original file |
column_names |
list[str] | Column headers for context |
original_format |
string | Source format: csv, xlsx, xls |
Example payload for a CSV row:
{
"domain": "car",
"document_type": "faq",
"language": "fa",
"source_file": "car_insurance_faq.csv",
"chunk_index": 0,
"content_hash": "a1b2c3...",
"created_at": "2025-01-15T10:30:00Z",
"row_number": 1,
"column_names": ["question", "answer", "category"],
"original_format": "csv"
}
Document Metadata (Optional, Future)
Reserved fields for future use:
policy_version: stringeffective_date: datedepartment: stringregion: string
Consequences
Positive
- Consistent filtering — all chunks have core fields for retrieval filtering
- Citation support — source file + chunk index enables precise citations
- Deduplication — content hash identifies duplicate content
- Tabular context — column names preserved for row-based chunks
Negative
- Manual tagging — domain and document_type require human input or heuristics at ingestion
- Storage overhead — metadata stored with each chunk
Neutral
- Schema is minimal but extensible
- Column names may vary across files — need normalization strategy
Implementation Notes
- Define metadata validator at ingestion time
- Auto-detect language using
langdetector similar - Derive domain from file path or manual mapping
- Compute content_hash before insertion
- Store
created_atin ISO 8601 format for consistency