82 lines
2.6 KiB
Markdown
82 lines
2.6 KiB
Markdown
# ADR 0008: Chunk Metadata Schema
|
|
|
|
## Status
|
|
|
|
Accepted
|
|
|
|
## Context
|
|
|
|
Each chunk in Qdrant needs metadata for filtering, citation, and debugging. The schema must balance flexibility with consistency, and handle both document-based (PDF, Word) and tabular (CSV, Excel) sources.
|
|
|
|
## Decision
|
|
|
|
### Core Metadata Schema
|
|
|
|
| Field | Type | Required | Description |
|
|
|-------|------|----------|-------------|
|
|
| `domain` | string | Yes | Insurance domain: `car`, `health`, `fire`, `life`, `travel`, etc. |
|
|
| `document_type` | string | Yes | Document category: `policy`, `faq`, `claim`, `terms`, `guidance`, etc. |
|
|
| `language` | string | Yes | Language code: `fa` (Persian), `en` (English) |
|
|
| `source_file` | string | Yes | Original filename with extension |
|
|
| `chunk_index` | int | Yes | Zero-based position in source document |
|
|
| `content_hash` | string | Yes | SHA-256 hash of chunk content |
|
|
| `created_at` | datetime | Yes | Ingestion timestamp |
|
|
|
|
### Tabular Data Extensions
|
|
|
|
For CSV and Excel sources, additional fields:
|
|
|
|
| Field | Type | Description |
|
|
|-------|------|-------------|
|
|
| `row_number` | int | Row index in original file |
|
|
| `column_names` | list[str] | Column headers for context |
|
|
| `original_format` | string | Source format: `csv`, `xlsx`, `xls` |
|
|
|
|
Example payload for a CSV row:
|
|
```json
|
|
{
|
|
"domain": "car",
|
|
"document_type": "faq",
|
|
"language": "fa",
|
|
"source_file": "car_insurance_faq.csv",
|
|
"chunk_index": 0,
|
|
"content_hash": "a1b2c3...",
|
|
"created_at": "2025-01-15T10:30:00Z",
|
|
"row_number": 1,
|
|
"column_names": ["question", "answer", "category"],
|
|
"original_format": "csv"
|
|
}
|
|
```
|
|
|
|
### Document Metadata (Optional, Future)
|
|
|
|
Reserved fields for future use:
|
|
- `policy_version`: string
|
|
- `effective_date`: date
|
|
- `department`: string
|
|
- `region`: string
|
|
|
|
## Consequences
|
|
|
|
### Positive
|
|
- **Consistent filtering** — all chunks have core fields for retrieval filtering
|
|
- **Citation support** — source file + chunk index enables precise citations
|
|
- **Deduplication** — content hash identifies duplicate content
|
|
- **Tabular context** — column names preserved for row-based chunks
|
|
|
|
### Negative
|
|
- **Manual tagging** — domain and document_type require human input or heuristics at ingestion
|
|
- **Storage overhead** — metadata stored with each chunk
|
|
|
|
### Neutral
|
|
- Schema is minimal but extensible
|
|
- Column names may vary across files — need normalization strategy
|
|
|
|
## Implementation Notes
|
|
|
|
- Define metadata validator at ingestion time
|
|
- Auto-detect language using `langdetect` or similar
|
|
- Derive domain from file path or manual mapping
|
|
- Compute content_hash before insertion
|
|
- Store `created_at` in ISO 8601 format for consistency
|