Backlog — PDF & ingestion follow-ups
Items deferred from PDF ingestion v1 (Text PDF + PyMuPDF + Heading Reconstruction).
| Item | When to pick up |
|---|---|
| Scanned PDF OCR | After Text PDF Experiments are stable; chatbot needs image PDFs |
| Persian / cloud OCR | Local OCR quality on Farsi sample set is too low |
| pdfplumber tables | PyMuPDF table flattening hurts Experiment scores |
| LibreOffice PDF→DOCX fallback | Heading Reconstruction false-positives dominate; want second opinion path |
| Multi-column layout | Reading order is wrong on multi-column policies |
| Page metadata in chunks | Users need “page N” citations in answers / Chunk Preview |