docs: capture pdf domain language and deferred ingestion backlog
Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
25
backlog/scanned-pdf-ocr.md
Normal file
25
backlog/scanned-pdf-ocr.md
Normal file
@@ -0,0 +1,25 @@
|
||||
# Scanned PDF OCR (v1.1+)
|
||||
|
||||
## Why deferred
|
||||
|
||||
v1 only accepts Text PDFs and applies a Text-layer Gate: Scanned PDFs are hard-rejected at upload (not stored, not processed). OCR is a second product surface: noisy text poisons every Strategy equally and makes Experiment comparisons misleading.
|
||||
|
||||
## Trigger
|
||||
|
||||
Pick this up when:
|
||||
|
||||
1. Text PDF path is trusted on real docs
|
||||
2. You have ~10–20 labeled Scanned PDFs + question JSONs matching chatbot traffic
|
||||
3. Production coverage of those docs matters more than clean benchmarks alone
|
||||
|
||||
## Approach to try
|
||||
|
||||
1. Detect empty / near-empty text layer at upload (same gate as v1 rejection)
|
||||
2. Run OCR as an **explicit pre-step** that still emits `ParseResult` (markdown + DocumentTree)
|
||||
3. Tag source as `ocr` so Experiments can filter Text PDF vs Scanned PDF runs
|
||||
4. Reuse Heading Reconstruction heuristics on OCR text (expect more false positives)
|
||||
|
||||
## Success criteria
|
||||
|
||||
- OCR’d markdown is readable enough that recursive Strategy headings aren’t nonsense
|
||||
- Side-by-side Experiment: Text PDF original vs OCR of a scan of the same doc — score delta understood, not mysterious
|
||||
Reference in New Issue
Block a user