docs: capture pdf domain language and deferred ingestion backlog

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
2026-08-02 16:59:42 +03:30
parent 987b0493ac
commit 118255acdc
8 changed files with 166 additions and 1 deletions

View File

@@ -0,0 +1,25 @@
# Scanned PDF OCR (v1.1+)
## Why deferred
v1 only accepts Text PDFs and applies a Text-layer Gate: Scanned PDFs are hard-rejected at upload (not stored, not processed). OCR is a second product surface: noisy text poisons every Strategy equally and makes Experiment comparisons misleading.
## Trigger
Pick this up when:
1. Text PDF path is trusted on real docs
2. You have ~10–20 labeled Scanned PDFs + question JSONs matching chatbot traffic
3. Production coverage of those docs matters more than clean benchmarks alone
## Approach to try
1. Detect empty / near-empty text layer at upload (same gate as v1 rejection)
2. Run OCR as an **explicit pre-step** that still emits `ParseResult` (markdown + DocumentTree)
3. Tag source as `ocr` so Experiments can filter Text PDF vs Scanned PDF runs
4. Reuse Heading Reconstruction heuristics on OCR text (expect more false positives)
## Success criteria
- OCR’d markdown is readable enough that recursive Strategy headings aren’t nonsense
- Side-by-side Experiment: Text PDF original vs OCR of a scan of the same doc — score delta understood, not mysterious