Files
chunking_strategies_evaluation/backlog/scanned-pdf-ocr.md

1.1 KiB
Raw Blame History

Scanned PDF OCR (v1.1+)

Why deferred

v1 only accepts Text PDFs and applies a Text-layer Gate: Scanned PDFs are hard-rejected at upload (not stored, not processed). OCR is a second product surface: noisy text poisons every Strategy equally and makes Experiment comparisons misleading.

Trigger

Pick this up when:

  1. Text PDF path is trusted on real docs
  2. You have ~10–20 labeled Scanned PDFs + question JSONs matching chatbot traffic
  3. Production coverage of those docs matters more than clean benchmarks alone

Approach to try

  1. Detect empty / near-empty text layer at upload (same gate as v1 rejection)
  2. Run OCR as an explicit pre-step that still emits ParseResult (markdown + DocumentTree)
  3. Tag source as ocr so Experiments can filter Text PDF vs Scanned PDF runs
  4. Reuse Heading Reconstruction heuristics on OCR text (expect more false positives)

Success criteria

  • OCR’d markdown is readable enough that recursive Strategy headings aren’t nonsense
  • Side-by-side Experiment: Text PDF original vs OCR of a scan of the same doc — score delta understood, not mysterious