Files
chatbot_v3/docs/adr/0018-docx-and-spreadsheet-parsing-with-fixed-size-chunking.md
Ali Zarinkolah 5cdfb70085 feat(ingestion): add DOCX/CSV/XLSX parsing and fixed-size chunking (ADR-0018)
Adds src/application/ingestion/ -- Persian normalization, DOCX body
walk with structural data/layout table classification, CSV/XLSX row
rendering, and fixed-size token chunking (cl100k_base, 400/60/512) --
as pure functions per ADR-0015, tested against real production
documents (asia_data_sample, kept out of the repo). ADR-0018 records
where this diverges from ADR-0004 (fixed-size default, no invented
headings/tree, structural table classification, header-provable
labeling only). Plan 001's scope line is corrected from CSV-only to
DOCX/XLSX/CSV, and CLAUDE.md's stale project-status paragraph is
updated to match current implementation state.
2026-08-18 10:22:17 +03:30

16 KiB
Raw Permalink Blame History

0018. DOCX and spreadsheet parsing with fixed-size chunking

Status

Proposed

Context

ADR-0004 specified the full parsing and chunking design: structural extraction before chunking (table rows and Q&A pairs kept atomic, prose chunked separately), semantic-aware chunking as the default, image captioning via a vision API, and legacy .doc conversion through headless LibreOffice. Plan 001's first vertical slice needs a working parser now, and that full design is substantially more work than the slice can absorb. This ADR records what v1 actually ships and why it differs, so the code does not silently contradict an Accepted ADR.

Four forces shaped the decision:

A working extractor already exists. The chunking_strategies_evaluation repository — the harness the user built to compare chunking strategies against this same Farsi corpus — contains a DOCX extractor that walks the document body in reading order and handles the "prose lives inside table cells" pattern common in Farsi documents exported from older Word versions. Roughly 200 lines of it are production-quality; the rest is evaluation scaffolding (five competing strategies, an LLM-as-judge benchmark, a dashboard, an HTML report generator). Porting it is cheaper and better-tested against real documents than writing a parser from scratch.

Semantic chunking does not fit the inline request. ADR-0004 chose semantic-aware chunking as the default on the strength of the user's own offline accuracy comparison, in which fixed-size was the "viable, simpler runner-up". That comparison measured retrieval accuracy, not ingestion cost. Semantic boundary detection requires embedding every sentence before chunk boundaries can be decided — a second network round-trip pass inside the request budget ADR-0017 bounds with INGESTION_TIMEOUT_SECONDS. ADR-0017's own cost table already assumes the cheaper strategy, listing "Chunk (fixed-size, ADR-0004) | Blocking CPU, pure Python | Negligible". ADR-0004 and ADR-0017 are therefore already in tension, and this ADR resolves it toward ADR-0017 for v1. Fixed-size is acceptable specifically because ADR-0001 gives every chunk previous_chunk_id/next_chunk_id pointers: a chunk boundary that cuts a thought in half is recoverable by expanding to neighbors at retrieval time (ADR-0003).

Nothing normalizes the text the dense embedders see. ADR-0005 resolved the sparse analyzer as bm25-fa-norm-stop — "normalization + stopword removal", computed in our own BM25 pipeline outside Qdrant. That covers the sparse vector only. Persian text authored on mixed Arabic/Persian keyboards contains both ک (U+06A9) and ك (U+0643), both ی (U+06CC) and ي (U+064A); these are distinct codepoints and therefore distinct tokens to nomic-embed-text-v2-moe and text-embedding-3-large alike, so the same Persian word can embed two different ways depending on which key the author pressed. Word documents in this corpus reliably contain both forms.

The corpus is spreadsheets more than it is CSVs. ADR-0004 inspected the real sample and found the tabular files are entirely .xlsx; no .csv exists in practice. Plan 001 and the source_files.source_type CHECK constraint both name csv. The row-to-chunk mapping is identical either way — only the reader differs — so v1 reads both rather than forcing a manual export step that would silently drop the merged-cell values ADR-0004 warns about.

Decision

1. Formats

v1 ingests .docx, .csv, and .xlsx.

.doc is rejected with 415 Unsupported Media Type. ADR-0004 specified conversion via soffice --headless --convert-to docx; that is a subprocess with a multi-second startup cost running inside the inline request ADR-0017 defines, and it adds a system binary to the container image. Conversion is deferred to an out-of-process HTTP conversion service, tracked in the backlog. source_files.source_type continues to allow doc so the row can be recorded once conversion lands.

2. Structural units before chunking

Every source file is decomposed into ordered structural units before any chunking runs, as ADR-0004 requires. There are two kinds:

  • PARAGRAPH — a run of flowing prose. Consecutive paragraphs accumulate into one unit rather than one unit each, so the splitter sees flowing text instead of a series of one-sentence fragments ("the whole remaining run of paragraphs", per ADR-0004).
  • TABLE_ROW — one row of a data table. Atomic; split only when a single row exceeds the model's sequence length.

Walk doc.element.body children in document order — not doc.paragraphs and doc.tables separately — so tables interleaved with paragraphs keep their position. This is ADR-0004's rule, unchanged.

No headings are invented. A real Heading N Word style becomes a # prefix on its paragraph; where a document declares none, none appear. The evaluation repository upgrades paragraphs to headings by text pattern (^بخش, ^\d+[-.]\d+, leading *); those patterns are tuned to a Farsi regulatory corpus, not an insurance one, and a wrongly-detected heading silently reshapes the document in a way that is hard to notice downstream. They are not adopted.

There is also no document tree. An earlier draft built a Document > Section > Article > Paragraph hierarchy from heading styles. Not one document in the sample corpus carries a single Heading style, so that tree was flat in every real case, and nothing consumed it — chunking works from the unit list, and ADR-0001's payload has no tree field. It is not built.

3. Data tables against layout tables

A docx table is either data or page furniture, and the two need opposite treatment. The classification is structural, never a reading of content: a data cell is by definition small enough to be a chunk, so a table containing a cell that alone exceeds chunk_size, or a cell containing nested tables, is a layout container. In the sample corpus this separates by two orders of magnitude — 24 to 171 tokens for the largest cell of each data table, against 54,007 tokens for a cell holding an entire sub-document across 738 paragraphs and 5 nested tables.

  • Data table → one TABLE_ROW unit per row, rendered by the same code that renders spreadsheet rows.
  • Layout table → its cells are prose, recursed into and folded into the surrounding prose block.

4. Table rows are labeled only when a header is provable

Applied to docx tables and spreadsheets alike:

  • A header is row 0 or nothing. Never scan further down for a header-shaped row. Scanning discarded every row above the match and then labeled the rest from a data row, turning a 30-row compensation table into chunks reading 80: 70.
  • Leading rows that are structurally a merged banner — fewer than two populated cells, or one value repeated across the row — are skipped first. That is a fact about the merge, not a guess about meaning.
  • Row 0 is accepted as a header only when it is inconsistent with the column beneath it: a text label above a numeric column, or a short label above much longer values. This is the test csv.Sniffer.has_header uses; it is a property of the table rather than a pattern borrowed from one document.
  • When no header is provable, cells are joined with " | " — unlabeled, but never mislabeled. Losing a label is recoverable at retrieval time; labeling every row from a data row is not.
  • A header merged vertically across two rows resolves to the same text in the row below it; that duplicate is skipped rather than emitted as data.
  • A cell merged across columns is reported once per grid position it spans; those repeats are collapsed.

is a pipeline invariant

5. Persian normalization is a pipeline invariant

Every extracted text block is normalized before chunking, for all formats:

  • Arabic to Persian letter folding: ك→ک, ي→ی, ى→ی, أ/إ→ا
  • unicodedata.normalize("NFKC")
  • Removal of harakat (diacritics) and tatweel
  • ¬ → space, then collapse runs of whitespace

Digits and punctuation are not rewritten. Persian digits (۱۲۳) and Persian punctuation (؛, ٬) are left as authored, because chunk content is what citations render back to the user and Western digits inside Persian prose read as wrong.

Normalization runs per text block, before the markdown is assembled — the whitespace-collapse step maps \n to a space, so applying it to an assembled document would flatten every heading and paragraph onto a single line.

This complements rather than replaces ADR-0005's sparse-side normalization, which additionally removes stopwords and is specific to the BM25 vector. It also stabilizes the text that feeds content_hash.

6. Spreadsheet handling

ADR-0004's rules stand, under the header discipline of section 4: each cell is rendered as "{column_header}: {cell_value}", merged cell ranges are forward-filled before rendering, and sheets with no non-empty data rows are skipped.

Forward-filling merges is not cosmetic. openpyxl stores a merged range's value only in its top-left cell, so in the branch directory the province is present on the first branch of each province and absent from every other one. Filling the range makes each row-chunk self-contained — a branch carries its province even though the source cell is blank.

One correction: .csv is read with the standard library's csv module, not pandas as ADR-0004 states. Adding pandas for delimiter handling and row iteration is not warranted. .xlsx uses openpyxl, as ADR-0004 assumed.

There is no sheet-shape sniffing. A two-column Q&A sheet and a branch-directory sheet go through the same generic renderer; a Q&A row renders as question: …\nanswer: … and a branch row as branch_name: …\ncity: …, both self-describing without a schema heuristic that could misfire.

7. Chunking

The fixed_size strategy, over tokens counted with tiktoken cl100k_base:

Setting Value
chunk_size 400 tokens
chunk_overlap 60 tokens
max_chunk_tokens 512 (hard cap)

These numbers are recorded here because they exist in no ADR today — ADR-0001 explicitly left "size/overlap are tunable config, not fixed by this ADR" open, and ADR-0004 gives only the 512 ceiling.

cl100k_base is a deliberate proxy. text-embedding-3-large has an 8191-token window and never binds; nomic-embed-text-v2-moe's 512-token sequence length is the only real constraint. cl100k tokenizes Persian inefficiently while nomic's multilingual tokenizer does not, so a cl100k count reliably over-estimates the nomic count — measuring with cl100k and capping at 512 is safe in the conservative direction, without shipping a second tokenizer and its model download into the ingestion path. The 400/512 gap leaves headroom for the mandatory search_document: task prefix (ADR-0004) and any heading text carried into a chunk.

Spreadsheet rows are atomic and bypass the splitter. A row that exceeds max_chunk_tokens falls through the fixed-size splitter in place, emitting several ordered chunks, rather than being silently truncated by the embedding model.

8. content_type values emitted

ADR-0004's four-value set is unchanged. v1 emits paragraph (DOCX prose and flattened tables) and table_row (spreadsheet rows). qa_pair and image_caption remain defined but are not produced.

9. Deferred, not rejected

Semantic-aware chunking; DOCX table_row and qa_pair structural detection; embedded-image captioning; .doc conversion; Matryoshka-256 truncation. Each remains ADR-0004's stated intent; this ADR only records that v1 does not ship them.

Consequences

Positive

  • Plan 001's ingestion slice is unblocked with a parser already proven against this specific Farsi corpus, rather than one written speculatively.
  • Chunking stays pure, synchronous, and cheap — it fits ADR-0017's inline request budget with no network round-trip, and runs safely under anyio.to_thread.run_sync.
  • Persian normalization closes a real defect that would otherwise degrade both dense vectors silently, with no error and no obvious symptom.
  • Concrete chunk-size numbers and their rationale are now recorded somewhere other than a config default, so a later change is a visible decision.
  • Q&A sheets, branch directories, and docx contact tables are all served by one renderer with no schema sniffing, so a new sheet shape needs no new code.
  • Refusing to label a table whose header is unprovable means the parser degrades to unlabeled rows instead of producing confidently wrong ones, which is the failure mode that is hard to notice downstream.

Negative

  • v1 ships the strategy the user's own comparison ranked second. Retrieval accuracy is expected to be measurably lower than semantic chunking would give. Neighbor expansion via ADR-0001's pointers is the mitigation, and it is unproven at this scale.
  • A table whose header cannot be proven — short text over short text, which is genuinely ambiguous — produces unlabeled " | " rows. A reader or model can still see the values but not which column each belongs to.
  • The layout-table rule keys on chunk_size, so changing that setting silently changes which tables are treated as data. The observed margin is two orders of magnitude, so this is unlikely to flip in practice, but it is a coupling.
  • A DOCX that alternates literal سوال:/پاسخ: paragraphs loses question/answer atomicity; a boundary can fall between a question and its answer, because qa_pair detection is deferred.
  • cl100k is a proxy for nomic's tokenizer. The relationship is safe in the conservative direction for Persian, but a document in another language could in principle tokenize the other way; the 512 assertion is what catches it.
  • Normalizing stored content means the text served in citations is not byte-identical to the source document. Letter folding was chosen over full normalization specifically to keep this difference invisible to a reader.
  • .doc files are rejected outright rather than converted, so any legacy document must be re-saved by hand until the conversion service lands.

Alternatives Considered

  • Implement ADR-0004 in full now (DOCX table-row detection with header inference, Q&A pair heuristics, image captioning, semantic boundary detection): rejected for v1 as roughly triple the work, none of which exists in the evaluation repository to port, and which would block the first ingestion slice on parser research.
  • Semantic chunking inside the inline request: rejected — it adds a per-sentence embedding pass to a request already bounded by INGESTION_TIMEOUT_SECONDS, and ADR-0017 chose inline ingestion on the assumption that chunking is negligible. Revisit when ingestion moves back off the request path.
  • The nomic-embed-text-v2-moe tokenizer for exact chunk sizing: rejected — it requires transformers/tokenizers and a model file download in the ingestion path to buy precision that the conservative cl100k over-estimate already provides.
  • Character-based splitting (no tokenizer at all): rejected — Persian characters-per-token varies enough that a character budget cannot guarantee the 512-token ceiling that actually matters.
  • Full Persian normalization including digit unification (۱۲۳→123) and punctuation mapping: rejected — it would improve lexical matching slightly when a query uses the other digit form, at the cost of rendering Persian citations with Western digits.
  • Storing raw and normalized text as separate payload fields: rejected — ADR-0001 fixes the payload field list, and doubling the stored text per point is not justified when letter folding alone is visually lossless.
  • Row-chunking DOCX tables like spreadsheets: rejected for v1 — the Farsi .doc exports in this corpus use tables as page layout, with ordinary prose inside cells, so treating every row as an atomic unit would shred paragraphs mid-sentence. Revisit once real chunk output from the table-heavy documents has been inspected.