feat(autoindex): fill vertically merged XLSX cells down the rows they cover - #1132
Open
thomas-villani wants to merge 1 commit into
Open
thomas-villani wants to merge 1 commit into
thomas-villani wants to merge 1 commit into
Conversation
… cover Motivation: a cell merged DOWN over several rows stores its value only in the top row. pandas writes exactly that for every MultiIndex export (`df.groupby([...]).sum().to_excel(...)`, merge_cells=True is the default), and hand-made reports do it for a category beside its line items. xerj-org#1124 read only the top cell, so for a pandas groupby export with region merged A2:A4 / A5:A6: before Sales!r3 {"product":"B","sales":20} (no region) after Sales!r3 {"region":"East","product":"B","sales":20} `region: East` matched 1 of East's 3 rows and a `terms` aggregation on region undercounted. Same for a hand-made report's merged Category column. Mechanism: `<mergeCells>` follows `<sheetData>`, but rows are streamed and emitted as read. So each sheet gets a pre-pass, `scan_merges`, that streams the part once WITHOUT parsing cells: it byte-searches (memchr::memmem) for the `<mergeCells` tag — unambiguous, since `<` cannot appear unescaped in XML text — and XML-parses only the tail from there (capped at 16 MB), which keeps a following `<hyperlink ref=...>` from being read as a merge. `FillDown` then gives each streamed row the value of any live vertical merge covering it, in the merge's own (top-left) column. Deliberately NOT done: - Merges across columns are not expanded: a title merged over a table's width must stay one cell, or it becomes a header-shaped row. A block merge (A3:B4) fills down column A only. - No row is invented: a covered row with no cells of its own stays absent. - A covered cell that has its own value (malformed file) keeps it. - Two-row headers (pandas MultiIndex columns: Q1 over Jan/Feb) are a separate change. Cost, measured (release, median of 3, rust:latest container; full time includes the pre-pass): 300k x 8 sheet, 112.6 MB XML full 6.274 s pre-pass 0.640 s (10.2%) pandas groupby, 1.5 MB, 1204 vertical merges full 0.067 s pre-pass 0.011 s (16.8%) The pre-pass is charged to the decompression budget like any read. A sampling run (phase A / --dry-run, 500 rows per sheet) must stay cheap, so it skips the pre-pass on sheets declared over 64 MB decompressed (~0.35 s of scanning at the measured rate); such a sample has no fill-down. Full runs always scan. Tests (extract/xlsx.rs): - a_vertical_merge_gives_its_value_to_every_row_it_covers (pandas shape, merges out of order, hyperlink ref after the list) - merges_across_columns_are_not_expanded (merged title, block merge) - fill_down_invents_no_rows_and_overwrites_no_values - the_merge_scan_finds_the_tag_and_only_the_tag (prefixed tag, tag split across 7-byte reads, "mergeCells" in cell text, tail cap) - a_sampling_run_skips_the_merge_scan_only_on_a_large_sheet - cell_references_parse_to_column_and_row With `fill.apply` disabled, the three behavior tests fail; with it, all pass. Real pandas/openpyxl workbooks (MultiIndex rows, merged-title report) checked with a temporary probe, removed before commit. Evidence: cargo test -p xerj-autoindex --lib: 1195 passed, 0 failed, 2 ignored. cargo build --release -p xerj-autoindex ok; clippy -D warnings clean; cargo fmt --check clean. Written by an AI agent (Claude Code) on behalf of the PR author.
4 of 7 tasks
thomas-villani
added a commit
to thomas-villani/xerj
that referenced
this pull request
Oct 4, 2026
…org#1132-xerj-org#1134) Motivation: xerj-org#1132 fills vertical merges down, xerj-org#1133 names columns from a two-row grouped header (Q1_Jan) and xerj-org#1134 adds a roff man(7) extractor. The format lists and the fact-check matrix should say so. This commit must merge AFTER those three code PRs, because it describes them as shipped. - README, landing/llms.txt, landing/llms-full.txt: man pages added to the format lists (llms-full: plain or gzipped, one record per section, titled NAME(SECT), mdoc(7) pages stay plain text). The XLSX entry in llms-full now says vertical merges are filled down, a two-row grouped header is combined, and merges across columns are not expanded. It replaces "merged cells not expanded". - scripts/seo/claims_rules.py: new GREEN THING row "Man pages (roff man(7))" citing extract/man.rs:1, and its gate says mdoc is not parsed. The .xlsx gate gets the same merged-cell wording as llms-full. - testdata/factcheck: new good_man_answer.md fixture (git check-ignore confirms the !scripts/seo/testdata/**/*.md re-include applies), and good_xlsx_answer.md loses the old "merged cells are not expanded" line. Verified: factcheck --self-test OK (35 THING rows); --fail-on error 0 ERROR; build_articles --check current; landing-constants-guard passes. --fixture-check reports 1 false positive (FC-EV-DANGLING) on this branch alone, because good_man_answer.md cites extract/man.rs, which only exists once xerj-org#1134 merges. With man.rs from xerj-org#1134 checked out it reports 0 false positives and 13/13 good fixtures clean.
2 of 7 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes, and why
A cell merged down over several rows stores its value only in the top row. pandas does this for every MultiIndex export (
df.groupby([...]).sum().to_excel(...), wheremerge_cells=Trueis the default), and hand-made reports do it for a category label beside its line items. #1124 read only the top cell, so the value was missing from every other row:region: Eastmatched 1 of East's 3 rows, and atermsaggregation onregionundercounted.Mechanism.
<mergeCells>comes after<sheetData>, but rows are streamed and emitted as they are read. So each sheet gets a pre-pass,scan_merges, that streams the part once without parsing cells:<mergeCellstag withmemchr::memmem. The search is unambiguous, because<cannot appear unescaped in XML text.<hyperlink ref=…>is not mistaken for a merge.Then
FillDowngives each streamed row the value of any vertical merge that covers it, in the merge's own top-left column.Deliberately not done:
Jan,Jan_2) are left for a follow-up PR.Sampling. A sampling run (phase A and
--dry-run, 500 rows per sheet) has to stay cheap, so it skips the pre-pass on sheets whose declared decompressed size is over 64 MB. A sample of such a sheet gets no fill-down. Full runs always scan. If you'd rather keep sampling and full runs identical regardless of cost, it's a one-line change.Evidence
Release build, median of 3 runs, in a
rust:latestcontainer. The full time includes the pre-pass.There are six new tests:
a_vertical_merge_gives_its_value_to_every_row_it_coversmerges_across_columns_are_not_expandedfill_down_invents_no_rows_and_overwrites_no_valuesthe_merge_scan_finds_the_tag_and_only_the_tag(a prefixed tag, a tag split across 7-byte reads,mergeCellsinside cell text, the tail cap)a_sampling_run_skips_the_merge_scan_only_on_a_large_sheetcell_references_parse_to_column_and_rowWith
fill.applydisabled, the three behavior tests fail:With the fix:
Checks
cargo fmt --all(--checkclean)cargo build --release -j 32 -p xerj-autoindexcargo test -p xerj-autoindex --libpassesNot run / notes:
failure_resume_http_tests::a_run_reports_progress_through_every_phase_and_closes_the_streamfails when run alone on this machine (3/3). It fails identically on cleanmain(3/3), and it passes inside the full suite. It asserts on wall-clock time ("a sub-2s run relays exactly one bar"), and the host was under heavy load from another job, so I'm treating it as environmental and haven't filed it.llms-full.txtand into the fact-check gate for.xlsx. Whichever of the two PRs lands second, I'll update that wording to "vertical merges filled down; merges across columns not expanded".Provenance
<mergeCells>block after<sheetData>. No Excel-authored file was tested. The 64 MB sampling threshold comes from one measured scan rate (about 176 MB/s on this host) and is not tuned beyond that.🤖 Generated with Claude Code