feat(autoindex): name XLSX columns from a two-row grouped header (stacked on #1132) - #1133
Open
thomas-villani wants to merge 2 commits into
Open
thomas-villani wants to merge 2 commits into
thomas-villani wants to merge 2 commits into
Conversation
… cover Motivation: a cell merged DOWN over several rows stores its value only in the top row. pandas writes exactly that for every MultiIndex export (`df.groupby([...]).sum().to_excel(...)`, merge_cells=True is the default), and hand-made reports do it for a category beside its line items. xerj-org#1124 read only the top cell, so for a pandas groupby export with region merged A2:A4 / A5:A6: before Sales!r3 {"product":"B","sales":20} (no region) after Sales!r3 {"region":"East","product":"B","sales":20} `region: East` matched 1 of East's 3 rows and a `terms` aggregation on region undercounted. Same for a hand-made report's merged Category column. Mechanism: `<mergeCells>` follows `<sheetData>`, but rows are streamed and emitted as read. So each sheet gets a pre-pass, `scan_merges`, that streams the part once WITHOUT parsing cells: it byte-searches (memchr::memmem) for the `<mergeCells` tag — unambiguous, since `<` cannot appear unescaped in XML text — and XML-parses only the tail from there (capped at 16 MB), which keeps a following `<hyperlink ref=...>` from being read as a merge. `FillDown` then gives each streamed row the value of any live vertical merge covering it, in the merge's own (top-left) column. Deliberately NOT done: - Merges across columns are not expanded: a title merged over a table's width must stay one cell, or it becomes a header-shaped row. A block merge (A3:B4) fills down column A only. - No row is invented: a covered row with no cells of its own stays absent. - A covered cell that has its own value (malformed file) keeps it. - Two-row headers (pandas MultiIndex columns: Q1 over Jan/Feb) are a separate change. Cost, measured (release, median of 3, rust:latest container; full time includes the pre-pass): 300k x 8 sheet, 112.6 MB XML full 6.274 s pre-pass 0.640 s (10.2%) pandas groupby, 1.5 MB, 1204 vertical merges full 0.067 s pre-pass 0.011 s (16.8%) The pre-pass is charged to the decompression budget like any read. A sampling run (phase A / --dry-run, 500 rows per sheet) must stay cheap, so it skips the pre-pass on sheets declared over 64 MB decompressed (~0.35 s of scanning at the measured rate); such a sample has no fill-down. Full runs always scan. Tests (extract/xlsx.rs): - a_vertical_merge_gives_its_value_to_every_row_it_covers (pandas shape, merges out of order, hyperlink ref after the list) - merges_across_columns_are_not_expanded (merged title, block merge) - fill_down_invents_no_rows_and_overwrites_no_values - the_merge_scan_finds_the_tag_and_only_the_tag (prefixed tag, tag split across 7-byte reads, "mergeCells" in cell text, tail cap) - a_sampling_run_skips_the_merge_scan_only_on_a_large_sheet - cell_references_parse_to_column_and_row With `fill.apply` disabled, the three behavior tests fail; with it, all pass. Real pandas/openpyxl workbooks (MultiIndex rows, merged-title report) checked with a temporary probe, removed before commit. Evidence: cargo test -p xerj-autoindex --lib: 1195 passed, 0 failed, 2 ignored. cargo build --release -p xerj-autoindex ok; clippy -D warnings clean; cargo fmt --check clean. Written by an AI agent (Claude Code) on behalf of the PR author.
Motivation: pandas writes a DataFrame with MultiIndex columns as a header
two rows deep, the outer level merged across its inner columns:
row 1: Q1 (B1:C1) Q2 (D1:E1)
row 2: Jan Feb Jan Feb
row 4: East 1 2 3 4
Header detection passes over row 1 (a narrower candidate above a wider
one is a title line), so the fields came out as Jan, Feb, Jan_2, Feb_2:
the quarter was lost and two columns were named by a dedup suffix.
Hand-built reports with "Actual | Budget" bands over the same month
columns have the same shape.
Change (extract/xlsx.rs):
- scan_merges keeps every multi-cell range as a `Merge` (it kept only
ranges more than one row tall). FillDown takes the vertical ones, so
fill-down is unchanged.
- group_prefixes: the row directly above the header counts as group
labels only when it holds only text, has a cell merged across two or
more header columns, and no merge spans every header column (that is a
title over the table, not a group). Each of its cells then prefixes the
header columns it spans; an unmerged cell prefixes its own column,
since pandas does not merge a one-column group. A cell whose merge runs
down into the header row (`Region` in A1:A2) is part of the header and
gives no prefix. The name is sanitize_field_name("Q1 Jan") = Q1_Jan.
- Sample/full consistency: the sampling run skips the merge scan on a
sheet over MAX_SAMPLE_MERGE_SCAN_BYTES (64 MB decompressed), and the
sample types the fields the full run indexes, so a column's name must
not depend on the run. Header merges are therefore used only on sheets
both runs scan; over that size a two-row header keeps its lower row,
as before. Documented in the module's known limits.
Evidence:
- 5 new/updated tests; cargo test -p xerj-autoindex --lib: 1200 passed,
0 failed, 2 ignored (1195 before). clippy -D warnings clean, fmt clean,
release build ok.
- Mutation checks, each run and then reverted: disabling the prefixing
fails the 4 positive tests; dropping the full-width-title rule fails 3
(the shared test fixture has always carried an A1:D1 title merge, so
every older header test now also guards that rule); dropping the
vertical-label skip fails 1.
- pandas 2.x output (`df.to_excel` with MultiIndex columns), extracted in
a temporary probe test: records are
{"col_A":"East","Q1_Jan":1,"Q1_Feb":2,"Q2_Jan":3,"Q2_Feb":4}, the same
with a 500-row sample and with no limit. A report with a title merged
over its width and the two fill-down workbooks from the previous PR
extract unchanged.
Stacked on feat/xlsx-merged-cells (the vertical fill-down PR).
thomas-villani
added a commit
to thomas-villani/xerj
that referenced
this pull request
Oct 4, 2026
…org#1132-xerj-org#1134) Motivation: xerj-org#1132 fills vertical merges down, xerj-org#1133 names columns from a two-row grouped header (Q1_Jan) and xerj-org#1134 adds a roff man(7) extractor. The format lists and the fact-check matrix should say so. This commit must merge AFTER those three code PRs, because it describes them as shipped. - README, landing/llms.txt, landing/llms-full.txt: man pages added to the format lists (llms-full: plain or gzipped, one record per section, titled NAME(SECT), mdoc(7) pages stay plain text). The XLSX entry in llms-full now says vertical merges are filled down, a two-row grouped header is combined, and merges across columns are not expanded. It replaces "merged cells not expanded". - scripts/seo/claims_rules.py: new GREEN THING row "Man pages (roff man(7))" citing extract/man.rs:1, and its gate says mdoc is not parsed. The .xlsx gate gets the same merged-cell wording as llms-full. - testdata/factcheck: new good_man_answer.md fixture (git check-ignore confirms the !scripts/seo/testdata/**/*.md re-include applies), and good_xlsx_answer.md loses the old "merged cells are not expanded" line. Verified: factcheck --self-test OK (35 THING rows); --fail-on error 0 ERROR; build_articles --check current; landing-constants-guard passes. --fixture-check reports 1 false positive (FC-EV-DANGLING) on this branch alone, because good_man_answer.md cites extract/man.rs, which only exists once xerj-org#1134 merges. With man.rs from xerj-org#1134 checked out it reports 0 false positives and 13/13 good fixtures clean.
2 of 7 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes, and why
pandas writes a DataFrame with MultiIndex columns as a header two rows deep, with each outer label merged across its inner columns:
Header detection passes over row 1, because a narrower candidate above a wider one is treated as a title line. The fields therefore came out as
Jan, Feb, Jan_2, Feb_2: the quarter was lost, and two columns were named only by a dedup suffix. Hand-built reports withActual | Budgetbands over the same month columns have the same shape.Now the row directly above the header names the columns under it:
Q1_Jan, Q1_Feb, Q2_Jan, Q2_Feb.When the upper row counts as group labels (
group_prefixes). All of these must hold:When the row qualifies:
RegioninA1:A2) belongs to the header, so it gives no prefix. feat(autoindex): fill vertically merged XLSX cells down the rows they cover #1132's fill-down already puts it into the header row.Sample/full consistency. The sampling run skips the merge scan on any sheet over 64 MB decompressed (#1132). The sample types the fields the full run indexes, so a column's name must not depend on which run read the sheet. Header merges are therefore used only on sheets that both runs scan. Above that size, a two-row header keeps only its lower row, as it does today; this is listed in the module's known limits.
scan_mergesnow returns every multi-cell range as aMerge; before, it kept only ranges more than one row tall.FillDowntakes the vertical ones, so fill-down behavior is unchanged.Evidence
Mutation checks. Each mutation was applied, the xlsx tests were run, and the mutation was reverted:
A1:D1title merge, so the older header tests also guard this rule.Real file: a pandas 2.x
df.to_excelwith MultiIndex columns, extracted with a temporary probe test that was removed before commit:The output is identical with a 500-row sample and with no limit. These still extract exactly as they do on #1132:
Checks
cargo fmt --all(--checkclean)cargo build --release -j 32 -p xerj-autoindexcargo test -p xerj-autoindex --libpasses<mergeCells>, scan failure, or a large sheet), names are exactly what they were before.Not run: no throughput benchmark. The change adds one pass over the merge list per sheet, run when the header is found, plus a filter when
FillDownis built. The per-row path is unchanged.Provenance
🤖 Generated with Claude Code