Skip to content

feat(autoindex): name XLSX columns from a two-row grouped header (stacked on #1132) - #1133

Open
thomas-villani wants to merge 2 commits into
xerj-org:mainfrom
thomas-villani:feat/xlsx-two-row-headers
Open

thomas-villani wants to merge 2 commits into
xerj-org:mainfrom
thomas-villani:feat/xlsx-two-row-headers

Conversation

@thomas-villani

Copy link
Copy Markdown
Contributor

Stacked on #1132. This branch builds on #1132's merge scan, so until #1132 merges, this diff also shows its commit (dc719492). The new change is the single commit ab19f6c3. Please review and merge #1132 first.

What this changes, and why

pandas writes a DataFrame with MultiIndex columns as a header two rows deep, with each outer label merged across its inner columns:

row 1:        Q1 (B1:C1)    Q2 (D1:E1)
row 2:        Jan   Feb     Jan   Feb
row 4: East   1     2       3     4

Header detection passes over row 1, because a narrower candidate above a wider one is treated as a title line. The fields therefore came out as Jan, Feb, Jan_2, Feb_2: the quarter was lost, and two columns were named only by a dedup suffix. Hand-built reports with Actual | Budget bands over the same month columns have the same shape.

Now the row directly above the header names the columns under it: Q1_Jan, Q1_Feb, Q2_Jan, Q2_Feb.

When the upper row counts as group labels (group_prefixes). All of these must hold:

  • It sits directly above the header.
  • It holds only text.
  • At least one of its cells is merged across two or more header columns.
  • No merge spans every header column. A title merged over the whole table labels the table, not a group.

When the row qualifies:

  • Each of its cells prefixes the header columns it spans.
  • An unmerged cell prefixes only its own column, because pandas does not merge a group that is one column wide.
  • A cell whose merge runs down into the header row (Region in A1:A2) belongs to the header, so it gives no prefix. feat(autoindex): fill vertically merged XLSX cells down the rows they cover #1132's fill-down already puts it into the header row.

Sample/full consistency. The sampling run skips the merge scan on any sheet over 64 MB decompressed (#1132). The sample types the fields the full run indexes, so a column's name must not depend on which run read the sheet. Header merges are therefore used only on sheets that both runs scan. Above that size, a two-row header keeps only its lower row, as it does today; this is listed in the module's known limits.

scan_merges now returns every multi-cell range as a Merge; before, it kept only ranges more than one row tall. FillDown takes the vertical ones, so fill-down behavior is unchanged.

Evidence

cargo test -p xerj-autoindex --lib     1200 passed; 0 failed; 2 ignored   (1195 before)
cargo clippy -p xerj-autoindex --all-targets -- -D warnings   clean
cargo build --release -j 32 -p xerj-autoindex                 ok

Mutation checks. Each mutation was applied, the xlsx tests were run, and the mutation was reverted:

mutation failing tests
prefixing disabled 4: all of the new positive tests
full-width-title rule removed 3. The shared test fixture has always carried an A1:D1 title merge, so the older header tests also guard this rule.
vertical-label skip removed 1

Real file: a pandas 2.x df.to_excel with MultiIndex columns, extracted with a temporary probe test that was removed before commit:

Quarters!r4 {"col_A":"East","Q1_Jan":1,"Q1_Feb":2,"Q2_Jan":3,"Q2_Feb":4}
Quarters!r5 {"col_A":"West","Q1_Jan":5,"Q1_Feb":6,"Q2_Jan":7,"Q2_Feb":8}

The output is identical with a 500-row sample and with no limit. These still extract exactly as they do on #1132:

  • an openpyxl report with a title merged over its width;
  • the MultiIndex-rows export;
  • the 14k-row groupby export.

Checks

  • cargo fmt --all (--check clean)
  • Scoped release build: cargo build --release -j 32 -p xerj-autoindex
  • cargo test -p xerj-autoindex --lib passes
  • ES-YAML conformance suite: not run. The change is confined to the autoindex XLSX extractor, a client-side crate.
  • New ES-compatible YAML case: N/A
  • Docs: the module doc describes the behavior and its limits. The public "merged cells" wording follow-up from feat(autoindex): fill vertically merged XLSX cells down the rows they cover #1132 still applies.
  • CONTRIBUTION_REVIEW audit: not run in full. Failure behavior: with no merge list (no <mergeCells>, scan failure, or a large sheet), names are exactly what they were before.

Not run: no throughput benchmark. The change adds one pass over the merge list per sheet, run when the header is found, plus a filter when FillDown is built. The per-row path is unchanged.

Provenance

  • Written by: AI agent (Claude Code, Claude Opus 5.5), run by @thomas-villani, who reviewed it.
  • Verified: every command and number above.
  • Assumed:
    • Excel-authored workbooks with banded headers use the same merged-label layout. Only pandas and hand-built fixtures were tested.
    • Headers three or more rows deep keep only their two lowest rows, and no file of that shape was tested.

🤖 Generated with Claude Code

… cover

Motivation: a cell merged DOWN over several rows stores its value only in
the top row. pandas writes exactly that for every MultiIndex export
(`df.groupby([...]).sum().to_excel(...)`, merge_cells=True is the default),
and hand-made reports do it for a category beside its line items. xerj-org#1124
read only the top cell, so for a pandas groupby export with region merged
A2:A4 / A5:A6:

  before  Sales!r3 {"product":"B","sales":20}           (no region)
  after   Sales!r3 {"region":"East","product":"B","sales":20}

`region: East` matched 1 of East's 3 rows and a `terms` aggregation on
region undercounted. Same for a hand-made report's merged Category column.

Mechanism: `<mergeCells>` follows `<sheetData>`, but rows are streamed and
emitted as read. So each sheet gets a pre-pass, `scan_merges`, that streams
the part once WITHOUT parsing cells: it byte-searches (memchr::memmem) for
the `<mergeCells` tag — unambiguous, since `<` cannot appear unescaped in
XML text — and XML-parses only the tail from there (capped at 16 MB),
which keeps a following `<hyperlink ref=...>` from being read as a merge.
`FillDown` then gives each streamed row the value of any live vertical
merge covering it, in the merge's own (top-left) column.

Deliberately NOT done:
- Merges across columns are not expanded: a title merged over a table's
  width must stay one cell, or it becomes a header-shaped row. A block
  merge (A3:B4) fills down column A only.
- No row is invented: a covered row with no cells of its own stays absent.
- A covered cell that has its own value (malformed file) keeps it.
- Two-row headers (pandas MultiIndex columns: Q1 over Jan/Feb) are a
  separate change.

Cost, measured (release, median of 3, rust:latest container; full time
includes the pre-pass):
  300k x 8 sheet, 112.6 MB XML   full 6.274 s   pre-pass 0.640 s  (10.2%)
  pandas groupby, 1.5 MB, 1204 vertical merges
                                 full 0.067 s   pre-pass 0.011 s  (16.8%)
The pre-pass is charged to the decompression budget like any read.
A sampling run (phase A / --dry-run, 500 rows per sheet) must stay cheap,
so it skips the pre-pass on sheets declared over 64 MB decompressed
(~0.35 s of scanning at the measured rate); such a sample has no
fill-down. Full runs always scan.

Tests (extract/xlsx.rs):
- a_vertical_merge_gives_its_value_to_every_row_it_covers (pandas shape,
  merges out of order, hyperlink ref after the list)
- merges_across_columns_are_not_expanded (merged title, block merge)
- fill_down_invents_no_rows_and_overwrites_no_values
- the_merge_scan_finds_the_tag_and_only_the_tag (prefixed tag, tag split
  across 7-byte reads, "mergeCells" in cell text, tail cap)
- a_sampling_run_skips_the_merge_scan_only_on_a_large_sheet
- cell_references_parse_to_column_and_row
With `fill.apply` disabled, the three behavior tests fail; with it, all
pass. Real pandas/openpyxl workbooks (MultiIndex rows, merged-title
report) checked with a temporary probe, removed before commit.

Evidence: cargo test -p xerj-autoindex --lib: 1195 passed, 0 failed,
2 ignored. cargo build --release -p xerj-autoindex ok; clippy -D warnings
clean; cargo fmt --check clean.

Written by an AI agent (Claude Code) on behalf of the PR author.
Motivation: pandas writes a DataFrame with MultiIndex columns as a header
two rows deep, the outer level merged across its inner columns:

    row 1:        Q1 (B1:C1)    Q2 (D1:E1)
    row 2:        Jan   Feb     Jan   Feb
    row 4: East   1     2       3     4

Header detection passes over row 1 (a narrower candidate above a wider
one is a title line), so the fields came out as Jan, Feb, Jan_2, Feb_2:
the quarter was lost and two columns were named by a dedup suffix.
Hand-built reports with "Actual | Budget" bands over the same month
columns have the same shape.

Change (extract/xlsx.rs):
- scan_merges keeps every multi-cell range as a `Merge` (it kept only
  ranges more than one row tall). FillDown takes the vertical ones, so
  fill-down is unchanged.
- group_prefixes: the row directly above the header counts as group
  labels only when it holds only text, has a cell merged across two or
  more header columns, and no merge spans every header column (that is a
  title over the table, not a group). Each of its cells then prefixes the
  header columns it spans; an unmerged cell prefixes its own column,
  since pandas does not merge a one-column group. A cell whose merge runs
  down into the header row (`Region` in A1:A2) is part of the header and
  gives no prefix. The name is sanitize_field_name("Q1 Jan") = Q1_Jan.
- Sample/full consistency: the sampling run skips the merge scan on a
  sheet over MAX_SAMPLE_MERGE_SCAN_BYTES (64 MB decompressed), and the
  sample types the fields the full run indexes, so a column's name must
  not depend on the run. Header merges are therefore used only on sheets
  both runs scan; over that size a two-row header keeps its lower row,
  as before. Documented in the module's known limits.

Evidence:
- 5 new/updated tests; cargo test -p xerj-autoindex --lib: 1200 passed,
  0 failed, 2 ignored (1195 before). clippy -D warnings clean, fmt clean,
  release build ok.
- Mutation checks, each run and then reverted: disabling the prefixing
  fails the 4 positive tests; dropping the full-width-title rule fails 3
  (the shared test fixture has always carried an A1:D1 title merge, so
  every older header test now also guards that rule); dropping the
  vertical-label skip fails 1.
- pandas 2.x output (`df.to_excel` with MultiIndex columns), extracted in
  a temporary probe test: records are
  {"col_A":"East","Q1_Jan":1,"Q1_Feb":2,"Q2_Jan":3,"Q2_Feb":4}, the same
  with a 500-row sample and with no limit. A report with a title merged
  over its width and the two fill-down workbooks from the previous PR
  extract unchanged.

Stacked on feat/xlsx-merged-cells (the vertical fill-down PR).
@cla-bot cla-bot Bot added the cla-signed label Oct 4, 2026
thomas-villani added a commit to thomas-villani/xerj that referenced this pull request Oct 4, 2026
…org#1132-xerj-org#1134)

Motivation: xerj-org#1132 fills vertical merges down, xerj-org#1133 names columns from a
two-row grouped header (Q1_Jan) and xerj-org#1134 adds a roff man(7) extractor.
The format lists and the fact-check matrix should say so. This commit
must merge AFTER those three code PRs, because it describes them as shipped.

- README, landing/llms.txt, landing/llms-full.txt: man pages added to the
  format lists (llms-full: plain or gzipped, one record per section, titled
  NAME(SECT), mdoc(7) pages stay plain text). The XLSX entry in llms-full
  now says vertical merges are filled down, a two-row grouped header is
  combined, and merges across columns are not expanded. It replaces
  "merged cells not expanded".
- scripts/seo/claims_rules.py: new GREEN THING row "Man pages (roff man(7))"
  citing extract/man.rs:1, and its gate says mdoc is not parsed. The .xlsx
  gate gets the same merged-cell wording as llms-full.
- testdata/factcheck: new good_man_answer.md fixture (git check-ignore
  confirms the !scripts/seo/testdata/**/*.md re-include applies), and
  good_xlsx_answer.md loses the old "merged cells are not expanded" line.

Verified: factcheck --self-test OK (35 THING rows); --fail-on error 0 ERROR;
build_articles --check current; landing-constants-guard passes.
--fixture-check reports 1 false positive (FC-EV-DANGLING) on this branch
alone, because good_man_answer.md cites extract/man.rs, which only exists
once xerj-org#1134 merges. With man.rs from xerj-org#1134 checked out it reports 0 false
positives and 13/13 good fixtures clean.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant