Skip to content

docs: list .pptx/.xlsx as extracted formats; retire "no XLSX/PPTX extractor" - #1130

Open
thomas-villani wants to merge 3 commits into
xerj-org:mainfrom
thomas-villani:docs/office-formats
Open

thomas-villani wants to merge 3 commits into
xerj-org:mainfrom
thomas-villani:docs/office-formats

Conversation

@thomas-villani

Copy link
Copy Markdown
Contributor

What this changes, and why

#1117 (PPTX) and #1124 (XLSX) shipped in v1.0.0-rc.81, but the site and the fact-check gate still describe both formats as missing. Two published pages now make false claims:

  • answers/search-word-documents-in-a-folder: "XERJ has no extractor for .xlsx, .pptx … and refuses a file in any of those formats".
  • compare/xerj-vs-localsynapse: "there is no XLSX or PPTX extractor at all". The TL;DR, description, FAQ and "choose LocalSynapse" section all concede Excel workbooks on that basis.

Changes:

  • scripts/seo/claims_rules.py: the RED "Excel / PowerPoint" THING row is split, the same way the mail row was for feat(autoindex): mbox and Google Takeout ingest #949:
    • "Excel workbooks (.xlsx)" is GREEN and cites extract/xlsx.rs. Its gate lists the known limits: one table per sheet, merged cells not expanded, and no value for a formula without a cached result.
    • "PowerPoint decks (.pptx)" is GREEN and cites extract/pptx.rs.
    • A new RED row covers what is still refused: legacy .xls/.ppt, .xlsb and OpenDocument.
    • The FC-THING-RED rewrite text no longer lists XLSX/PPTX as missing.
  • Fixtures: bad_thing_red.md now targets legacy .xls. It still expects FC-THING-RED, which now comes from the new RED row. A new good_xlsx_answer.md covers the GREEN row.
  • LocalSynapse comparison: rewritten, not just flipped. XERJ now reads workbooks as typed tables. LocalSynapse still keeps cell coordinates and merged ranges, so a workbook that depends on its layout still fits LocalSynapse better. The mail and desktop concessions are unchanged.
  • README, llms.txt, llms-full.txt: the extractor lists now include PPTX and XLSX. In llms.txt only the "Then, from an unknown folder…" list changed. Line 3 is left alone on purpose: it already goes past the ~300-character first-screen limit (rule 7 in docs/research/llms-txt-2026-09). No other open PR edits llms.txt (checked before branching).
  • landing/answers and landing/compare were regenerated with build_articles.py, and the sitemap was regenerated in a separate commit after the content commit (own-commit rule).

Evidence

factcheck.py --self-test       self-test OK: 56 rules, 55/37/32 tiered numbers, 34 THING rows
factcheck.py --fixture-check   58 TP, 0 FN, 0 FP; true negatives 12 (of 12); rules with no fixture: 0
factcheck.py --fail-on error   92 file(s), 0 ERROR, 93 WARN   (same 0 ERROR / 93 WARN as main)
build_articles.py --check      ok (189 files; 92 selected article(s))
gen_sitemap.py --check         ok (183 URLs)
fix_heads.py / fix_links.py / mk_og_card.py --check, test_article_schema.py: ok
landing-constants-guard.sh     all checks passed

Checks

  • cargo fmt --all: N/A, no Rust changed
  • Scoped release build: N/A
  • cargo test -p <crate>: N/A
  • ES-YAML conformance suite: not applicable, this change is docs/landing/fact-check-only
  • New ES-compatible behavior YAML case: N/A
  • Docs updated
  • CONTRIBUTION_REVIEW audit: N/A, no operational document or runnable recipe changed

Not run:

  • seo_lint.py fails locally with sitemap.stale, and it fails the same way on a clean main checkout. This Windows clone has core.autocrlf=true, so sitemap.xml is checked out with CRLF line endings. gen_sitemap.py --check passes.
  • The first-time-agent harness (llms-txt rule 11) was not re-run, because this change only adds two names to an existing list.
  • ste_check.py is advisory. On the LocalSynapse page it reports one fewer ERROR than on main. The remaining ERROR is the existing mail sentence.

Provenance

🤖 Generated with Claude Code

…ractor"

Motivation: xerj-org#1117 (PPTX) and xerj-org#1124 (XLSX) shipped in v1.0.0-rc.81, but the
site and the fact-check gate still described both as missing. Two published
pages made claims that are now false:

- answers/search-word-documents-in-a-folder: "XERJ has no extractor for
  `.xlsx`, `.pptx` ... and refuses a file in any of those formats".
- compare/xerj-vs-localsynapse: "there is no XLSX or PPTX extractor at all",
  and the TL;DR, description, FAQ and "choose LocalSynapse" section all
  conceded Excel workbooks on that basis.

Changes:

- scripts/seo/claims_rules.py: the RED "Excel / PowerPoint" THING row is
  split the way the mail row was split for xerj-org#949 — "Excel workbooks (.xlsx)"
  GREEN (cites extract/xlsx.rs, gate carries the known limits: one table per
  sheet, merged cells not expanded, formulas without a cached value absent),
  "PowerPoint decks (.pptx)" GREEN (cites extract/pptx.rs), and a RED row for
  what is still refused: legacy .xls/.ppt, .xlsb, OpenDocument. The
  FC-THING-RED rewrite text no longer lists XLSX/PPTX as missing.
- Fixtures: bad_thing_red.md retargeted from xlsx to legacy .xls (still
  expects FC-THING-RED, now via the new RED row); new good_xlsx_answer.md
  covers the GREEN row with zero ERRORs.
- LocalSynapse comparison rewritten honestly rather than flipped: XERJ now
  reads workbooks as typed tables; LocalSynapse still keeps cell coordinates
  and merged ranges, so layout-heavy workbooks remain its shape. Mail and
  desktop concessions unchanged. `updated:` bumped to 2026-10-03.
- README, llms.txt (extractor list on the "Then, from an unknown folder"
  line) and llms-full.txt (§ formats) list PPTX and XLSX. llms.txt line 3 is
  deliberately untouched: it already exceeds the ~300-char first-screen limit
  (docs/research/llms-txt-2026-09 rule 7) and this change should not grow it.
- landing/answers + landing/compare regenerated with build_articles.py.

Verified (commands run locally):
- factcheck.py --self-test OK (34 THING rows); --fixture-check: 58 TP,
  0 FN, 0 FP, 12/12 good fixtures clean; --fail-on error: 0 ERROR.
- build_articles.py --check, gen_sitemap.py --check, fix_heads.py --check,
  fix_links.py --check, mk_og_card.py --check, test_article_schema.py: ok.
- landing-constants-guard.sh: all checks passed.
- seo_lint.py reports sitemap.stale on this Windows checkout, identically on
  clean main: core.autocrlf=true gives sitemap.xml CRLF endings. Not caused
  by this change; CI runs on Linux.
- ste_check (advisory): the localsynapse page has one fewer STE ERROR than
  on main; the remaining one is the pre-existing mail sentence.

Not done: the first-time-agent harness re-run (llms-txt rule 11) — this
change only adds two format names to an existing list.

Written by an AI agent (Claude Code) on behalf of the PR author.
…org#1132-xerj-org#1134)

Motivation: xerj-org#1132 fills vertical merges down, xerj-org#1133 names columns from a
two-row grouped header (Q1_Jan) and xerj-org#1134 adds a roff man(7) extractor.
The format lists and the fact-check matrix should say so. This commit
must merge AFTER those three code PRs, because it describes them as shipped.

- README, landing/llms.txt, landing/llms-full.txt: man pages added to the
  format lists (llms-full: plain or gzipped, one record per section, titled
  NAME(SECT), mdoc(7) pages stay plain text). The XLSX entry in llms-full
  now says vertical merges are filled down, a two-row grouped header is
  combined, and merges across columns are not expanded. It replaces
  "merged cells not expanded".
- scripts/seo/claims_rules.py: new GREEN THING row "Man pages (roff man(7))"
  citing extract/man.rs:1, and its gate says mdoc is not parsed. The .xlsx
  gate gets the same merged-cell wording as llms-full.
- testdata/factcheck: new good_man_answer.md fixture (git check-ignore
  confirms the !scripts/seo/testdata/**/*.md re-include applies), and
  good_xlsx_answer.md loses the old "merged cells are not expanded" line.

Verified: factcheck --self-test OK (35 THING rows); --fail-on error 0 ERROR;
build_articles --check current; landing-constants-guard passes.
--fixture-check reports 1 false positive (FC-EV-DANGLING) on this branch
alone, because good_man_answer.md cites extract/man.rs, which only exists
once xerj-org#1134 merges. With man.rs from xerj-org#1134 checked out it reports 0 false
positives and 13/13 good fixtures clean.
@thomas-villani

Copy link
Copy Markdown
Contributor Author

Merge order: I've added commit docs: list man pages; describe XLSX merged-cell handling to this PR. It describes the code in #1132 (vertical merges filled down), #1133 (two-row grouped headers → Q1_Jan) and #1134 (man pages) as shipped. Please merge this PR after those three. I put it here rather than in a new PR because only one open PR should edit landing/llms.txt at a time.

Until #1134 lands, factcheck.py --fixture-check reports one FC-EV-DANGLING on the new good_man_answer.md, because it cites extract/man.rs. That error also flags a merge in the wrong order. With man.rs from #1134 present it reports 0 false positives and 13/13 good fixtures clean. The other gates (--self-test, --fail-on error, build_articles --check, gen_sitemap --check, landing-constants-guard) pass on the branch as it is.

This comment and commit were written by an AI agent (Claude Code) on behalf of @thomas-villani.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant