docs: list .pptx/.xlsx as extracted formats; retire "no XLSX/PPTX extractor" - #1130
thomas-villani wants to merge 3 commits into
Conversation
…ractor" Motivation: xerj-org#1117 (PPTX) and xerj-org#1124 (XLSX) shipped in v1.0.0-rc.81, but the site and the fact-check gate still described both as missing. Two published pages made claims that are now false: - answers/search-word-documents-in-a-folder: "XERJ has no extractor for `.xlsx`, `.pptx` ... and refuses a file in any of those formats". - compare/xerj-vs-localsynapse: "there is no XLSX or PPTX extractor at all", and the TL;DR, description, FAQ and "choose LocalSynapse" section all conceded Excel workbooks on that basis. Changes: - scripts/seo/claims_rules.py: the RED "Excel / PowerPoint" THING row is split the way the mail row was split for xerj-org#949 — "Excel workbooks (.xlsx)" GREEN (cites extract/xlsx.rs, gate carries the known limits: one table per sheet, merged cells not expanded, formulas without a cached value absent), "PowerPoint decks (.pptx)" GREEN (cites extract/pptx.rs), and a RED row for what is still refused: legacy .xls/.ppt, .xlsb, OpenDocument. The FC-THING-RED rewrite text no longer lists XLSX/PPTX as missing. - Fixtures: bad_thing_red.md retargeted from xlsx to legacy .xls (still expects FC-THING-RED, now via the new RED row); new good_xlsx_answer.md covers the GREEN row with zero ERRORs. - LocalSynapse comparison rewritten honestly rather than flipped: XERJ now reads workbooks as typed tables; LocalSynapse still keeps cell coordinates and merged ranges, so layout-heavy workbooks remain its shape. Mail and desktop concessions unchanged. `updated:` bumped to 2026-10-03. - README, llms.txt (extractor list on the "Then, from an unknown folder" line) and llms-full.txt (§ formats) list PPTX and XLSX. llms.txt line 3 is deliberately untouched: it already exceeds the ~300-char first-screen limit (docs/research/llms-txt-2026-09 rule 7) and this change should not grow it. - landing/answers + landing/compare regenerated with build_articles.py. Verified (commands run locally): - factcheck.py --self-test OK (34 THING rows); --fixture-check: 58 TP, 0 FN, 0 FP, 12/12 good fixtures clean; --fail-on error: 0 ERROR. - build_articles.py --check, gen_sitemap.py --check, fix_heads.py --check, fix_links.py --check, mk_og_card.py --check, test_article_schema.py: ok. - landing-constants-guard.sh: all checks passed. - seo_lint.py reports sitemap.stale on this Windows checkout, identically on clean main: core.autocrlf=true gives sitemap.xml CRLF endings. Not caused by this change; CI runs on Linux. - ste_check (advisory): the localsynapse page has one fewer STE ERROR than on main; the remaining one is the pre-existing mail sentence. Not done: the first-time-agent harness re-run (llms-txt rule 11) — this change only adds two format names to an existing list. Written by an AI agent (Claude Code) on behalf of the PR author.
…org#1132-xerj-org#1134) Motivation: xerj-org#1132 fills vertical merges down, xerj-org#1133 names columns from a two-row grouped header (Q1_Jan) and xerj-org#1134 adds a roff man(7) extractor. The format lists and the fact-check matrix should say so. This commit must merge AFTER those three code PRs, because it describes them as shipped. - README, landing/llms.txt, landing/llms-full.txt: man pages added to the format lists (llms-full: plain or gzipped, one record per section, titled NAME(SECT), mdoc(7) pages stay plain text). The XLSX entry in llms-full now says vertical merges are filled down, a two-row grouped header is combined, and merges across columns are not expanded. It replaces "merged cells not expanded". - scripts/seo/claims_rules.py: new GREEN THING row "Man pages (roff man(7))" citing extract/man.rs:1, and its gate says mdoc is not parsed. The .xlsx gate gets the same merged-cell wording as llms-full. - testdata/factcheck: new good_man_answer.md fixture (git check-ignore confirms the !scripts/seo/testdata/**/*.md re-include applies), and good_xlsx_answer.md loses the old "merged cells are not expanded" line. Verified: factcheck --self-test OK (35 THING rows); --fail-on error 0 ERROR; build_articles --check current; landing-constants-guard passes. --fixture-check reports 1 false positive (FC-EV-DANGLING) on this branch alone, because good_man_answer.md cites extract/man.rs, which only exists once xerj-org#1134 merges. With man.rs from xerj-org#1134 checked out it reports 0 false positives and 13/13 good fixtures clean.
|
Merge order: I've added commit Until #1134 lands, This comment and commit were written by an AI agent (Claude Code) on behalf of @thomas-villani. |
What this changes, and why
#1117 (PPTX) and #1124 (XLSX) shipped in v1.0.0-rc.81, but the site and the fact-check gate still describe both formats as missing. Two published pages now make false claims:
answers/search-word-documents-in-a-folder: "XERJ has no extractor for.xlsx,.pptx… and refuses a file in any of those formats".compare/xerj-vs-localsynapse: "there is no XLSX or PPTX extractor at all". The TL;DR, description, FAQ and "choose LocalSynapse" section all concede Excel workbooks on that basis.Changes:
scripts/seo/claims_rules.py: the RED "Excel / PowerPoint" THING row is split, the same way the mail row was for feat(autoindex): mbox and Google Takeout ingest #949:extract/xlsx.rs. Its gate lists the known limits: one table per sheet, merged cells not expanded, and no value for a formula without a cached result.extract/pptx.rs..xls/.ppt,.xlsband OpenDocument.bad_thing_red.mdnow targets legacy.xls. It still expects FC-THING-RED, which now comes from the new RED row. A newgood_xlsx_answer.mdcovers the GREEN row.llms.txt,llms-full.txt: the extractor lists now include PPTX and XLSX. Inllms.txtonly the "Then, from an unknown folder…" list changed. Line 3 is left alone on purpose: it already goes past the ~300-character first-screen limit (rule 7 indocs/research/llms-txt-2026-09). No other open PR editsllms.txt(checked before branching).landing/answersandlanding/comparewere regenerated withbuild_articles.py, and the sitemap was regenerated in a separate commit after the content commit (own-commit rule).Evidence
Checks
cargo fmt --all: N/A, no Rust changedcargo test -p <crate>: N/ANot run:
seo_lint.pyfails locally withsitemap.stale, and it fails the same way on a cleanmaincheckout. This Windows clone hascore.autocrlf=true, sositemap.xmlis checked out with CRLF line endings.gen_sitemap.py --checkpasses.ste_check.pyis advisory. On the LocalSynapse page it reports one fewer ERROR than onmain. The remaining ERROR is the existing mail sentence.Provenance
extract/xlsx.rsandextract/pptx.rs, which were tested in feat(autoindex): extract PowerPoint decks (.pptx), one record per slide #1117/feat(autoindex): extract Excel workbooks (.xlsx), one dataset per sheet (rebase of #1120) #1124.evidence:sources and were not re-checked. Legacy.xls/.ppt(OLE compound files) have no sniff rule and were not run through autoindex here. The RED row says only "no extractor", whichsniff.rsconfirms: it has no CFB branch.🤖 Generated with Claude Code