From b4bd15bf7aea8b37ec8a573aed095a37c5e8b313 Mon Sep 17 00:00:00 2001 From: Tom Villani Date: Sat, 3 Oct 2026 21:16:54 -0400 Subject: [PATCH 1/3] docs: list .pptx/.xlsx as extracted formats; retire "no XLSX/PPTX extractor" MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Motivation: #1117 (PPTX) and #1124 (XLSX) shipped in v1.0.0-rc.81, but the site and the fact-check gate still described both as missing. Two published pages made claims that are now false: - answers/search-word-documents-in-a-folder: "XERJ has no extractor for `.xlsx`, `.pptx` ... and refuses a file in any of those formats". - compare/xerj-vs-localsynapse: "there is no XLSX or PPTX extractor at all", and the TL;DR, description, FAQ and "choose LocalSynapse" section all conceded Excel workbooks on that basis. Changes: - scripts/seo/claims_rules.py: the RED "Excel / PowerPoint" THING row is split the way the mail row was split for #949 — "Excel workbooks (.xlsx)" GREEN (cites extract/xlsx.rs, gate carries the known limits: one table per sheet, merged cells not expanded, formulas without a cached value absent), "PowerPoint decks (.pptx)" GREEN (cites extract/pptx.rs), and a RED row for what is still refused: legacy .xls/.ppt, .xlsb, OpenDocument. The FC-THING-RED rewrite text no longer lists XLSX/PPTX as missing. - Fixtures: bad_thing_red.md retargeted from xlsx to legacy .xls (still expects FC-THING-RED, now via the new RED row); new good_xlsx_answer.md covers the GREEN row with zero ERRORs. - LocalSynapse comparison rewritten honestly rather than flipped: XERJ now reads workbooks as typed tables; LocalSynapse still keeps cell coordinates and merged ranges, so layout-heavy workbooks remain its shape. Mail and desktop concessions unchanged. `updated:` bumped to 2026-10-03. - README, llms.txt (extractor list on the "Then, from an unknown folder" line) and llms-full.txt (§ formats) list PPTX and XLSX. llms.txt line 3 is deliberately untouched: it already exceeds the ~300-char first-screen limit (docs/research/llms-txt-2026-09 rule 7) and this change should not grow it. - landing/answers + landing/compare regenerated with build_articles.py. Verified (commands run locally): - factcheck.py --self-test OK (34 THING rows); --fixture-check: 58 TP, 0 FN, 0 FP, 12/12 good fixtures clean; --fail-on error: 0 ERROR. - build_articles.py --check, gen_sitemap.py --check, fix_heads.py --check, fix_links.py --check, mk_og_card.py --check, test_article_schema.py: ok. - landing-constants-guard.sh: all checks passed. - seo_lint.py reports sitemap.stale on this Windows checkout, identically on clean main: core.autocrlf=true gives sitemap.xml CRLF endings. Not caused by this change; CI runs on Linux. - ste_check (advisory): the localsynapse page has one fewer STE ERROR than on main; the remaining one is the pre-existing mail sentence. Not done: the first-time-agent harness re-run (llms-txt rule 11) — this change only adds two format names to an existing list. Written by an AI agent (Claude Code) on behalf of the PR author. --- README.md | 3 +- .../search-word-documents-in-a-folder.md | 2 +- content/compare/xerj-vs-localsynapse.md | 18 +++++----- .../search-word-documents-in-a-folder.html | 2 +- .../search-word-documents-in-a-folder.md | 2 +- landing/compare/index.html | 4 +-- landing/compare/index.json | 4 +-- landing/compare/xerj-vs-localsynapse.html | 26 +++++++------- landing/compare/xerj-vs-localsynapse.md | 16 ++++----- landing/llms-full.txt | 5 ++- landing/llms.txt | 2 +- scripts/seo/claims_rules.py | 34 +++++++++++++++---- .../seo/testdata/factcheck/bad_thing_red.md | 10 +++--- .../testdata/factcheck/good_xlsx_answer.md | 22 ++++++++++++ 14 files changed, 99 insertions(+), 51 deletions(-) create mode 100644 scripts/seo/testdata/factcheck/good_xlsx_answer.md diff --git a/README.md b/README.md index 6cff0f422..94f838746 100644 --- a/README.md +++ b/README.md @@ -229,7 +229,8 @@ done in 158.1s, 593 datasets, 83103 records live, 790 junk records ``` Source files go through tree-sitter, so code arrives with its symbols and line numbers -instead of as flat text. CSV, JSON, JSONL, XML, YAML, SQLite, PDF, DOCX, HTML, mail (`.eml` and mbox mailboxes, +instead of as flat text. CSV, JSON, JSONL, XML, YAML, SQLite, PDF, DOCX, PowerPoint +(`.pptx`, one record per slide), Excel (`.xlsx`, one dataset per sheet), HTML, mail (`.eml` and mbox mailboxes, including a Google Takeout export) and common log formats are all handled. Unity projects get first-class treatment: text-serialized scenes, prefabs and assets become one record per GameObject/Component, `.meta` files become a GUID-to-path table, and MonoBehaviour records carry `script_class`/`script_path` so "which diff --git a/content/answers/search-word-documents-in-a-folder.md b/content/answers/search-word-documents-in-a-folder.md index 1cef9c5da..6609a0738 100644 --- a/content/answers/search-word-documents-in-a-folder.md +++ b/content/answers/search-word-documents-in-a-folder.md @@ -112,6 +112,6 @@ We publish this as a finding, not as a feature. The guard does the correct thing The measurement is a single-node run of 2 small files on one host, so it shows behavior and not throughput. XERJ has no replication and no failover in this configuration. -The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. XERJ has no extractor for `.xlsx`, `.pptx`, `.rtf` or `.odt`, and refuses a file in any of those formats rather than parsing part of it. +The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. Excel workbooks (`.xlsx`) and PowerPoint decks (`.pptx`) have their own extractors, so they can sit in the same folder. XERJ has no extractor for `.rtf` or `.odt`, and refuses a file in either format rather than parsing part of it. Ranking is BM25 over the extracted paragraph text. The default embedder in XERJ is lexical feature hashing and cannot connect a query to a synonym; neural embeddings are opt-in through `--embed-mode neural`. diff --git a/content/compare/xerj-vs-localsynapse.md b/content/compare/xerj-vs-localsynapse.md index 627ae7328..c59a184fe 100644 --- a/content/compare/xerj-vs-localsynapse.md +++ b/content/compare/xerj-vs-localsynapse.md @@ -1,13 +1,13 @@ --- title: "XERJ vs LocalSynapse for MCP local file search" h1: "What's a local MCP search engine for Claude?" -description: "Two local MCP search engines. LocalSynapse wins the desktop experience, Excel workbooks and mail shapes beyond mbox and eml. XERJ wins data files and the Elasticsearch REST API." +description: "Two local MCP search engines. LocalSynapse wins the desktop experience and mail shapes beyond mbox and eml. XERJ wins data files, Excel sheets read as typed tables, and the Elasticsearch REST API." slug: "xerj-vs-localsynapse" cluster: "Comparison: local MCP search" question: "What's a local MCP search engine for Claude?" intent: "comparison" published: "2026-08-22" -updated: "2026-09-26" +updated: "2026-10-03" author: "XERJ documentation team" reviewer: "XERJ engineering team" schema_type: "TechArticle" @@ -39,13 +39,13 @@ evidence: source: "https://github.com/LocalSynapse/LocalSynapse" faq: - q: "What does LocalSynapse do better?" - a: "The desktop experience. It ships a search window, indexes whole drives automatically, and reads Excel workbooks that XERJ has no extractor for. On mail the two overlap — XERJ reads mbox and eml — while LocalSynapse also reads .msg." + a: "The desktop experience. It ships a search window, indexes whole drives automatically, and keeps Excel cell coordinates and merged ranges, which XERJ does not. On mail the two overlap — XERJ reads mbox and eml — while LocalSynapse also reads .msg." - q: "LocalSynapse vs something that also indexes CSV and SQLite?" a: "That is the split. XERJ takes the data files and the query surface: SQLite databases, hostile CSV dialects, SQL exports and code, answered over the Elasticsearch REST API." - q: "Can XERJ read my mail archive?" a: "Partly, since v1.0.0-rc.75. XERJ reads .mbox and .eml, detecting mbox by content — a From-separator plus RFC 5322 headers — so the extension is irrelevant. PST, OST and Maildir are still unhandled. LocalSynapse reads .eml, .msg and .mbox." - q: "Can XERJ read a spreadsheet?" - a: "Not an Excel workbook. A CSV export is read with dialect detection, but there is no XLSX or PPTX extractor at all." + a: "Yes, since v1.0.0-rc.81. Each .xlsx sheet becomes its own dataset with one record per row and typed fields, and a .pptx deck becomes one record per slide. A sheet is read as one table, merged cells are not expanded, and legacy .xls must be saved as .xlsx first." - q: "How many MCP tools does each side expose?" a: "LocalSynapse documents four. XERJ serves 10 through xerj mcp, including memory tools." - q: "Is either one open source?" @@ -56,7 +56,7 @@ faq: a: "Give it a local MCP server over an index of those files. Both sides here run on your machine, they index different things, and an agent can hold a tool list from each." --- -**TL;DR** — LocalSynapse wins the desktop experience: a search window you double-click, whole-drive indexing, and Excel workbooks — and mail shapes beyond mbox and eml — that XERJ has no extractor for. XERJ wins data files and the Elasticsearch query surface. No head-to-head was run for this page. +**TL;DR** — LocalSynapse wins the desktop experience: a search window you double-click, whole-drive indexing, and mail shapes beyond mbox and eml that XERJ has no extractor for. XERJ wins data files — Excel sheets included, read as typed tables — and the Elasticsearch query surface. No head-to-head was run for this page. ## Concede the desktop @@ -70,13 +70,13 @@ XERJ has none of that. There is no window, no installer with a tray icon, and no LocalSynapse reads PDF, Word, PowerPoint, Hangul, CSV, Markdown and plain text. It also reads mail files as `.eml`, `.msg` and `.mbox`, and it reads Excel workbooks with cell coordinates and merged ranges up to 25 MB. -One of those is a real gap on the XERJ side, and it has no workaround worth publishing: there is no XLSX or PPTX extractor, so a spreadsheet has to become CSV first. Mail is no longer a flat gap — XERJ reads .mbox and .eml, detecting mbox by content rather than by extension ([#949](https://github.com/xerj-org/xerj/pull/949), shipped in v1.0.0-rc.75) — but PST, OST and Maildir are still unhandled, and LocalSynapse adds .msg. +Office files are no longer a gap on the XERJ side. Since v1.0.0-rc.81 it reads `.pptx` decks one document per slide ([#1117](https://github.com/xerj-org/xerj/pull/1117)) and `.xlsx` workbooks one dataset per sheet ([#1124](https://github.com/xerj-org/xerj/pull/1124)). The two read a workbook differently: XERJ turns a sheet into a typed table under its header row, while LocalSynapse keeps cell coordinates and merged ranges. XERJ reads one table per sheet. It does not expand merged cells, so a workbook built around its layout suits LocalSynapse better. Mail is no longer a flat gap — XERJ reads .mbox and .eml, detecting mbox by content rather than by extension ([#949](https://github.com/xerj-org/xerj/pull/949), shipped in v1.0.0-rc.75) — but PST, OST and Maildir are still unhandled, and LocalSynapse adds .msg. -If your corpus is spreadsheets, or mail in a shape XERJ does not parse, stop here and use LocalSynapse. No measurement on this page would change that answer. +If your corpus is mail in a shape XERJ does not parse, stop here and use LocalSynapse. No measurement on this page would change that answer. ## What XERJ reads instead -Thirteen families are covered. The list holds JSON and JSONL, CSV with dialect detection, structured logs, SQL exports and SQLite. It also holds PDF, DOCX, HTML, XML, YAML, plain text, code and gzip variants. +The list holds JSON and JSONL, CSV with dialect detection, structured logs, SQL exports, SQLite and Excel workbooks. It also holds PDF, DOCX, PowerPoint, HTML, XML, YAML, plain text, mail, code and gzip variants. The data end of that list is where the two products part company. A SQLite database, a semicolon CSV with a decimal comma, and a multi-gigabyte SQL export are shapes a document-first indexer usually skips. @@ -118,7 +118,7 @@ XERJ does no optical character recognition. A page image with no text layer stay Choose LocalSynapse when a person wants a search window on Windows or macOS. That is its job. -Choose LocalSynapse when the corpus is Excel workbooks, or mail XERJ does not parse. XERJ reads .mbox and .eml; PST, OST and Maildir stay out, and there is no workbook extractor. +Choose LocalSynapse for mail XERJ does not parse, or for workbooks that depend on layout. XERJ reads .mbox and .eml, but not PST, OST or Maildir. It reads a sheet as one table and does not expand merged cells. Choose LocalSynapse when you want whole-drive coverage with no folder list to maintain. Its neural embedder is on by default. diff --git a/landing/answers/search-word-documents-in-a-folder.html b/landing/answers/search-word-documents-in-a-folder.html index 631969a29..4a4f38976 100644 --- a/landing/answers/search-word-documents-in-a-folder.html +++ b/landing/answers/search-word-documents-in-a-folder.html @@ -309,7 +309,7 @@

The reason string does not name the guard

We publish this as a finding, not as a feature. The guard does the correct thing, and the message it leaves behind is too generic to diagnose from.

What the capture does not cover

The measurement is a single-node run of 2 small files on one host, so it shows behavior and not throughput. XERJ has no replication and no failover in this configuration.

-

The extractor reads the zipped OpenXML format that .docx uses. Convert a legacy .doc file to .docx first. XERJ has no extractor for .xlsx, .pptx, .rtf or .odt, and refuses a file in any of those formats rather than parsing part of it.

+

The extractor reads the zipped OpenXML format that .docx uses. Convert a legacy .doc file to .docx first. Excel workbooks (.xlsx) and PowerPoint decks (.pptx) have their own extractors, so they can sit in the same folder. XERJ has no extractor for .rtf or .odt, and refuses a file in either format rather than parsing part of it.

Ranking is BM25 over the extracted paragraph text. The default embedder in XERJ is lexical feature hashing and cannot connect a query to a synonym; neural embeddings are opt-in through --embed-mode neural.

FAQ

How do I search through a folder of contracts in .docx?

diff --git a/landing/answers/search-word-documents-in-a-folder.md b/landing/answers/search-word-documents-in-a-folder.md index 8e2c66868..e45362d08 100644 --- a/landing/answers/search-word-documents-in-a-folder.md +++ b/landing/answers/search-word-documents-in-a-folder.md @@ -111,7 +111,7 @@ We publish this as a finding, not as a feature. The guard does the correct thing The measurement is a single-node run of 2 small files on one host, so it shows behavior and not throughput. XERJ has no replication and no failover in this configuration. -The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. XERJ has no extractor for `.xlsx`, `.pptx`, `.rtf` or `.odt`, and refuses a file in any of those formats rather than parsing part of it. +The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. Excel workbooks (`.xlsx`) and PowerPoint decks (`.pptx`) have their own extractors, so they can sit in the same folder. XERJ has no extractor for `.rtf` or `.odt`, and refuses a file in either format rather than parsing part of it. Ranking is BM25 over the extracted paragraph text. The default embedder in XERJ is lexical feature hashing and cannot connect a query to a synonym; neural embeddings are opt-in through `--embed-mode neural`. diff --git a/landing/compare/index.html b/landing/compare/index.html index d63151d04..ab5103e12 100644 --- a/landing/compare/index.html +++ b/landing/compare/index.html @@ -63,7 +63,7 @@ "url": "https://xerj.org/compare/", "image": "https://xerj.org/og/xerj-card.png", "inLanguage": "en", - "dateModified": "2026-09-26", + "dateModified": "2026-10-03", "isPartOf": { "@id": "https://xerj.org/#website" }, @@ -244,7 +244,7 @@

Comparison: desktop search

Comparison: local MCP search

Comparison: local search engines