Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -229,7 +229,8 @@ done in 158.1s, 593 datasets, 83103 records live, 790 junk records
```

Source files go through tree-sitter, so code arrives with its symbols and line numbers
instead of as flat text. CSV, JSON, JSONL, XML, YAML, SQLite, PDF, DOCX, HTML, mail (`.eml` and mbox mailboxes,
instead of as flat text. CSV, JSON, JSONL, XML, YAML, SQLite, PDF, DOCX, PowerPoint
(`.pptx`, one record per slide), Excel (`.xlsx`, one dataset per sheet), man pages (roff, one record per section), HTML, mail (`.eml` and mbox mailboxes,
including a Google Takeout export) and common log formats are all handled. Unity projects get first-class treatment: text-serialized scenes,
prefabs and assets become one record per GameObject/Component, `.meta` files become a
GUID-to-path table, and MonoBehaviour records carry `script_class`/`script_path` so "which
Expand Down
2 changes: 1 addition & 1 deletion content/answers/search-word-documents-in-a-folder.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,6 +112,6 @@ We publish this as a finding, not as a feature. The guard does the correct thing

The measurement is a single-node run of 2 small files on one host, so it shows behavior and not throughput. XERJ has no replication and no failover in this configuration.

The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. XERJ has no extractor for `.xlsx`, `.pptx`, `.rtf` or `.odt`, and refuses a file in any of those formats rather than parsing part of it.
The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. Excel workbooks (`.xlsx`) and PowerPoint decks (`.pptx`) have their own extractors, so they can sit in the same folder. XERJ has no extractor for `.rtf` or `.odt`, and refuses a file in either format rather than parsing part of it.

Ranking is BM25 over the extracted paragraph text. The default embedder in XERJ is lexical feature hashing and cannot connect a query to a synonym; neural embeddings are opt-in through `--embed-mode neural`.
18 changes: 9 additions & 9 deletions content/compare/xerj-vs-localsynapse.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: "XERJ vs LocalSynapse for MCP local file search"
h1: "What's a local MCP search engine for Claude?"
description: "Two local MCP search engines. LocalSynapse wins the desktop experience, Excel workbooks and mail shapes beyond mbox and eml. XERJ wins data files and the Elasticsearch REST API."
description: "Two local MCP search engines. LocalSynapse wins the desktop experience and mail shapes beyond mbox and eml. XERJ wins data files, Excel sheets read as typed tables, and the Elasticsearch REST API."
slug: "xerj-vs-localsynapse"
cluster: "Comparison: local MCP search"
question: "What's a local MCP search engine for Claude?"
intent: "comparison"
published: "2026-08-22"
updated: "2026-09-26"
updated: "2026-10-03"
author: "XERJ documentation team"
reviewer: "XERJ engineering team"
schema_type: "TechArticle"
Expand Down Expand Up @@ -39,13 +39,13 @@ evidence:
source: "https://github.com/LocalSynapse/LocalSynapse"
faq:
- q: "What does LocalSynapse do better?"
a: "The desktop experience. It ships a search window, indexes whole drives automatically, and reads Excel workbooks that XERJ has no extractor for. On mail the two overlap — XERJ reads mbox and eml — while LocalSynapse also reads .msg."
a: "The desktop experience. It ships a search window, indexes whole drives automatically, and keeps Excel cell coordinates and merged ranges, which XERJ does not. On mail the two overlap — XERJ reads mbox and eml — while LocalSynapse also reads .msg."
- q: "LocalSynapse vs something that also indexes CSV and SQLite?"
a: "That is the split. XERJ takes the data files and the query surface: SQLite databases, hostile CSV dialects, SQL exports and code, answered over the Elasticsearch REST API."
- q: "Can XERJ read my mail archive?"
a: "Partly, since v1.0.0-rc.75. XERJ reads .mbox and .eml, detecting mbox by content — a From-separator plus RFC 5322 headers — so the extension is irrelevant. PST, OST and Maildir are still unhandled. LocalSynapse reads .eml, .msg and .mbox."
- q: "Can XERJ read a spreadsheet?"
a: "Not an Excel workbook. A CSV export is read with dialect detection, but there is no XLSX or PPTX extractor at all."
a: "Yes, since v1.0.0-rc.81. Each .xlsx sheet becomes its own dataset with one record per row and typed fields, and a .pptx deck becomes one record per slide. A sheet is read as one table, merged cells are not expanded, and legacy .xls must be saved as .xlsx first."
- q: "How many MCP tools does each side expose?"
a: "LocalSynapse documents four. XERJ serves 10 through xerj mcp, including memory tools."
- q: "Is either one open source?"
Expand All @@ -56,7 +56,7 @@ faq:
a: "Give it a local MCP server over an index of those files. Both sides here run on your machine, they index different things, and an agent can hold a tool list from each."
---

**TL;DR** — LocalSynapse wins the desktop experience: a search window you double-click, whole-drive indexing, and Excel workbooks — and mail shapes beyond mbox and eml — that XERJ has no extractor for. XERJ wins data files and the Elasticsearch query surface. No head-to-head was run for this page.
**TL;DR** — LocalSynapse wins the desktop experience: a search window you double-click, whole-drive indexing, and mail shapes beyond mbox and eml that XERJ has no extractor for. XERJ wins data files — Excel sheets included, read as typed tables — and the Elasticsearch query surface. No head-to-head was run for this page.

## Concede the desktop

Expand All @@ -70,13 +70,13 @@ XERJ has none of that. There is no window, no installer with a tray icon, and no

LocalSynapse reads PDF, Word, PowerPoint, Hangul, CSV, Markdown and plain text. It also reads mail files as `.eml`, `.msg` and `.mbox`, and it reads Excel workbooks with cell coordinates and merged ranges up to 25 MB.

One of those is a real gap on the XERJ side, and it has no workaround worth publishing: there is no XLSX or PPTX extractor, so a spreadsheet has to become CSV first. Mail is no longer a flat gap — XERJ reads .mbox and .eml, detecting mbox by content rather than by extension ([#949](https://github.com/xerj-org/xerj/pull/949), shipped in v1.0.0-rc.75) — but PST, OST and Maildir are still unhandled, and LocalSynapse adds .msg.
Office files are no longer a gap on the XERJ side. Since v1.0.0-rc.81 it reads `.pptx` decks one document per slide ([#1117](https://github.com/xerj-org/xerj/pull/1117)) and `.xlsx` workbooks one dataset per sheet ([#1124](https://github.com/xerj-org/xerj/pull/1124)). The two read a workbook differently: XERJ turns a sheet into a typed table under its header row, while LocalSynapse keeps cell coordinates and merged ranges. XERJ reads one table per sheet. It does not expand merged cells, so a workbook built around its layout suits LocalSynapse better. Mail is no longer a flat gap — XERJ reads .mbox and .eml, detecting mbox by content rather than by extension ([#949](https://github.com/xerj-org/xerj/pull/949), shipped in v1.0.0-rc.75) — but PST, OST and Maildir are still unhandled, and LocalSynapse adds .msg.

If your corpus is spreadsheets, or mail in a shape XERJ does not parse, stop here and use LocalSynapse. No measurement on this page would change that answer.
If your corpus is mail in a shape XERJ does not parse, stop here and use LocalSynapse. No measurement on this page would change that answer.

## What XERJ reads instead

Thirteen families are covered. The list holds JSON and JSONL, CSV with dialect detection, structured logs, SQL exports and SQLite. It also holds PDF, DOCX, HTML, XML, YAML, plain text, code and gzip variants.
The list holds JSON and JSONL, CSV with dialect detection, structured logs, SQL exports, SQLite and Excel workbooks. It also holds PDF, DOCX, PowerPoint, HTML, XML, YAML, plain text, mail, code and gzip variants.

The data end of that list is where the two products part company. A SQLite database, a semicolon CSV with a decimal comma, and a multi-gigabyte SQL export are shapes a document-first indexer usually skips.

Expand Down Expand Up @@ -118,7 +118,7 @@ XERJ does no optical character recognition. A page image with no text layer stay

Choose LocalSynapse when a person wants a search window on Windows or macOS. That is its job.

Choose LocalSynapse when the corpus is Excel workbooks, or mail XERJ does not parse. XERJ reads .mbox and .eml; PST, OST and Maildir stay out, and there is no workbook extractor.
Choose LocalSynapse for mail XERJ does not parse, or for workbooks that depend on layout. XERJ reads .mbox and .eml, but not PST, OST or Maildir. It reads a sheet as one table and does not expand merged cells.

Choose LocalSynapse when you want whole-drive coverage with no folder list to maintain. Its neural embedder is on by default.

Expand Down
2 changes: 1 addition & 1 deletion landing/answers/search-word-documents-in-a-folder.html
Original file line number Diff line number Diff line change
Expand Up @@ -309,7 +309,7 @@ <h2>The reason string does not name the guard</h2>
<p>We publish this as a finding, not as a feature. The guard does the correct thing, and the message it leaves behind is too generic to diagnose from.</p>
<h2>What the capture does not cover</h2>
<p>The measurement is a single-node run of 2 small files on one host, so it shows behavior and not throughput. XERJ has no replication and no failover in this configuration.</p>
<p>The extractor reads the zipped OpenXML format that <code class="inline">.docx</code> uses. Convert a legacy <code class="inline">.doc</code> file to <code class="inline">.docx</code> first. XERJ has no extractor for <code class="inline">.xlsx</code>, <code class="inline">.pptx</code>, <code class="inline">.rtf</code> or <code class="inline">.odt</code>, and refuses a file in any of those formats rather than parsing part of it.</p>
<p>The extractor reads the zipped OpenXML format that <code class="inline">.docx</code> uses. Convert a legacy <code class="inline">.doc</code> file to <code class="inline">.docx</code> first. Excel workbooks (<code class="inline">.xlsx</code>) and PowerPoint decks (<code class="inline">.pptx</code>) have their own extractors, so they can sit in the same folder. XERJ has no extractor for <code class="inline">.rtf</code> or <code class="inline">.odt</code>, and refuses a file in either format rather than parsing part of it.</p>
<p>Ranking is BM25 over the extracted paragraph text. The default embedder in XERJ is lexical feature hashing and cannot connect a query to a synonym; neural embeddings are opt-in through <code class="inline">--embed-mode neural</code>.</p>
<h2>FAQ</h2>
<h3>How do I search through a folder of contracts in .docx?</h3>
Expand Down
2 changes: 1 addition & 1 deletion landing/answers/search-word-documents-in-a-folder.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ We publish this as a finding, not as a feature. The guard does the correct thing

The measurement is a single-node run of 2 small files on one host, so it shows behavior and not throughput. XERJ has no replication and no failover in this configuration.

The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. XERJ has no extractor for `.xlsx`, `.pptx`, `.rtf` or `.odt`, and refuses a file in any of those formats rather than parsing part of it.
The extractor reads the zipped OpenXML format that `.docx` uses. Convert a legacy `.doc` file to `.docx` first. Excel workbooks (`.xlsx`) and PowerPoint decks (`.pptx`) have their own extractors, so they can sit in the same folder. XERJ has no extractor for `.rtf` or `.odt`, and refuses a file in either format rather than parsing part of it.

Ranking is BM25 over the extracted paragraph text. The default embedder in XERJ is lexical feature hashing and cannot connect a query to a synonym; neural embeddings are opt-in through `--embed-mode neural`.

Expand Down
4 changes: 2 additions & 2 deletions landing/compare/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@
"url": "https://xerj.org/compare/",
"image": "https://xerj.org/og/xerj-card.png",
"inLanguage": "en",
"dateModified": "2026-09-26",
"dateModified": "2026-10-03",
"isPartOf": {
"@id": "https://xerj.org/#website"
},
Expand Down Expand Up @@ -244,7 +244,7 @@ <h2>Comparison: desktop search</h2>
</ul>
<h2>Comparison: local MCP search</h2>
<ul>
<li><a href="/compare/xerj-vs-localsynapse">What's a local MCP search engine for Claude?</a> — Two local MCP search engines. LocalSynapse wins the desktop experience, Excel workbooks and mail shapes beyond mbox and eml. XERJ wins data files and the Elasticsearch REST API.</li>
<li><a href="/compare/xerj-vs-localsynapse">What's a local MCP search engine for Claude?</a> — Two local MCP search engines. LocalSynapse wins the desktop experience and mail shapes beyond mbox and eml. XERJ wins data files, Excel sheets read as typed tables, and the Elasticsearch REST API.</li>
</ul>
<h2>Comparison: local search engines</h2>
<ul>
Expand Down
4 changes: 2 additions & 2 deletions landing/compare/index.json
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,8 @@
"canonical_url": "https://xerj.org/compare/xerj-vs-localsynapse",
"markdown_url": "https://xerj.org/compare/xerj-vs-localsynapse.md",
"cluster": "Comparison: local MCP search",
"updated": "2026-09-26",
"summary": "Two local MCP search engines. LocalSynapse wins the desktop experience, Excel workbooks and mail shapes beyond mbox and eml. XERJ wins data files and the Elasticsearch REST API."
"updated": "2026-10-03",
"summary": "Two local MCP search engines. LocalSynapse wins the desktop experience and mail shapes beyond mbox and eml. XERJ wins data files, Excel sheets read as typed tables, and the Elasticsearch REST API."
},
{
"title": "XERJ vs Elasticsearch run on one machine",
Expand Down
Loading
Loading