Command-line tools that let AI agents search, read, and cite sources from your personal library.
Created for my own philosophical research.
Python and ts wrappers included.
- Who This Is For
- What's Included
- Quick Start
- Requirements
- Installation
- Usage
- Example Workflows
- Status
- Limitations
- Customization
- For Human Users
- Testing
- Python and TypeScript Wrappers
- Contributing
- Background
- Contact
This toolkit is designed for researchers who:
- Have a collection of academic papers as markdown files
- Maintain a BibTeX bibliography for their research
- Want to use claude, codex, opencode or other AI agents to help with fact-checking, citation verification, and research tasks
- Need programmatic access to their personal research library
Eleven command-line tools for working with your research library:
| Tool | Purpose |
|---|---|
cite2md |
Convert a BibTeX key or citation to a markdown file path (or print contents) |
cite2abs |
Convert a BibTeX key or citation to an abstract Markdown file path, or print the abstract |
cite2bib |
Get the BibTeX entry for a citation or key |
cite2pdf |
Locate the PDF file for a citation or key |
draft2keys |
Extract all BibTeX keys from a draft document |
find-bib |
Filter the complete canonical bibliography catalogue by author, title, year, DOI, or abstract |
path2key |
Extract the BibTeX key from a filename or path |
rg-sources |
Search full text of all papers using ripgrep |
rg-abstracts |
Search Markdown abstracts using ripgrep |
fd-sources |
Find papers by filename pattern |
cat-sources |
Print contents of source files from filenames or paths |
All tools follow a consistent pattern: list → then read one. This helps manage context limits when working with AI agents.
# Clone the repository
git clone https://github.com/butterfill/agent-tools-for-philosophy.git
cd agent-tools-for-philosophy
# Run install script (copies tools to your PATH and runs tests)
./install.shBefore using the tools, set these environment variables (add to your .bashrc, .zshrc, etc.):
# Required primary/cited CSL-JSON bibliography
export BIB_JSON="$HOME/endnote/phd_biblio.json"
# Optional complete Zotero CSL-JSON export
export ZOTERO_JSON="$HOME/endnote/zotero-export.json"
# Path to your BibTeX file (used for BibTeX artifact commands)
export BIB_FILE="$HOME/documents/research/my-bibliography.bib"
# Path to your directory of markdown papers
# Each file should have the BibTeX key (without colons) in its filename
export PAPERS_DIR="$HOME/papers"
# Optional: path to your directory of markdown abstracts
export ABSTRACTS_DIR="$HOME/syncthing/db/papers-abstracts"# Copy instructions to your working directory
cp ~/path/to/agent-tools-for-philosophy/agent-tool-instructions.md .
# start your agent
codex # or claude, opencode, ...
# in the agent:
"We are aiming to achieve X. Please review agent-tool-instructions.md \
to understand the available tools. Your task is to ..."
# Search for papers by an author
find-bib --author "Steward"
# Open markdown file for a specific paper in VS Code (assumes vs code is installed)
cite2md --vs vesper:2012_jumping
# Read the full text (assumes glow is installed)
cite2md --cat vesper:2012_jumping | glow- macOS or Linux
jq— JSON processing (required forcite2bibandcite2md; installation guide)- Python 3 + pip — used to install/run the canonical
ReferenceCatalogbackingfind-bib fd— Fast file finding; alias tofdif it’sfdfindon your platform (installation guide)rg(ripgrep) — Fast text search (installation guide)
Install on macOS:
brew install jq fd ripgrep pythonInstall on Ubuntu/Debian:
sudo apt install jq fd-find ripgrep python3 python3-pip
ln -s $(command -v fdfind) ~/.local/bin/fdgawkon macOS (brew install gawk) — will be used if available for better performance
shellcheck— used by the test suite to lint every shell script (install withbrew install shellcheckorsudo apt install shellcheck)- On macOS, upgrade Bash via
brew install bashso that built-in utilities likemapfileare available when running the test suite (./run-tests.sh).
$BIB_JSON— Required primary/cited CSL-JSON bibliography; defaults to$HOME/endnote/phd_biblio.json$ZOTERO_JSON— Optional complete-library CSL-JSON export; defaults to$HOME/endnote/zotero-export.json$BIB_FILE— Must point to your BibTeX file$PAPERS_DIR— Must point to your directory of.mdfiles- Each markdown filename should include the BibTeX key with colons removed
- Example:
vesper2012_jumping.mdfor keyvesper:2012_jumping
$ABSTRACTS_DIR— Optional path to Markdown abstract files- Defaults to
$HOME/syncthing/db/papers-abstracts - Abstract filenames should be BibTeX keys with
.md, with either colon form or colons removed
- Defaults to
If you do not want BibTeX keys in .md file names, you can create a bibtex-index.jsonl file in your $PAPERS_DIR:
{"key": "liu:2022_facial", "filename": "Liu et al 2022 - Facial expressions elicit multiplexed perceptions of emotion categories.md"}Run the install script:
./install.shThis will:
- Install the canonical Python
ReferenceCatalogruntime and dependencies beside the CLI tools - Copy all executable tools to a directory on your PATH (tries
~/syncthing/bin,~/.local/bin, or~/bin) - Copy help text files to
help-text/subdirectory - Run the test suite to verify everything works
If the installer reports that the target directory is not on your PATH, add it to your shell profile:
# For bash
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
# For zsh
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.zshrcCopy agent-tool-instructions.md to your agent's working directory:
cp /path/to/agent-tools-for-philosophy/agent-tool-instructions.md .Then instruct your AI agent (via your agent framework) to read this file and use the tools. The file contains concise, agent-friendly documentation for each tool.
Example with an AI agent CLI (e.g., aider, codex, or similar):
# Copy instructions to your working directory
cp ~/path/to/agent-tools-for-philosophy/agent-tool-instructions.md .
# Start your AI agent and give it instructions
# (this attempts everything at once; unless it's a very small draft, better to break into steps)
your-agent-cli "Please read agent-tool-instructions.md and help me fact-check draft.md
against the sources it cites. For each citation, verify the quotes and claims are accurate."For extended human-friendly features, use --human instead of --help:
cite2md --humanHuman-specific extensions (like opening files in VS Code, revealing in Finder) are documented in agent-tool-instructions-FOR-HUMANS-ONLY.md but are hidden from agents to save tokens.
Scenario: You've written a draft that cites several sources. You want to verify each citation is accurate.
# Copy agent instructions to your working directory
cp ~/agent-tools-for-philosophy/agent-tool-instructions.md .
# Ask your AI agent to extract citation keys
your-agent "Please identify the most important sources cited in draft.md.
For each source, find its BibTeX key using the tools in agent-tool-instructions.md.
Create a file called found-keys.txt with the keys, one per line."
# Review the keys file
cat found-keys.txt
# Ask the agent to fact-check each source
your-agent "For each key in `found-keys.txt`, read the corresponding source and check
whether `draft.md` represents it accurately.
Use the tools in agent-tool-instructions.md to find and read sources.
Create a report in `checking/[key-without-colon].md`."Scenario: You want to check if your draft accurately represents a specific source.
# Copy agent instructions
cp ~/agent-tools-for-philosophy/agent-tool-instructions.md .
# Ask agent to verify
your-agent "Review agent-tool-instructions.md to understand the available tools.
Check draft.md against the source with key mylopoulos:2019_intentions.
Assess whether draft.md represents the source accurately, fairly, and charitably.
If there are mistakes, they are of first importance. If there is relevant material
in the source that draft.md overlooks, note that too.
Create a file checking/mylopoulos2019_intentions.md with headings for each issue.
Under each heading include:
1. A concise statement of the issue
2. Verbatim quotes from draft.md (in \"double quotes\")
3. Verbatim quotes from the source (in \"double quotes\")
Use double quotes only for quotations."Scenario: The agent is researching a topic and wants to see what's in your library.
# Search by author
find-bib --author "Steward" --author "Velleman"
# Search abstracts for key terms
find-bib --abstract "motor representation"
# Search Markdown abstract files for candidate sources
rg-abstracts -l -i "joint action"
# Full-text search across all papers
rg-sources "bayesian prior" -i -C 2
# Find papers by filename pattern
fd-sources "intention" | head -5Status: Experimental
(I'm making the repo public mainly to share the approach rather than the code.)
These tools are not designed to provide security isolation. Some tools work with or return absolute paths (cite2md, cite2pdf, path2key). The tools scope searches to $PAPERS_DIR for convenience, not security.
I use these tools in a disposable VPS. Unless you are a great babysitter, the tools should not be used by an AI agent in an environment where giving the agent access to arbitrary file paths could be a problem. These tools are unsuitable if you're running an agent with access to sensitive files or systems (but see Customization below).
These tools work well for me, but you may want to adapt them for your own workflow.
Copy just the specs/ folder, the 'Contributing' section from this README, and agent-tool-instructions.md, modify the specifications to suit your needs, then ask an agent to implement your own versions.
I initially intended these tools to be for agents only but ended up using them myself.
Human-specific features (hidden from agents):
- Open files in VS Code:
cite2md --vs <key> - Reveal in Finder:
cite2md --reveal <key> - ...
See agent-tool-instructions-FOR-HUMANS-ONLY.md for full details.
Run any tool with --human to see human-friendly help:
cite2md --humanThis is hidden from agents to save tokens.
I have not read the tests carefully.
./run-tests.sh# Unit tests
bash tests/cite2bib.test.sh
bash tests/draft2keys.test.sh
# End-to-end tests
bash tests/fd-sources-cat.e2e.test.sh
bash tests/rg-sources.e2e.test.sh- Tests call scripts via
./prefix (assuming current directory) - Some tests use fixtures in
tests/fixtures/ tests/shellcheck.test.shruns shellcheck across every shell script; installshellchecklocally so it doesn't skip- The
find-bibtests use primary and secondary CSL-JSON fixtures to protect canonical union/precedence behavior - The
cite2bibtests usetests/fixtures/sample.bib
uv add "agent-tools @ git+ssh://git@github.com/butterfill/agent-tools.git#subdirectory=python"
pnpm install "git+ssh://git@github.com/butterfill/agent-tools.git#subdirectory=typescript"
If you want to contribute a new tool:
- Create a spec in
specs/describing the tool's purpose, inputs, outputs, and behavior - Pick a clear, short command name (hyphenated if necessary). Extensions should not be used.
- Implement the tool as an executable script in the root directory
- Follow the baseline pattern:
- Start with
#!/usr/bin/env bashandset -euo pipefail - Provide
--helpflag with concise usage - Return consistent exit codes: 0 (success), 1 (not found), 2 (usage/config error)
- Use environment variables for configuration (e.g.,
PAPERS_DIR)
- Start with
- Externalize help text in
help-text/<tool-name>-help.txtandhelp-text/<tool-name>-human.txt - Write tests in
tests/<tool-name>.test.sh - Update documentation:
- Add entry to
agent-tool-instructions.md - Mention in this README if it's a major addition
- Add entry to
- Linting: All shell scripts must pass
shellcheck(runtests/shellcheck.test.sh) - Prefer standard Unix tools:
fd,rg,jq,yq,ast-grep - Output discipline
- Keep stdout for the tool’s primary data. Send diagnostics to stderr.
- Avoid extra wrapper noise; let underlying tools (e.g.,
rg) control formatting.
- Fail fast: Check dependencies and show clear error messages
- Minimal CLI surface: Long flags only, strong defaults over configurability, avoid short aliases and decorative output (e.g., no headers/line numbers).
- List → read pattern: Encourage workflows that list first, then read one item at a time
- Environment and scope
- Prefer environment variables for configurable roots (e.g.,
PAPERS_DIRwith a safe default). - Check environment variables already in use before requiring any new ones.
- If a tool must access outside the repo, scope carefully and validate inputs (e.g., reject absolute
paths).
- Prefer environment variables for configurable roots (e.g.,
See the "Contributing Tools" section in the current README for full guidelines.
I used to copy sources that I wanted an agent to read for each task into a directory. This was quite slow, and it became harder and harder to ensure that the agent could only see the sources I wanted it to see.
I first thought about using MCP, but this is simpler. I was inspired by Armin Ronacher’s Tools: Code Is All You Need and Cameron’s My Take on the MCP vs CLI Debate as well as some things Simon Willison wrote. Thank you!
MIT License. Attribution appreciated.
See LICENSE.md for full text.