Status: Living document. Retrofitted at v0.4 to anchor the project's expanding scope. Owner: John (@gatewaynode) Last reviewed: 2026-04-25
This document describes what llm_context_shield is, who it serves, and which deployment contexts it must support. It is the source of truth for scope decisions. Phase-level how and when live in tasks/todo.md; architecture lives in tasks/ARCHITECTURE.md; installation and CLI usage live in README.md. This PRD links to them and does not duplicate them.
A general-purpose, UNIX-philosophy threat scanner for LLM context-injection attacks. The same binary and library serve every layer of an LLM stack — from a single-shot pipe filter feeding one prompt, to a long-running gateway watching thousands of concurrent sessions. The scanner reports; it never rewrites. It runs locally; nothing leaves the process unless the operator opts in. It is small, fast, deterministic, cross-platform, and embeddable.
| Persona | Description | Primary use cases |
|---|---|---|
| Skill author | Writes Claude Code skills, MCP servers, or agent tools that fetch external content. Wants a one-line lcs scan -p filter between fetch and model. |
UC-1 |
| Application developer | Builds a Rust application (CLI tool, batch processor, internal service) that needs to scan untrusted text before passing it to an LLM. Embeds the library, customises the engine, ignores the CLI. | UC-2, UC-3 |
| Security engineer | Audits stored content — prompt logs, document corpora, ingested datasets — for injection attempts. Runs the tool over batches of files and reviews aggregated results. | UC-3, UC-4 |
| Platform operator | Runs an LLM-backed service (chatbot, agent platform, RAG endpoint, API gateway). Needs per-user threat tracking across requests and pluggable storage that survives across processes. | UC-5 |
The five contexts that scope this project. Every roadmap phase must trace back to one of these.
A skill or shell pipeline routes external content through lcs scan -p before the content reaches a model. Clean content flows through; threats block the pipe.
- Primary user: Skill author, application developer
- Status: ✅ Shipped
- Reference implementation:
skill/safe-fetch.md,README.md"Recommended workflow" section - Key NFR: Latency. The scan must add negligible overhead to a
curl | modelpipeline. Single-process, single-input, no state.
A user or script invokes lcs scan <file> (or Shield::scan(text) from a Rust binary) to check one input. Output is consumed directly — exit codes for shell scripts, JSON for tooling, text for humans.
- Primary user: Application developer, security engineer
- Status: ✅ Shipped
- Reference implementation:
README.md"Usage" section,examples/embed.rs - Key NFR: Determinism. Repeated scans of the same input produce byte-identical output (BTreeMap-ordered class scores; stable rule-name sorting).
A user scans many files as a related set — a downloaded corpus, a PR diff with multiple touched documents, an entire prompt directory. Per-file results are useful, but so are group-level aggregations: which threat classes appeared across the batch, whether suspicious patterns cluster across multiple files, what the worst-offender file is.
- Primary user: Security engineer, application developer
- Status: 📅 Roadmap — Phase 13 (scan groups)
- Key NFR: Aggregation. Group-level threat scoreboards and cross-input correlation must work without persistence or temporal semantics. A scan group is a one-shot snapshot.
Scan a captured chat history (system prompt + N user turns + N assistant turns) as a related set of inputs to identify which turn introduced an injection. Orderless replay only — treat the history as a scan group.
- Primary user: Security engineer, application developer
- Status: 📅 Roadmap — Phase 13 (scan groups). The temporal half — replaying the history through a stateful session to detect crescendo / accumulation patterns — was originally Phase 12 in this project; transferred to the aegis orchestrator project on 2026-04-30 to keep lcs UNIX-composable as a single-shot scanner.
- Key NFR: Aggregation. Group-level threat scoreboards and cross-input correlation must work without persistence or temporal semantics. A scan group is a one-shot snapshot.
This use case (long-running service with per-user / per-API-key / per-conversation state surviving across requests and restarts) was originally Phase 12 in this project. As of 2026-04-30 it lives in the aegis orchestrator project. lcs's contract from this side is "single-shot scan with stable JSON output and rule-set fingerprint" — aegis composes lcs into the WAF / gateway shape via subprocess (decided 2026-05-01; the library wrap option is deferred to a future productized aegis where in-process latency matters).
- Primary user: Platform operator
- Status: Out of scope for lcs; owned by aegis.
The scanner detects the categories listed in README.md "Scanner Categories" — prompt injection, instruction override, jailbreak, delimiter manipulation, data exfiltration, hidden content, refusal suppression, response steering, secret probing, context shift, ICL exploitation, coercion, refusal bypass, session protocol, and obfuscation.
Coverage is documented per category against the CrowdStrike Prompt Injection Attack Taxonomy in data/prompt-injection-attack-taxonomy.md.
Three interchangeable engines, selected at build time and runtime, documented in detail in README.md "Scan Engines":
simple— hardcoded Rust regex. Fastest. No optional features.yara— YARA-X (VirusTotal's pure-Rust YARA). Editable.yarfiles. Build flag--features yara.syara— SYARA-X (Super YARA), adds semantic matchers in three tiers: regex (always on),similarity:via local ONNX MiniLM (--features syara-sbert), andllm:via an OpenAI-compatible endpoint (--features syara-llm).
The library exposes the Engine trait as the extension point. Custom engines (see examples/custom_engine.rs) can be plugged in without modifying the library.
Capabilities that combine signals across rules, engines, scans, or sessions. These live in the orchestrator (Shield), not in individual engines.
| Capability | Phase | Status | What it adds |
|---|---|---|---|
| Heuristic threat scoring | 7 | ✅ Shipped | Per-class and cumulative threat accumulators. Threshold-gated rules can stay silent until cheaper rules raise suspicion. |
| Cross-rule correlation | 11 | ✅ Shipped | Rules that fire only when two findings co-occur (proximate, ordered, combined, or cross-engine). Composite scores feed back into the scoreboard. |
| Session tracking | 12 | Transferred to aegis on 2026-04-30. Per-session history, multi-turn patterns (crescendo, frequency, spread, spike), and any signals that require state across scans live there, not in lcs. | |
| Scan groups | 13 | ✅ Shipped | Orderless multi-input correlation. Per-input results plus group-level aggregations. Cross-input correlation reuses CorrelationEngine with each input bucketed as a synthetic engine labelled input:<label> — that label appears as the engine field on cross-input correlation findings. |
| Confidence calibration | 14 | 📅 Roadmap | Calibrated probability estimates combining evidence from string matches, semantic similarity, LLM verdicts, and correlation into a unified per-scan confidence score. Single-scan ensemble only (decided 2026-05-01); the multi-scan ensemble that adds session-signal evidence is a separate layer in aegis. |
- Exit codes —
0= clean,1= findings detected,2= error. Stable contract across all subcommands. - Output formats —
text(human-readable to stderr, summary to stdout),json(full report to stdout,jq-composable),quiet(exit code only). - Severity filtering —
--severity low|medium|high|criticalfilters before reporting. - Passthrough —
lcs scan -pwrites the original input to stdout when clean, suppressing the scan summary so the content can flow through a pipeline. Optional-o <file>redirects to a file.
Configuration is opt-in. CLI flags always override config; config always overrides built-in defaults. On first run, lcs init (or any lcs scan invocation with no config dir present) creates ~/.config/llm_context_shield/config.toml with all options commented out.
Config sections, all optional: [scan], [rules], [syara], [scoring], [correlation], and (📅 Phase 13) [scan_group], (📅 Phase 14) [confidence]. Session-related config ([session]) is owned by the aegis orchestrator project, not lcs.
The yara and syara engines load user-authored rules from $XDG_DATA_HOME/llm_context_shield/rules/{yara,syara}/. The correlation engine loads custom correlation rules from a TOML path specified in [correlation] custom_rules.
Rule authoring is documented in docs/rule-authoring.md. The simple→YARA migration guide lives at docs/migration-from-simple.md. The semantic-rule guide lives at docs/semantic-rules.md.
- Latency. A single-input scan on the
simpleengine is well under 100ms p99 on commodity hardware for inputs up to 1 MiB. Theyaraengine is comparable. Semantic tiers (syara-sbert,syara-llm) degrade gracefully — they add latency but only when their respective rule types are enabled. - Memory. Single-process resident set is bounded by input size + rule set size + (when enabled) one MiniLM model. No unbounded growth in long-running embeddings; session storage is caller-controlled (opt-in).
- Input size cap.
read_inputenforces a 100 MiB cap on stdin and file input.
No scan content leaves the process unless the operator has explicitly opted in to LLM-backed semantic rules (syara-llm). Even then, the configured endpoint is operator-chosen — it can be a fully local server (LMStudio, Ollama, vLLM) or a remote API. The default-build CLI (no LLM features) makes no network calls.
GroupReport (Phase 13) carries only metadata — categories, threat classes, severity histograms, scores. It does not store input text or finding details. (Phase 12's ScanSummary is owned by aegis as of 2026-04-30; the same minimisation posture applies there but is enforced in that project.)
Output across runs is byte-identical for the same input and configuration. Class scores use BTreeMap for alphabetical ordering. Rule iteration is stable. Exit codes are reproducible.
Eight targets ship in the release matrix: macOS arm64, macOS x86-64, Linux x86-64 (musl), Linux aarch64 (musl), Windows x86-64 (GNU), Windows ARM64, FreeBSD x86-64, WebAssembly (WASI). OpenBSD and NetBSD build natively but are not cross-compilable. See README.md "Release Builds" for the full matrix.
Per CLAUDE.md: never use the latest dependency version (target N-1), never use packages less than 30 days old, pin and verify hashes when possible, audit transitive dependencies on every bump.
Library consumers must be able to depend on llm_context_shield without pulling in clap, tracing-subscriber, or tracing-appender. The cli feature is default-on for binary builds and opt-out for library consumers (default-features = false).
The five use cases above translate to four concrete embedding surfaces. Each has a stable API contract from v0.4 onward.
Stable from v0.4. The lcs scan subcommand, its flags, and its three exit codes are part of the public contract. New flags may be added; existing ones do not change semantics. New output fields (in text and json) are additive.
Documented in README.md "Usage" and "Options" sections.
Stable from v0.4 with the Shield builder API at the surface. Engine implementations and the Engine trait are stable extension points. Send + Sync bounds on Engine are required for multi-threaded embeddings. (Session storage and cross-scan state are owned by the aegis orchestrator project — see §6.4.)
The Engine trait exposes per-rule introspection via rule_metadata() -> Vec<RuleMeta> (default impl returns empty for source compatibility — custom engines opt in by overriding). RuleMeta carries name, category, severity, threat_class, version, threat_level, and threshold — covering both rule identity and the scoring metadata the engine actually uses at scan time. Embedding hosts can call Shield::rule_set_fingerprint() to obtain a SHA-256 over the canonical-JSON sort of the loaded rule set, suitable for audit-trail attribution. The fingerprint is sensitive to scoring-metadata changes (bumping any rule's threshold shifts the fingerprint), which is the intended audit signal. See docs/rule-introspection.md for the full contract.
Crate features:
cli(default) — pulls inclap,tracing-subscriber,tracing-appender.yara— YARA-X engine.syara— SYARA-X engine (string-only by default).syara-sbert— adds local ONNX MiniLM similarity matching.syara-classifier— adds local ONNX classifier matching.syara-llm— adds OpenAI-compatible LLM matching (network-using).
Library consumers should use default-features = false and opt into only the engines they need. See README.md "Library Usage" and examples/embed.rs.
The wasm32-wasip1 target ships in the release matrix. Use cases:
- In-browser scanning (via WASI runtime in browser).
- In-runtime scanning embedded in a host application that uses Wasmtime / Wasmer / WasmEdge.
Constraints:
- LLM-backed rules (
syara-llm) are not usable in WASM today (no HTTP client in pure WASI preview-1). - Filesystem rule discovery requires WASI preview-2 capability grants; bundled rules work without filesystem access.
- The
simpleandyaraengines are the recommended starting points for WASM embeddings.
Long-running services, session-aware scanning, multi-step state, multi-tenancy, and any "MCP / HTTP server in front of the scanner" surface live in the separate aegis project. aegis is the orchestrator/wrapper that composes lcs into the gateway / WAF / multi-turn shapes. The transfer happened on 2026-04-30 — keeping lcs UNIX-composable as a single-shot scanner was preferred over bundling the full operator-grade stack into one binary.
Integration mode (decided 2026-05-01): aegis composes lcs as a subprocess, not a library dependency. aegis spawns lcs, pipes input on stdin, parses JSON from stdout. This decouples release cadences, sidesteps lcs's optional-feature combinatorics (yara, syara, syara-llm), and matches the contract surface listed below — which was deliberately built for subprocess composability. The library wrap (linking llm_context_shield as a Cargo dep) is deferred to a future productized aegis where ~ms-scale CLI startup + JSON parse cost would matter; for session-aware orchestration where one scan = one user turn, that cost is noise.
lcs's contracts that aegis composes against:
lcs scanJSON output (stable from v0.4).lcs rules --all --jsonper-instance rule schema export (Phase 11.6b).lcs rules --all --fingerprintcross-engine rule-set fingerprint (Phase 11.6b).- Three exit codes (
0clean,1findings,2error).
The following are explicitly not part of this project's scope. Each appears here because someone has asked or might ask, and we want a single canonical "no" with reasoning.
- Inline LLM rewriting / sanitisation. The scanner reports findings and lets the caller decide what to do. It does not mutate input. Sanitisation is a different problem with its own tradeoffs (false rewrites, semantic loss); coupling it to detection would conflate two concerns.
- Real-time streaming scans. All scans are single-shot per input. Callers wanting to chunk a large stream must do so themselves and either submit each chunk separately or assemble a complete input first.
- Authentication, rate-limiting, multi-tenant isolation. These belong to the embedding host. The library does not enforce them; the CLI does not provide them.
- Outbound LLM-output scanning. The taxonomy targets inputs to the model. Scanning model outputs for harmful content, hallucinations, or data leaks is a related but distinct problem with different threat models, ground truths, and rule shapes.
- Adversarial robustness guarantees. No detector is bulletproof. The project commits to detecting patterns from the documented taxonomy at the documented engine tiers. It does not commit to defeating adaptive attackers who craft inputs specifically to evade these rules. Adversarial-robustness testing is captured as a future research direction in Phase 14.
- Replacement for trust boundaries. This is a defence-in-depth layer. It does not replace input validation, output sanitisation, principle-of-least-privilege tool design, human review, or any of the other layers an LLM-backed system needs.
The active phase plan, sub-phases, checklists, and review notes live in tasks/todo.md. High-level phase status as of this revision:
| Phase | Topic | Status |
|---|---|---|
| 1–6 | Foundation, engines, library packaging, releases | ✅ Shipped |
| 7 | Heuristic threat scoring | ✅ Shipped |
| 8 | High-confidence rule expansion | ✅ Shipped (8d deferred) |
| 9 | Threshold-gated behavioural rules | ✅ Shipped |
| 10 | SYARA-only semantic rules | ✅ Shipped |
| 11 | Cross-rule correlation | ✅ Shipped |
| 12 | Session-aware scanning | |
| 13 | Scan groups | 📅 Roadmap — confirmed in lcs (2026-05-01); cross-input correlation reuse keeps it here. Boundary: lcs takes paths/strings, no input-collection plumbing |
| 14 | Confidence calibration and ensemble scoring | 📅 Roadmap — single-scan ensemble in lcs (2026-05-01); aegis owns a separate multi-scan ensemble layer over lcs's per-scan output + session signals |
The PRD owns what and why. tasks/todo.md owns how and when. When a roadmap phase ships, this table moves the row from 📅 to ✅; the use-case mapping table in §3 also gets updated.
- 2026-05-01 — Three open architectural questions from the 2026-04-30 transfer resolved. Q1: aegis composes lcs via subprocess (library option deferred to a future productized aegis). Q2: Phase 13 stays in lcs (cross-input correlation reuse argues against transfer; boundary held at "lcs takes paths/strings, no input-collection plumbing"). Q3: Phase 14 single-scan ensemble stays in lcs; aegis owns a separate multi-scan ensemble layer that wraps lcs's per-scan
ConfidenceScorewith session-signal evidence. Follow-on PRD fixes: §3 UC-5 pinned subprocess; §4.3 capabilities table updated (Session tracking marked transferred; Phase 14 description scoped to single-scan, multi-scan layer noted as aegis-owned); §5.2 dropped staleScanSummaryreference; §6.4 added "Integration mode" paragraph stating subprocess composition with rationale. - 2026-04-30 — Phase 12 (session-aware scanning) transferred to the aegis orchestrator project. lcs scope tightened to single-shot scanning; UC-5 dropped from lcs's use-case set, UC-4 reframed as orderless-only. §6.4 rewritten from "future MCP / HTTP server" speculation to a concrete reference to aegis. Roadmap table updated.
- 2026-04-25 — Initial PRD created. Retrofitted at v0.4 after Phase 11 shipped. Anchored UC-1 through UC-5; revised Phase 12 scope (session tracking only) and split out Phase 13 (scan groups) based on the use-case framing.