Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,5 @@ dist
*.log
.e2e-profile
.e2e-shots/
vendor/
.e2e-cache/
36 changes: 36 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,42 @@ so an update never silently changes what an existing user sees; the bridge to ma
userData), never shared. `settings/options.ts` lists every terminal option by id; a unit test holds
the list and the sections together, and each app's e2e (`options`) asserts its page shows that
list and no terminal-looking row of its own. A row outside the list is a fork.
- **THIS IS A PRODUCT FOR OTHER PEOPLE** (owner, 2026-09-19: "this isn't an app for just me. keep
that in mind with all things you implement"). A feature bundles or fetches what it needs and works
on a fresh Windows install: never lean on the owner's GPU, tools, caches or installed runtimes.
- **DICTATION** (#13, 2026-09-19; spec in `docs/superpowers/specs/2026-09-19-dictation-design.md`).
Built ONCE in `core/`, for Prism too. Hold Right Alt, speak, release: the FINAL text is pasted at the
cursor of the shell that was in front when you started. The rules that must not regress:
- **It is the FOURTH written exception to "the app never types into your shell"** (with the agent
resume, Prism's `Set-Location`, and drop-to-type-path): the user spoke it on purpose, it arrives as
one bracketed paste, and it NEVER carries Enter or a newline (`cleanTranscript` and the panel's
`pasteSpokenText` both strip them; the e2e asserts no new prompt appears).
- **Off means off.** Default off. With it off no key listener is armed, no process exists, nothing
downloads. Switching it off kills the server. The e2e counts `whisper-server.exe` processes.
- **Local only.** Audio goes renderer -> main -> `whisper-server` on 127.0.0.1 and is never written to
disk. Official whisper.cpp binaries only, nothing we compile; every download (engine, models, the
NVIDIA pack) is pinned by SHA-256 in `core/shared/dictationCatalog.ts` and an unverified file is
deleted, never loaded. Model urls are pinned to a Hugging Face COMMIT, not `main`.
- **Right Alt is AltGr** on Norwegian and most European keyboards (Windows sends Ctrl+RightAlt, then
the key). So a bare-modifier hotkey is a SOLO HOLD of 200 ms; any other key during that window
means it was typing. `core/renderer/lib/dictationKey.ts` is the pure reducer, heavily tested.
- **Live text is provisional and only SHOWN** (in the pill); only the final pass is pasted. MEASURED:
a phrase heard wrong at 6 s was corrected by 8 s, and text typed into a prompt cannot be unsent.
- **The engine ships its own C++ runtime.** whisper.cpp's exe and DLLs import MSVCP140 /
VCRUNTIME140(_1) / VCOMP140, which neither official zip carries and a fresh Windows may lack.
`core/tools/fetch-whisper.mjs` copies them app-local from the build machine, only if validly
signed by Microsoft, plus the MIT notice. The GPU pack finds them through PATH (the engine sets it).
- **Shared files, per-app values.** Models and the GPU pack live in `%LOCALAPPDATA%\PrismDictation`
(both apps; never removed by an uninstall); every setting VALUE is per app.
- **GPU:** the official CUDA 12.4 pack (643 MB) is an optional download, offered only when an NVIDIA
adapter is found. MEASURED on an RTX 5090: it runs (first run 9 s compiling kernels), Large-v3 then
answers in 0.36-0.5 s. AMD/Intel stay on CPU (Base/Small) until there is an official build.
- **Testing is REAL:** the `dictation` e2e feeds Chromium's fake microphone a WAV
(`PT_E2E_MIC`), runs the real bundled engine with the Tiny model (cached in `.e2e-cache/`), and
asserts the sentence lands on the prompt line. Not machine-testable, so on the hands-on list: how
it sounds, AltGr on a physical keyboard, pause-media against a real player.
- Re-pinning the engine: `fetch-whisper.mjs` and `ENGINE` in the catalog must agree (a test holds
them together), and the GPU pack must be the SAME release tag.
- **A PAGE THAT WORKS IS NOT A PAGE THAT LOOKS RIGHT** (#20, 2026-09-19). Moving the settings into
`core/` dropped every Tailwind class used only there (`core/` is outside the scanned root; the fix
is the `@source` line at the top of `index.css`, do not remove it). All 14 e2e scenarios passed over
Expand Down
13 changes: 12 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,14 @@ and want to know, at a glance, which one has finished.
milliseconds of Enter and clears the instant the answer lands, where a terminal that scores output
is a second or two late at both ends. It wears the theme's accent, as a quiet line under the tab
or, turned up, the whole tab filled, which then also holds a "finished" colour until you visit it.
- **Dictation, on your own machine.** Turn it on in Settings, hold `Right Alt` and speak: the words
land at the cursor when you let go, in any tab and any shell. It is
[whisper.cpp](https://github.com/ggml-org/whisper.cpp) running locally, so your voice never leaves
the PC and it works offline. A pill shows a live level meter and the text as you speak, so a muted
microphone is obvious at once, not after a lost sentence. It never presses Enter for you. Pick a
model in Settings (Base runs well on any CPU; with an NVIDIA card one click adds the GPU engine and
the large models answer in under a second), a language or auto-detect, and optionally pause your
music while you talk. Off until you switch it on: nothing listens and nothing downloads before that.
- **Tabs that come back.** Close the app and reopen it: every tab returns in the folder its shell
was in, and a tab that hosted Claude or Codex resumes that conversation (`claude --resume <id>`,
`codex resume --last`).
Expand Down Expand Up @@ -76,6 +84,7 @@ lands on a start screen with the folders you were last in, each one press from a
| `Ctrl+,` | Settings |
| `Ctrl+scroll` | Zoom this tab's text |
| `F11` | Fullscreen |
| hold `Right Alt` | Dictate (when switched on; rebindable, or press-to-toggle) |

Everything else belongs to the shell: `Escape` is still vim's, and `Ctrl+Backspace` deletes a word
(`Ctrl+W` closes the tab, as it does in a browser).
Expand All @@ -90,14 +99,16 @@ prompt, which is what names the tab; a WSL tab keeps the folder it opened in.

```bash
npm install
npm run fetch:bin # the speech engine, once (pinned by SHA-256; e2e and package run it too)
npm run dev # run it
npm test # unit tests (vitest)
npm run e2e # drives the built app, parked offscreen and unfocused
npm run package # dist/PrismTerminal-Setup-x64-<version>.exe
```

Electron, React 19, TypeScript, Vite, Tailwind v4, [xterm.js](https://xtermjs.org) and
[node-pty](https://github.com/microsoft/node-pty). Design notes live in
[node-pty](https://github.com/microsoft/node-pty). Dictation is [whisper.cpp](https://github.com/ggml-org/whisper.cpp)'s
official Windows build (MIT), fetched at build time and shipped beside the app. Design notes live in
[`docs/superpowers`](docs/superpowers).

## License
Expand Down
41 changes: 31 additions & 10 deletions core/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,14 @@ ask me."* And: *"why can't this repo be the core?"* It can, and this is it.
2. **What is an app's shell stays in the app.** Prism: roots and the wall, the
sidebar, the split dock, several shells per tab, app styles. Prism Terminal:
folder tabs, the start screen, window chrome from the theme, the lifecycle.
Dictation (#13) is in here too, whole: the key, the microphone, the engine,
the model store, the pill, the tab mark and its Settings page. A host wires
it with four lines: `registerDictationIpc` in main, `createDictationApi` in
the preload, the `dictation` field of its host config, and
`useDictationArm` + `<DictationPill>` where its terminal is drawn. It also
runs `core/tools/fetch-whisper.mjs <dir>` at build time and ships that folder
as `resources/bin/whisper`, and grants its own window the `media` permission
(audio only).
3. **A difference between the apps is DECLARED, never forked.** Every place the
two legitimately differ is a field of `TermHostConfig` in
[`renderer/host.ts`](renderer/host.ts): the default each untouched setting
Expand Down Expand Up @@ -97,16 +105,29 @@ lockfile but ships against Prism's.

## Releasing the core

1. Land the change in Prism Terminal (PR, the app's full gate).
2. Bump `core/package.json`'s version. Publish: `git subtree split --prefix=core -b core-dist`,
tag `core-v<version>`, push the branch and the tag. Tags are `core-v*` for the
core and `v*` for the app, since this repo is both.
3. In Prism, and only after ASKING THE OWNER (2026-09-19: a core release is
never pulled into Prism unasked): bump the pin, run `npm run e2e:terminal`
(Prism's terminal gate: every scenario the terminal can break, required
green for any pin bump, because Prism has the larger footprint and so more
ways to break), then Prism's usual gate, install, PR. The core's lint and
unit tests run only here; Prism's gate on it is its compiler and its e2e.
Nobody does, by hand (PrismTerminal #23; owner, 2026-09-19: "that compiled copy
needs to be auto bumped when a new Prism Terminal release or merge to main
happens"). Land the change through a PR in Prism Terminal, WITH a bump of
`core/package.json`'s version (CI fails the PR otherwise: a released tag is never
moved). On merge, `.github/workflows/core-release.yml`:

1. re-splits this folder into `core-dist` and tags `core-v<version>`. Tags are
`core-v*` for the core and `v*` for the app, since this repo is both;
2. opens a PR in Prism bumping the pin and Prism's own version;
3. waits for that PR's checks, which include Prism's TERMINAL GATE on a GitHub
runner (`terminal-gate.yml` there): the built app driven through every
scenario the terminal can break, real dictation included. Prism has the larger
footprint and so more ways to break, which is why the proof runs there;
4. merges it, if the owner has set the repo variable `PRISM_AUTO_MERGE` to
`true`, and Prism releases itself; otherwise the green PR waits for a person.
A red check never merges.

The one thing still cut by hand is a RELEASE CANDIDATE for an open PR
(`core-v0.2.0-rc.N`, split from the PR's branch), so Prism's half can be built
and tested before the core's half merges. The core's lint and unit tests run
only here; Prism's gate on it is its compiler, its unit suite and that e2e.
**Keep Prism's gate runner-safe:** nothing in those scenarios may assume one
particular machine.

## Status (2026-09-19)

Expand Down
116 changes: 116 additions & 0 deletions core/main/__fixtures__/fakeMediaHelper.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
// A stand-in for the resident PowerShell media helper (#13), so mediaPause's
// tests never touch the owner's real media sessions and never start a real
// powershell.exe. It speaks the same line protocol on stdin and stdout:
//
// pause <n> -> <n> <comma-separated ids that WERE Playing>
// resume <n> <ids> -> <n> ok
// list <n> -> <n> <id>=<status>,...
//
// and keeps a tiny world of sessions in memory, so a second pause finds nothing
// left to pause and a resume really does put a session back to Playing.
//
// Arguments (all optional):
// --sessions a=Playing,b=Paused the world it starts with
// --log <file> every request line is appended here
// --mode normal|silent|garbage|reverse|crash-once
// --marker <file> crash-once: exit on the first request unless
// this file exists, creating it on the way out
// --delay <ms> answer each request this late
//
// Everything is imported rather than taken off a global, because eslint gives a
// bare .mjs under core/ no Node globals.
import { appendFileSync, existsSync, writeFileSync } from 'node:fs'
import process from 'node:process'
import { createInterface } from 'node:readline'
import { setTimeout as later } from 'node:timers'

function arg(name, fallback) {
const at = process.argv.indexOf(`--${name}`)
return at >= 0 && at + 1 < process.argv.length ? process.argv[at + 1] : fallback
}

const mode = arg('mode', 'normal')
const logFile = arg('log', '')
const marker = arg('marker', '')
const delay = Number(arg('delay', '0'))

/** id -> status, in the order given. */
const sessions = new Map(
arg('sessions', '')
.split(',')
.filter(Boolean)
.map((pair) => {
const cut = pair.lastIndexOf('=')
return [pair.slice(0, cut), pair.slice(cut + 1)]
})
)

function answerTo(line) {
const [verb, n, rest] = splitThree(line)
if (verb === 'pause') {
const ids = []
for (const [id, status] of sessions) {
if (status === 'Playing') {
sessions.set(id, 'Paused')
ids.push(id)
}
}
return `${n} ${ids.join(',')}`
}
if (verb === 'resume') {
for (const id of rest.split(',')) if (sessions.get(id) === 'Paused') sessions.set(id, 'Playing')
return `${n} ok`
}
if (verb === 'list') return `${n} ${[...sessions].map(([id, status]) => `${id}=${status}`).join(',')}`
return null
}

/** "verb n rest of the line", where the rest may itself hold spaces. */
function splitThree(line) {
const first = line.indexOf(' ')
if (first < 0) return [line, '', '']
const second = line.indexOf(' ', first + 1)
if (second < 0) return [line.slice(0, first), line.slice(first + 1), '']
return [line.slice(0, first), line.slice(first + 1, second), line.slice(second + 1)]
}

function say(text) {
if (text === null) return
if (delay > 0) later(() => process.stdout.write(`${text}\n`), delay)
else process.stdout.write(`${text}\n`)
}

/** reverse mode holds the first answer until the second exists, then swaps them. */
let held = null

const lines = createInterface({ input: process.stdin })
lines.on('line', (line) => {
if (logFile) appendFileSync(logFile, `${line}\n`)
if (mode === 'crash-once' && marker && !existsSync(marker)) {
writeFileSync(marker, 'crashed once')
process.exit(1)
}
if (mode === 'silent') return
const answer = answerTo(line)
if (mode === 'garbage') {
// What a real PowerShell can put on stdout around an answer: a stray
// warning, a blank line, and an answer to a request nobody made.
say('WARNING: something nobody asked for')
say('')
say('999999 not-yours')
}
if (mode === 'reverse') {
if (held === null) {
held = answer
return
}
say(answer)
say(held)
held = null
return
}
say(answer)
})
// Exactly what the real helper does when the app goes away: stdin closes, the
// read loop ends, the process exits. It is why no orphan can outlive its parent.
lines.on('close', () => process.exit(0))
123 changes: 123 additions & 0 deletions core/main/__fixtures__/fakeWhisperServer.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,123 @@
// A stand-in for whisper-server.exe, for dictationEngine.test.ts. It speaks the
// two routes the engine uses (GET / and POST /inference, measured against the
// real b5130 server on 2026-09-19) and answers with a description of what it
// RECEIVED, so a test can read the argv, the multipart parts and the order of
// the passes back out of `text` instead of trusting the engine's word for them.
//
// Started the way the real one is: -m <model> --host 127.0.0.1 --port <n> -l <lang>.
// Switches, as argv or as the environment variable beside each:
// --fake-exit FAKE_WHISPER_EXIT=1 say why on stderr and exit 3 at once
// (a GPU engine that cannot start)
// --fake-never-ready FAKE_WHISPER_NEVER_READY=1 stay alive and never listen
// --fake-delay <ms> FAKE_WHISPER_DELAY=<ms> hold every /inference answer this long
// --fake-die-on-inference FAKE_WHISPER_DIE=1 read the upload, say why on stderr, exit 4
//
// The globals are imported by name because the repo's eslint config gives Node
// globals to tools/ only, and this file may not change that config.
import { Buffer } from 'node:buffer'
import http from 'node:http'
import process from 'node:process'
import { setInterval, setTimeout } from 'node:timers'

const argv = process.argv.slice(2)
const flag = (name) => argv.includes(name)
const value = (name) => {
const i = argv.indexOf(name)
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : null
}

const model = value('-m')
const language = value('-l')
const host = value('--host') ?? '127.0.0.1'
const port = Number(value('--port') ?? 0)
const delay = Number(value('--fake-delay') ?? process.env.FAKE_WHISPER_DELAY ?? 0)
const exitNow = flag('--fake-exit') || process.env.FAKE_WHISPER_EXIT === '1'
const neverReady = flag('--fake-never-ready') || process.env.FAKE_WHISPER_NEVER_READY === '1'
const dieOnInference = flag('--fake-die-on-inference') || process.env.FAKE_WHISPER_DIE === '1'

/** The parts of a multipart/form-data body: name, filename, type and the bytes. */
function parseMultipart(body, contentType) {
const m = /boundary=(?:"([^"]+)"|([^;]+))/i.exec(contentType ?? '')
if (!m) return []
const mark = Buffer.from(`--${m[1] ?? m[2]}`)
const parts = []
let at = body.indexOf(mark)
while (at >= 0) {
const next = body.indexOf(mark, at + mark.length)
if (next < 0) break
// Between two marks: CRLF, the headers, a blank line, the bytes, CRLF.
const chunk = body.subarray(at + mark.length + 2, next - 2)
const split = chunk.indexOf('\r\n\r\n')
if (split >= 0) {
const head = chunk.subarray(0, split).toString('utf8')
parts.push({
name: /name="([^"]*)"/i.exec(head)?.[1] ?? null,
filename: /filename="([^"]*)"/i.exec(head)?.[1] ?? null,
type: /content-type:\s*(.+)/i.exec(head)?.[1]?.trim() ?? null,
data: chunk.subarray(split + 4)
})
}
at = next
}
return parts
}

if (exitNow) {
process.stderr.write('fake: CUDA error: no kernel image is available for execution on the device\n')
process.exit(3)
} else if (neverReady) {
// Alive, and deaf: what a server stuck loading its model looks like from outside.
setInterval(() => {}, 1000)
} else {
let count = 0
let active = 0
let maxActive = 0
const server = http.createServer((req, res) => {
if (req.method === 'GET') {
res.writeHead(200, { 'content-type': 'text/html' })
res.end('<html>fake whisper server</html>')
return
}
if (req.method !== 'POST' || req.url !== '/inference') {
res.writeHead(404)
res.end()
return
}
const chunks = []
req.on('data', (c) => chunks.push(c))
req.on('end', () => {
if (dieOnInference) {
// The whole message goes out before the exit, so the engine's stderr tail has it.
process.stderr.write('fake: ggml_cuda: out of memory\n', () => process.exit(4))
return
}
active += 1
maxActive = Math.max(maxActive, active)
count += 1
const n = count
const parts = parseMultipart(Buffer.concat(chunks), req.headers['content-type'])
const file = parts.find((p) => p.name === 'file')
const format = parts.find((p) => p.name === 'response_format')
setTimeout(() => {
const seen = {
n,
maxActive,
bytes: file ? file.data.length : -1,
filename: file?.filename ?? null,
type: file?.type ?? null,
format: format ? format.data.toString('utf8') : null,
model,
language,
host,
port,
pid: process.pid,
cwd: process.cwd()
}
active -= 1
res.writeHead(200, { 'content-type': 'application/json' })
res.end(JSON.stringify({ text: JSON.stringify(seen) }))
}, delay)
})
})
server.listen(port, host)
}
Loading
Loading