Repository navigation
Dictation when using the agents #13
Description
Activity
Owner, 2026-09-19:
- Win+H is NOT good enough: the model cannot be chosen, and it is limited to the user's Windows language. Both are requirements here: a choosable model, and not tied to the OS language.
- The owner uses OpenWhispr today, almost only in the terminal, so building dictation INTO Prism Terminal is wanted.
- 'I think I'm using a model from OpenAI maybe, but there are so many options.'
What the owner's OpenWhispr setup actually is (read off this machine, 2026-09-19; secrets not read):
- Transcription is LOCAL, not a cloud API:
LOCAL_TRANSCRIPTION_PROVIDER=whisper(whisper.cpp, GGML models),WHISPER_CUDA_ENABLED=true,DICTATION_LANGUAGE=auto. - Selected model:
LOCAL_WHISPER_MODEL=base(ggml-base.bin, 141 MB). Also downloaded and unused:ggml-large-v3.bin(2.9 GB) and NVIDIA Parakeet (int8 ONNX, ~630 MB). Models live in~/.cache/openwhispr/{whisper-models,parakeet-models}. - Hardware: RTX 5090, so a large Whisper model is near-instant locally.
Implication for the design: the closest thing to what the owner already likes is a bundled local whisper.cpp (a fetched binary, the way Prism bundles ffmpeg, so no new npm/native dependency) with a model picker. A cloud engine is the alternative or an addition. To be decided with the owner before building.
Decided (owner, 2026-09-19):
- Engine: local Whisper (whisper.cpp), the kind the owner already uses, with a choosable model. It must be BUNDLED with the app: 'this isn't an app for just me. keep that in mind with all things you implement.' So nothing may depend on OpenWhispr, its model cache, the owner's GPU or CUDA. Design for a stranger's PC: a build that runs on any GPU vendor or on the CPU, and a model the app fetches itself.
- Trigger: the user chooses in Settings between hold-to-talk and press-to-start / press-to-stop.
- Output: the recognised text goes straight into the prompt at the cursor as a bracketed paste. It NEVER presses Enter. This is a written exception to the 'never types into your shell' rule, the same kind as drop-to-type-path.
- Scope: any tab and any shell, not only while an agent is present.
Standing rule to add to CLAUDE.md with this PR: Prism Terminal is a product for other users; every feature bundles or fetches what it needs and works on a fresh Windows install.
Correction to the findings above: the owner's OpenWhispr screenshot shows Large (2952 MB) as the ACTIVE model, not
base. TheLOCAL_WHISPER_MODEL=baseline in its.envis stale; the app keeps its real choice elsewhere. So the accuracy the owner is used to is large-v3 on an RTX 5090.Decided (owner, 2026-09-19), round two:
- Model manager in Settings, under its own Dictation tab, in the shape of OpenWhispr's (owner screenshot): a list of models, each with its size and a Download button, a 'Recommended' badge, an 'Active' marker on the one in use. The app downloads models itself; nothing is installed by hand.
- A few models are shown, one marked (Recommended); the user picks. (Which one gets the badge per machine is a design detail: see the hardware note below.)
- Off by default. Nothing listens and nothing downloads until the user turns dictation on.
- Hotkey: Right Alt. Hold-to-talk is the DEFAULT; press-to-start / press-to-stop is the other mode, chosen in Settings. Rebindable.
- Sub-option: pause media while dictating (and resume it after).
Design notes these decisions force (for the spec):
- Right Alt is AltGr on Norwegian and most European layouts, where it types
@ { } [ ] \ | ~ €. Windows reports it as LeftCtrl+RightAlt (code: AltRight,key: AltGraph). Since the app is for other people too, the key must be read as a SOLO hold: recording starts only after Right Alt has been down ~200ms with no other key, and any other key pressed during the hold cancels it (that was AltGr typing, not dictation). Its keyup is swallowed so Alt does not open the window's system menu. - Pause media cannot be the media play/pause key: that key TOGGLES, so it would START music that was not playing. It has to ask Windows what is playing (the system media transport controls,
GlobalSystemMediaTransportControlsSessionManager), pause only the sessions that were playing, and resume exactly those. Reachable from a persistent PowerShell helper, the patterndwmHelper.tsalready uses, so no new dependency. - GPU: whisper.cpp's Vulkan build (any vendor) with a CPU fallback, so there is no CUDA-only path and ideally no 'Enable GPU' step at all.
Decided (owner, 2026-09-19), round three:
- Live feedback is a requirement. "I like whisper but its main issue is that you see no text before you stop. So if your mic's off you won't know until you finish. A model that could write as you speak would be nice." The owner has had issues with NVIDIA's models, but is open to them or to a better alternative.
- Indicator: a "Listening..." pill over the terminal with a live level meter, turning to "Transcribing..." on release, plus a mic mark on the tab, plus a microphone sound effect on start and stop.
- Cleanup: light touch only, no second LLM: Whisper's own punctuation and capitalisation, plus simple deterministic cleanup (doubled words, fillers, stray whitespace, engine artifacts like
[BLANK_AUDIO]). What OpenWhispr does when its AI cleanup is off. - Next step: write the spec + plan for dictation.
Measured 2026-09-19 (spike, throwaway; owner's PC, Ryzen 9 9950X3D, CPU ONLY, no GPU pack):
- Official whisper.cpp Windows builds (release
b5130):whisper-bin-x64.zipis 8 MB (sha256f9ec6c52a2e949b62ab51fa21d0d497958f9e41c3010c157c4e42932d5316f3c), CUDA builds are 260-640 MB, and there is NO official Vulkan build. "One Vulkan build for every GPU" would mean compiling it ourselves. - That 8 MB zip contains
whisper-cli.exe,whisper-server.exe(resident HTTP server: the model loads once),whisper-stream.exe, VAD tools, andparakeet-cli.exe+parakeet.dll: NVIDIA Parakeet now ships inside whisper.cpp, so one small bundle could serve both families. whisper-cli, base, 11 s clip, cold (model load included): 1.46 s.whisper-server, base, a GROWING clip re-sent each time (what live text needs): 2 s -> 0.87 s, 4 s -> 0.87 s, 6 s -> 0.89 s, 8 s -> 0.90 s, 11 s -> 0.93 s. About 0.9 s per pass regardless of length, so text can update roughly once a second with no GPU at all.- Partials are PROVISIONAL: at 6 s it heard "ask not what you are coming from", corrected by 8 s. So live text belongs in the pill, and only the FINAL text is pasted into the shell (text typed into a prompt cannot be cleanly taken back).
whisper-cli, large-v3, CPU only: 20.4 s for the 11 s clip. Large and Turbo are GPU-only in practice; Base and Small are the CPU models. This was a very fast desktop CPU; a laptop will be several times slower.- OpenWhispr's own packaging matches: CPU engine in the installer, a CUDA pack downloaded by its "Enable GPU" button (
%APPDATA%\open-whispr\bin\whisper-cuda), sherpa-onnx bundled for the NVIDIA models.
What this means for the design: bundle the 8 MB official CPU build (pinned tag + SHA-256, fetched at build time like Prism's ffmpeg); run
whisper-serverresident while dictation is on (idle stand-down); live text by re-sending the growing clip, shown in the pill; an optional GPU pack download in the model manager (the owner's "Enable GPU" shape); recommend Base or Small on CPU and Turbo once a GPU pack is in. Parakeet via the same bundle is worth a second spike before promising it.Sequencing (owner, 2026-09-19): dictation waits until the shared core (#15) exists, so it is built once and lands in both Prism and Prism Terminal. The decisions and measurements above stand and feed the spec then.
Decided (owner, 2026-09-19), round four, after the shared core landed:
- GPU: CPU engine bundled in the installer; an 'Enable GPU' button in the model manager downloads the OFFICIAL NVIDIA (CUDA) pack, which unlocks Turbo and Large. AMD/Intel users stay on CPU (Base/Small) in v1. No engine we compile ourselves. Checked the same day: whisper.cpp's latest build release (
b5130) still has no any-vendor (Vulkan) Windows x64 build. - Model files: ONE shared folder for both apps (
%LOCALAPPDATA%\PrismDictation), so a model is downloaded once. Settings VALUES stay per app (active model, hotkey, on/off). Uninstalling one app leaves the models. - Language: Auto-detect by default, with a Language picker to pin one. Multilingual models only.
- Prism: its own 'Dictation' page in the Settings rail under Behaviour (Prism Terminal: its own tab, as decided).
- GPU: CPU engine bundled in the installer; an 'Enable GPU' button in the model manager downloads the OFFICIAL NVIDIA (CUDA) pack, which unlocks Turbo and Large. AMD/Intel users stay on CPU (Base/Small) in v1. No engine we compile ourselves. Checked the same day: whisper.cpp's latest build release (
Spec and plan are written, on branch
feat/13-dictation, waiting for the owner's single approval:docs/superpowers/specs/2026-09-19-dictation-design.md,docs/superpowers/plans/2026-09-19-dictation.md.Measured 2026-09-19 (GPU spike, owner's RTX 5090): the official
whisper-cublas-12.4.0-bin-x64.zip(643 MB, sha256af520ddd034d985b55dfeea3e465ed93653ba2aee1a55e865033edc548c272a7) RUNS on a 5090 although the card is newer than CUDA 12.4: first run 9.1 s (one-time kernel compile), then Large-v3 through a residentwhisper-serveris 0.36 to 0.50 s per pass on an 11 s clip. So one pack covers old and new NVIDIA cards, and live text works with the largest model.- added 4 commits that reference this issue
on Sep 19, 2026 Built, PR #22. Spikes (2026-09-19): Chromium's fake microphone works under Electron (peak 0.78 from the WAV), so the e2e is real end to end; Windows' media sessions are reachable from Windows PowerShell 5.1 via WinRT, so pause-media shipped as designed. Found while building: the engine needs Microsoft's C++ runtime, which no official zip carries; it now ships app-local beside the engine (Microsoft-signed copies only). Next: Prism, after asking.
2 remaining items
- added 11 commits that reference this issue
on Sep 19, 2026 - added a commit that references this issue
on Sep 19, 2026 - added a commit that references this issue
on Sep 25, 2026 - added a commit that references this issue
on Sep 25, 2026
Owner, 2026-09-19. Planned, not started. Before building: ask the owner the open questions below and agree the design (spec + plan, one approval).
Goal
Dictation when using the agents: speak a prompt instead of typing it.
What the owner asked for
Open questions to ask before designing
Constraints already known