Skip to content

Dictation when using the agents #13

Description

@Maxaubert

Owner, 2026-09-19. Planned, not started. Before building: ask the owner the open questions below and agree the design (spec + plan, one approval).

Goal

Dictation when using the agents: speak a prompt instead of typing it.

What the owner asked for

  • A dictation feature for use with the agents (Claude Code, Codex).
  • Optional: a setting decides whether the user sees it at all.

Open questions to ask before designing

  • Which speech engine? Options differ a lot: Windows' own dictation (Win+H already works in any text field, including this terminal, with nothing to build), Chromium's Web Speech API (sends audio to a cloud service, and may not be available in Electron at all), a local model such as Whisper (private and offline, but a large download and a new native/runtime dependency), or a hosted API (needs a key and sends audio out). The owner already has local speech tooling on this machine (Sonara, OpenWhispr), which may be a reason to integrate rather than rebuild.
  • If Win+H already does the job, what should this add: a visible button, push-to-talk, better accuracy, punctuation, a preview?
  • How is it triggered: a hold-to-talk hotkey, a toggle button in the chrome, or both? Which key, given the shell keeps most of them?
  • Where does the text go: straight into the agent's prompt as it is recognised, or into a preview you can correct before it is sent? Does it ever press Enter? (CLAUDE.md's standing rule is that the app never types into a user's shell; text the user dictated on purpose is a new, owner-visible exception to decide, not to assume.)
  • Only when an agent is present in the tab, or in any shell?
  • Languages: English only, or Norwegian too?
  • Privacy: is audio leaving the machine acceptable?

Constraints already known

  • The renderer is sandboxed; microphone access needs an explicit permission path.
  • No new runtime dependency without a reason (eight today).

Activity

  1. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Owner, 2026-09-19:

    • Win+H is NOT good enough: the model cannot be chosen, and it is limited to the user's Windows language. Both are requirements here: a choosable model, and not tied to the OS language.
    • The owner uses OpenWhispr today, almost only in the terminal, so building dictation INTO Prism Terminal is wanted.
    • 'I think I'm using a model from OpenAI maybe, but there are so many options.'

    What the owner's OpenWhispr setup actually is (read off this machine, 2026-09-19; secrets not read):

    • Transcription is LOCAL, not a cloud API: LOCAL_TRANSCRIPTION_PROVIDER=whisper (whisper.cpp, GGML models), WHISPER_CUDA_ENABLED=true, DICTATION_LANGUAGE=auto.
    • Selected model: LOCAL_WHISPER_MODEL=base (ggml-base.bin, 141 MB). Also downloaded and unused: ggml-large-v3.bin (2.9 GB) and NVIDIA Parakeet (int8 ONNX, ~630 MB). Models live in ~/.cache/openwhispr/{whisper-models,parakeet-models}.
    • Hardware: RTX 5090, so a large Whisper model is near-instant locally.

    Implication for the design: the closest thing to what the owner already likes is a bundled local whisper.cpp (a fetched binary, the way Prism bundles ffmpeg, so no new npm/native dependency) with a model picker. A cloud engine is the alternative or an addition. To be decided with the owner before building.

  2. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Decided (owner, 2026-09-19):

    • Engine: local Whisper (whisper.cpp), the kind the owner already uses, with a choosable model. It must be BUNDLED with the app: 'this isn't an app for just me. keep that in mind with all things you implement.' So nothing may depend on OpenWhispr, its model cache, the owner's GPU or CUDA. Design for a stranger's PC: a build that runs on any GPU vendor or on the CPU, and a model the app fetches itself.
    • Trigger: the user chooses in Settings between hold-to-talk and press-to-start / press-to-stop.
    • Output: the recognised text goes straight into the prompt at the cursor as a bracketed paste. It NEVER presses Enter. This is a written exception to the 'never types into your shell' rule, the same kind as drop-to-type-path.
    • Scope: any tab and any shell, not only while an agent is present.

    Standing rule to add to CLAUDE.md with this PR: Prism Terminal is a product for other users; every feature bundles or fetches what it needs and works on a fresh Windows install.

  3. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Correction to the findings above: the owner's OpenWhispr screenshot shows Large (2952 MB) as the ACTIVE model, not base. The LOCAL_WHISPER_MODEL=base line in its .env is stale; the app keeps its real choice elsewhere. So the accuracy the owner is used to is large-v3 on an RTX 5090.

    Decided (owner, 2026-09-19), round two:

    • Model manager in Settings, under its own Dictation tab, in the shape of OpenWhispr's (owner screenshot): a list of models, each with its size and a Download button, a 'Recommended' badge, an 'Active' marker on the one in use. The app downloads models itself; nothing is installed by hand.
    • A few models are shown, one marked (Recommended); the user picks. (Which one gets the badge per machine is a design detail: see the hardware note below.)
    • Off by default. Nothing listens and nothing downloads until the user turns dictation on.
    • Hotkey: Right Alt. Hold-to-talk is the DEFAULT; press-to-start / press-to-stop is the other mode, chosen in Settings. Rebindable.
    • Sub-option: pause media while dictating (and resume it after).

    Design notes these decisions force (for the spec):

    • Right Alt is AltGr on Norwegian and most European layouts, where it types @ { } [ ] \ | ~ €. Windows reports it as LeftCtrl+RightAlt (code: AltRight, key: AltGraph). Since the app is for other people too, the key must be read as a SOLO hold: recording starts only after Right Alt has been down ~200ms with no other key, and any other key pressed during the hold cancels it (that was AltGr typing, not dictation). Its keyup is swallowed so Alt does not open the window's system menu.
    • Pause media cannot be the media play/pause key: that key TOGGLES, so it would START music that was not playing. It has to ask Windows what is playing (the system media transport controls, GlobalSystemMediaTransportControlsSessionManager), pause only the sessions that were playing, and resume exactly those. Reachable from a persistent PowerShell helper, the pattern dwmHelper.ts already uses, so no new dependency.
    • GPU: whisper.cpp's Vulkan build (any vendor) with a CPU fallback, so there is no CUDA-only path and ideally no 'Enable GPU' step at all.
  4. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Decided (owner, 2026-09-19), round three:

    • Live feedback is a requirement. "I like whisper but its main issue is that you see no text before you stop. So if your mic's off you won't know until you finish. A model that could write as you speak would be nice." The owner has had issues with NVIDIA's models, but is open to them or to a better alternative.
    • Indicator: a "Listening..." pill over the terminal with a live level meter, turning to "Transcribing..." on release, plus a mic mark on the tab, plus a microphone sound effect on start and stop.
    • Cleanup: light touch only, no second LLM: Whisper's own punctuation and capitalisation, plus simple deterministic cleanup (doubled words, fillers, stray whitespace, engine artifacts like [BLANK_AUDIO]). What OpenWhispr does when its AI cleanup is off.
    • Next step: write the spec + plan for dictation.

    Measured 2026-09-19 (spike, throwaway; owner's PC, Ryzen 9 9950X3D, CPU ONLY, no GPU pack):

    • Official whisper.cpp Windows builds (release b5130): whisper-bin-x64.zip is 8 MB (sha256 f9ec6c52a2e949b62ab51fa21d0d497958f9e41c3010c157c4e42932d5316f3c), CUDA builds are 260-640 MB, and there is NO official Vulkan build. "One Vulkan build for every GPU" would mean compiling it ourselves.
    • That 8 MB zip contains whisper-cli.exe, whisper-server.exe (resident HTTP server: the model loads once), whisper-stream.exe, VAD tools, and parakeet-cli.exe + parakeet.dll: NVIDIA Parakeet now ships inside whisper.cpp, so one small bundle could serve both families.
    • whisper-cli, base, 11 s clip, cold (model load included): 1.46 s.
    • whisper-server, base, a GROWING clip re-sent each time (what live text needs): 2 s -> 0.87 s, 4 s -> 0.87 s, 6 s -> 0.89 s, 8 s -> 0.90 s, 11 s -> 0.93 s. About 0.9 s per pass regardless of length, so text can update roughly once a second with no GPU at all.
    • Partials are PROVISIONAL: at 6 s it heard "ask not what you are coming from", corrected by 8 s. So live text belongs in the pill, and only the FINAL text is pasted into the shell (text typed into a prompt cannot be cleanly taken back).
    • whisper-cli, large-v3, CPU only: 20.4 s for the 11 s clip. Large and Turbo are GPU-only in practice; Base and Small are the CPU models. This was a very fast desktop CPU; a laptop will be several times slower.
    • OpenWhispr's own packaging matches: CPU engine in the installer, a CUDA pack downloaded by its "Enable GPU" button (%APPDATA%\open-whispr\bin\whisper-cuda), sherpa-onnx bundled for the NVIDIA models.

    What this means for the design: bundle the 8 MB official CPU build (pinned tag + SHA-256, fetched at build time like Prism's ffmpeg); run whisper-server resident while dictation is on (idle stand-down); live text by re-sending the growing clip, shown in the pill; an optional GPU pack download in the model manager (the owner's "Enable GPU" shape); recommend Base or Small on CPU and Turbo once a GPU pack is in. Parakeet via the same bundle is worth a second spike before promising it.

  5. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Sequencing (owner, 2026-09-19): dictation waits until the shared core (#15) exists, so it is built once and lands in both Prism and Prism Terminal. The decisions and measurements above stand and feed the spec then.

  6. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Decided (owner, 2026-09-19), round four, after the shared core landed:

    • GPU: CPU engine bundled in the installer; an 'Enable GPU' button in the model manager downloads the OFFICIAL NVIDIA (CUDA) pack, which unlocks Turbo and Large. AMD/Intel users stay on CPU (Base/Small) in v1. No engine we compile ourselves. Checked the same day: whisper.cpp's latest build release (b5130) still has no any-vendor (Vulkan) Windows x64 build.
    • Model files: ONE shared folder for both apps (%LOCALAPPDATA%\PrismDictation), so a model is downloaded once. Settings VALUES stay per app (active model, hotkey, on/off). Uninstalling one app leaves the models.
    • Language: Auto-detect by default, with a Language picker to pin one. Multilingual models only.
    • Prism: its own 'Dictation' page in the Settings rail under Behaviour (Prism Terminal: its own tab, as decided).
  7. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Spec and plan are written, on branch feat/13-dictation, waiting for the owner's single approval: docs/superpowers/specs/2026-09-19-dictation-design.md, docs/superpowers/plans/2026-09-19-dictation.md.

    Measured 2026-09-19 (GPU spike, owner's RTX 5090): the official whisper-cublas-12.4.0-bin-x64.zip (643 MB, sha256 af520ddd034d985b55dfeea3e465ed93653ba2aee1a55e865033edc548c272a7) RUNS on a 5090 although the card is newer than CUDA 12.4: first run 9.1 s (one-time kernel compile), then Large-v3 through a resident whisper-server is 0.36 to 0.50 s per pass on an 11 s clip. So one pack covers old and new NVIDIA cards, and live text works with the largest model.

  8. Maxaubert commented on Sep 19, 2026

    @Maxaubert
    OwnerAuthor

    Built, PR #22. Spikes (2026-09-19): Chromium's fake microphone works under Electron (peak 0.78 from the WAV), so the e2e is real end to end; Windows' media sessions are reachable from Windows PowerShell 5.1 via WinRT, so pause-media shipped as designed. Found while building: the engine needs Microsoft's C++ runtime, which no official zip carries; it now ships app-local beside the engine (Microsoft-signed copies only). Next: Prism, after asking.

  9. 2 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions