Repository navigation
Push-to-talk dictation: hold a key, talk, release — the interaction none of the voice PRs implement #7417
Replies: 6 comments 1 reply
|
I support this one. I'm currently evaluating T3 Code, and I'm a heavy Wispr Flow user, but I can't use Wispr Flow with T3 Code. It simply does not recognize it. And I'm not going back to typing everything, so this is a showstopper. |
|
+1 for push-to-talk on desktop. One requirement worth building in from the start: transcription in the speaker's language, not English only. I dictate in Russian, and the iPhone app already shows what happens otherwise: non-English speech gets transcribed with the English model and comes out as unrelated English words (#12743 for Korean, discussion #13313, fix in PR #13312). Whatever backend lands on desktop, it would help to pick the language from the OS or a setting and to let the transcription model run multilingual. Until then, macOS's own dictation (double Fn) works in the composer and handles Russian, so a desktop version would mostly need to beat that on speed and on the hold-to-talk flow. |
|
+1 for speech-to-text prompt input across both the desktop and mobile apps. Push-to-talk on desktop, and on mobile a hands-free mode that sends when i stop talking, so i can steer agents while driving (#14394). Paired with read-aloud answers (#6762) that would make T3 fully usable by voice. |
|
Linking my related discussion #15422 where I propose adding a cloud speech-to-text input on mobile. I think the main blocker and the reason why other attempts were blocked is that the subscriptions don't allow voice transcription yet. The way I solved it on my fork, is by adding an option in settings so users can add their own transcription endpoints and key (securely stored on iPhone keychain). This approach may be useful for a desktop voice input too. |
|
I built this: hold a key, speak, and T3 Code carries out short spoken orders (switch thread by number, write and send, switch model or effort, scroll, take and attach a screenshot, chain several in one sentence). It works in English and Spanish without a language setting, and you bring your own Groq or OpenAI key, so it does not depend on a subscription. Details and a fully local option are in #17317. |
|
Hey @juliusmarminge, I'd like to help get push-to-talk working on desktop. Before I build anything: is this something you'd merge, and which direction do you want? I see #14882 and #14865 are both open, and I don't want to add a third competing approach. If one of those is the base you want, I'm happy to help land it (tests, fixes, whatever's blocking) and build hold-to-talk on top. I use T3 Code with voice daily, so I'll stick around to maintain it. |
Uh oh!
There was an error while loading. Please reload this page.
This is not another "add voice support" request — it is about one specific interaction model that none of the existing voice work implements, and that I think is the reason the earlier attempts keep stalling.
What I mean
Hold a key. Talk. Release. The transcript lands in the composer. You never leave the composer, never click anything, never review a panel.
That is how Wispr Flow, superwhisper, macOS dictation and the Claude Code desktop app do it, and it is the thing that makes voice actually usable for short bursts — "run the tests again", "revert that last edit", "check the second file too".
Why this is a distinct ask
I went through the existing work, and every implementation is a modal panel, not a key:
CONFLICTINGCONFLICTINGCONFLICTINGCONFLICTINGIn #6625 — the most current one — the only keyboard handling in the entire diff is
window.addEventListener("keydown", cancelOnEscape). Recording starts fromonClick={() => void startVoiceTranscription()}and ends fromonClick={onStop}/onClick={onSend}. There is no held-key path anywhere in any of these branches.A panel is a destination: open it, record, watch a waveform, stop, review, insert, close. Push-to-talk is modeless: the held key is the recording state, so there is no panel to open, no timer to watch, no cancel affordance to design (releasing early with no speech is the cancel), and no lifecycle to clean up on navigation. It is a meaningfully smaller surface than what has been attempted four times.
On the earlier decision
#653 was closed
NOT_PLANNEDin March with "out of scope, there are apps that do this already." I would gently push back on that specific reasoning: people reach for those third-party apps because the in-app options are panels. Nobody installs Wispr Flow to get a waveform — they install it to get a key. The panel-shaped version is the one that loses to external tools; the key-shaped version is the one that does not.Also worth noting
@akhmerov's unanswered follow-up on that thread: codex app-server has supported voice input since v0.147.Smallest useful scope
Known blocker
On signed macOS builds this cannot work until #5321 lands — the desktop build has no
com.apple.security.device.audio-inputentitlement, so the OS never even shows a microphone prompt. That PR is+3lines andMERGEABLE, with an open question about also settingentitlementsInheritfor Electron helper processes (which is exactly the renderer path any dictation feature needs). See also #728 and #7268.All reactions