Blog / Dictating to Claude Code

Dictating to Claude Code: what /voice sends, and a local way

Talking to a coding agent is the first kind of dictation many developers have ever wanted. Prompts are long, conversational and full of words like "refactor the auth middleware", and typing them is slower than saying them. Claude Code now ships a voice mode, and it is good. This article explains exactly what it does with your audio, where it stops, and gives measured numbers for the alternative: a dictation key that works in every terminal, chat box and editor on the machine, with the speech model running on your own GPU.

What Claude Code's voice mode does

You enable it with /voice. In the default hold mode you hold the space bar while you speak; in tap mode you tap once to start and once to send. The transcript appears in the prompt as you talk, dimmed until it is final, and you can mix typing and speech in the same message. Recording stops on its own after fifteen seconds of silence or two minutes total. Transcription is tuned for code vocabulary, and Claude Code quietly adds your project name and git branch as recognition hints, which is a clever touch. It costs nothing against your usage limits.

Now the part Anthropic's own documentation states in its first requirement, verbatim: "Voice dictation streams your recorded audio to Anthropic's servers for transcription. Audio is not processed locally." Three consequences follow from that sentence and are spelled out on the same page.

  • It needs a claude.ai account. The speech service is not available when Claude Code runs on an API key, Amazon Bedrock, Google Cloud's Agent Platform or Microsoft Foundry. Enterprise setups that chose those routes for data-residency reasons get no voice mode, which is consistent, if inconvenient.
  • It needs a local microphone. No voice in SSH sessions, in Claude Code on the web, in VS Code Remote, Dev Containers or Codespaces, and not in WSL1.
  • It speaks twenty languages, one at a time. Dictation follows the language setting and defaults to English. Set it to Polish and you dictate in Polish; speak Polish while it is set to English and the transcript is garbage, as the troubleshooting section notes. There is no automatic detection.

None of that is a criticism of the engineering. It is a description of an architecture: a very good cloud transcription service, embedded in one application, available to one kind of account.

Where it stops for a working developer

The account and network conditions matter to some people and not others. Two limits bite everyone.

It only works inside Claude Code's prompt. The moment you switch to a Codex session, a Cursor chat, a GitHub issue in the browser, a commit message in your editor or a Slack thread, voice is gone. Developers who talk to agents rarely talk to only one of them.

The audio leaves the machine. For a personal side project this is nobody's concern. For a client's codebase under NDA, a regulated employer, or simply a developer who chose local models for a reason, it is a line. Some organizations have already turned the feature off by policy, which Claude Code reports as "Voice mode is disabled by your organization's policy".

The local alternative, measured

A system-wide dictation key sidesteps both limits: press it in any window, talk, press it again, and the text lands where the cursor was, transcribed by a Whisper model on your own graphics card. The obvious question is whether local transcription is fast enough to feel like a conversation. We measured our own engine, the one inside AlphaDictate, on the laptop this article was written on, which has a discrete NVIDIA RTX 5070 Ti laptop GPU and an integrated AMD Radeon 890M. The clip is 77 seconds of synthesized English speech about a product meeting, sliced to typical take lengths; each number is the median of three runs through the shipped engine, and excludes the fraction of a second the paste itself takes.

Take lengthFast model, discreteBest model, discreteFast model, integratedBest model, integrated
5 seconds0.13 s0.29 s0.31 s1.2 s
15 seconds0.30 s0.77 s0.64 s2.0 s
60 seconds1.6 s3.0 s2.7 s9.3 s

Loading a model into the GPU happens once, and took between one and two and a half seconds depending on size. A typical prompt to an agent is ten to twenty seconds of speech, so on a discrete GPU the text appears under a second after you stop talking with the most accurate model, and on an integrated GPU in about two seconds, or well under a second if you pick the Fast model. Cloud dictation is not meaningfully quicker than that once the round trip is counted, and it does not get slower on a bad connection.

Longer takes rarely wait even that long: AlphaDictate transcribes what you have said at each pause while you keep talking, so after a minute of dictation with pauses only the last stretch is left, about half a second on the discrete GPU and two seconds on the integrated one with the most accurate model. Our GPU article has those numbers.

Two honest caveats. Synthesized speech is cleaner than a human mumbling at a laptop microphone, so accuracy is not what this test measures; latency is. And the integrated-GPU column is a laptop from 2025; older integrated graphics will be slower, which is why the app offers three model sizes and tells you at first run which one your hardware supports.

The terminal problem nobody mentions

Every dictation tool on Windows, ours included, gets text into a terminal the same way: it puts the transcript on the clipboard and sends a paste. Windows Terminal, the VS Code terminal and the console Claude Code runs in all accept Ctrl+V, so this works. Two things break it, and they break it for every tool.

The first is an elevated terminal. If you run your shell as administrator, Windows silently discards keystrokes injected by a normal-privilege program, and the paste appears to do nothing, with no error anywhere. AlphaDictate checks for this before pasting and tells you the text is on the clipboard so you can press Ctrl+V yourself, rather than losing the take. The second is the clipboard itself: the paste replaces whatever you had copied. We chose to leave the transcript on the clipboard deliberately, so that a paste that failed for any reason is one Ctrl+V away, and every take is also kept in the app's local history. If you dictate in the middle of copying something precious, copy it again afterwards. That is the trade, and we would rather say it here than have you discover it.

Setting up a local dictation key for agent work

  1. Install a local dictation app and let it download a model. Pick the Fast model if you are on integrated graphics and want sub-second results; pick Best on a discrete GPU.
  2. Bind it to something you can hit without looking. A side mouse button is ideal for agent work: hand on the mouse, press, talk, press, the prompt fills.
  3. Leave the language on automatic detection if you think in more than one language. Whisper will follow you mid-sentence.
  4. Add your project's odd words to the vocabulary hint once: internal service names, the library you keep saying, your co-founder's surname.
  5. Keep Claude Code's own voice mode enabled too if you like its live-preview typing. The two do not conflict; one is a key inside one app, the other is a key everywhere.

We make AlphaDictate, so read the recommendation with that in mind, and note that the free alternative in the same category is Handy, an open-source local dictation app with fewer features and no cost. Our earlier note on Windows voice typing covers the built-in option, which also sends audio to a server and also stops at one language.


Frequently asked

Does Claude Code voice mode process audio locally? No. Anthropic's documentation states that audio is streamed to Anthropic's servers for transcription and is not processed locally.

Why does /voice say it requires a Claude.ai account? The speech-to-text service is only offered to claude.ai logins. Claude Code sessions authenticated with an API key, Bedrock, Google Cloud's Agent Platform or Microsoft Foundry cannot use it.

Can I dictate to Claude Code over SSH? Not with the built-in voice mode, which needs a local microphone. A system-wide dictation app running on your local machine types into the SSH terminal like any other window.

Which languages does Claude Code voice support? Twenty, including English, German, French, Spanish, Portuguese, Polish, Japanese, Korean and Ukrainian, chosen by the language setting rather than detected.

How fast is local dictation on a laptop? On the laptop we measured, a fifteen-second take transcribed in under a second on the discrete GPU with the most accurate model, and in about two seconds on the integrated GPU, or under a second with the Fast model.

Does dictation work in an administrator terminal? Injected pastes are ignored by elevated windows on Windows, for every tool. AlphaDictate detects this and leaves the text on the clipboard for a manual Ctrl+V.