Blog / Which GPU do you need for local dictation? Probably the one you have

Which GPU do you need for local dictation? Probably the one you have

"Runs on your GPU" sounds like it needs a gaming PC. It does not. We timed our own dictation engine on the two graphics chips in one ordinary 2025 laptop: a discrete NVIDIA RTX 5070 Ti laptop GPU and the AMD Radeon 890M integrated into its processor, the kind of graphics that ships inside the CPU of a mid-range laptop. Both turn a typical spoken paragraph into text faster than you can move your hand from the key to the mouse, and for long dictation the app hides most of the remaining wait by working while you talk. This page has the numbers, the memory each model needs, and the one kind of machine where local dictation does not work.

How we measured

The audio is a 77-second clip of synthesized English speech about a product meeting, cut into the lengths people actually dictate: 5 seconds for a chat reply, 15 for a paragraph or a prompt to an AI agent, 60 for a long email. Each clip went through the exact engine AlphaDictate ships, Whisper running through whisper.cpp on Vulkan, three times per point; the tables show the median. The timer covers transcription only, from the moment you stop talking to the moment the text exists, and leaves out the fraction of a second the paste takes. Every run returned the full transcript, so the fast numbers are not the engine skipping audio.

The three models are the app's three tiers: Fast is Whisper Small, Balanced is Whisper Medium, and Best is Whisper Large v3, all compressed to 5-bit weights.

A spoken paragraph: 15 seconds of speech

GraphicsFastBalancedBest
RTX 5070 Ti laptop GPU, discrete0.30 s0.55 s0.77 s
Radeon 890M, integrated0.64 s1.2 s2.0 s

On the discrete card every model finishes in under a second, so there is no reason to pick anything but Best. On integrated graphics the Best model takes about two seconds, which you notice but do not mind for an email; the Fast model is under a second and still handles clear speech well.

A long take, and why you rarely wait for all of it

Transcribed in one piece, a minute of speech is where the two chips part ways:

60 seconds, all at onceFastBalancedBest
RTX 5070 Ti laptop GPU, discrete1.6 s2.3 s3.0 s
Radeon 890M, integrated2.7 s5.4 s9.3 s

Nine seconds after a minute of talking would be irritating ten times a day. It rarely comes to that, because AlphaDictate does not wait for you to finish. Once it has eight seconds of audio and you pause, five seconds by default, it hands what you have said so far to the GPU and keeps listening. When you press stop, only the part since your last pause is left to do. Uninterrupted speech is never split, so a word is never cut in half.

To measure that, we built a take the way people actually dictate: about a minute of speech in four stretches of 13 to 18 seconds, with a five-second pause to think between them, 74 seconds in all. We replayed it through the app's own splitting logic at its real timing and timed the wait after stop:

A minute of dictation with pausesFast, all at onceFast, split at pausesBest, all at onceBest, split at pauses
RTX 5070 Ti laptop GPU, discrete1.4 s0.24 s2.8 s0.53 s
Radeon 890M, integrated3.1 s0.57 s11.0 s2.1 s

Splitting cuts the wait about fivefold on both chips. With the most accurate model, integrated graphics answer a minute of dictation in about two seconds, the same as a single paragraph, because by the time you stop only your last paragraph is left.

Our own use of the app on this laptop tells the same story. Over the last week of dictating with it, about half of the takes longer than 20 seconds had a pause long enough to split on. For those, a median of 13 seconds of speech was left to transcribe after stop, and the median wait was 0.8 seconds. Long takes spoken without such a pause waited a median of 1.7 seconds. The pause length is a setting, Pause that ends a part under Accuracy & speed, and can go as low as two seconds if you want results sooner and speak without hesitating.

A recording: a 10-minute file

GraphicsFastBest
RTX 5070 Ti laptop GPU, discrete16 s28 s
Radeon 890M, integrated36 s1 min 54 s

Files get no head start, since the whole recording is there before transcription begins, so this is where a discrete GPU earns its keep. A one-hour meeting recording on the discrete card with the best model takes about three minutes; on integrated graphics, about eleven. Both beat sending an hour of a confidential meeting to a cloud service and waiting for it to come back.

How much memory each model needs

ModelGPU memory used
Fastabout 0.6 GB
Balancedabout 1.2 GB
Bestabout 2.1 GB

We read these from Windows' own per-process GPU counters while each model was loaded, and they were the same on both chips. Two practical consequences. Any discrete card with 4 GB of memory holds even the largest model with room to spare, so the card's age matters more than its size. And integrated graphics have no memory of their own: they borrow system RAM, so on a laptop with 8 GB the Best model takes a quarter of it. On such a machine Balanced is the considerate choice; with 16 GB or more it does not matter.

Why there is no CPU version

The one machine local dictation does not serve well is a machine with no usable GPU at all: a remote desktop session, a virtual machine, a server. We built a version that ran on the processor alone and measured it at five to twenty times slower than integrated graphics, which turns a two-second wait into half a minute. We removed it the same day. A dictation tool that quietly takes thirty seconds is worse than one that tells you plainly it needs a graphics chip, so AlphaDictate checks for a GPU with a Vulkan driver at first run and says so if it cannot find one.

In practice nearly every Windows PC sold in the last several years has one, integrated or discrete. What usually stands in the way is an old graphics driver, and updating it from the manufacturer's site fixes that.

What to pick

  • Gaming PC or laptop with any NVIDIA or AMD discrete card: Best. Everything is under a second, and you get the most accurate transcripts.
  • Recent laptop with integrated graphics, 16 GB of RAM: Best. Short passages take about two seconds and long ones barely more, thanks to splitting at pauses. Balanced if you transcribe many long recordings.
  • Laptop with 8 GB of RAM: Balanced, or Fast if the machine is older.
  • Older or low-power integrated graphics: start with Fast. We have not measured older chips yet, so treat the integrated numbers above as the upper end of what integrated graphics does.
  • A laptop with both chips: you do not have to choose. AlphaDictate uses the discrete GPU when it is available and the integrated one when the discrete chip is switched off to save power, and it moves between them in the pause between takes.

AlphaDictate's setup screen names the graphics it found and suggests a starting model: Best on machines with NVIDIA graphics, Balanced elsewhere. The model is one click to change later, and all three are included in the price, $39 a year after a free hour. You can check what your own PC has in Task Manager under Performance, where each GPU is listed with its memory.

If you have a GPU we have not measured, especially an older Intel or AMD integrated chip, we would like your numbers at hello@alphadictate.com. Our note on dictating to Claude Code uses the same measurements from a developer's point of view.


Frequently asked

Do I need a gaming GPU for local speech recognition? No. On the integrated graphics of a 2025 laptop, a 15-second spoken paragraph transcribes in about two seconds with Whisper Large v3 and in under a second with the smallest model.

Do I have to wait for a long dictation to be transcribed from the start? No. AlphaDictate transcribes what you have said at each pause while you keep talking, so after a minute of dictation with pauses the wait was about two seconds on integrated graphics and half a second on a discrete GPU with the most accurate model.

How much VRAM does Whisper Large v3 need? About 2.1 GB in the compressed form AlphaDictate uses. Whisper Medium needs about 1.2 GB and Whisper Small about 0.6 GB.

Can Whisper run on integrated graphics? Yes, through a Vulkan driver. Integrated graphics use system memory, so the model's memory comes out of your RAM.

Does AlphaDictate work without a GPU? No. It needs a graphics chip with a Vulkan driver, integrated or discrete. A processor-only version was five to twenty times slower than integrated graphics, so we do not ship one.

How long does it take to transcribe an hour of audio locally? In our measurement, about three minutes on a laptop RTX 5070 Ti and about eleven minutes on Radeon 890M integrated graphics with the most accurate model.