Which small local LLM cleans up dictation best on Apple Silicon?
We benchmarked four Gemma 4 checkpoints on an M4 Max for dictation cleanup: real latency, memory, and the prompt bug hiding underneath.
Saydrop is a push-to-talk dictation app for Apple Silicon. You double-tap a hotkey, talk, and the text lands in whatever app has focus. Transcription is mlx-whisper (Whisper medium by default), running entirely on your Mac — we wrote about that part here.
Optionally, a second stage runs the raw transcript through a small local LLM to clean it up. That stage is the subject of this post: which model should it be, and how much does a bigger one buy you?
We ran the sweep in August 2026 expecting a legible trade-off curve. We did not get one. Here is the data, including the parts that make us look bad.
Why a cleanup step exists at all
Whisper transcribes what you said, which is not what you meant to write. Real dictation output looks like this (an excerpt from our long fixture):
…i think the allocation for marketing is maybe a bit low we should probably bump it to like fifty thousand instead of thirty five and then for operations i think we can actually cut some of the training budget…
No punctuation, no capitalisation, self-corrections mid-sentence. Pasting that into Slack is fine; pasting it into a commit message or a customer email is not.
So a small instruction-tuned model gets one job: fix grammar, spelling and punctuation, drop the “um”s and “ähm”s, change nothing else. The shipped system prompt says it in exactly those words, including “Keep the wording, tone, and meaning exactly as they are.” Remember that sentence — it turns out to be the whole story.
The constraint: this stage sits on the critical path — the user is staring at a status pill waiting for text — and it shares unified memory with a resident Whisper model on a laptop that is also running Xcode.
What we tested
Four 4-bit Gemma 4 checkpoints, all through mlx-lm:
mlx-community/gemma-4-e2b-it-4bit— the shipped defaultmlx-community/gemma-4-e4b-it-4bit— the obvious next size upmlx-community/gemma-4-e4b-it-OptiQ-4bit— an “optimised” quantisation of E4Bmlx-community/gemma-4-E4B-it-qat-4bit— quantisation-aware training variant
Two third-party conversions (FakeRockert543/gemma-4-e2b-it-MLX-4bit and the E4B equivalent) were in the plan and never produced a number, because neither one loads. More on that below.
The 12B and 31B tiers were dropped deliberately rather than measured — flagging that up front: this benchmark has no data above E4B.
Methodology
Hardware: Apple M4 Max, 48 GB unified memory. One number, one machine — no cross-hardware claims.
Harness: app_osx/bin/bench_cleanup.py, driving MLXGemmaBackend directly. The model is loaded and warmed once per process, so load time is paid outside the measured window; one OS process per model, so peak memory is attributable to that model alone. Median of 3 runs.
Fixtures: three synthetic raw-ASR transcripts in app_osx/bin/fixtures/cleanup/ — short (25 words, 1 checkable fact), medium (187 words, 8 facts), long (640 words, 32 facts), carrying fillers, hedges, self-corrections and run-ons in the shape real dictation produces. Synthetic on purpose: no user speech anywhere in this benchmark.
Scoring is deliberately dumb and positive-only: are the expected facts still present, are the forbidden strings (um, uh) gone, what fraction of words survived, what is the length ratio. Fact matching is word-boundary anchored and hyphen-insensitive. Peak memory comes from mlx.core.get_peak_memory() — MLX’s own allocator high-water mark, weights plus KV cache, not OS RSS.
One parity caveat to carry through every number below: production uses SubprocessMLXGemmaBackend, which runs the model in an isolated child process, while this harness measures the in-process backend. That is the right instrument for comparing models — it isolates model cost from IPC noise — but it under-reports what a user feels, which also pays subprocess spawn, model load in the child, and a round trip. Do not read these as end-to-end cleanup latency.
Results
Median of 3 runs, full-text mode, uncontended machine.
| model | short (25 w) | medium (187 w) | long (640 w) | peak mem | facts | fillers | word retention |
|---|---|---|---|---|---|---|---|
gemma-4-e2b-it-4bit (shipped) |
0.31 s | 1.27 s | 4.00 s | 3.33 GB | 32/32 | 0 | 93.9% |
gemma-4-e4b-it-4bit |
0.45 s | 2.10 s | 11.10 s ¹ | 4.96 GB | 32/32 | 0 | 99.0% |
gemma-4-e4b-it-OptiQ-4bit |
0.52 s | 2.52 s | 8.23 s | 7.28 GB | 32/32 | 0 | 98.6% |
gemma-4-E4B-it-qat-4bit |
0.57 s | 2.79 s | 9.27 s | 6.52 GB | 32/32 | 0 | 98.6% |
¹ That cell drifted 8.87 → 11.10 → 12.53 s across three consecutive runs. The long fixture carries thermal noise the short one does not. Treat differences under ~20% in the long column as noise, not signal.
Two things fall out immediately.
The quality axis is saturated. Every model, including the smallest, scores 32/32 on fact retention and zero filler violations on the hardest fixture. A metric on which everything ties does no ranking work — that is evidence our proxies are too easy, not that the models are equal.
The latency axis is not saturated. E2B is roughly 2× faster than every E4B variant at every length, on half to a third of the memory. And note that “optimised” quantisation bought nothing: OptiQ is both slower than stock E4B on short and medium input and uses 47% more memory.
Chunking does not help
We tested splitting the transcript into sentences and cleaning them sequentially (--mode chunked), on the theory that shorter prompts return sooner. On clean measurements it is indistinguishable from full-text: E2B medium is 1.27 s full vs 1.29 s chunked; OptiQ long is 8.23 s in both modes, to three decimal places.
The reason is boring once you see it: cleanup latency is dominated by tokens generated, and chunking does not reduce the token count — it only changes how the same tokens are batched. Chunking is a real win against a parallel cloud API. Against local sequential inference it is a no-op at best.
Two checkpoints that do not load
Both third-party conversions fail identically at load:
Received 140 parameters not in model:
language_model.model.layers.15.self_attn.k_norm.weight, ...k_proj..., ...v_proj...
They were converted against a different KV-sharing layout than our pinned mlx-lm expects — with no warning until load time. An availability result, not a quality one, but a concrete argument for preferring mlx-community checkpoints, which track the runtime you actually ship.
The real finding: the small model is the misbehaving one
The proxies say every model is equivalent. Reading the actual output says otherwise — in the opposite direction from what the headline implies. Given that raw fixture above, E2B produces:
…I think the allocation for marketing is a bit low. We should probably bump it to fifty thousand instead of thirty-five. For operations, I think we can cut some of the training budget…
E4B, OptiQ and QAT all produce:
…I think the allocation for marketing is maybe a bit low. We should probably bump it to like fifty thousand instead of thirty-five. And then for operations, I think we can actually cut some of the training budget…
The prompt said keep the wording exactly as it is and remove filler words (for example ‘um’, ‘uh’). By that specification the E4B family is correct and E2B is over-editing. E2B is deleting hedges — maybe, like, actually, and then — that it was explicitly told to preserve.
This also means the word-retention column reads backwards: E2B’s lower 93.9% is hedge deletion, not better filler removal. Do not cite it as a quality score without this paragraph attached.
So the axis that actually separates these models is instruction adherence, not capability — and adherence is far cheaper to buy with prompt engineering than with 2× latency and 2× memory.
What we picked, and why
We kept mlx-community/gemma-4-e2b-it-4bit. Fastest at every length, smallest footprint, and its one demonstrated weakness is a prompt-adherence gap. Paying double latency and double memory to fix a prompt bug is a bad trade for a menu-bar app that must also keep Whisper resident.
There is an open product question underneath that the benchmark cannot answer: is E2B’s punchier output a defect, or the thing we actually want? If a defect, tighten the prompt against hedge deletion and re-measure. If not, the prompt is wrong and should sanction light rewording. Either way, the model does not change.
Limitations, stated plainly
- One machine. M4 Max, 48 GB. No data for M1/M2 or 8–16 GB configurations from this sweep.
- No model above E4B. 12B and 31B were dropped: the measurable quality axis is already saturated, and the 640-word fixture costs 4 s on the smallest model. A reasoned exclusion, not a tested one — we would rather say so than imply coverage we do not have.
- Synthetic fixtures, three of them. No real-user speech, no accent variety, no multi-speaker audio.
- Positive-only scoring. The proxies confirm content survived and listed fillers left. They cannot express “did not invent anything” beyond a forbidden-string list. Reading the dumped output is still mandatory — here, reading is what produced the actual finding.
- In-process, not end-to-end. See the parity caveat above.
- Cloud comparison is out of scope. Saydrop’s optional Gemini backend is not in this table.
Three measurement traps we walked into
Each produced a confident, wrong reading before it was caught. Together they cost more time than the benchmark itself.
1. ru_maxrss is unusable for this. The harness first reported peak memory via resource.getrusage, which claimed 86.42 GB for a model on a 48 GB machine — resident set size cannot exceed physical RAM. macOS’s high-water mark counts pages later compressed or reclaimed, overstating by up to 12×. All memory numbers above come from mlx.core.get_peak_memory() instead.
2. Substring matching poisons a filler check. "um" in output matches document, documentation, jump and numbers — all present in the long fixture. Every run reported a filler violation regardless of what the model did. Matching is now word-boundary anchored.
3. Fact notation has to match the fixture. Expected facts were first written in digits (september 15, 22) while raw-ASR fixtures spell them out (the fifteenth of september, twenty two percent). This scored a perfectly faithful model at 6/32 — an apparent 80% content loss that was pure notation mismatch.
And the one that nearly shipped as a finding: the first sweep ran while a second copy of the benchmark was still running. Contention inflated results by up to 2.6× (E2B medium: 3.36 s contended vs 1.27 s clean) and manufactured two fictitious results — a 1.9× chunking speedup, and OptiQ being 3.6× faster than stock E4B. Both evaporated on re-measurement. Every row above was re-run uncontended.
Reproducing this
The harness and fixtures ship in the repo:
cd app_osx
uv run python bin/bench_cleanup.py \
--model mlx-community/gemma-4-e2b-it-4bit \
--fixture bin/fixtures/cleanup/short.txt \
--fixture bin/fixtures/cleanup/medium.txt \
--fixture bin/fixtures/cleanup/long.txt \
--mode full --mode chunked --runs 3 --json /tmp/cleanup-bench.json
Run one model per process, and make sure nothing else is touching the GPU. If your long-fixture numbers differ from ours by more than 20%, check contention and thermals before concluding anything.
Saydrop is a one-time CHF 39 licence for Apple Silicon Macs — transcription on-device via mlx-whisper, optional local Gemma cleanup in an isolated subprocess, optional Gemini cleanup only if you explicitly turn it on. Download the 14-day trial, or read the MLX Whisper deep dive for how the transcription half works.