I fine-tuned a 0.8B local model for dictation cleanup. It matched a hosted frontier model on this narrow task
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I built SpeakoFlow Mini, an Apache-2.0 fine-tune of Qwen3.5-0.8B for dictation cleanup. It is not a chat model or a general rewriter. It takes speech-to-text output, applies corrections the speaker actually made, and leaves everything else alone. That last part is harder than it sounds. General models often polish text that was already correct, but in dictation an unnecessary improvement is still a wrong edit. For example: On my specialized English-only benchmark, SpeakoFlow Mini scored 70.7% and GPT-5.6 Luna scored 65.0%. Both used the same fixed short prompt, with reasoning disabled. The gap was +5.8 points, but the 95% interval was [-1.5, +12.9], so this is a statistical tie, not a win. That setup matters. It is the low-overhead configuration this model was trained for. With a longer prompt and a reasoning budget, Luna does better. It is also the better model for unusual cases outside the narrow patterns covered here. The claim is narrower. Under one fixed short prompt with no reasoning budget, a 0.8B model running locally matched a hosted frontier model on this task. The controlled comparison is against the untuned Qwen3.5-0.8B base. Fine-tuning moved the score from 47.3% to 70.7%, a gain of +23.4 points with a 95% interval of [+16.3, +30.3]. The second image shows that comparison. The Q8_0 build is 833 MB and runs fully offline. The Hugging Face card has the model files, run commands, benchmark method, quantization results, limitations and public examples: https://huggingface.co/SpeakoFlow/speakoflow-mini The evaluation set is not part of the public release. If you try it, I am most interested in cases where it changes something that should have been left alone. Disclosure: I built SpeakoFlow Mini and SpeakoFlow. [link] [comments] |
More from r/LocalLLaMA
-
Demo of local document extraction (52 pages) using Arctic Embed and Bonsai 8B on an Iphone 16 (KernelAI app)
Aug 30
-
Will apple still release devices with mobile HbM in 2027 ?
Aug 30
-
Whatever happened to OpenClaw and its derivatives?
Aug 30
-
Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.