Theme 1: Accuracy and accent/language limitations
Users report that the model works well for clear, western accents but struggles with non‑standard speech, regional accents, and languages other than English.
- “Spanish is not good (seems to write non existing words and/or with terrible typos…) but English seem to work good even with my (Spanish) accent…” — tecleandor
- “the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke” — INTPenis
- “data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.” — yu3zhou4
Theme 2: LLM‑based post‑processing and sharing concerns
Many commenters pair the STT output with a lightweight LLM to clean up transcripts (removing filler words, correcting errors), while expressing uncertainty about the ethics of sharing AI‑edited text.
- “And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words.” — cgbur
- “+1 for handy and then using LLM's for the cleanup pass… I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated” — Imustaskforhelp
- “I prompt it to: ‘Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings.’” — flockonus
Theme 3: Preference for local/on‑device, private, small‑footprint STT
The discussion highlights appreciation for models that run entirely on the CPU, need no GPU or cloud, and keep data private—often citing tiny size and offline capability as key advantages.
- “love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud)…” — mrkn1
- “FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model.” — rpdillon
- “For speech to text UX I currently use https://handy.computer/ as it's cross platform and open source… I am currently using Parakeet Unified EN 0.6B and finding it excellent.” — contingencies