Project ideas from Hacker News discussions.

Whistle: Speech to Text in 16.9 MB

📝 Discussion Summary (Click to expand)

Theme 1: Accuracy and accent/language limitations
Users report that the model works well for clear, western accents but struggles with non‑standard speech, regional accents, and languages other than English.
- “Spanish is not good (seems to write non existing words and/or with terrible typos…) but English seem to work good even with my (Spanish) accent…” — tecleandor
- “the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke” — INTPenis
- “data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.” — yu3zhou4

Theme 2: LLM‑based post‑processing and sharing concerns
Many commenters pair the STT output with a lightweight LLM to clean up transcripts (removing filler words, correcting errors), while expressing uncertainty about the ethics of sharing AI‑edited text.
- “And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words.” — cgbur
- “+1 for handy and then using LLM's for the cleanup pass… I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated” — Imustaskforhelp
- “I prompt it to: ‘Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings.’” — flockonus

Theme 3: Preference for local/on‑device, private, small‑footprint STT
The discussion highlights appreciation for models that run entirely on the CPU, need no GPU or cloud, and keep data private—often citing tiny size and offline capability as key advantages.
- “love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud)…” — mrkn1
- “FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model.” — rpdillon
- “For speech to text UX I currently use https://handy.computer/ as it's cross platform and open source… I am currently using Parakeet Unified EN 0.6B and finding it excellent.” — contingencies


🚀 Project Ideas

Handy++: Open‑Source Cross‑Platform Dictation App

Summary

  • Cross‑platform dictation tool with push‑to‑talk input and optional clipboard‑mode paste.
  • Uses lightweight local STT models (Whistle/Parakeet) plus speaker‑adaptation for accents.
  • Provides raw transcript and optional LLM‑based cleanup (run locally) with a toggle to share original vs cleaned version.
  • Designed for privacy‑first, hands‑free writing and accessibility.

Details

Key Value
Target Audience Professionals, writers, RSI sufferers, and users with speech impairments who need hands‑free typing.
Core Feature Push‑to‑talk dictation with local STT, accent adaptation, and optional LLM cleanup.
Tech Stack Electron/Tauri UI, Whisper.cpp or Whistle (ONNX) model, Llama.cpp for local LLM cleanup, SQLite for user profiles.
Difficulty Medium
Monetization Hobby

Notes

  • HN users praised Handy’s push‑to‑talk and clipboard mode but asked for better accent handling and local processing (tecleandor, saturn8601, cgbur).
  • Accent adaptation directly addresses Mexican/Venezuelan Spanish and heavy‑accent issues raised by chilicuil and INTPenis.
  • Offering a raw‑transcript toggle lets users decide whether to share AI‑cleaned text, responding to ethics concerns from Imustaskforhelp and asa123.
  • Could spark discussion on open‑source dictation pipelines, speaker‑adaptation techniques, and privacy‑preserving AI assistance.

ScribeGuard: Personalized, On‑Device Spell Checker

Summary

  • System‑level spell checker that learns your personal vocabulary and writing style to reduce false positives.
  • Works as a text‑input extension across macOS, Windows, Linux, and mobile platforms.
  • Uses a small on‑device transformer fine‑tuned on your own text (via local federated learning) to suggest only truly needed corrections.
  • Includes a review UI and exportable personal dictionary.

Details

Key Value
Target Audience Writers, programmers, multilingual users annoyed by aggressive, generic spell checkers (e.g., Apple, Google).
Core Feature Adaptive, privacy‑first spell correction that minimizes annoying false corrections.
Tech Stack Rust core with platform text‑service bindings (NSTextInput, Text Services Framework), DistilBERT‑base transformer fine‑tuned locally via ONNX Runtime, encrypted local dictionary storage.
Difficulty High
Monetization Revenue-ready: Freemium (free base, premium for cloud sync of personal dictionary across devices).

Notes

  • Andy_ppp’s complaint about spell checkers “annoy[ing] the crap out of everyone” matches the need for a less intrusive checker.
  • Personal adaptation helps with technical jargon, code, and mixed‑language inputs noted by tecleandor and chilicuil.
  • Fully local processing addresses privacy worries about sending text to servers.
  • Could generate discussion on balancing prescriptive correctness with user intent in language tools.

AccentKit: Developer SDK for Accent‑Adaptive Streaming STT

Summary

  • SDK that lets developers embed accent‑personalized, streaming speech‑to‑text into any app.
  • End‑users upload a short voice sample (~30 s) to adapt the model to their accent/dialect.
  • Provides low‑latency streaming transcription with filler‑word removal, punctuation insertion, and optional LLM post‑processing.
  • Returns versioned transcripts (raw vs cleaned) for transparency and ethical sharing.

Details

Key Value
Target Audience App developers building voice‑driven tools (note‑taking, accessibility, dictation, voice control).
Core Feature Accent‑personalized streaming STT with filler suppression and optional LLM cleanup.
Tech Stack C++ core using Whisper.cpp backend, lightweight speaker‑adaptation adapter layers, WebAssembly/Wasm web port, gRPC API for edge deployment, optional Llama.cpp for cleanup.
Difficulty Medium‑High
Monetization Revenue-ready: Pay‑per‑use API (free tier ≤10k utterances/month, then $0.0005 per utterance).

Notes

  • Commenters reported accent‑related failures (tecleandor, chilicuil, saturn8601) and desire for models that ignore umms/ahhs (ComputerGuru).
  • Streaming and low latency were highlighted as crucial by nicksaroha, solarkraft, and wkcheng.
  • Providing raw vs cleaned transcript options addresses ethics worries raised by Imustaskforhelp and asa123.
  • Could stimulate discussion on speaker‑adaptation techniques, privacy‑preserving personalization, and real‑time STT trade‑offs.

Read Later