Whistle ships a 16.9 MB on-device speech recognition model
What happened
On October 9, Hacker News' front page covered the release of Whistle, an on-device speech recognition model. It runs on CPU as a single 16.9 MB file, supports seven languages, and shares a C++ engine with Needle. It transcribes up to 30 seconds of audio at a time and provides word-level timestamps and speech embeddings. The report also describes how the two work together: loaded jointly, one program can turn audio directly into tool calls. The coverage lays out the model's runtime requirements, language support, transcription ability and its integration with Needle.
Written by AI from the coverage · updated 2 hours ago
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageWhistle: Speech to Text in 16.9 MB
Whistle 发布端侧语音识别模型,以单个 16.9 MB 文件在 CPU 上运行,支持七种语言并与 Needle 共用 C++ 引擎。模型可一次转写最长 30 秒的音频,提供词级时间戳和语音嵌入向量;与 Needle 联合加载时,同一程序可直接将音频转为工具调用。
Heat over time
Not enough continuous observations to draw a trend yet.