Today we release Whistle, a speech recognition model for mobiles, wearables, robots, smart home, automotive and microcontrollers. It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle, from the same container and the same quantisation.

Whistle does three jobs, all of them on the device:

  • Transcription. 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish. The language is detected unless you name it.
  • Word timestamps. Every word with its start, end and probability, aligned from the decoder's attention.
  • Speech embedding. The encoder output, one row per 80 ms frame, without decoding a transcript.