Cactus Compute has released Whistle, a speech recognition model that fits in a single 16.9-megabyte file, runs on a CPU with no dependencies, and covers seven languages: English, German, French, Spanish, Italian, Dutch and Polish. It was published on October 2 and is climbing Hacker News now.

The headline numbers come from the company's own benchmarks, run on an Apple M4 Pro CPU. At ten seconds of audio, Whistle reaches its first output token in 11.1 milliseconds; Whisper base takes 73.2 ms. It decodes roughly 1,319 tokens per second against Whisper base's 266, from a file 8.6 times smaller — 145.3 MB. On word error rate, the company says Whistle leads on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average, while Whisper base stays ahead on TED-LIUM, AMI and the MLS average. Where a model has never published numbers for a benchmark, the comparison is left blank rather than filled in.

Three capabilities matter for devices. Whistle transcribes up to 30 seconds of 16 kHz mono audio in one pass and detects the language on its own; it emits word-level timestamps; and it can return the encoder's speech embedding, one row per 80 ms frame, without decoding a transcript.

The most strategic detail is architectural. Whistle loads into the same C++ engine as Needle, Cactus's function-calling model, so a single small binary can transcribe a clip and immediately turn it into tool calls — the pattern behind local voice control for wearables, robots and smart-home hardware. The engine ships prebuilt for 17 targets, from macOS, Linux, Android, iOS and watchOS to Windows on ARM, RISC-V, MIPS, the browser and a WASI component.

These are the vendor's own measurements, so the usual caveat applies: independent replication is the real test. But the direction of travel is the notable part. Sub-20-megabyte, dependency-free speech recognition on an ordinary CPU turns voice into something you can ship inside a microcontroller-class footprint instead of routing audio to a cloud subscription.