What Fits in 16.9 Megabytes

By Siam Sukkhee Trading Co., Ltd — 2026-10-09 — Hacker News

A model that transcribes speech in 11 milliseconds. A file smaller than most app icons. This is Whistle, and it works entirely on your device's CPU.

The constraint here is the whole point.

Cactus Compute released Whistle in early October as a single 16.9 MB package that runs speech recognition on anything—phones, wearables, automotive systems, microcontrollers. No internet. No dependencies. The audio stays where it came from. It detects language automatically or you force one yourself, handles up to 30 seconds of 16 kHz mono audio per pass, and covers seven languages: English, German, French, Spanish, Italian, Dutch, Polish.

Compare that footprint to what's out there. Whisper base weighs 145.3 MB. Moonshine tiny v2 is 41.9 MB. More than eight copies of Whistle fit in the space Whisper base needs. On an Apple M4 Pro CPU, Whistle reaches the first token in 11.1 milliseconds; Whisper base takes 73.2 ms. Decoding speed is 1,319 tokens per second against 266 for Whisper base.

For a microcontroller or a watch, that's the entire argument.

The architecture itself is not complicated—actually, that's not quite right. The architecture is elegant precisely because it avoids overcomplicating things. An audio encoder with eight attention blocks reads a log-mel front end. That output feeds into a decoder through gated cross-attention. The decoder runs eight more blocks, selects from five beam candidates, and applies keyword biasing via an Aho-Corasick automaton. It outputs not just text but word-level timestamps with confidence scores, and optionally speech embeddings without transcription at all.

The quantization is what makes the size work. That's an engineering choice, and it's a deliberate one—the authors trained every depth from 2 layers upward as a standalone model, so you can load lighter or heavier configurations at runtime without retraining.

Word-level timestamps matter more than they sound. Every word comes back with start time, end time, and probability. The decoder's own attention provides the alignment, so you're not bolting on a separate alignment step later. If you need speech embeddings instead of text—for search, matching, clustering—you get one row per 80 milliseconds of audio and never decode a transcript.

That silence detector deserves mention. Before the beam search even starts, Whistle measures the clip's loudness range. Below threshold, it returns empty text and empty language and stops. You're not burning tokens on quiet intervals.

The benchmarks show Whistle ahead on five of eight word-error-rate comparisons against both Whisper base and Moonshine, even though it's a tenth the size or smaller. LibriSpeech test-clean, test-other, SPGISpeech, Earnings-22, FLEURS average—all Whistle's. Whisper base wins on TED-LIUM, AMI, and MLS average, which are the bigger datasets. The tradeoff is explicit: smaller model, faster inference, less data to beat.

This isn't a cloud transcription API that happens to work offline. It's a model built from first principles for devices that have 16 MB of space and a CPU that needs to wake up, transcribe, and sleep again. One binary can load speech and text together and turn audio into structured tool calls directly—a voice interface that never phones home.

The release came with source on GitHub and weights on Hugging Face, both open. Python users get pip install cactus-needle and a three-line script. The C++ engine ships prebuilt for seventeen targets: macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser, WASI. You point it at any .cact file and it runs.

What matters here isn't the innovation in attention blocks or the logarithmic mel-frequency scaling. It's the ruthlessness about fitting something functional into a box so small that cloud transcription starts to look wasteful by comparison.

Source: "Cactus-Compute has released Whistle, an open-weight speech recognition model optimized for edge devices and embedded systems." — Hyper

Tags: metals recycling Thailand