TL;DR
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Cactus Compute has released Whistle, a 16.9 MB speech-recognition model designed to run locally on CPUs without external dependencies. The company reports support for seven languages and fast results in its own tests, but performance depends on the hardware and comparison setup; the supplied material does not include independent validation.
Cactus Compute announced Whistle on October 2, 2026, describing it as a 16.9 MB speech-recognition model that can transcribe audio locally on a CPU. The release targets devices including phones, wearables, robots and smart-home systems, and says Whistle supports seven languages while returning word-level timestamps and speech embeddings.
According to the company, Whistle processes 16 kHz mono audio in clips of up to 30 seconds. It supports English, German, French, Spanish, Italian, Dutch and Polish, with language detection unless a user specifies one. The model can also provide the start and end times and a probability for each word, or return speech embeddings without generating a transcript.
Cactus Compute says the model runs in its C++ CPU engine with no dependencies and uses the same engine as Needle. Its browser demonstration downloads the 16.9 MB model on first use; the company says audio stays on the device. The release does not specify a device list or a minimum memory requirement, so the file size alone does not establish how well it will run across different hardware.
The company reports that, on an Apple M4 Pro CPU with 10 seconds of audio, Whistle reached its first token in 11.1 milliseconds and decoded at 1,319 tokens per second. The stated comparison uses each model’s official runtime and defaults: Whistle at five beams, OpenAI Whisper and Moonshine Voice in non-streaming mode. Those timings describe the company’s test setup, not results guaranteed on other CPUs.
Local Transcription on Small Devices
A 16.9 MB model could make speech input practical in products that cannot or should not send audio to a remote service. Local processing can reduce reliance on network access and keep audio on-device, if an implementation follows the behavior described in Cactus Compute’s browser demo. This may be relevant for wearables, embedded systems and other devices with limited storage or connectivity.
Whistle also bundles transcription, word timestamps and embeddings in one model. That may help developers build interfaces that need aligned captions or speech-based features without assembling separate model components. However, the announcement does not report power use, memory consumption, or latency across a range of phones and embedded processors. Those factors will determine whether the model’s small file size translates into useful performance in real products.
on-device speech recognition software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Whistle Fits With Needle
Cactus Compute presents Whistle as a speech model built on the same CPU engine and shared components as Needle, its existing model. The release describes an eight-block encoder and an eight-block decoder, with the decoder attending to audio features produced by the encoder. Developers can select a decoder depth at load time, but the company says all eight encoder blocks still run.
The company compares Whistle with Whisper base and Moonshine tiny v2 on model size, decoding speed and word error rate. It reports Whistle at 16.9 MB, compared with 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. Its report says Whistle scores better on some listed speech benchmarks, while Whisper base scores better on others. The benchmark coverage is not identical: Moonshine is English-only, some results were not published by competing model authors, and the AMI figures cited for Whisper and the other models refer to different subsets.
For latency, the company says Whistle’s time to first token changes with clip length: 5.9 milliseconds for five seconds of audio, 11.1 milliseconds for 10 seconds and 36.3 milliseconds for 30 seconds. It notes that Whisper pads inputs to 30 seconds, affecting that comparison. The figures are therefore tied to the test methods and runtimes described in the announcement.
““It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.””
— Cactus Compute
As an affiliate, we earn on qualifying purchases.
Independent Tests Still Needed
The benchmark and speed figures in the announcement are reported by Cactus Compute; the provided material does not include independent test results. Results may vary with hardware, software versions, settings and audio conditions. The source also does not give word-error-rate values in the supplied text, limiting readers’ ability to judge the size of the reported accuracy differences.
It is not clear which devices beyond the Apple M4 Pro were tested, how much memory Whistle requires at runtime, or how much energy it uses. The release describes a browser demonstration and a CPU engine but does not specify availability terms, licensing, or a full list of supported operating systems. Those details matter for developers considering deployment.
voice transcription app for mobile
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deployment Details to Watch
The next useful evidence will be independent evaluations of accuracy, latency and resource use across the devices Whistle is designed to serve. Developers will also need clear documentation on licensing, supported platforms, runtime memory and integration requirements. Cactus Compute’s announcement does not provide dates for those details or name a separate validation effort.
For now, the company’s browser demo offers a way to try short clips in the listed languages after downloading the model. Whether Whistle moves from a compact release and company-reported benchmarks into wider device use will depend on developer access, real-world performance and support for the hardware targeted by the announcement.
embedded speech recognition hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model released by Cactus Compute. The company says it runs on a CPU and has a model file size of 16.9 MB.
Which languages does Whistle support?
Cactus Compute lists English, German, French, Spanish, Italian, Dutch and Polish. It says the language is detected automatically unless the user specifies one.
Does Whistle send audio to the cloud?
The company says audio in its browser demonstration stays on the device. The announcement does not establish how every product or third-party integration using Whistle will process audio.
How fast is Whistle?
In Cactus Compute’s test on an Apple M4 Pro CPU using 10 seconds of audio, Whistle reached its first token in 11.1 milliseconds and decoded at 1,319 tokens per second. These are company-reported figures from a specified test setup, not a guarantee for other devices.
Has Whistle’s performance been independently verified?
The supplied announcement reports the company’s own benchmarks. It does not provide independent validation or full word-error-rate figures in the source material provided here.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
