AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Cactus Compute has released Whistle, a 16.9 MB speech-recognition model designed to run locally on CPUs without external dependencies. The company reports support for seven languages and fast results in its own tests, but performance depends on the hardware and comparison setup; the supplied material does not include independent validation.

Cactus Compute announced Whistle on October 2, 2026, describing it as a 16.9 MB speech-recognition model that can transcribe audio locally on a CPU. The release targets devices including phones, wearables, robots and smart-home systems, and says Whistle supports seven languages while returning word-level timestamps and speech embeddings.

According to the company, Whistle processes 16 kHz mono audio in clips of up to 30 seconds. It supports English, German, French, Spanish, Italian, Dutch and Polish, with language detection unless a user specifies one. The model can also provide the start and end times and a probability for each word, or return speech embeddings without generating a transcript.

Cactus Compute says the model runs in its C++ CPU engine with no dependencies and uses the same engine as Needle. Its browser demonstration downloads the 16.9 MB model on first use; the company says audio stays on the device. The release does not specify a device list or a minimum memory requirement, so the file size alone does not establish how well it will run across different hardware.

The company reports that, on an Apple M4 Pro CPU with 10 seconds of audio, Whistle reached its first token in 11.1 milliseconds and decoded at 1,319 tokens per second. The stated comparison uses each model’s official runtime and defaults: Whistle at five beams, OpenAI Whisper and Moonshine Voice in non-streaming mode. Those timings describe the company’s test setup, not results guaranteed on other CPUs.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact, CPU-based speech recognition model that runs on-device and shares an engine with its Needle model.

Local Transcription on Small Devices

A 16.9 MB model could make speech input practical in products that cannot or should not send audio to a remote service. Local processing can reduce reliance on network access and keep audio on-device, if an implementation follows the behavior described in Cactus Compute’s browser demo. This may be relevant for wearables, embedded systems and other devices with limited storage or connectivity.

Whistle also bundles transcription, word timestamps and embeddings in one model. That may help developers build interfaces that need aligned captions or speech-based features without assembling separate model components. However, the announcement does not report power use, memory consumption, or latency across a range of phones and embedded processors. Those factors will determine whether the model’s small file size translates into useful performance in real products.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Fits With Needle

Cactus Compute presents Whistle as a speech model built on the same CPU engine and shared components as Needle, its existing model. The release describes an eight-block encoder and an eight-block decoder, with the decoder attending to audio features produced by the encoder. Developers can select a decoder depth at load time, but the company says all eight encoder blocks still run.

The company compares Whistle with Whisper base and Moonshine tiny v2 on model size, decoding speed and word error rate. It reports Whistle at 16.9 MB, compared with 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. Its report says Whistle scores better on some listed speech benchmarks, while Whisper base scores better on others. The benchmark coverage is not identical: Moonshine is English-only, some results were not published by competing model authors, and the AMI figures cited for Whisper and the other models refer to different subsets.

For latency, the company says Whistle’s time to first token changes with clip length: 5.9 milliseconds for five seconds of audio, 11.1 milliseconds for 10 seconds and 36.3 milliseconds for 30 seconds. It notes that Whisper pads inputs to 30 seconds, affecting that comparison. The figures are therefore tied to the test methods and runtimes described in the announcement.

““It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.””

— Cactus Compute

Amazon

portable speech to text device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests Still Needed

The benchmark and speed figures in the announcement are reported by Cactus Compute; the provided material does not include independent test results. Results may vary with hardware, software versions, settings and audio conditions. The source also does not give word-error-rate values in the supplied text, limiting readers’ ability to judge the size of the reported accuracy differences.

It is not clear which devices beyond the Apple M4 Pro were tested, how much memory Whistle requires at runtime, or how much energy it uses. The release describes a browser demonstration and a CPU engine but does not specify availability terms, licensing, or a full list of supported operating systems. Those details matter for developers considering deployment.

Amazon

voice transcription app for mobile

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Deployment Details to Watch

The next useful evidence will be independent evaluations of accuracy, latency and resource use across the devices Whistle is designed to serve. Developers will also need clear documentation on licensing, supported platforms, runtime memory and integration requirements. Cactus Compute’s announcement does not provide dates for those details or name a separate validation effort.

For now, the company’s browser demo offers a way to try short clips in the listed languages after downloading the model. Whether Whistle moves from a compact release and company-reported benchmarks into wider device use will depend on developer access, real-world performance and support for the hardware targeted by the announcement.

Amazon

embedded speech recognition hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. The company says it runs on a CPU and has a model file size of 16.9 MB.

Which languages does Whistle support?

Cactus Compute lists English, German, French, Spanish, Italian, Dutch and Polish. It says the language is detected automatically unless the user specifies one.

Does Whistle send audio to the cloud?

The company says audio in its browser demonstration stays on the device. The announcement does not establish how every product or third-party integration using Whistle will process audio.

How fast is Whistle?

In Cactus Compute’s test on an Apple M4 Pro CPU using 10 seconds of audio, Whistle reached its first token in 11.1 milliseconds and decoded at 1,319 tokens per second. These are company-reported figures from a specified test setup, not a guarantee for other devices.

Has Whistle’s performance been independently verified?

The supplied announcement reports the company’s own benchmarks. It does not provide independent validation or full word-error-rate figures in the source material provided here.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Work By Valve’s Timur Kristóf On Improving Old AMD GPUs On Linux

Timur Kristóf presented work improving AMDGPU support for decade-old GCN 1.0 and 1.1 cards, including display, power management and reset fixes.

Breaking Up With Google Play: Why Conversations Is Now Free

A reported rise in interest in the phrase “Breaking Up with Google Play: Why Conversations Is Now Free” has put the app Conversations and Google Play in fo

How to Choose Automated Testing Tools

Set up automated testing with Jest, Playwright, and GitHub Actions: write unit tests, run them in CI, and verify every commit automatically.

Book Review: Is Parallel Programming Hard, And, If So, What Can You Do About It?

A book review asking whether parallel programming is inherently hard is drawing renewed interest. The review’s author and outlet remain unconfirmed.