TL;DR
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on CPUs and support seven languages. The company reports fast processing and competitive word-error rates against selected models, but the benchmark results come from its own tests and do not cover every model on every dataset.
Cactus Compute released Whistle on October 2, describing it as a 16.9 MB speech-recognition model that runs on a CPU without additional dependencies and transcribes speech on the device. The company says it supports English, German, French, Spanish, Italian, Dutch and Polish, positioning it for phones, wearables and other devices where sending audio to a remote service may be undesirable.
According to Cactus Compute, Whistle processes 16 kHz mono audio clips of up to 30 seconds. It can return a transcript, word-level start and end times with probabilities, or a speech embedding containing one row per 80-millisecond frame. The language is detected automatically unless the user specifies one. The company says audio remains on the device in its browser demonstration after the model is downloaded.
The model uses the same C++ CPU engine, container and quantisation as Needle, Cactus Compute’s other model. Its encoder has eight attention blocks, and the decoder also has eight layers; a load-time setting can select a shallower decoder, while the encoder continues to run all eight blocks. The company says its silence check can return an empty transcript without running beam search when a clip is below a loudness threshold.
In Cactus Compute’s test on an Apple M4 Pro CPU using 10 seconds of audio, Whistle’s reported time to first token was 11.1 milliseconds, compared with 73.2 ms for Whisper base and 22.8 ms for Moonshine tiny v2. It reported 1,319 decoded tokens per second, versus 266 and 262, respectively. The company lists model sizes of 16.9 MB for Whistle, 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2.
Smaller Models for Local Transcription
A model that fits in 16.9 MB could make speech recognition more practical on devices with limited storage, computing capacity or network access. Local processing can also reduce the need to send recorded speech to a cloud service, though the company’s description of on-device operation does not by itself establish how every deployment handles storage, permissions or privacy.
The reported speed figures suggest Whistle may suit applications that need quick responses, including voice interfaces in wearables, smart-home products, cars and robots. Those figures are vendor test results on one stated CPU, not a guarantee of performance on other processors or in real-world use. Battery consumption, accuracy in noisy environments and performance on longer or overlapping speech are not established in the supplied report.
Accuracy matters as much as size and speed. Cactus Compute says Whistle scored better than the listed competitors on several datasets, including LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. It also reports Whisper base ahead on TED-LIUM, AMI and the MLS average. Those mixed results make the model’s suitability dependent on the language, recording conditions and task.
on-device speech recognition software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Cactus Compute Tested Whistle
The release compares Whistle with Whisper base and Moonshine tiny v2 on model size, time to first token, decode speed and word error rate. The company says each model ran in its official runtime with default settings; Whistle used its C++ engine at five beams, while Whisper used openai-whisper and Moonshine used moonshine-voice in non-streaming mode. These runtime differences are relevant when interpreting the timings.
Cactus Compute says Whisper pads every input to 30 seconds, making its first-token timing flat across clip lengths. By contrast, Whistle’s reported time to first token varies with clip duration: 5.9 ms at five seconds, 11.1 ms at 10 seconds and 36.3 ms at 30 seconds. The report defines decode speed as tokens divided by the wall time after the first token, excluding the encoder from that calculation.
The benchmark table has gaps because some model authors did not publish results for particular datasets. The company notes that Moonshine is English-only and that Whisper has no reported figures for SPGISpeech, Earnings-22 or AMI cleaned. It also cautions that Whisper’s AMI result uses the AMI-IHM subset, which differs from the AMI version used for the other models.
““It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.””
— Cactus Compute, in the Whistle release
voice transcription app for smartphones
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Published Results
The supplied release does not provide independent verification of the accuracy, speed or memory measurements. It does not give enough detail here to assess repeat counts, variation between runs, power use, memory consumption during inference or performance across different CPU architectures. The reported M4 Pro timings should not be treated as representative of every supported target device.
It is also unclear how Whistle performs beyond the listed datasets, across accents and background-noise conditions, or on speech involving multiple speakers. The stated language list covers seven languages, but the release summary does not give language-by-language accuracy figures. No results are supplied for longer continuous recordings, since the described one-pass clip limit is 30 seconds.
The company says it released the model, but the source material does not specify the full licensing terms, distribution channels, hardware requirements or a schedule for updates. It also does not establish that the browser demo’s privacy behavior automatically applies to every application built with the model.
As an affiliate, we earn on qualifying purchases.
Availability and Independent Testing
The immediate next step for developers is to check the model package, license and runtime instructions before integrating Whistle into a product; those details are not fully stated in the supplied release text. Cactus Compute’s browser demo offers a way to try transcription, but its first use downloads the 16.9 MB model.
Further evaluation would help establish whether the published accuracy and latency figures hold on target devices outside the M4 Pro test. Useful follow-up comparisons would report results by language and dataset, disclose resource use and repeatability, and test noisy and multi-speaker recordings. No independent evaluation or additional release milestone is identified in the source material.
privacy-focused speech recognition device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model released by Cactus Compute. The company says it can transcribe audio locally and also provide word timings or speech embeddings.
Which languages does Whistle support?
The release lists English, German, French, Spanish, Italian, Dutch and Polish. It says the language can be detected automatically or specified by the user.
How large is the model, and how long can a clip be?
Cactus Compute gives the model size as 16.9 MB and says it processes 16 kHz mono clips of up to 30 seconds in one pass.
Are the speed and accuracy results independently verified?
Not in the supplied source. The results are Cactus Compute’s own comparisons; the release describes the runtimes and test conditions but does not cite an independent evaluation.
Does the model work on every device?
The company says Whistle runs on CPUs and targets devices including phones, wearables and robots. The source does not provide performance measurements for every device or CPU, so actual speed and resource use may vary.
Source: hn
Columbus Day / Indigenous Peoples' Day Picks
long weekend sales
As an affiliate, we earn on qualifying purchases.
