Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to transcribe audio locally on CPUs without external dependencies. The company reports an 11.1 ms time to first token on an Apple M4 Pro and better word-error rates than selected competitors on several tests, while Whisper base performed better on other benchmarks.

Cactus Compute released Whistle, a speech-recognition model packaged as a 16.9 MB file and designed to run on-device using a CPU, the company said on October 2. The model transcribes speech in seven languages and is built to share the C++ engine used by the company’s Needle model, positioning it for devices where network access, memory or power may be limited.

Whistle accepts 16 kHz mono audio and can process clips of up to 30 seconds in one pass. Cactus says it supports English, German, French, Spanish, Italian, Dutch and Polish, with language detection unless a user specifies one. The model can also return word-level start and end times, probability estimates, and speech embeddings rather than a decoded transcript.

The company says audio is processed on the device and does not leave it when using the browser demonstration. The first use of that demo downloads the 16.9 MB model. Cactus also describes Whistle as having no runtime dependencies and running in the same container and quantisation as Needle; these are company statements, and the report does not provide independent verification.

In tests reported by Cactus, Whistle reached its first output token in 11.1 milliseconds and decoded at 1,319 tokens per second on an Apple M4 Pro CPU using a 10-second audio clip. The company compared official runtimes at their default settings: Whistle used five-beam decoding, alongside OpenAI Whisper base and Moonshine tiny v2. These figures describe that specific setup, not performance across other processors or devices.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute has released Whistle, a compact speech-to-text model for on-device CPU use, with company-published latency and accuracy comparisons.

Speech Recognition Without a Server

A model small enough to fit in a 16.9 MB package could make speech transcription more practical on phones, wearables, robots and other devices with limited resources. On-device processing can also reduce reliance on a continuous internet connection and keep recorded audio from being sent to a remote service, though actual privacy depends on how an application integrates the model and handles audio.

The release matters beyond file size because the company reports low first-token latency and a CPU-based runtime. If those results hold on target hardware, developers could use voice input where sending each clip to a cloud service would add delay, cost or connectivity requirements. Cactus’s comparison also shows that compactness does not settle the accuracy question: Whistle scored better on some listed benchmarks, while Whisper base led on others.

Potential users should treat the figures as vendor-reported results, not a broad independent evaluation. The published numbers are tied to a particular chip, clip length, runtime and decoding configuration. Performance on lower-powered mobile CPUs, microcontrollers or noisy real-world audio is not established by the supplied report.

Amazon

on-device speech to text software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Processes Audio

Cactus describes a pipeline that turns audio into log-mel features, compresses the audio frames through a convolutional stem, then uses an encoder and decoder to produce text. For a 30-second clip, the report says the front end generates 3,000 frames and the convolutional stage reduces that to 375 frames, one for every 80 milliseconds. The encoder uses eight attention blocks, while the decoder uses a separate set of eight blocks and reads the encoded audio through gated cross-attention.

The model shares components with Needle, Cactus’s existing model, rather than using a wholly separate engine. Developers can select decoder depth at load time, according to the report, but the encoder still runs all eight blocks. Whistle uses five-beam search, supports keyword biasing, and caps transcripts at 320 tokens. The company says it checks clip loudness before decoding and returns an empty transcript for audio below its silence threshold.

Cactus’s benchmark summary says Whistle had lower word-error rates than Whisper base on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base scored better on TED-LIUM, AMI and the MLS average. The report warns that some comparisons are unavailable because other model authors did not publish results, and that Whisper’s AMI figure uses AMI-IHM, a different subset from the AMI results reported for the other models.

““audio never leaves your device.””

— Cactus Compute

Amazon

low latency speech recognition app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests Still Needed

The published report does not establish how Whistle performs across a wide range of devices, accents, recording conditions or background-noise levels. Its latency and decoding figures are from an Apple M4 Pro test, and no independent replication or third-party benchmark is included in the source material.

Some comparison points are missing because the respective model authors did not publish results, Cactus says. The benchmark summary also identifies a different AMI subset for Whisper base, making that particular comparison less direct. The source does not provide enough detail here to independently assess every dataset, scoring choice or the impact of different runtime defaults.

The release announcement does not specify licensing terms, minimum hardware requirements, memory use during inference, battery consumption, or whether the model can be redistributed freely. It also does not report the exact silence threshold or quantify how keyword biasing affects recognition errors. Those details matter to developers evaluating deployment, but remain unconfirmed in the supplied report.

Amazon

multilingual speech transcription device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Deployment Details to Watch

Developers considering Whistle will need to check Cactus’s model and engine documentation for availability, licensing and integration instructions, as well as hardware requirements and memory use. The company’s release material describes a browser demo and a C++ engine, but does not provide a timeline for additional languages, longer audio limits or further platform support.

The next useful evidence would be independent testing across mobile and embedded CPUs, with consistent benchmark datasets and clearly stated runtime settings. Until then, Whistle’s size and reported speed are measurable claims from its maker, while suitability for a particular application will depend on device performance and transcription accuracy in that application’s audio conditions.

Amazon

privacy-focused voice recognition hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. It transcribes audio locally using a CPU-based engine, according to the company.

Which languages does Whistle support?

Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. It says the model detects the language unless the user specifies one.

How large is the model, and what audio can it process?

The model is distributed as a 16.9 MB file. The company says it accepts 16 kHz mono audio clips of up to 30 seconds in one pass.

How did Whistle compare with Whisper base?

In Cactus Compute’s reported tests, Whistle had lower word-error rates on some listed benchmarks, including LibriSpeech and SPGISpeech. Whisper base scored better on TED-LIUM, AMI and the MLS average, and the company notes that some benchmark comparisons are missing or use different dataset subsets.

Are the performance claims independently verified?

The supplied report presents results from Cactus Compute, including measurements on an Apple M4 Pro. It does not include independent replication or a broad evaluation across device types.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can Codex Costs Fall While Development Stays Fast? LegalOn’s Approach

LegalOn says it cut Codex costs by half without slowing development, but the available report details no baseline, timeframe or measurement method.

What’s New In AI? Access SpaceXAI’s Grok Bot Before Its Official Launch

SpaceXAI’s Grok chatbot is now available in an early beta phase, offering select users a first look ahead of broader release plans. Access details are still emerging.

Anthropic Unveils Claude Docs, Elevating Competition With Microsoft In AI

Anthropic introduces Claude Docs, a new document-focused AI feature, challenging Microsoft’s dominance in workplace AI tools and expanding its enterprise offerings.

The Future Is Here: 12 AI-Powered Note Apps For 2026

Discover the leading AI note-taking apps of 2026, featuring transcription, summaries, and seamless device integration to boost productivity.