SpaceXAI has released Grok Voice Transcribe 2.0, its newest speech-to-text (STT) model. The development team claims it to be twice as accurate as Grok Voice Transcribe 1.0 at the same price. The model targets hard audio: noisy phone lines, competing voices, local accents, and spoken credentials. It runs in batch and real-time streaming modes through the Speech to Text API.
Is it deployable? Yes, as a hosted API. It is live today under the model ID grok-voice-transcribe-2.0. SpaceXAI has not announced open weights, so self-hosting is not an option.
What is Grok Voice Transcribe 2.0?
Grok Voice Transcribe 2.0 is built on the audio foundation model behind Grok Voice. SpaceXAI team states that Grok Voice already handles tens of thousands of customer-support calls a day. It also transcribes millions of hours of video narration and runs the Grok assistant in Tesla vehicles.
The training data is live, noisy, multilingual audio recorded across diverse environments. SpaceXAI then refined the model with post-training.
Benchmarks: What SpaceXAI Reports
SpaceXAI reports a first-place accuracy rank among 32 streaming models on the public Artificial Analysis leaderboard. That benchmark, AA-WER Streaming, uses about 8 hours of audio. It weights AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%. See the methodology for details.
SpaceXAI also measures word error rate (WER) on 4 internal sets drawn from production traffic:
- Telephony (8 kHz): English customer support calls
- Conversational: English conversations with Grok
- Credentials: phone numbers, emails, and addresses in English
- Short phrases: voice-assistant utterances in 19 languages
Version 2.0 improves on 1.0 across all 4 sets. On telephony, SpaceXAI says it leads every model the company tested. These internal results are vendor-reported and not independently reproduced.
Multilingual Transcription and Language Switching
The model transcribes dozens of languages and detects the language automatically. It also follows mid-recording language switches in a single pass. SpaceXAI team calls multilingual accuracy the largest improvement over 1.0.
Short phrases, such as in-car commands, give a model little context to identify the language. On that set, WER drops from 20.6% to 6.8%. That works out to roughly 67% fewer word errors.
The docs list 25 languages for written-form formatting of numbers, currencies, and units.
