Back to all stories

SpaceXAI Launches Grok Voice Transcribe 2.0 with Double the Accuracy

SpaceXAI has launched Grok Voice Transcribe 2.0, a speech transcription model that doubles the accuracy of its predecessor while keeping pricing unchanged. The new model introduces improved multilingual support, speaker labeling, and other features designed to enhance transcription quality for batch and streaming use cases.

LA

LazyFounders

·5 min read
SpaceXAI Launches Grok Voice Transcribe 2.0 with Double the Accuracy
Image: Image Credit: SpaceXAI via THE BRIDGE

SpaceXAI has launched Grok Voice Transcribe 2.0, a speech transcription model that doubles the accuracy of its predecessor while keeping pricing unchanged. The new model introduces improved multilingual support, speaker labeling, and other features designed to enhance transcription quality for batch and streaming use cases.

30 SEC SUMMARY

  • SpaceXAI launched Grok Voice Transcribe 2.0, a speech transcription model with twice the accuracy of its predecessor.
  • The new model supports batch and streaming modes, with improved multilingual accuracy and no price increase.
  • Pricing remains at $0.10 per hour for batch processing and $0.20 per hour for streaming.
  • SpaceXAI plans to deprecate the older model within weeks, making 2.0 the default API option.
  • Features include speaker labeling, timestamps, and support for up to 8 channels at no extra cost.

TABLE OF CONTENTS

  • Launch and Accuracy Improvements
  • Features and Pricing
  • Transition Plan for Users
  • What this means
  • Key takeaways
  • FAQ
  • Sources

KEY HIGHLIGHTS

  • SpaceXAI released Grok Voice Transcribe 2.0, a speech transcription model with double the accuracy of its predecessor.
  • The model supports both batch and streaming modes, with significant improvements in multilingual accuracy.
  • Pricing remains at $0.10 per hour for batch processing and $0.20 per hour for streaming.
  • Features include speaker labeling, timestamps, and support for up to 8 channels at no additional cost.
  • SpaceXAI will deprecate Grok Voice Transcribe 1.0 within weeks, making 2.0 the default model for its API.

Launch and Accuracy Improvements

According to THE BRIDGE, SpaceXAI has released Grok Voice Transcribe 2.0, a new version of its speech transcription model. The company claims the model achieves twice the accuracy of its predecessor, Grok Voice Transcribe 1.0, while maintaining the same pricing.

The new model supports both batch and streaming transcription and is built on the same speech foundation model as SpaceXAI’s Grok Voice voice assistant. It also introduces improvements in multilingual accuracy, which SpaceXAI highlights as the most significant enhancement over the previous version.

Features and Pricing

Grok Voice Transcribe 2.0 includes several features designed to improve transcription quality and flexibility. It assigns start and end times, as well as confidence levels, for each transcribed word. Speaker labeling, which identifies individual speakers in a recording, is available at no additional cost. The model can also transcribe up to 8 channels individually.

Users can specify up to 100 unique terms per request to improve recognition accuracy. The model formats numbers, dates, currencies, phone numbers, and email addresses, while removing speech hesitations and detecting utterance breaks for voice agents.

Multilingual support is a key feature, with the model capable of handling dozens of languages. It automatically detects the language and tracks switches during a recording in a single process. According to THE BRIDGE, external metrics rank the model first in accuracy among 32 streaming-capable models on the Artificial Analysis leaderboard.

Pricing for the new model remains unchanged: $0.10 per hour of audio for batch processing and $0.20 per hour for streaming. Speaker labeling, timestamps, and term specification are included in this price.

Transition Plan for Users

SpaceXAI announced on September 18 that Grok Voice Transcribe 2.0 will soon become the default model for its Speech-to-Text API. The company plans to deprecate Grok Voice Transcribe 1.0 within a few weeks, though the exact timeline has not been disclosed.

Users who wish to continue using the older model during the transition period can do so by specifying the model name as "grok-voice-transcribe-1.0" in their API requests. For most users, integrating 2.0 requires no code changes to achieve improved accuracy.

What this means

LazyFounders analysis — our interpretation, not reported fact.

SpaceXAI’s launch of Grok Voice Transcribe 2.0 reflects a strategic push to improve accuracy and multilingual support without raising costs. For founders and operators relying on transcription services, this update could reduce the need for manual edits or third-party tools, particularly in multilingual use cases. The decision to deprecate the older model quickly may accelerate adoption but could also create urgency for teams needing to test and migrate.

The inclusion of features like speaker labeling and timestamping at no additional cost sets a new benchmark for competitors, potentially forcing them to reevaluate their pricing or feature sets. Startups in the AI and transcription space should monitor how this impacts customer expectations and pricing models.

Key takeaways

  • Grok Voice Transcribe 2.0 offers double the accuracy of its predecessor, with a focus on multilingual support.
  • Pricing remains unchanged at $0.10/hour for batch processing and $0.20/hour for streaming.
  • SpaceXAI will deprecate the older model within weeks, making migration to 2.0 necessary for most users.
  • Features like speaker labeling, timestamps, and multichannel support are included at no extra cost.

FAQ

What is Grok Voice Transcribe 2.0?

Grok Voice Transcribe 2.0 is SpaceXAI’s latest speech transcription model. It offers twice the accuracy of its predecessor, Grok Voice Transcribe 1.0, and includes features like multilingual support, speaker labeling, and timestamps.

How much does Grok Voice Transcribe 2.0 cost?

The model is priced at $0.10 per hour of audio for batch processing and $0.20 per hour for streaming. Features like speaker labeling and term specification are included at no extra cost.

What are the key improvements in Grok Voice Transcribe 2.0?

The most significant improvement is accuracy, particularly for multilingual transcription. The model also supports up to 8 channels, automatic language detection, and formatting for structured data like numbers and dates.

Will Grok Voice Transcribe 1.0 be discontinued?

Yes, SpaceXAI plans to deprecate Grok Voice Transcribe 1.0 within a few weeks and make 2.0 the default model for its Speech-to-Text API.

Do users need to change their code to adopt Grok Voice Transcribe 2.0?

No, users integrating with the existing Speech-to-Text API can achieve improved accuracy without changing their code. However, those who wish to continue using the older model must specify the model name in their requests.

Related on LazyFounders

Sources

  1. THE BRIDGEJapanese · 2026-09-22
    SpaceXAIが音声書き起こしモデル「Grok Voice Transcribe 2.0」を公開——価格据え置きで1.0比2倍の精度と説明

This story is an original summary and analysis written by LazyFounders from the reporting listed above. Facts are attributed to their original publishers; sections marked as analysis are LazyFounders's opinion. Where a source is in another language, facts were machine-translated and quotations are reported, not reproduced. Read the original coverage via the links.

Lazy Founder - Powered by Blogy.in