Skip to content
DMarketer Tayeeb – Digital Marketing Expert in Bangalore | SEO, SEM & SMM Expert
Contact

Grok Voice Transcribe 2.0 API: Pricing, Modes and Benchmarks

Short answer, checked September 23, 2026: Grok Voice Transcribe 2.0 is available through xAI’s batch REST and streaming WebSocket Speech-to-Text API. Current xAI docs list grok-voice-transcribe-2.0 as the default, but the September 17 release note still lists 1.0 and the September 18 launch post says the default switch is coming soon. Pin the model ID in production while those snapshots differ; xAI has not posted an exact 1.0 retirement date. Batch costs $0.10 per audio hour and streaming costs $0.20. Artificial Analysis currently shows a batch WER of 2.3% for 2.0 versus 4.0% for 1.0, but that batch result is separate from xAI’s launch-day claim that 2.0 ranked first on the streaming board.

Choose batch when the recording can be transcribed after upload and you want the lower rate. Choose streaming when your app needs partial words while someone is still speaking. Before switching a live workload, test your own accents, noise, names, numbers, speaker mix, and latency needs: a leaderboard average is a comparison tool, not a promise for your calls or meetings.

Availability and the model-default mismatch

xAI’s developer release notes date API availability for grok-voice-transcribe-2.0 to September 17. The public Grok Voice Transcribe 2.0 announcement followed on September 18. That post said 2.0 would soon become the default and that 1.0 would be deprecated in the coming weeks, without a calendar deadline. The current Speech-to-Text documentation, last updated September 18, now names 2.0 as the default when the model field is omitted. The earlier September 17 release note still says 1.0 is the default.

The practical fix is simple: send the intended slug on every request instead of inheriting a changing default. The API docs list both 1.0 and 2.0. xAI says existing integrations receive the 2.0 improvement without other code changes, but pinning the slug makes your own test and rollback repeatable. Keep 1.0 only if you need to hold the old behavior while you compare outputs; check the current docs and release notes again before setting a retirement date in your migration plan.

Batch REST or streaming WebSocket?

Both modes use audio-duration pricing. Streaming costs twice the published batch rate, so it earns its place when an application must show interim text, respond to a live caller, or coordinate speech turns in real time.

ModeEndpoint and inputBest fitPublished price
BatchPOST https://api.x.ai/v1/stt; multipart audio file or a server-side audio URLRecorded meetings, uploaded calls, podcasts, and video audio that can wait for a completed response$0.10 per audio hour
Streamingwss://api.x.ai/v1/stt; send binary audio frames and receive transcript eventsLive captions, dictation, call assistance, or a voice app that needs partial text before the person finishes speaking$0.20 per audio hour

At the published rates, 1,000 audio hours cost $100 in batch mode or $200 in streaming mode. That is $0.00167 or $0.00333 per audio minute, respectively. This arithmetic uses audio duration, not the wall-clock time your app spends waiting. It excludes retries, storage, network transfer, and any application-side processing. See the current xAI API pricing page before budgeting a production rollout.

Try a source-based batch request

The documented REST endpoint accepts a file upload or a URL for xAI to fetch. This example follows the current docs and has not been executed by DMT:

curl -X POST https://api.x.ai/v1/stt \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -F model=grok-voice-transcribe-2.0 \
  -F format=true \
  -F language=en \
  -F "keyterm=product name" \
  -F file=@meeting.mp3

Keep the API key in a server-side secret. The uploaded file can be at most 500 MB, and the file field must come after the other multipart fields; xAI warns that fields sent after it may be ignored. The example’s format=true applies inverse text normalization, such as converting spoken numbers or currency into written form, and requires a language value. For a server-held audio URL, pass the documented url field instead of file.

Use streaming when the transcript must arrive during speech

The WebSocket endpoint accepts model and audio settings as query parameters. A minimal source-based configuration is:

wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0&sample_rate=16000&encoding=pcm&interim_results=true

Open the connection from your backend with the authorization header, wait for transcript.created, then send binary audio frames in the documented encoding. With interim_results=true, partial text can arrive about every 500 ms; send audio.done when the stream ends and handle the final transcript event. The docs recommend 16 kHz PCM and roughly 100 ms audio chunks. Do not put the xAI key in browser JavaScript: proxy the WebSocket through your backend. These are documented protocol steps, not a DMT-tested integration.

Features and limits to check before switching

xAI says 2.0 adds accuracy improvements without changing the existing API shape. The launch announcement lists word-level timestamps and confidence, speaker diarization, up to eight independently transcribed channels, key-term biasing, text formatting, filler-word controls, and Smart Turn. The current docs specify up to 100 key terms, each no longer than 50 characters. Multichannel streaming supports up to eight channels but is not available with Opus encoding.

There is a response-schema detail worth checking if confidence is essential to your product: the launch post mentions a confidence score for each word, while the current API response example shows word text and start/end timestamps (plus a speaker field when diarization is enabled) without a per-word confidence property. Confirm the field in a real response or current API reference before designing downstream logic around it.

The docs describe 25 language codes for formatting numbers, currencies, and units. They say speech transcription works regardless of the optional language parameter; xAI’s launch post describes automatic language detection, mid-recording language switches, and recognition across dozens of languages. Do not treat the formatting-code list as a complete list of languages the model can recognize, or assume a particular language pair works for your audio without testing it.

For a full-duplex voice application that listens and speaks, the transcription endpoint is only one component. DMT’s Gemini Live voice-app architecture guide covers a different provider and broader voice-app design; it does not document Grok’s model IDs or pricing.

Read the accuracy claims by test and mode

The headline that 2.0 is “twice as accurate” is xAI’s summary of its own evaluations, not a universal multiplier for every transcript. xAI describes four internal sets based on its production traffic: 8 kHz English customer-support calls, English conversations with Grok, spoken credentials such as phone numbers and email addresses, and short voice-assistant phrases across 19 languages. The post says 2.0 improves on 1.0 in all four. The only exact internal pair stated in readable prose is short-phrase WER falling from 20.6% to 6.8%. That is a result for this particular set, not a guaranteed reduction for your audio.

Artificial Analysis is a separate evaluation. Its current non-streaming leaderboard, checked September 23, lists Grok Voice Transcribe 2.0 at 2.3% AA-WER, a median batch speed factor of 157.7, and an estimated $1.67 per 1,000 audio minutes. Its 1.0 row shows 4.0% WER and a 279.8 speed factor at the same normalized price. The speed factor measures audio seconds processed per second; it is not a live-response latency measurement. On this batch snapshot, 2.0 has lower WER while 1.0 has a higher measured batch speed factor.

Do not turn that batch table into a current streaming rank. xAI’s September 18 post said 2.0 ranked first for accuracy among 32 streaming models at launch. Artificial Analysis’ current streaming page now lists 37 models in total and 31 plotted, but the accessible page text does not show Grok’s current row. The launch-day #1/32 statement is therefore attributed here to xAI; a current streaming position is not independently confirmed from the readable chart.

Artificial Analysis says its WER index is an approximately eight-hour sample tested directly against provider APIs, weighted across AA-AgentTalk (50%), VoxPopuli (25%), and Earnings22 (25%). WER counts substitutions, insertions, and deletions against a human reference transcript; lower is better. That makes it useful for comparing this defined sample. It does not replace a pilot with your accent mix, microphone quality, background noise, channel layout, domain vocabulary, and output requirements. Also, the live AA batch page’s model-count summary and FAQ ranking currently disagree, so this guide reports its visible v2 and v1 row values but no ordinal batch rank.

A safe migration check

  • Set model=grok-voice-transcribe-2.0 explicitly in batch forms and WebSocket query parameters. Keep a separately configured 1.0 path only if you need a controlled rollback.
  • Run both versions on the same representative audio. Include accents, numbers, product names, interruptions, multiple speakers, and the languages you support.
  • Compare the exact output your application consumes: WER or human correction rate, speaker labels, timestamps, formatting, partial/final timing, and failure/retry rate.
  • Estimate spend from hours of submitted audio. At the published prices, streaming is twice the audio-hour rate of batch; retries and application infrastructure are outside this calculation.
  • Recheck the current xAI docs and release notes before removing 1.0. The sources reviewed for this guide do not give a firm retirement date.

Disclosure: This guide synthesizes vendor documentation and the cited third-party benchmark. DMT did not run an API test, transcribe audio, or independently audit xAI’s production evaluations. The examples are unexecuted documentation guidance.

Sources

Share this article

Published by

Tayeeb Khan

Tayeeb Khan is the founder of DMarketer Tayeeb, covering digital marketing, SEO and AI. Articles may draw on professional experience, source-based research and AI-assisted or automated production. Firsthand tests are identified in the relevant article; a byline does not imply personal testing or human review of every claim.

Leave a Comment

Your email address will not be published. Required fields are marked *

Stay ahead of the curve

Get actionable digital marketing, SEO, and AI insights delivered to your inbox. No fluff, just value.

No spam. Unsubscribe anytime.