Skip to content
DMarketer Tayeeb – Digital Marketing Expert in Bangalore | SEO, SEM & SMM Expert
Contact

Qwen3.8 LiveTranslate API: Languages, Events and Pricing

Qwen3.8-LiveTranslate is a dedicated real-time translation model, not the general-purpose Omni assistant endpoint. Its Model Studio ID is qwen3.8-livetranslate-flash-realtime. The current model card lists 60 languages for audio input and translated text, but speech output covers 29; check the target language before choosing text-plus-audio mode. The documented Model Studio path uses a workspace-specific WebSocket endpoint in Beijing or Singapore, then streams text, translated audio, and source-transcription events. Qwen’s announcement is dated September 18, 2026; the current Model Studio guide lists the model as stable. Qwen’s announcement · Model Studio model information

Use LiveTranslate when the product needs simultaneous interpretation or captions. If the job is a broader real-time assistant with tools or MCP, compare the separate Qwen3.8 Omni Realtime API guide: the LiveTranslate model page lists function calling, structured outputs, web search, batch inference, and fine-tuning as unsupported.

Check speech output for your language first

Alibaba’s current model information and translation guide distinguish the 60-language input/text coverage from spoken output. The docs list audio plus text for 29 languages and text-only output for the remaining 31. “Understands 60 languages” therefore does not mean the model can speak a translation in all 60. The provider documents language lists, but does not publish a controlled quality result for every source-to-target pair; test your own language direction, accents, code-switching, names, and specialist terms before promising live interpretation.

DecisionWhat the current docs sayWhat to verify
Input and translated text60 supported languagesConfirm both ends of the actual language direction in the supported-language table.
Translated speech29 supported output languagesConfirm the target language is in the speech-output subset; otherwise plan for text output.
Output choiceText only, or text plus audioChoose with session.output_modalities; the docs do not present an audio-only mode.

Qwen’s September 22 Alibaba Cloud release note reports average lagging (LAAL) moving from 2.8 to 2.3 seconds. That is a vendor-reported evaluation result, not an independent measurement or a response-time guarantee for your pair, room, network, or app. Use it as a reason to test, not as a service-level target.

Connect to the right region and workspace

For Model Studio, use the WebSocket Realtime endpoint and replace {WorkspaceId} with the ID for the workspace that owns the API key. The selected host, workspace, model ID, and key must match. The current Model Studio model card and 3.8 setup guide document WebSocket. A separate QwenCloud protocol matrix shows a broader protocol list; because the Model Studio guide assigns extra AOQ/WebRTC support to the 3.5 model, verify transport support in the exact product and region before designing around anything else.

Beijing:
wss://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-livetranslate-flash-realtime

Singapore:
wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-livetranslate-flash-realtime

WebSocket handshake:
Authorization: Bearer <Model-Studio-API-key>

These are the provider’s documented Model Studio endpoint patterns, not a test connection. Store the API key on a trusted server rather than in a browser bundle or mobile app. The API guide uses a bearer key from Model Studio; check the selected region’s console for workspace access before connecting.

Configure the Qwen3.8 session and stream events

After the WebSocket handshake, send a session.update event. This documented fragment requests translated text and speech in English; change the target language and output mode to fit the product.

{
  "type": "session.update",
  "session": {
    "output_modalities": ["text", "audio"],
    "translation": {
      "language": "en"
    }
  }
}

Wait for the server’s session update acknowledgement before sending audio. For Qwen3.8, the documented target-language field is session.translation.language; the default is English. The source transcript is emitted in the same session. The client-events page says Qwen3.8 ASR stays enabled, but avoid copying the older 3.5-only session.modalities or optional ASR model configuration into a 3.8 request.

StageQwen3.8 event or fieldClient action
Send speechinput_audio_buffer.appendAppend base64 audio chunks; with default speaker detection, the server segments speech and starts translation.
Read source speechconversation.item.input_audio_transcription.delta and .completedAppend source transcript deltas; use the completed event for the final source text.
Text-only translationresponse.text.deltaAppend each delta in arrival order.
Text plus speechresponse.audio_transcript.delta and response.audio.deltaAppend text deltas and Base64-decode audio deltas into audio chunks.
Finish safelysession.finish, then session.finishedSend finish after the last input and wait for the server acknowledgement before closing.

The exact event names matter. Qwen3.5 examples on the same help page use fields such as modalities, response.text.text, and response.audio_transcript.text. Qwen3.8 instead uses output_modalities, response.text.delta for text-only output, and response.audio_transcript.delta when text and audio are requested. Do not mix the two event schemas.

The 3.8 client-events guide lists speaker_detection as the default turn-detection type; it enables real-time speaker diarization. It also documents server_vad. The server can return speaker IDs with speech events, which lets the client associate source text and translated items. The manual commit flow shown for 3.5 is not the default 3.8 recipe.

To add visual context, send video as sampled image frames through input_image_buffer.append, not as a single opaque video upload. The client-events reference says audio must be appended before the first image; images must be JPEG/JPG, under 500 KB before Base64, no more than two per second, with 480p or 720p recommended and 1080p the maximum. Images are optional; the realtime model still requires audio input.

One important close-path failure is losing the last utterance: the provider warns that closing the socket without session.finish can prevent the server from completing the final speech segment. Wait for session.finished before disconnecting. These examples are derived from Alibaba’s current documentation and have not been run against a live workspace.

Estimate cost by region and modality

The model page lists original API rates per million tokens in CNY and excludes limited-time offers. Audio, image, translated text, and generated speech have separate meters:

Model Studio regionAudio inputImage inputText outputAudio output
China (Beijing)¥40¥3.30¥100¥160
Singapore (International)¥54.688¥4.01¥145.835¥218.752

For Qwen3.8, the Model Studio billing guide documents input audio at 7 tokens per second and output audio at 12.5 tokens per second. As an audio-only component illustration, 60 seconds of input is 420 tokens: about ¥0.0168 in Beijing or ¥0.0230 in Singapore. Ten seconds of generated speech is 125 tokens: ¥0.0200 in Beijing or about ¥0.0273 in Singapore. The combined audio components are about ¥0.0368 and ¥0.0503, respectively, before any translated-text tokens, image frames, or other usage. The provider’s image rule is 0.5 token per 32-by-32-pixel tile.

Do not treat that example as a universal cost per minute. It assumes exactly one minute of input audio and ten seconds of generated audio. Output text has its own rate and depends on the actual translated-text token count; frames add image tokens. For a real session, use the returned usage record and the price for the workspace’s region rather than guessing tokens from words or assuming translated audio lasts as long as the source.

The current official Chinese-language pricing table lists a Beijing-only 1-million-token free quota for this exact model, valid for 90 days from the latest applicable Model Studio activation, model release, or application approval. The Singapore table has no free-quota entry. This is an account and region benefit, not a permanent list price; verify the selected workspace’s live console before including it in a budget. The English model-specific page shows standard rates and points readers to the console for promotions.

Include model limits in the proof of concept

Model Studio lists a 53,248-token context window, with up to 49,152 input tokens and 4,096 output tokens. Its model card lists 10 requests per minute and 100,000 tokens per minute for both Beijing and Singapore. Confirm the workspace’s effective limits before load testing; the public model card is not an account-specific capacity guarantee.

  • Measure time to first translated text and first audio chunk on the networks and devices you plan to support.
  • Test each real source/target language direction, especially proper names, technical terms, code-switching, accents, interruptions, and overlapping speakers.
  • Compare text-only and text-plus-audio output using actual token usage; confirm whether the 29-language speech subset covers the target.
  • For video, vary frame rate and resolution and inspect whether the visual context changes the translation enough to justify its image-token cost.
  • Record the session close behavior, reconnect path, rate limits, and final-segment handling before production.

Qwen3.8 LiveTranslate is worth evaluating when live translated captions or speech are the core product. Use the general Omni Realtime guide when you need a broader audio/video assistant and tools. For either model, published language counts and vendor-reported latency are starting points; your test should decide whether a particular language pair and interaction pattern are reliable enough to ship.

Official sources

Share this article

Published by

Tayeeb Khan

Tayeeb Khan is the founder of DMarketer Tayeeb, covering digital marketing, SEO and AI. Articles may draw on professional experience, source-based research and AI-assisted or automated production. Firsthand tests are identified in the relevant article; a byline does not imply personal testing or human review of every claim.

Leave a Comment

Your email address will not be published. Required fields are marked *

Stay ahead of the curve

Get actionable digital marketing, SEO, and AI insights delivered to your inbox. No fluff, just value.

No spam. Unsubscribe anytime.