{"id":3144,"date":"2026-09-23T14:42:32","date_gmt":"2026-09-23T14:42:32","guid":{"rendered":"https:\/\/dmarketertayeeb.com\/blog\/qwen3-8-livetranslate-api-languages-pricing\/"},"modified":"2026-09-24T03:47:54","modified_gmt":"2026-09-24T03:47:54","slug":"qwen3-8-livetranslate-api-languages-pricing","status":"publish","type":"post","link":"https:\/\/dmarketertayeeb.com\/blog\/qwen3-8-livetranslate-api-languages-pricing\/","title":{"rendered":"Qwen3.8 LiveTranslate API: Languages, Events and Pricing"},"content":{"rendered":"\n<p><strong>Qwen3.8-LiveTranslate is a dedicated real-time translation model, not the general-purpose Omni assistant endpoint.<\/strong> Its Model Studio ID is <code>qwen3.8-livetranslate-flash-realtime<\/code>. The current model card lists 60 languages for audio input and translated text, but speech output covers 29; check the target language before choosing text-plus-audio mode. The documented Model Studio path uses a workspace-specific WebSocket endpoint in Beijing or Singapore, then streams text, translated audio, and source-transcription events. Qwen\u2019s announcement is dated September 18, 2026; the current Model Studio guide lists the model as stable. <a href=\"https:\/\/qwen.ai\/blog?id=qwen3.8-livetranslate\">Qwen\u2019s announcement<\/a> \u00b7 <a href=\"https:\/\/help.aliyun.com\/en\/model-studio\/qwen3-8-livetranslate-flash-realtime\">Model Studio model information<\/a><\/p>\n\n\n\n<p>Use LiveTranslate when the product needs simultaneous interpretation or captions. If the job is a broader real-time assistant with tools or MCP, compare the separate <a href=\"https:\/\/dmarketertayeeb.com\/blog\/qwen3-8-omni-flash-realtime-api\">Qwen3.8 Omni Realtime API guide<\/a>: the LiveTranslate model page lists function calling, structured outputs, web search, batch inference, and fine-tuning as unsupported.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Check speech output for your language first<\/h2>\n\n\n\n<p>Alibaba\u2019s current model information and translation guide distinguish the 60-language input\/text coverage from spoken output. The docs list audio plus text for 29 languages and text-only output for the remaining 31. \u201cUnderstands 60 languages\u201d therefore does not mean the model can speak a translation in all 60. The provider documents language lists, but does not publish a controlled quality result for every source-to-target pair; test your own language direction, accents, code-switching, names, and specialist terms before promising live interpretation.<\/p>\n\n\n\n<figure class=\"wp-block-table\" style=\"overflow-x:auto;\"><table style=\"min-width:640px;\"><thead><tr><th>Decision<\/th><th>What the current docs say<\/th><th>What to verify<\/th><\/tr><\/thead><tbody><tr><td>Input and translated text<\/td><td>60 supported languages<\/td><td>Confirm both ends of the actual language direction in the supported-language table.<\/td><\/tr><tr><td>Translated speech<\/td><td>29 supported output languages<\/td><td>Confirm the target language is in the speech-output subset; otherwise plan for text output.<\/td><\/tr><tr><td>Output choice<\/td><td>Text only, or text plus audio<\/td><td>Choose with <code>session.output_modalities<\/code>; the docs do not present an audio-only mode.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>Qwen\u2019s September 22 Alibaba Cloud release note reports average lagging (LAAL) moving from 2.8 to 2.3 seconds. That is a vendor-reported evaluation result, not an independent measurement or a response-time guarantee for your pair, room, network, or app. Use it as a reason to test, not as a service-level target.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Connect to the right region and workspace<\/h2>\n\n\n\n<p>For Model Studio, use the WebSocket Realtime endpoint and replace <code>{WorkspaceId}<\/code> with the ID for the workspace that owns the API key. The selected host, workspace, model ID, and key must match. The current Model Studio model card and 3.8 setup guide document WebSocket. A separate <a href=\"https:\/\/docs.qwencloud.com\/api-reference\/realtime-api\/overview\">QwenCloud protocol matrix<\/a> shows a broader protocol list; because the Model Studio guide assigns extra AOQ\/WebRTC support to the 3.5 model, verify transport support in the exact product and region before designing around anything else.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Beijing:\nwss:\/\/{WorkspaceId}.cn-beijing.maas.aliyuncs.com\/api-ws\/v1\/realtime?model=qwen3.8-livetranslate-flash-realtime\n\nSingapore:\nwss:\/\/{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com\/api-ws\/v1\/realtime?model=qwen3.8-livetranslate-flash-realtime\n\nWebSocket handshake:\nAuthorization: Bearer &lt;Model-Studio-API-key&gt;<\/code><\/pre>\n\n\n\n<p>These are the provider\u2019s documented Model Studio endpoint patterns, not a test connection. Store the API key on a trusted server rather than in a browser bundle or mobile app. The API guide uses a bearer key from Model Studio; check the selected region\u2019s console for workspace access before connecting.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Configure the Qwen3.8 session and stream events<\/h2>\n\n\n\n<p>After the WebSocket handshake, send a <code>session.update<\/code> event. This documented fragment requests translated text and speech in English; change the target language and output mode to fit the product.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>{\n  \"type\": \"session.update\",\n  \"session\": {\n    \"output_modalities\": [\"text\", \"audio\"],\n    \"translation\": {\n      \"language\": \"en\"\n    }\n  }\n}<\/code><\/pre>\n\n\n\n<p>Wait for the server\u2019s session update acknowledgement before sending audio. For Qwen3.8, the documented target-language field is <code>session.translation.language<\/code>; the default is English. The source transcript is emitted in the same session. The client-events page says Qwen3.8 ASR stays enabled, but avoid copying the older 3.5-only <code>session.modalities<\/code> or optional ASR model configuration into a 3.8 request.<\/p>\n\n\n\n<figure class=\"wp-block-table\" style=\"overflow-x:auto;\"><table style=\"min-width:680px;\"><thead><tr><th>Stage<\/th><th>Qwen3.8 event or field<\/th><th>Client action<\/th><\/tr><\/thead><tbody><tr><td>Send speech<\/td><td><code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">input_audio_buffer.append<\/code><\/td><td>Append base64 audio chunks; with default speaker detection, the server segments speech and starts translation.<\/td><\/tr><tr><td>Read source speech<\/td><td><code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">conversation.item.input_audio_transcription.delta<\/code> and <code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">.completed<\/code><\/td><td>Append source transcript deltas; use the completed event for the final source text.<\/td><\/tr><tr><td>Text-only translation<\/td><td><code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">response.text.delta<\/code><\/td><td>Append each <code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">delta<\/code> in arrival order.<\/td><\/tr><tr><td>Text plus speech<\/td><td><code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">response.audio_transcript.delta<\/code> and <code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">response.audio.delta<\/code><\/td><td>Append text deltas and Base64-decode audio deltas into audio chunks.<\/td><\/tr><tr><td>Finish safely<\/td><td><code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">session.finish<\/code>, then <code style=\"white-space:normal;overflow-wrap:anywhere;word-break:break-word;\">session.finished<\/code><\/td><td>Send finish after the last input and wait for the server acknowledgement before closing.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>The exact event names matter. Qwen3.5 examples on the same help page use fields such as <code>modalities<\/code>, <code>response.text.text<\/code>, and <code>response.audio_transcript.text<\/code>. Qwen3.8 instead uses <code>output_modalities<\/code>, <code>response.text.delta<\/code> for text-only output, and <code>response.audio_transcript.delta<\/code> when text and audio are requested. Do not mix the two event schemas.<\/p>\n\n\n\n<p>The 3.8 client-events guide lists <code>speaker_detection<\/code> as the default turn-detection type; it enables real-time speaker diarization. It also documents <code>server_vad<\/code>. The server can return speaker IDs with speech events, which lets the client associate source text and translated items. The manual commit flow shown for 3.5 is not the default 3.8 recipe.<\/p>\n\n\n\n<p>To add visual context, send video as sampled image frames through <code>input_image_buffer.append<\/code>, not as a single opaque video upload. The client-events reference says audio must be appended before the first image; images must be JPEG\/JPG, under 500 KB before Base64, no more than two per second, with 480p or 720p recommended and 1080p the maximum. Images are optional; the realtime model still requires audio input.<\/p>\n\n\n\n<p>One important close-path failure is losing the last utterance: the provider warns that closing the socket without <code>session.finish<\/code> can prevent the server from completing the final speech segment. Wait for <code>session.finished<\/code> before disconnecting. These examples are derived from Alibaba\u2019s current documentation and have not been run against a live workspace.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Estimate cost by region and modality<\/h2>\n\n\n\n<p>The model page lists original API rates per million tokens in CNY and excludes limited-time offers. Audio, image, translated text, and generated speech have separate meters:<\/p>\n\n\n\n<figure class=\"wp-block-table\" style=\"overflow-x:auto;\"><table style=\"min-width:680px;\"><thead><tr><th>Model Studio region<\/th><th>Audio input<\/th><th>Image input<\/th><th>Text output<\/th><th>Audio output<\/th><\/tr><\/thead><tbody><tr><td>China (Beijing)<\/td><td>\u00a540<\/td><td>\u00a53.30<\/td><td>\u00a5100<\/td><td>\u00a5160<\/td><\/tr><tr><td>Singapore (International)<\/td><td>\u00a554.688<\/td><td>\u00a54.01<\/td><td>\u00a5145.835<\/td><td>\u00a5218.752<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p>For Qwen3.8, the <a href=\"https:\/\/help.aliyun.com\/en\/model-studio\/qwen3-5-livetranslate-flash-realtime\">Model Studio billing guide<\/a> documents input audio at 7 tokens per second and output audio at 12.5 tokens per second. As an audio-only component illustration, 60 seconds of input is 420 tokens: about \u00a50.0168 in Beijing or \u00a50.0230 in Singapore. Ten seconds of generated speech is 125 tokens: \u00a50.0200 in Beijing or about \u00a50.0273 in Singapore. The combined audio components are about \u00a50.0368 and \u00a50.0503, respectively, before any translated-text tokens, image frames, or other usage. The provider\u2019s image rule is 0.5 token per 32-by-32-pixel tile.<\/p>\n\n\n\n<p>Do not treat that example as a universal cost per minute. It assumes exactly one minute of input audio and ten seconds of generated audio. Output text has its own rate and depends on the actual translated-text token count; frames add image tokens. For a real session, use the returned usage record and the price for the workspace\u2019s region rather than guessing tokens from words or assuming translated audio lasts as long as the source.<\/p>\n\n\n\n<p>The current official Chinese-language pricing table lists a Beijing-only 1-million-token free quota for this exact model, valid for 90 days from the latest applicable Model Studio activation, model release, or application approval. The Singapore table has no free-quota entry. This is an account and region benefit, not a permanent list price; verify the selected workspace\u2019s live console before including it in a budget. The English model-specific page shows standard rates and points readers to the console for promotions.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Include model limits in the proof of concept<\/h2>\n\n\n\n<p>Model Studio lists a 53,248-token context window, with up to 49,152 input tokens and 4,096 output tokens. Its model card lists 10 requests per minute and 100,000 tokens per minute for both Beijing and Singapore. Confirm the workspace\u2019s effective limits before load testing; the public model card is not an account-specific capacity guarantee.<\/p>\n\n\n\n<ul class=\"wp-block-list\"><li>Measure time to first translated text and first audio chunk on the networks and devices you plan to support.<\/li><li>Test each real source\/target language direction, especially proper names, technical terms, code-switching, accents, interruptions, and overlapping speakers.<\/li><li>Compare text-only and text-plus-audio output using actual token usage; confirm whether the 29-language speech subset covers the target.<\/li><li>For video, vary frame rate and resolution and inspect whether the visual context changes the translation enough to justify its image-token cost.<\/li><li>Record the session close behavior, reconnect path, rate limits, and final-segment handling before production.<\/li><\/ul>\n\n\n\n<p>Qwen3.8 LiveTranslate is worth evaluating when live translated captions or speech are the core product. Use the general Omni Realtime guide when you need a broader audio\/video assistant and tools. For either model, published language counts and vendor-reported latency are starting points; your test should decide whether a particular language pair and interaction pattern are reliable enough to ship.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Official sources<\/h2>\n\n\n\n<ul class=\"wp-block-list\"><li><a href=\"https:\/\/qwen.ai\/blog?id=qwen3.8-livetranslate\">Qwen\u2019s September 18 LiveTranslate announcement<\/a><\/li><li><a href=\"https:\/\/help.aliyun.com\/en\/model-studio\/qwen3-8-livetranslate-flash-realtime\">Model Studio model information, stable status, capabilities, regional rates, and limits<\/a><\/li><li><a href=\"https:\/\/help.aliyun.com\/en\/model-studio\/qwen3-5-livetranslate-flash-realtime\">Current real-time translation guide, including the explicit Qwen3.8 sections<\/a><\/li><li><a href=\"https:\/\/help.aliyun.com\/en\/model-studio\/live-translator-client-events\">Qwen3.8 client configuration and input events<\/a><\/li><li><a href=\"https:\/\/help.aliyun.com\/en\/model-studio\/live-translator-server-events\">Qwen3.8 source, translation, and audio server events<\/a><\/li><li><a href=\"https:\/\/help.aliyun.com\/zh\/model-studio\/model-pricing\">Official Chinese pricing table and region-specific free-quota terms<\/a><\/li><li><a href=\"https:\/\/www.alibabacloud.com\/en\/press-room\/alibaba-unveils-roadmap-on-full-stack-ai-strategy?_p_lc=1\">Alibaba Cloud\u2019s September 22 release note and vendor-reported LAAL result<\/a><\/li><\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Set up Qwen3.8 LiveTranslate over WebSocket, compare its 60-language input with 29 speech outputs, parse 3.8 events, and estimate regional CNY prices.<\/p>\n","protected":false},"author":1,"featured_media":3143,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[180,178],"tags":[393,300,315,441],"class_list":["post-3144","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news","category-artificial-intelligence","tag-ai-model-releases","tag-ai-models","tag-developer-tools","tag-llm-inference","has-featured-image"],"_links":{"self":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3144","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/comments?post=3144"}],"version-history":[{"count":2,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3144\/revisions"}],"predecessor-version":[{"id":3148,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3144\/revisions\/3148"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media\/3143"}],"wp:attachment":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media?parent=3144"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/categories?post=3144"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/tags?post=3144"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}