Skip to content
DMarketer Tayeeb – Digital Marketing Expert in Bangalore | SEO, SEM & SMM Expert
Contact

Qwen3.8-Omni-Flash-Realtime API: Regions, Protocols and Pricing

Alibaba Cloud Model Studio listed qwen3.8-omni-flash-realtime on September 21, 2026 as a real-time audio and video model that returns text and speech. It supports WebSocket, WebRTC, and AOQ, plus multichannel audio and remote MCP tools. Start by choosing the region and transport, then use that region’s workspace-specific endpoint and API key. This is a separate endpoint from the general qwen3.8-omni-flash model listed on September 17.

This guide uses Alibaba Cloud’s published setup and billing documentation. The examples are source-based and have not been run against a live Model Studio account; confirm current workspace access, console settings, and billing before deployment.

Choose the transport for your client

The three transports serve different integration paths. For video, the model receives consecutive image frames; it does not require you to treat a video file as one opaque input.

TransportBest fit in the provider docsTrade-off to plan for
WebRTCBrowser-based interaction and traditional audio/video callsNative browser support and built-in echo/noise handling; more connection setup than WebSocket
WebSocketServer integration, prototypes, or direct event handlingSimplest integration, but weaker resilience on poor networks and no built-in echo/noise cancellation; the client must handle it
AOQ (AI over QUIC)Native multimodal real-time sessions where weak-network behavior mattersProvider describes strong weak-network resilience and low integration difficulty; it is not browser-supported and uses a server-side token flow

These are documented fit descriptions, not a latency or quality ranking from an independent test. If your product is a browser call, begin with WebRTC. If your service already manages the session and needs low-friction event access, start with WebSocket. Consider AOQ for a native client and unstable networks, then test it on the networks and devices your users actually have.

If your decision is specifically whether to evaluate Qwen3.8-Flash-Next locally or through a hosted service, see the separate Qwen3.8-Flash-Next local-versus-hosted guide. That post covers Flash-Next evaluation; it does not cover this Realtime endpoint.

Match the region, workspace, endpoint, and key

The current model page lists China (Beijing) and Singapore. Your key must belong to the selected region and workspace. Replace {WorkspaceId} with the workspace ID shown in the Model Studio console and pass the exact model ID in the query string.

RegionWebSocket endpointWebRTC signaling endpoint
China (Beijing)wss://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtimehttps://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/webrtc/realtime?model=qwen3.8-omni-flash-realtime
Singaporewss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtimehttps://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime?model=qwen3.8-omni-flash-realtime

For WebSocket, authenticate the handshake with a bearer API key. For WebRTC, the API key is used during the SDP exchange. AOQ keeps the long-lived key on your application server: the server requests a gateway token and gives the temporary token to the client. For public clients, keep long-lived keys out of source code, browser bundles, and mobile apps; use a trusted backend, and use the documented temporary-token flow specifically for AOQ.

Before connecting, activate Model Studio, create a key for the selected workspace, and check that the Realtime model is available in that region. The SDK documentation says Qwen3.8 Realtime requires the workspace-specific domains shown above; older generic DashScope domains are not the recommended endpoint for this model.

Configure audio, video, and turns before streaming

Audio input and response mode

The Model Studio Python guide’s single-channel example uses 16 kHz, 16-bit, mono PCM input and 24 kHz, 16-bit, mono PCM output. Check the SDK and selected protocol’s current audio format before wiring your capture and playback devices. The Realtime model also supports WebSocket input with 1, 2, or 4 channels; multichannel input uses 16 kHz PCM. Configure the channel/spatial format before sending the first audio segment because the client-event guide says it cannot be changed after audio starts.

Choose automatic turn detection (VAD) for conversational back-and-forth. Use manual turn control for push-to-talk or voice-message flows: append audio to the input buffer, commit it, then send response.create. In the documented manual mode, session.turn_detection is set to null. The provider’s interaction-flow guide shows the client sending tool results and requesting another response after a tool call.

Video frames and aggregation

Model Studio describes video as a sequence of image frames and recommends sampling at about one frame per second in its Realtime guide. For this model, the session setting session.video.input.representation_compact defaults to none; set it to normal to enable aggregation. The billing guide says normal aggregation reduces video-token use to one quarter of the count without aggregation. That is a token-count statement, not a guarantee that every scene or task will retain the same useful visual detail. Validate frame rate and aggregation against your own video task.

{
  "type": "session.update",
  "session": {
    "video": {
      "input": {
        "representation_compact": "normal"
      }
    }
  }
}

This is the documented setting expressed as a minimal session-update fragment. It is illustrative and unexecuted; send configuration before the first audio segment and confirm the server’s session update event before relying on it.

Add functions or remote MCP with an explicit handoff

Custom function calling and remote MCP tools can be configured in the same session.tools collection. They cannot be combined with built-in web search through enable_search in this Realtime flow. Model Studio says MCP tool invocation has no extra tool fee, but the model inference used to choose and process tools is still billed.

  • For function calling, read the completed argument event, run the function in your application, send the tool result back as a conversation item, then request the next model response.
  • For MCP, wait for mcp_list_tools.completed (or a failure event) before assuming discovery finished. A session-configuration acknowledgement does not mean the remote tools are ready.
  • Keep approval and execution on the application side. Treat tool arguments as untrusted input, enforce your own permissions, and avoid granting an agent broad access just because the session can call a tool.

Function calls are a control handoff, not a finished spoken answer. The documented flow returns arguments, your application executes the function and returns its result, and you send another response request. The model can then produce the user-facing text or audio response.

Budget per modality, including both parts of spoken output

Model Studio meters input and output tokens across modalities. The current CNY table lists these standard rates per million tokens:

RegionInput text, image, or videoInput audioOutput textOutput audio
China (Beijing)¥1.50¥6.00¥4.50¥12.00
Singapore¥1.677¥6.781¥5.104¥13.636

Speech output is billed twice across its modalities: audio tokens use the output-audio rate and the corresponding generated text uses the output-text rate. Do not price a spoken answer as audio-only. The provider’s current rate guide gives these conversion rules for this model:

  • Input audio: seconds × 7 tokens.
  • Output audio: seconds × 12.5 tokens.
  • Audio clips shorter than one second are counted as one second.
  • Spatial input audio uses twice the regular input-audio tokens for both two- and four-channel input.
  • Images use one token per 32 × 32 pixel tile; video uses the image-frame rules, with normal aggregation at one quarter of unaggregated video tokens.

For a checkable Beijing illustration, one minute of incoming audio is 420 input-audio tokens: 60 × 7 × ¥6 per million, or about ¥0.00252. Ten seconds of generated speech is 125 output-audio tokens: 10 × 12.5 × ¥12 per million, or about ¥0.0015. If that spoken response also contains 50 generated text tokens as an explicitly assumed example (not a typical token mix), its text costs 50 × ¥4.50 per million, or ¥0.000225. Audio plus text for the generated response is therefore about ¥0.001725; the illustrated audio input plus that response is about ¥0.004245. These totals exclude video/images, other text and all retained history.

History makes a live session’s cost grow differently from a single-turn estimate. On each response, retained audio, image/video, and text from earlier turns are processed again as input. Model Studio documents limits of 100 audio turns or 600 seconds and 50 video turns or 240 seconds; these are retained-history limits, not a total lifetime cap. A WebSocket session itself can run up to 120 minutes. Older history is discarded when the relevant turn or cumulative-media limit is exceeded, so test your own turn lengths and monitor usage rather than multiplying only the newest utterance.

Check the free-quota notice in your workspace

Alibaba’s official pages do not agree on which region receives the listed 1-million-token quota. The model-specific page and help.aliyun.com CNY pricing table say Beijing only; Alibaba Cloud’s separate international USD pricing page says Singapore only. Each describes a 90-day validity period tied to the later of service activation, model release, or application approval. Because the region wording conflicts, check the selected workspace’s console and billing view before counting on any free quota.

Limits to include in a proof of concept

The model page lists a maximum of 196,608 input tokens and 65,536 output tokens. Realtime audio history is capped at 100 turns and 600 seconds; video history at 50 turns and 240 seconds. These history limits refer to media retained in context, not the duration of the entire conversation. The same page lists 60 requests per minute and 2,000,000 tokens per minute for both Beijing and Singapore. Check the live console for your account’s effective limits before load testing.

A small evaluation should compare the exact transport and media path you intend to ship: microphone format, camera frame sampling, channel count, network conditions, interruption behavior, tool latency, and the model’s response quality on representative tasks. The published feature and pricing pages do not establish performance for your hardware, network, geography, or application.

Official documentation

Provider documentation describes the supported interfaces, settings, and prices. This guide does not report an independent benchmark, live account test, or production deployment.

Share this article

Published by

Tayeeb Khan

Tayeeb Khan is the founder of DMarketer Tayeeb, covering digital marketing, SEO and AI. Articles may draw on professional experience, source-based research and AI-assisted or automated production. Firsthand tests are identified in the relevant article; a byline does not imply personal testing or human review of every claim.

Leave a Comment

Your email address will not be published. Required fields are marked *

Stay ahead of the curve

Get actionable digital marketing, SEO, and AI insights delivered to your inbox. No fluff, just value.

No spam. Unsubscribe anytime.