Alibaba Cloud Model Studio listed qwen3.8-omni-flash-realtime on September 21, 2026 as a real-time audio and video model that returns text and speech. It supports WebSocket, WebRTC, and AOQ, plus multichannel audio and remote MCP tools. Start by choosing the region and transport, then use that region’s workspace-specific endpoint and API key. This is a separate endpoint from the general qwen3.8-omni-flash model listed on September 17.
This guide uses Alibaba Cloud’s published setup and billing documentation. The examples are source-based and have not been run against a live Model Studio account; confirm current workspace access, console settings, and billing before deployment.
Choose the transport for your client
The three transports serve different integration paths. For video, the model receives consecutive image frames; it does not require you to treat a video file as one opaque input.
| Transport | Best fit in the provider docs | Trade-off to plan for |
|---|---|---|
| WebRTC | Browser-based interaction and traditional audio/video calls | Native browser support and built-in echo/noise handling; more connection setup than WebSocket |
| WebSocket | Server integration, prototypes, or direct event handling | Simplest integration, but weaker resilience on poor networks and no built-in echo/noise cancellation; the client must handle it |
| AOQ (AI over QUIC) | Native multimodal real-time sessions where weak-network behavior matters | Provider describes strong weak-network resilience and low integration difficulty; it is not browser-supported and uses a server-side token flow |
These are documented fit descriptions, not a latency or quality ranking from an independent test. If your product is a browser call, begin with WebRTC. If your service already manages the session and needs low-friction event access, start with WebSocket. Consider AOQ for a native client and unstable networks, then test it on the networks and devices your users actually have.
If your decision is specifically whether to evaluate Qwen3.8-Flash-Next locally or through a hosted service, see the separate Qwen3.8-Flash-Next local-versus-hosted guide. That post covers Flash-Next evaluation; it does not cover this Realtime endpoint.
Match the region, workspace, endpoint, and key
The current model page lists China (Beijing) and Singapore. Your key must belong to the selected region and workspace. Replace {WorkspaceId} with the workspace ID shown in the Model Studio console and pass the exact model ID in the query string.
| Region | WebSocket endpoint | WebRTC signaling endpoint |
|---|---|---|
| China (Beijing) | wss://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime | https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/webrtc/realtime?model=qwen3.8-omni-flash-realtime |
| Singapore | wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/realtime?model=qwen3.8-omni-flash-realtime | https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime?model=qwen3.8-omni-flash-realtime |
For WebSocket, authenticate the handshake with a bearer API key. For WebRTC, the API key is used during the SDP exchange. AOQ keeps the long-lived key on your application server: the server requests a gateway token and gives the temporary token to the client. For public clients, keep long-lived keys out of source code, browser bundles, and mobile apps; use a trusted backend, and use the documented temporary-token flow specifically for AOQ.
Before connecting, activate Model Studio, create a key for the selected workspace, and check that the Realtime model is available in that region. The SDK documentation says Qwen3.8 Realtime requires the workspace-specific domains shown above; older generic DashScope domains are not the recommended endpoint for this model.
Configure audio, video, and turns before streaming
Audio input and response mode
The Model Studio Python guide’s single-channel example uses 16 kHz, 16-bit, mono PCM input and 24 kHz, 16-bit, mono PCM output. Check the SDK and selected protocol’s current audio format before wiring your capture and playback devices. The Realtime model also supports WebSocket input with 1, 2, or 4 channels; multichannel input uses 16 kHz PCM. Configure the channel/spatial format before sending the first audio segment because the client-event guide says it cannot be changed after audio starts.
Choose automatic turn detection (VAD) for conversational back-and-forth. Use manual turn control for push-to-talk or voice-message flows: append audio to the input buffer, commit it, then send response.create. In the documented manual mode, session.turn_detection is set to null. The provider’s interaction-flow guide shows the client sending tool results and requesting another response after a tool call.
Video frames and aggregation
Model Studio describes video as a sequence of image frames and recommends sampling at about one frame per second in its Realtime guide. For this model, the session setting session.video.input.representation_compact defaults to none; set it to normal to enable aggregation. The billing guide says normal aggregation reduces video-token use to one quarter of the count without aggregation. That is a token-count statement, not a guarantee that every scene or task will retain the same useful visual detail. Validate frame rate and aggregation against your own video task.
{
"type": "session.update",
"session": {
"video": {
"input": {
"representation_compact": "normal"
}
}
}
}
This is the documented setting expressed as a minimal session-update fragment. It is illustrative and unexecuted; send configuration before the first audio segment and confirm the server’s session update event before relying on it.
Add functions or remote MCP with an explicit handoff
Custom function calling and remote MCP tools can be configured in the same session.tools collection. They cannot be combined with built-in web search through enable_search in this Realtime flow. Model Studio says MCP tool invocation has no extra tool fee, but the model inference used to choose and process tools is still billed.
- For function calling, read the completed argument event, run the function in your application, send the tool result back as a conversation item, then request the next model response.
- For MCP, wait for
mcp_list_tools.completed(or a failure event) before assuming discovery finished. A session-configuration acknowledgement does not mean the remote tools are ready. - Keep approval and execution on the application side. Treat tool arguments as untrusted input, enforce your own permissions, and avoid granting an agent broad access just because the session can call a tool.
Function calls are a control handoff, not a finished spoken answer. The documented flow returns arguments, your application executes the function and returns its result, and you send another response request. The model can then produce the user-facing text or audio response.
Budget per modality, including both parts of spoken output
Model Studio meters input and output tokens across modalities. The current CNY table lists these standard rates per million tokens:
| Region | Input text, image, or video | Input audio | Output text | Output audio |
|---|---|---|---|---|
| China (Beijing) | ¥1.50 | ¥6.00 | ¥4.50 | ¥12.00 |
| Singapore | ¥1.677 | ¥6.781 | ¥5.104 | ¥13.636 |
Speech output is billed twice across its modalities: audio tokens use the output-audio rate and the corresponding generated text uses the output-text rate. Do not price a spoken answer as audio-only. The provider’s current rate guide gives these conversion rules for this model:
- Input audio: seconds × 7 tokens.
- Output audio: seconds × 12.5 tokens.
- Audio clips shorter than one second are counted as one second.
- Spatial input audio uses twice the regular input-audio tokens for both two- and four-channel input.
- Images use one token per 32 × 32 pixel tile; video uses the image-frame rules, with
normalaggregation at one quarter of unaggregated video tokens.
For a checkable Beijing illustration, one minute of incoming audio is 420 input-audio tokens: 60 × 7 × ¥6 per million, or about ¥0.00252. Ten seconds of generated speech is 125 output-audio tokens: 10 × 12.5 × ¥12 per million, or about ¥0.0015. If that spoken response also contains 50 generated text tokens as an explicitly assumed example (not a typical token mix), its text costs 50 × ¥4.50 per million, or ¥0.000225. Audio plus text for the generated response is therefore about ¥0.001725; the illustrated audio input plus that response is about ¥0.004245. These totals exclude video/images, other text and all retained history.
History makes a live session’s cost grow differently from a single-turn estimate. On each response, retained audio, image/video, and text from earlier turns are processed again as input. Model Studio documents limits of 100 audio turns or 600 seconds and 50 video turns or 240 seconds; these are retained-history limits, not a total lifetime cap. A WebSocket session itself can run up to 120 minutes. Older history is discarded when the relevant turn or cumulative-media limit is exceeded, so test your own turn lengths and monitor usage rather than multiplying only the newest utterance.
Check the free-quota notice in your workspace
Alibaba’s official pages do not agree on which region receives the listed 1-million-token quota. The model-specific page and help.aliyun.com CNY pricing table say Beijing only; Alibaba Cloud’s separate international USD pricing page says Singapore only. Each describes a 90-day validity period tied to the later of service activation, model release, or application approval. Because the region wording conflicts, check the selected workspace’s console and billing view before counting on any free quota.
Limits to include in a proof of concept
The model page lists a maximum of 196,608 input tokens and 65,536 output tokens. Realtime audio history is capped at 100 turns and 600 seconds; video history at 50 turns and 240 seconds. These history limits refer to media retained in context, not the duration of the entire conversation. The same page lists 60 requests per minute and 2,000,000 tokens per minute for both Beijing and Singapore. Check the live console for your account’s effective limits before load testing.
A small evaluation should compare the exact transport and media path you intend to ship: microphone format, camera frame sampling, channel count, network conditions, interruption behavior, tool latency, and the model’s response quality on representative tasks. The published feature and pricing pages do not establish performance for your hardware, network, geography, or application.
Official documentation
- Model Studio lifecycle and Sep21 release entry
- Qwen3.8-Omni-Flash-Realtime model information, regions, modalities, limits, and CNY prices
- Realtime protocol overview for AOQ, WebRTC, and WebSocket
- Region-specific authentication and temporary-token flow
- Workspace-specific WebSocket URLs and SDK audio formats
- Turn handling and function/MCP tool flow
- Realtime client events, session configuration and multichannel audio
- Realtime setup, token conversion, billing, and limits
- Model Studio CNY pricing and quota notes
- Alibaba Cloud international USD pricing table; its quota region note differs from the help.aliyun pages
Provider documentation describes the supported interfaces, settings, and prices. This guide does not report an independent benchmark, live account test, or production deployment.