[Explanation] Alibaba’s “Wan-Streamer” Real-Time Interaction

IT

A research team at Alibaba has announced a foundational model called “Wan-Streamer,” which integrates audio and video to enable real-time interaction at response speeds equivalent to humans. The underlying goal is to overcome the fatal processing delay of conventional voice AI and build next-generation AI agents that interact seamlessly with the physical world.

Alibaba’s response to the ‘sub-second’ challenge

Alibaba Group’s research division is making waves in the global tech industry by releasing “Wan-Streamer v0.1” in June 2026, followed by the enhanced version “v0.2” on July 8, 2026. Wan-Streamer is an “end-to-end (E2E) interactive infrastructure model” that simultaneously understands users’ facial expressions, gestures, and voice, and can respond instantly.

The biggest feature of this model is its extremely short streaming unit, with a minimum of 160ms, to reproduce the “pause” humans feel in natural conversation. Even considering the average network latency of 350ms, the total interaction latency of the entire system is kept to about 550ms, meaning a response time of “sub-seconds” significantly less than one second.

Alibaba is currently rushing to break away from being an “e-commerce giant” in pursuit of leaps in cloud and AI fields, planning a massive investment of about 380 billion yuan (approximately 8 trillion yen) over the next three years. The announcement of Wan-Streamer is part of this large-scale investment plan and has once again demonstrated its presence globally as a cutting-edge AI development company originating from China. The following diagram compares the delay generation process between humans and Wan-Streamers.

Figure 1

Breaking away from traditional cascade models

Before Wan-Streamer appeared, voice AI systems mainly relied on a “cascade” architecture that connected multiple independent models. Specifically, it went through three steps: speech recognition (ASR), which converts users’ voices into text; large language models (LLMs), which analyze the text to generate responses; and speech synthesis (TTS), which converts the generated text back into speech.

However, since this method involves information conversion at each step, minor delays accumulate at each stage, resulting in several seconds of time for a final response. With this, it was impossible to achieve full-duplex communication like humans do by overlapping words, or to interact based on real-time visual feedback. By integrating these functions within a single transformer, Wan-Streamer eliminates duplication and dramatically reduces information gaps and errors.

This technological innovation is rooted in the achievements of Alibaba’s “Qwen” series, which is being integrated into platforms like Taobao, and the advanced inference capabilities of “Qwen3-Max,” announced in 2025, have been directly transformed into real-time interactive capabilities. As a result, AI has evolved beyond just a “search tool” into a “partner” that understands physical context and acts accordingly.

End-to-end technology delivering astonishingly low latency

Sensory integration with a single transformer

The core strength of Wan-Streamer lies in converting different information from language, voice, and video into a common unit called a “token,” modeling it within a single transformer. Instead of combining different modules for each modality (sense) as traditionally, it adopts a “native streaming” structure where visual, audio, and text tokens are inserted into each other.

This design enables AI to simultaneously analyze the user’s voice tone, facial expressions, and surrounding environment. For example, if a user says sadly “It’s okay” but their expression is clouded, Wan-Streamer can instantly pick up on the contradiction and respond with more appropriate words of concern.

Alibaba Cloud has achieved high growth of +26% year-on-year, solidifying its position as the “Chinese AWS” riding the wave of AI demand. Advanced multimodal processing like Wan-Streamer requires enormous computing resources, but by combining Alibaba’s powerful cloud infrastructure with its natively integrated architecture, it achieves near-theoretical high-speed inference.

Full-Duplex Communication Brought by Causal Encoders

In real-time conversations, turn-based communication, where processing starts after the user finishes, is a major factor that undermines naturalness. To solve this problem, Wan-Streamer introduced two innovative technologies: the Causal Encoder and the Multimodal Token Scheduling.

The causal encoder starts sequential analysis from the order data is received, without waiting for the information to be fully entered. This allows AI to predict the user’s intent the moment they start speaking and to prepare for a response. Furthermore, token scheduling determines in nanoseconds which information should be prioritized from vast data sources such as visual and audio, optimizing resource allocation.

This technology enables advanced full-duplex communication, where AI continuously listens to the user’s words without interruption while simultaneously generating its own responses—”listening, speaking, and thinking simultaneously.” This serves as a powerful weapon to counter competing models like OpenAI’s “GPT-Live-1,” boasting world-class performance in interactions down to the subsecond level.

High resolution and spatial awareness evolved in version 0.2

Deepening Environmental Understanding through Scene Grounding

The latest “Wan-Streamer v0.2,” released in July 2026, has succeeded in significantly improving video resolution without compromising the low latency performance of the initial version. A particularly noteworthy feature in this update is the evolution of the ability called “Scene Grounding.”

Scene grounding refers to the ability of AI to physically link the layout of objects and spaces within a video to the context of its own conversation. The latest models can accurately recognize the user’s posture, gaze ahead, hand movements, and the position of nearby objects as if they were “happening right in front of you.”

For example, if a user instructs you to “take the cup over there,” AI can instantly identify which cup is being pointed to based on their gaze and hand movements, and proceed the conversation based on that location. This level of spatial awareness suggests that AI is beginning to “understand” the three-dimensional world beyond mere two-dimensional image analysis, making it an essential element for future robot integration.

Improved accuracy as a mid-shot agent

In Wan-Streamer v0.2, the agent feature was especially strengthened for “mid-shot” (the angle of view capturing the upper body). This sense of distance is the most commonly used angle for people at reception, customer support, and conversations at home, aiming to maximize practicality in business and daily life.

With this update, not only will the subtle gestures and facial expressions of AI characters be rendered in higher definition, but the AI can also respond appropriately to users’ subtle movements, reflecting on its own “body.” Alibaba is outlining a strategy to link this advanced visual and audio integration technology to the same-day delivery service developed by its logistics subsidiary, Cainiao, as well as AI applications in education and healthcare.

Achieving both high resolution and maintaining low latency was a very high hurdle from a scientific and technical standpoint, but Alibaba overcame this by leveraging its proprietary compression algorithms and parallel processing technologies. The following image visualizes how AI recognizes objects within a scene in v0.2 and reflects them in dialogue.

Figure 2

Prospects for Social Implementation and the Path to Embodied AI

The Revolution of ‘Empathy’ in Education and Customer Service

The real-time conversational technology brought by Wan-Streamer is driving disruptive innovation in areas where “face-to-face” communication is highly valued, such as customer service and education. Until now, AI tutors and reception robots excelled at information accuracy, but struggled with ’emotional interactions’ such as reading users’ emotions and responding at the right moment.

However, if it can respond to users’ facial expressions and tone of voice at speeds around 500ms, like Wan-Streamer, it becomes possible to gently ask learners, “Is there something troubling you?” when they are stumbled. In this way, AI responding to users’ subtle signals without missing them goes beyond mere convenience improvements, supporting the realization of “empathetic AI” that builds trust with users.

Alibaba aims to leverage these technologies in new hardware such as smart glasses and in digital transformation (DX) in education and healthcare settings. It is no longer just a supplementary tool for e-commerce sites, but is accelerating its evolution into a tech company that supports people in every aspect of daily life, supporting them as a whole-of-life company.

Alibaba’s AI strategy that integrates into every aspect of daily life

Wan-Streamer’s ultimate goal is to realize “Embodied AI,” where AI operates in our world with a physical body. Alibaba is strongly advancing the integration of digital and physical spaces through the promotion of voice shopping on Taobao and the advancement of its logistics network.

The Wan-Streamer’s ability to integrate visual information, voice, and action is truly demonstrated when installed in home robots and automated delivery units. AI, which not only “listens” to user instructions but also “sees” the situation on the spot and can “act” without delay, is at the core of Alibaba’s next-generation flagship business, the AI and cloud strategy.

Entering 2026, Alibaba’s AI strategy is undergoing a dramatic shift under the leadership of Chen Yusen, described as a “34-year-old tech geek.” With the integration of enterprise AI agents and the shift to a development system focused on practicality, research results like those of Wan-Streamer are expected to be returned to our daily lives in an extremely short cycle.

Overcoming Regulatory Risks and the Future of Next-Generation Agents

Regulatory Trends on Emotional Interaction AI by the Chinese Government

With advances in advanced conversational technologies like Wan-Streamer, regulations on AI use within China are also becoming stricter. In January 2026, China’s Market Regulatory Authority (SAMR) and Cyberspace Regulatory Authority (CAC) finalized new measures regarding the use of AI in live streaming in e-commerce.

Particularly noteworthy is the strengthening of regulations on “emotional interactions” with AI. The Chinese government is concerned that AI behaving like humans and building excessive intimacy with users could pose addictive and ethical risks. In response, ByteDance and Alibaba have taken steps to suspend or disable anthropomorphic AI companion features one after another.

The real-time, intimate interactive experience provided by Wan-Streamer is highly likely to be subject to regulations regarding this “anthropomorphic interaction,” forcing Alibaba to face an extremely difficult balance between technological development and regulatory compliance. The interim measures implemented on July 15 strictly prohibit emotional manipulation by AI and behaviors that exploit user vulnerabilities, making the establishment of highly transparent safeguards essential for the commercial deployment of Wan-Streamer.

A society where humans and AI coexist through massive investments

Despite numerous challenges such as regulation and competition, Alibaba’s ambitions for AI remain undiminished. The Alibaba Partner Committee states, “Friendship and growth make Alibaba culture possible,” and while eliminating organizational impatience and aggressive management, it continues to pursue a blueprint for AI to change the world. The development of Wan-Streamer is also one step toward making those ideals a reality.

The future focus will be whether “native multimodal agents” like Wan-Streamer will be released as models for Open Weight. If Alibaba opens this model to the community like the earlier “Qwen” series, developers worldwide will use it to create a wide variety of social problem-solving apps such as support for people with disabilities and telemedicine.

In its latest financial results for 2025 and beyond, Alibaba clearly defined AI and cloud as the “second growth engine” to offset the slowdown in the e-commerce business. There is no doubt that the technology called Wan-Streamer holds the key to overcoming China’s strict regulatory environment and becoming a “winner in the AI and cloud era” in the global market.

[#アリババ #WanStreamer #リアルタイム対話 #マルチモーダルAI #低遅延 #中国テック #科学技術 #Qwen]

コメント

Copied title and URL