By using this site, you agree to our Privacy Policy and Terms of Use.
Accept
VellaTimesVellaTimesVellaTimes
  • News
    NewsShow More
    Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
    Espresso Extraction Science: The Finer Grind Flaw
    May 18, 2026
    A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
    Amazon Alexa for Shopping Replaces Rufus AI Assistant
    May 18, 2026
    Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
    OpenAI acquires Weights.gg to boost voice AI tools
    May 18, 2026
    Federal agents standing outside a modern university biology laboratory building at dusk during an active investigation.
    US Arrests Chinese Scientists for Smuggling Biological Materials
    May 18, 2026
    A dramatically lit modern corporate courtroom with futuristic technology elements, representing a high-stakes artificial intelligence legal trial.
    Elon Musk OpenAI Lawsuit Exposes Clashes Over AI Safety
    May 18, 2026
  • Technology
    TechnologyShow More
    Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
    OpenAI acquires Weights.gg to boost voice AI tools
    May 18, 2026
    A polished silicon wafer rests on a surface inside a modern semiconductor manufacturing facility.
    Samsung Strike Threatens Global AI Chip Production
    May 18, 2026
    A glowing computer screen displaying the text GPT-5.5 Instant in a modern, high-tech office environment with soft blue and purple lighting.
    GPT-5.5 Instant: OpenAI’s New Default ChatGPT Model
    May 10, 2026
    Wide view of a modern AI data center with server racks, glowing fiber-optic cables, and semiconductor hardware in the foreground.
    AI Infrastructure Spending Drives Nvidia, AMD Shares
    May 10, 2026
    A glowing computer monitor displaying lines of code and digital network graphics in a modern tech office setting.
    Airbnb AI Coding: 60% of New Software Now Generated by AI
    May 9, 2026
  • AI
    AIShow More
    A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
    Amazon Alexa for Shopping Replaces Rufus AI Assistant
    May 18, 2026
    A dramatically lit modern corporate courtroom with futuristic technology elements, representing a high-stakes artificial intelligence legal trial.
    Elon Musk OpenAI Lawsuit Exposes Clashes Over AI Safety
    May 18, 2026
    A high-tech global map visualization showing glowing digital connections across different continents, representing the worldwide adoption of artificial intelligence.
    Global AI Adoption in 2026: Trends and Growing Divide
    May 10, 2026
    A modern smartphone displaying an artificial intelligence chat interface used for online shopping and product comparison.
    Alibaba Qwen AI Taobao Integration Launches Agentic Shopping
    May 10, 2026
    A split-screen illustration showing a high-tech modern office using advanced AI tools contrasted against an older, dimly lit workspace.
    Global AI Adoption Surges But Rich-Poor Divide Widens
    May 9, 2026
  • Science
    ScienceShow More
    Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
    Espresso Extraction Science: The Finer Grind Flaw
    May 18, 2026
    Federal agents standing outside a modern university biology laboratory building at dusk during an active investigation.
    US Arrests Chinese Scientists for Smuggling Biological Materials
    May 18, 2026
    Header image of a quantum communication lab setup with fiber-optic equipment, a telecom quantum dot device, and interferometer components used for long-distance quantum key distribution.
    Quantum Key Distribution Reaches 120 km With Quantum Dots
    May 10, 2026
    Abstract geometric representation of glowing quantum paraparticles interacting within a three-dimensional mathematical grid in deep blue and gold tones.
    Quantum Paraparticles Exist: New Math Challenges Physics
    May 10, 2026
    A large expedition cruise ship is navigating rough ocean waters under a cloudy sky.
    Global Authorities Respond to Andes Hantavirus Outbreak on MV Hondius Cruise Ship
    May 9, 2026
  • World
    WorldShow More
    Allu Arjun Commitment to Ethical Brand Partnerships
    Exploring Allu Arjun’s Commitment to Ethical Brand Partnerships
    December 18, 2023
    Orry aka Orhan Awatramani
    Orhan Awatramani ‘Orry’ Biography, Lifestyle and Rise to Fame
    December 8, 2023
    Alia Bhatt Latest Deepake Video Victim
    Alia Bhatt becomes latest victim of Deepfake Videos, Obscene Video goes Viral
    November 28, 2023
    Napoleon Movie Review
    Napoleon Movie Review: A Historical Epic by Ridley Scott Reviewed
    November 25, 2023
  • Bookmarks
Search
Category
  • News
  • Technology
  • AI
  • Science
  • World
Company
  • About Us
  • Contact Us
  • Fact Checking Policy
  • Terms & Conditions
  • Privacy Policy
  • Copyright Policy
Resources
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
© 2022 VellaTimes • All Rights Reserved.
Reading: Alibaba Launches Qwen3.5-Omni Multimodal AI to Rival Gemini
Share
Notification Show More
Font ResizerAa
VellaTimesVellaTimes
Font ResizerAa
  • News
  • Technology
  • AI
  • Science
  • World
Search
  • Explore
    • News
    • Technology
    • AI
    • Science
    • World
  • Useful Links
    • About Us
    • Contact Us
    • Fact Checking Policy
    • Terms & Conditions
    • Privacy Policy
    • Copyright Policy
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
© 2022 VellaTimes • All Rights Reserved.
News

Alibaba Launches Qwen3.5-Omni Multimodal AI to Rival Gemini

Rakesh Paul
Last updated: 31/03/2026
Rakesh Paul
Share
6 Min Read
A futuristic server room screen displaying a glowing AI neural network merging soundwaves, text code, and video frames, representing the native multimodal capabilities of an advanced AI model.

On March 30, 2026, Alibaba introduced Qwen3.5-Omni, a native multimodal AI model designed to process text, images, audio, and video simultaneously. Moving away from older systems that simply stitch together separate text and vision tools, this new release uses a unified computational pipeline to handle all data types natively. The model aims to compete directly with major industry players, delivering real-time interaction, complex problem solving, and advanced reasoning capabilities for both enterprise and everyday users.

Contents
Unified Architecture and Context CapacityOutperforming Gemini in Audio BenchmarksAudio-Visual Vibe Coding and Real-Time VoiceConflicting Reports on Open-Source AvailabilityLeadership Changes at Alibaba

The Alibaba Qwen3.5-Omni series is available in three sizes to balance performance and cost. The Plus tier focuses on maximum accuracy and complex reasoning, the Flash version prioritizes high-throughput and low-latency interactions, and the Light variant is built for efficiency.

Unified Architecture and Context Capacity

All three models share a massive 256,000-token context window. This large data capacity allows the system to process over ten hours of continuous audio input or more than 400 seconds of 720p video at one frame per second. The system relies on a specialized Thinker-Talker architecture powered by a Hybrid-Attention Mixture of Experts framework. The Thinker component manages all multimodal reasoning and text generation, analyzing everything from visual cues to spoken words. Meanwhile, the Talker component seamlessly transforms those internal representations into streaming speech outputs for real-time conversations.

Outperforming Gemini in Audio Benchmarks

Pre-trained on over 100 million hours of native audio-visual data, the new model sets several performance milestones. The flagship Plus version achieved state-of-the-art results across 215 audio and audio-visual subtasks. It outright outperforms Google’s Gemini 3.1 Pro in general audio understanding, reasoning, speech recognition, and translation, while matching the Google flagship in overall audio-visual comprehension. In addition to audio dominance, the Kursol blog reports that the model matches GPT-5.4 in many core reasoning domains, making it a highly competitive alternative in the broader AI market.

The system brings significant upgrades to language support. Speech recognition now handles 113 languages and dialects, including 74 languages and 39 Chinese dialects. This is a massive jump from the previous generation, which only supported 11 languages and eight dialects. The model also generates speech in 36 languages and dialects, offering 55 different voices. In tests evaluating multilingual voice stability across 20 languages, it outperformed competitors like ElevenLabs, GPT-Audio, and Minimax.

Audio-Visual Vibe Coding and Real-Time Voice

A unique emergent capability discovered during the model’s training process is a feature the team calls audio-visual vibe coding. Without relying on traditional text prompts, developers can use a camera to show a software interface or a physical object, speak their instructions out loud, and the model will generate functional code to address the request. By processing the visual evidence and spoken intent simultaneously in a single pass, the system seamlessly writes code directly from video and voice inputs. In one demonstration, the model built a working snake game from a brief verbal description and a video clip.

To handle real-time voice interactions smoothly, Alibaba developed the Adaptive Rate Interleave Alignment technique. Because text and speech tokens process at different speeds, streaming voice AI often suffers from dropped words or stuttering. This new alignment method dynamically synchronizes text and speech units, improving the naturalness of the voice output without increasing delay or sacrificing performance.

The update also introduces native semantic interruption for voice assistants. The AI can intelligently distinguish between harmless background noise, simple listener feedback, and an actual attempt by the user to interrupt the conversation. This allows for more natural, human-like turn-taking without the AI stopping its thought process prematurely. Additionally, the system includes built-in live web search to answer current questions without relying on separate pipelines, alongside custom voice cloning capabilities that allow users to generate custom voices from short reference clips.

Conflicting Reports on Open-Source Availability

Reports conflict regarding the model’s availability to the public. According to the news outlet The Decoder and a briefing by The Information, Alibaba has not released the model weights openly, making Qwen3.5-Omni accessible only as a paid API service. However, the Build Fast with AI blog states that while the Plus and Flash versions are limited to Alibaba Cloud’s DashScope API, the Light variant is available as open weights on Hugging Face.

Leadership Changes at Alibaba

This major technical release arrives during a period of internal change at Alibaba. The Decoder reports that Junyang Lin, the chief AI developer behind the Qwen series, recently announced his sudden departure alongside other key team leaders. The exits reportedly stem from a management shakeup involving a new researcher hired from Google’s Gemini team. In response, Alibaba CEO Eddie Wu announced a new Foundation Model Task Force to maintain the company’s strategic focus on AI development.

TAGGED: AI voice assistants, Alibaba, Artificial Intelligence, Gemini 3.1 Pro, machine learning, multimodal AI, Qwen3.5-Omni
Share This Article
Facebook Twitter Whatsapp Whatsapp Telegram Copy Link
By Rakesh Paul
I'm the Co-Founder of VellaTimes and an experienced digital marketer. With substantial experience in the blogging industry, I love crafting insightful and engaging news articles on technology, sports, and automobiles.
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *


Most Read

Quantum Gravity Experiments Challenge Classical Physics

March 11, 2026

Microsoft AI Investment in Australia Reaches Record $25 Billion

April 23, 2026

Microsoft 365 Copilot Update: Cowork and Multi-Model AI

April 1, 2026

PixVerse real-time AI video tool launches, rivals Sora

January 14, 2026

Agentic AI in Insurance: How Tech Is Reshaping the Industry

February 12, 2026

Sun Unleashes Powerful X8.1 Flare as Giant Sunspot Region AR4366 Erupts

February 4, 2026

Related News

Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
News

Espresso Extraction Science: The Finer Grind Flaw

Nisha Pradhan Nisha Pradhan May 18, 2026
A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
News

Amazon Alexa for Shopping Replaces Rufus AI Assistant

Sameer Katoch Sameer Katoch May 18, 2026
Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
News

OpenAI acquires Weights.gg to boost voice AI tools

Rakesh Paul Rakesh Paul May 18, 2026

About Us

VellaTimesVellaTimesVellaTimes

VellaTimes is a leading news portal that covers the latest trending news in technology, lifestyle, entertainment, automobiles, travel, and sports.

Explore

  • News
  • Technology
  • AI
  • Science
  • World

Useful Links

  • About Us
  • Contact Us
  • Fact Checking Policy
  • Terms & Conditions
  • Privacy Policy
  • Copyright Policy

Subscribe Us

Subscribe to our newsletter for the Latest News and Top Stories!

© 2022 VellaTimes • All Rights Reserved.
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
adbanner
AdBlocker Detected
Our site is an advertising supported site. Please whitelist us to support our work.
Okay, I'll Whitelist