By using this site, you agree to our Privacy Policy and Terms of Use.
Accept
VellaTimesVellaTimesVellaTimes
  • News
    NewsShow More
    Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
    Espresso Extraction Science: The Finer Grind Flaw
    May 18, 2026
    A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
    Amazon Alexa for Shopping Replaces Rufus AI Assistant
    May 18, 2026
    Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
    OpenAI acquires Weights.gg to boost voice AI tools
    May 18, 2026
    Federal agents standing outside a modern university biology laboratory building at dusk during an active investigation.
    US Arrests Chinese Scientists for Smuggling Biological Materials
    May 18, 2026
    A dramatically lit modern corporate courtroom with futuristic technology elements, representing a high-stakes artificial intelligence legal trial.
    Elon Musk OpenAI Lawsuit Exposes Clashes Over AI Safety
    May 18, 2026
  • Technology
    TechnologyShow More
    Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
    OpenAI acquires Weights.gg to boost voice AI tools
    May 18, 2026
    A polished silicon wafer rests on a surface inside a modern semiconductor manufacturing facility.
    Samsung Strike Threatens Global AI Chip Production
    May 18, 2026
    A glowing computer screen displaying the text GPT-5.5 Instant in a modern, high-tech office environment with soft blue and purple lighting.
    GPT-5.5 Instant: OpenAI’s New Default ChatGPT Model
    May 10, 2026
    Wide view of a modern AI data center with server racks, glowing fiber-optic cables, and semiconductor hardware in the foreground.
    AI Infrastructure Spending Drives Nvidia, AMD Shares
    May 10, 2026
    A glowing computer monitor displaying lines of code and digital network graphics in a modern tech office setting.
    Airbnb AI Coding: 60% of New Software Now Generated by AI
    May 9, 2026
  • AI
    AIShow More
    A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
    Amazon Alexa for Shopping Replaces Rufus AI Assistant
    May 18, 2026
    A dramatically lit modern corporate courtroom with futuristic technology elements, representing a high-stakes artificial intelligence legal trial.
    Elon Musk OpenAI Lawsuit Exposes Clashes Over AI Safety
    May 18, 2026
    A high-tech global map visualization showing glowing digital connections across different continents, representing the worldwide adoption of artificial intelligence.
    Global AI Adoption in 2026: Trends and Growing Divide
    May 10, 2026
    A modern smartphone displaying an artificial intelligence chat interface used for online shopping and product comparison.
    Alibaba Qwen AI Taobao Integration Launches Agentic Shopping
    May 10, 2026
    A split-screen illustration showing a high-tech modern office using advanced AI tools contrasted against an older, dimly lit workspace.
    Global AI Adoption Surges But Rich-Poor Divide Widens
    May 9, 2026
  • Science
    ScienceShow More
    Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
    Espresso Extraction Science: The Finer Grind Flaw
    May 18, 2026
    Federal agents standing outside a modern university biology laboratory building at dusk during an active investigation.
    US Arrests Chinese Scientists for Smuggling Biological Materials
    May 18, 2026
    Header image of a quantum communication lab setup with fiber-optic equipment, a telecom quantum dot device, and interferometer components used for long-distance quantum key distribution.
    Quantum Key Distribution Reaches 120 km With Quantum Dots
    May 10, 2026
    Abstract geometric representation of glowing quantum paraparticles interacting within a three-dimensional mathematical grid in deep blue and gold tones.
    Quantum Paraparticles Exist: New Math Challenges Physics
    May 10, 2026
    A large expedition cruise ship is navigating rough ocean waters under a cloudy sky.
    Global Authorities Respond to Andes Hantavirus Outbreak on MV Hondius Cruise Ship
    May 9, 2026
  • World
    WorldShow More
    Allu Arjun Commitment to Ethical Brand Partnerships
    Exploring Allu Arjun’s Commitment to Ethical Brand Partnerships
    December 18, 2023
    Orry aka Orhan Awatramani
    Orhan Awatramani ‘Orry’ Biography, Lifestyle and Rise to Fame
    December 8, 2023
    Alia Bhatt Latest Deepake Video Victim
    Alia Bhatt becomes latest victim of Deepfake Videos, Obscene Video goes Viral
    November 28, 2023
    Napoleon Movie Review
    Napoleon Movie Review: A Historical Epic by Ridley Scott Reviewed
    November 25, 2023
  • Bookmarks
Search
Category
  • News
  • Technology
  • AI
  • Science
  • World
Company
  • About Us
  • Contact Us
  • Fact Checking Policy
  • Terms & Conditions
  • Privacy Policy
  • Copyright Policy
Resources
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
© 2022 VellaTimes • All Rights Reserved.
Reading: Microsoft Releases Phi-4-Reasoning-Vision-15B: A Compact Multimodal AI
Share
Notification Show More
Font ResizerAa
VellaTimesVellaTimes
Font ResizerAa
  • News
  • Technology
  • AI
  • Science
  • World
Search
  • Explore
    • News
    • Technology
    • AI
    • Science
    • World
  • Useful Links
    • About Us
    • Contact Us
    • Fact Checking Policy
    • Terms & Conditions
    • Privacy Policy
    • Copyright Policy
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
© 2022 VellaTimes • All Rights Reserved.
News

Microsoft Releases Phi-4-Reasoning-Vision-15B: A Compact Multimodal AI

Sameer Katoch
Last updated: 09/03/2026
Sameer Katoch
Share
6 Min Read
A glowing holographic display showing mathematical equations and scientific charts in a modern, brightly lit server room.

Microsoft has officially launched Phi-4-reasoning-vision-15B, an open-weight multimodal artificial intelligence model featuring 15 billion parameters. This new system is specifically tailored for vision-language applications, allowing the AI to effectively process and analyze both text and images. The model demonstrates exceptional capabilities in generating image captions, analyzing complex documents, and performing mathematical and scientific reasoning based on visual inputs.

Contents
A New Approach to Hybrid ReasoningMid-Fusion Architecture and Hardware EfficiencyTraining Process and Data RefinementOutperforming Larger Models on BenchmarksPowering AI Agents and Visual Analysis

By introducing a compact, hardware-efficient system, Microsoft aims to provide a powerful alternative to larger, resource-heavy AI models. Phi-4-reasoning-vision-15B stands out for its unique hybrid reasoning capabilities. This architecture allows the model to actively decide when a task requires a multi-step thought process and when a direct, straightforward answer is sufficient, saving valuable computing time.

A New Approach to Hybrid Reasoning

One of the most defining features of this model is its mixed reasoning and non-reasoning training strategy. Instead of forcing the AI to use a complex chain-of-thought process for every single prompt, Microsoft trained the system to alternate seamlessly between two distinct modes.

For complicated challenges, such as mathematical problems or scientific chart evaluations, the model activates a “think” mode. This generates structured, multi-step reasoning traces to ensure high accuracy. For simpler, perception-focused tasks like basic image captioning or optical character recognition, it relies on a “no-think” mode to provide immediate, low-latency responses.

This hybrid setup was achieved by making reasoning data approximately 20 percent of the overall training mixture. While the AI learns the boundary between these modes implicitly, users retain full control. Developers can override the default behavior by using specific prompt tags to force the model into either state depending on their exact needs.

Mid-Fusion Architecture and Hardware Efficiency

To successfully balance performance with compute costs, Microsoft researchers built the model using a mid-fusion architecture. The system physically combines two existing algorithms: the SigLIP-2 vision encoder and the previously released Phi-4-Reasoning language model.

SigLIP-2 operates by compressing images into a numerical format, generating visual tokens that neural networks can easily understand. These tokens are then projected into the language model’s embedding space. In a mid-fusion setup, only some of the model’s layers support multimodal processing, unlike early-fusion designs where every layer handles multimodal data.

This strategic design trades a minimal amount of output quality for a massive reduction in hardware usage. To lower the infrastructure footprint even further, users can completely disable the reasoning feature via prompts if they want to prioritize pure speed. The model’s dynamic resolution vision encoder supports up to 3,600 visual tokens, ensuring detailed high-resolution perception without the sluggish latency often found in larger models.

Training Process and Data Refinement

Microsoft managed to train the model efficiently over just four days using 240 B200 GPUs. The model processed 200 billion multimodal tokens during its training phase. This is a mere fraction of the trillion-plus tokens required to train other recent multimodal models currently on the market.

The training data primarily consisted of open-source image and text collections, but Microsoft heavily refined this data through a multi-step process. High-quality datasets were preserved, while images featuring inaccurate captions were given entirely new, corrected descriptions generated by GPT-4o and o4-mini. The researchers also enriched the training mix with internally created data, targeted acquisitions, and specific safety datasets designed to prevent harmful outputs.

Outperforming Larger Models on Benchmarks

Despite its highly compact size, the model achieved impressive results across numerous open-source evaluations using testing frameworks like Eureka ML Insights and VLMEvalKit. On the MathVista_Mini benchmark, which specifically tests multimodal mathematics, the model scored 75.2, outperforming Google’s gemma-3-12b-it by a significant 17 percent margin.

The model also recorded notable scores on several other comprehensive tests. It achieved an 84.8 on the AI2D_TEST, an 83.3 on ChartQATEST, an 88.2 on ScreenSpotv2, and a 76.0 on OCRBench. Microsoft researchers note that the model delivers better accuracy than similarly fast models and offers highly competitive performance against slower models that require ten times more computing power.

Powering AI Agents and Visual Analysis

With its ability to accurately detect graphical user interface elements, the model is exceptionally well-suited for computer-use agents. It can interpret screen content, deduce the exact functions of buttons and menus from standard screenshots, and provide precise click coordinates for automation. This functionality makes it an ideal base model for navigating web, mobile, and desktop interfaces.

The system also excels at analyzing highly complicated visual assets. In a demonstration shared by Microsoft, a user uploaded a photograph of a tilted Saturn. The model accurately explained that the planet’s orientation depended entirely on the time of year and the specific position of the telescope used to capture the image. Developers can now access the model’s code directly through Hugging Face, GitHub, and Azure AI Foundry.

TAGGED: AI agents, Artificial Intelligence, machine learning, Microsoft AI, multimodal AI, open-source AI, tech news
Share This Article
Facebook Twitter Whatsapp Whatsapp Telegram Copy Link
By Sameer Katoch
As the Founder of VellaTimes and an avid traveler, I'm passionate about the daily news events happening globally. With over five years of experience in the writing field, I am committed to delivering top-notch news that satisfies your daily news intake.
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *


Most Read

Starlink Privacy Policy Update Allows AI Training on Customer Data

February 2, 2026

Nvidia Reaches 5 Trillion Market Cap as AI Chip Rally Accelerates

April 27, 2026

Spain Storms Push Rainfall to Worst Levels in 47 Years

March 13, 2026

Tesla Cybertruck Price Cut: New Base Trim Saves Buyers $20,000

February 22, 2026

Meta Corning fiber optic deal: up to $6B agreement

January 28, 2026

US-UAE AI campus deal faces security hurdles in Abu Dhabi

January 28, 2026

Related News

Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
News

Espresso Extraction Science: The Finer Grind Flaw

Nisha Pradhan Nisha Pradhan May 18, 2026
A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
News

Amazon Alexa for Shopping Replaces Rufus AI Assistant

Sameer Katoch Sameer Katoch May 18, 2026
Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
News

OpenAI acquires Weights.gg to boost voice AI tools

Rakesh Paul Rakesh Paul May 18, 2026

About Us

VellaTimesVellaTimesVellaTimes

VellaTimes is a leading news portal that covers the latest trending news in technology, lifestyle, entertainment, automobiles, travel, and sports.

Explore

  • News
  • Technology
  • AI
  • Science
  • World

Useful Links

  • About Us
  • Contact Us
  • Fact Checking Policy
  • Terms & Conditions
  • Privacy Policy
  • Copyright Policy

Subscribe Us

Subscribe to our newsletter for the Latest News and Top Stories!

© 2022 VellaTimes • All Rights Reserved.
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
adbanner
AdBlocker Detected
Our site is an advertising supported site. Please whitelist us to support our work.
Okay, I'll Whitelist