Back to Frontiers

Multimodal AI

High75%

Vision-language-audio unification, cross-modal understanding, and unified sequence modeling. Making AI see, hear, and understand the world.

vision-languageGPT-4VGeminiCLIPmultimodalvideo understandingaudio
178
Papers
35
Milestones
$1.0B
Funding
2
Benchmarks

Key Benchmarks

Video-MME

Comprehensive video understanding benchmark

81.6%Human: 92%
Leader: video-SALMONN 2+medium saturation

MMMU

Massive Multi-discipline Multimodal Understanding benchmark

65.9%Human: 88.6%
Leader: KiQ-v0low saturation

Recent Papers

Recent Milestones

MiniMax hits near real-time AI video for streams

Chinese lab MiniMax’s H3 Max video model can generate a five second, 768p video clip in under three seconds, according to a September 1, 2026 report from QbitAI. Developers have already used the model to run AI-generated livestreams where content is created on the fly.

Sep 1, 2026releaseImpact: 70/100

DeepSeek V4 Flash Vision 305B weights go fully open

DeepSeek published MIT-licensed open weights for its 305B-parameter multimodal model DeepSeek‑V4‑Flash‑Vision‑Exp on August 31, 2026. Jiufeng’s September 1, 2026 briefing highlights that this is the first V4 vision model and that the checkpoint, tokenizer and reference inference code are now downloadable from Hugging Face.

Sep 1, 2026releaseImpact: 90/100

Alibaba launches HappyShrimp AI music platform

On August 30, 2026 Chinese media report that Alibaba has launched HappyShrimp 1.0, an AI music generation model that brings back the retired Xiami brand as an AI‑native service. The product went live on PC web in China and overseas and announced a same‑day strategic partnership with Taihe Music Group covering licensing and musician co‑creation.

Aug 30, 2026releaseImpact: 70/100

Chinese labs seize lead in AI video benchmarks

On August 29, 2026, the South China Morning Post reported that Chinese AI video models from Alibaba, MiniMax and ByteDance now hold eight of the top 10 spots on benchmark platform Artificial Analysis. Alibaba’s Wan 3.0 leads the text‑to‑video leaderboard with audio, ahead of Google’s Gemini Omni Flash, with MiniMax H3 variants and ByteDance’s Seedance 2.0 also scoring highly.

Aug 29, 2026benchmarkImpact: 70/100

Stability AI Gets $76M From Music & Gaming Majors

Stability AI announced on August 26, 2026 that it closed a $76 million Series B round led by major entertainment and tech investors including Electronic Arts, Sony Music Group, Universal Music Group and Warner Music Group. The Los Angeles based generative AI company said the new capital lifts its total funding to $232 million and will fund creative AI products for music, gaming and entertainment professionals.

Aug 26, 2026fundingImpact: 70/100

Apple Intelligence Lands in China via Alibaba, Baidu

On July 15, 2026, Techlusive reported that China’s Cyberspace Administration has approved Apple Intelligence for rollout on iPhones, iPads, Macs and Vision Pro in mainland China. Apple will rely on Alibaba’s Qwen model and Baidu’s AI services to meet local generative AI rules, after months of delay in the Chinese market.

Jul 15, 2026releaseImpact: 70/100

PixVerse Owner AIsphere Lands $4.4B for AI Video

Chinese AI video startup AIsphere (PixVerse) said on July 14, 2026 it has completed a cumulative RMB 2.98 billion Series C round, including a new C+ tranche led by Alibaba. The funding will deepen its video generation foundation models and real‑time “world model” research while accelerating global product growth and commercial deployments.

Jul 14, 2026fundingImpact: 80/100

OpenAI rolls out GPT‑Live real‑time voice

OpenAI launched GPT-Live on July 8, 2026, introducing new GPT‑Live‑1 and GPT‑Live‑1 mini voice models that power a more natural real-time ‘ChatGPT Live’ experience across its apps. The update lets ChatGPT listen and speak simultaneously, handle interruptions, and shows visual responses while adding safeguards to prevent voice impersonation.

Jul 8, 2026releaseImpact: 80/100

Meta debuts Muse Image for social apps

Meta introduced Muse Image on July 7, 2026, its first in‑house image‑generation model built by Meta Superintelligence Labs. The system now powers image creation in Meta AI, Instagram Stories and WhatsApp chats, with more than 30 new AI effects and a preview of a coming Muse Video model. ([tech.yahoo.com](https://tech.yahoo.com/ai/meta-ai/articles/meta-launches-muse-image-first-185347805.html?utm_source=openai))

Jul 7, 2026releaseImpact: 70/100

Tripo AI Secures $150M for 3D World Models

On July 6, 2026, Tripo AI announced it had raised $150 million in a Series A3 round to deepen work on 3D foundation models and world-model technologies. The financing drew strategic investment from Geely Capital and several Chinese gaming companies, including 4399 Network, Tanwan and Giant Network, alongside existing backers INCE Capital and Genesis Capital.

Jul 6, 2026fundingImpact: 70/100

xAI completes Grok Imagine image+video suite

Elon Musk said “Done with Grok Imagine” on July 5, signaling that xAI has completed development of its Grok Imagine image and video generation feature. The multimodal tool, already in beta, is now positioned as a fully built‑out creative suite integrated across Grok and the X app.

Jul 5, 2026releaseImpact: 70/100

Kling AI raises $3B to challenge Sora

Pandaily reports that Chinese short-video giant Kuaishou is spinning off its Kling AI video generation unit into an independent company and backing it with a funding round of up to $3 billion. Tencent, Alibaba and Baidu are named as key investors, and the new entity reportedly faces a mandatory IPO deadline in 2031.

Jul 4, 2026fundingImpact: 80/100

China’s Kling AI Scores $2B for Video Models

Kuaishou’s video-generation unit Kling AI has closed an initial $2 billion funding round at around a $15 billion pre-money valuation. The round includes Alibaba, Tencent, Baidu and several major Chinese and Gulf investment firms, and could expand to $3 billion, cutting Kuaishou’s stake to about 68%.

Jul 3, 2026fundingImpact: 70/100

Google Drives AI Image Costs Toward Pennies

Google has rolled out its Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) model, a fast low-cost image generator now available via Gemini API, Google AI Studio and the Gemini Enterprise Agent Platform. The model generates 1K-resolution images in about four seconds at roughly $0.034 per image, with rollouts and coverage reported across India and Japan on July 1, 2026.([gadgets360.com](https://www.gadgets360.com/ai/news/google-personalised-gemini-intelligence-ai-image-creation-nano-banana-photos-free-11711303?utm_source=openai))

Jul 1, 2026releaseImpact: 70/100

Gemini Makes Personal Image AI Free in US

Google made Gemini's personalized Nano Banana-powered image generation free for eligible U.S. users via the Gemini app on June 29, 2026.([techcrunch.com](https://techcrunch.com/2026/06/29/geminis-personalized-ai-image-generation-is-now-free-for-u-s-users/)) The feature, previously limited to paid tiers, uses data from services like Gmail, Photos and Search to create images tailored to each user’s interests and photos.([techcrunch.com](https://techcrunch.com/2026/06/29/geminis-personalized-ai-image-generation-is-now-free-for-u-s-users/))

Jun 29, 2026releaseImpact: 70/100

L’Oréal turns ChatGPT into AI beauty counter

L’Oréal announced a strategic partnership with OpenAI at VivaTech 2026, enabling Maybelline New York’s ModiFace-powered virtual makeup try-on directly inside ChatGPT. The deal also includes promoting L’Oréal brands via product ads in ChatGPT in the US, with further integrations for Lancôme and Kérastase to follow.

Jun 19, 2026releaseImpact: 70/100

Gemma 4 12B puts open multimodal AI on laptops

Regional coverage in Chinese and Korean tech media highlights Google’s new Gemma 4 12B model, a 12B‑parameter unified multimodal model released under Apache 2.0 that can run on laptops with 16GB of memory.([blog.google](https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/?utm_source=openai)) The model removes heavy vision and audio encoders, enabling local image, audio and video understanding on consumer hardware.

Jun 4, 2026releaseImpact: 80/100

Reactor raises $59M for real‑time AI world models

On May 28, 2026 AWS announced that startup Reactor emerged from stealth with $59 million in funding led by Lightspeed Venture Partners. The company is building a platform for real‑time “world model”–based generative video and interactive experiences, with early adopters like Overworld already building on it.

May 28, 2026fundingImpact: 70/100

Microsoft’s MAI-Image-2.5 hits No.3 on Arena

On May 26–27, 2026, Microsoft announced MAI‑Image‑2.5, a new in‑house image generation model that ranks third on Arena’s text‑to‑image leaderboard. Japanese outlet GIGAZINE reported on May 27 that the model will roll out to the MAI Playground within about two weeks.([gigazine.net](https://gigazine.net/gsc_news/en/20260527-mai-image-25-generation-ai/))

May 26, 2026benchmarkImpact: 70/100

Microsoft debuts in-house MAI speech & image AIs

On April 4, 2026, Tech Insider detailed how Microsoft has launched three in‑house MAI models — MAI‑Transcribe‑1, MAI‑Voice‑1 and MAI‑Image‑2 — following their April 2 release. The models target speech recognition, voice generation and image creation, and are being rolled out via Microsoft’s Foundry and MAI Playground platforms.

Apr 4, 2026releaseImpact: 70/100

Leading Organizations

Google
OpenAI
Meta
Anthropic

ArXiv Categories

cs.CVcs.CLcs.LGcs.MM

Related Frontiers