Multimodal AI
Vision-language-audio unification, cross-modal understanding, and unified sequence modeling. Making AI see, hear, and understand the world.
Key Benchmarks
Recent Papers
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Kang Liao, Yihang Luo, Xiao-Ming Wu +7 more
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
Puneet Mathur, Dinesh Manocha
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
Howard Qian, Yiting Chen, Yunfei Xie +6 more
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng +7 more
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Hanyang Wang, Yimo Cai, Weiliang Chen +14 more
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Hanyang Wang, Yimo Cai, Weiliang Chen +14 more
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
Suhwan Choi, Jaeyoon Jung, Sungkyung Kim +2 more
Beyond Retrieval: Analytic Memory for Multimodal Agents
Zhoujin Tian, Yao Tian, Hao Zhang +4 more
A Human-Centered Validation of the Explainability-Performance Coefficient
Christian Oliva, Luis F. Lago-Fernández
TerraNova: A Foundation Model for the Anthropocene
Carlos Rodriguez-Pardo, Massimo Tavoni
Recent Milestones
MiniMax hits near real-time AI video for streams
Chinese lab MiniMax’s H3 Max video model can generate a five second, 768p video clip in under three seconds, according to a September 1, 2026 report from QbitAI. Developers have already used the model to run AI-generated livestreams where content is created on the fly.
DeepSeek V4 Flash Vision 305B weights go fully open
DeepSeek published MIT-licensed open weights for its 305B-parameter multimodal model DeepSeek‑V4‑Flash‑Vision‑Exp on August 31, 2026. Jiufeng’s September 1, 2026 briefing highlights that this is the first V4 vision model and that the checkpoint, tokenizer and reference inference code are now downloadable from Hugging Face.
Alibaba launches HappyShrimp AI music platform
On August 30, 2026 Chinese media report that Alibaba has launched HappyShrimp 1.0, an AI music generation model that brings back the retired Xiami brand as an AI‑native service. The product went live on PC web in China and overseas and announced a same‑day strategic partnership with Taihe Music Group covering licensing and musician co‑creation.
Chinese labs seize lead in AI video benchmarks
On August 29, 2026, the South China Morning Post reported that Chinese AI video models from Alibaba, MiniMax and ByteDance now hold eight of the top 10 spots on benchmark platform Artificial Analysis. Alibaba’s Wan 3.0 leads the text‑to‑video leaderboard with audio, ahead of Google’s Gemini Omni Flash, with MiniMax H3 variants and ByteDance’s Seedance 2.0 also scoring highly.
Stability AI Gets $76M From Music & Gaming Majors
Stability AI announced on August 26, 2026 that it closed a $76 million Series B round led by major entertainment and tech investors including Electronic Arts, Sony Music Group, Universal Music Group and Warner Music Group. The Los Angeles based generative AI company said the new capital lifts its total funding to $232 million and will fund creative AI products for music, gaming and entertainment professionals.
Apple Intelligence Lands in China via Alibaba, Baidu
On July 15, 2026, Techlusive reported that China’s Cyberspace Administration has approved Apple Intelligence for rollout on iPhones, iPads, Macs and Vision Pro in mainland China. Apple will rely on Alibaba’s Qwen model and Baidu’s AI services to meet local generative AI rules, after months of delay in the Chinese market.
PixVerse Owner AIsphere Lands $4.4B for AI Video
Chinese AI video startup AIsphere (PixVerse) said on July 14, 2026 it has completed a cumulative RMB 2.98 billion Series C round, including a new C+ tranche led by Alibaba. The funding will deepen its video generation foundation models and real‑time “world model” research while accelerating global product growth and commercial deployments.
OpenAI rolls out GPT‑Live real‑time voice
OpenAI launched GPT-Live on July 8, 2026, introducing new GPT‑Live‑1 and GPT‑Live‑1 mini voice models that power a more natural real-time ‘ChatGPT Live’ experience across its apps. The update lets ChatGPT listen and speak simultaneously, handle interruptions, and shows visual responses while adding safeguards to prevent voice impersonation.
Meta debuts Muse Image for social apps
Meta introduced Muse Image on July 7, 2026, its first in‑house image‑generation model built by Meta Superintelligence Labs. The system now powers image creation in Meta AI, Instagram Stories and WhatsApp chats, with more than 30 new AI effects and a preview of a coming Muse Video model. ([tech.yahoo.com](https://tech.yahoo.com/ai/meta-ai/articles/meta-launches-muse-image-first-185347805.html?utm_source=openai))
Tripo AI Secures $150M for 3D World Models
On July 6, 2026, Tripo AI announced it had raised $150 million in a Series A3 round to deepen work on 3D foundation models and world-model technologies. The financing drew strategic investment from Geely Capital and several Chinese gaming companies, including 4399 Network, Tanwan and Giant Network, alongside existing backers INCE Capital and Genesis Capital.
xAI completes Grok Imagine image+video suite
Elon Musk said “Done with Grok Imagine” on July 5, signaling that xAI has completed development of its Grok Imagine image and video generation feature. The multimodal tool, already in beta, is now positioned as a fully built‑out creative suite integrated across Grok and the X app.
Kling AI raises $3B to challenge Sora
Pandaily reports that Chinese short-video giant Kuaishou is spinning off its Kling AI video generation unit into an independent company and backing it with a funding round of up to $3 billion. Tencent, Alibaba and Baidu are named as key investors, and the new entity reportedly faces a mandatory IPO deadline in 2031.
China’s Kling AI Scores $2B for Video Models
Kuaishou’s video-generation unit Kling AI has closed an initial $2 billion funding round at around a $15 billion pre-money valuation. The round includes Alibaba, Tencent, Baidu and several major Chinese and Gulf investment firms, and could expand to $3 billion, cutting Kuaishou’s stake to about 68%.
Google Drives AI Image Costs Toward Pennies
Google has rolled out its Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) model, a fast low-cost image generator now available via Gemini API, Google AI Studio and the Gemini Enterprise Agent Platform. The model generates 1K-resolution images in about four seconds at roughly $0.034 per image, with rollouts and coverage reported across India and Japan on July 1, 2026.([gadgets360.com](https://www.gadgets360.com/ai/news/google-personalised-gemini-intelligence-ai-image-creation-nano-banana-photos-free-11711303?utm_source=openai))
Gemini Makes Personal Image AI Free in US
Google made Gemini's personalized Nano Banana-powered image generation free for eligible U.S. users via the Gemini app on June 29, 2026.([techcrunch.com](https://techcrunch.com/2026/06/29/geminis-personalized-ai-image-generation-is-now-free-for-u-s-users/)) The feature, previously limited to paid tiers, uses data from services like Gmail, Photos and Search to create images tailored to each user’s interests and photos.([techcrunch.com](https://techcrunch.com/2026/06/29/geminis-personalized-ai-image-generation-is-now-free-for-u-s-users/))
L’Oréal turns ChatGPT into AI beauty counter
L’Oréal announced a strategic partnership with OpenAI at VivaTech 2026, enabling Maybelline New York’s ModiFace-powered virtual makeup try-on directly inside ChatGPT. The deal also includes promoting L’Oréal brands via product ads in ChatGPT in the US, with further integrations for Lancôme and Kérastase to follow.
Gemma 4 12B puts open multimodal AI on laptops
Regional coverage in Chinese and Korean tech media highlights Google’s new Gemma 4 12B model, a 12B‑parameter unified multimodal model released under Apache 2.0 that can run on laptops with 16GB of memory.([blog.google](https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/?utm_source=openai)) The model removes heavy vision and audio encoders, enabling local image, audio and video understanding on consumer hardware.
Reactor raises $59M for real‑time AI world models
On May 28, 2026 AWS announced that startup Reactor emerged from stealth with $59 million in funding led by Lightspeed Venture Partners. The company is building a platform for real‑time “world model”–based generative video and interactive experiences, with early adopters like Overworld already building on it.
Microsoft’s MAI-Image-2.5 hits No.3 on Arena
On May 26–27, 2026, Microsoft announced MAI‑Image‑2.5, a new in‑house image generation model that ranks third on Arena’s text‑to‑image leaderboard. Japanese outlet GIGAZINE reported on May 27 that the model will roll out to the MAI Playground within about two weeks.([gigazine.net](https://gigazine.net/gsc_news/en/20260527-mai-image-25-generation-ai/))
Microsoft debuts in-house MAI speech & image AIs
On April 4, 2026, Tech Insider detailed how Microsoft has launched three in‑house MAI models — MAI‑Transcribe‑1, MAI‑Voice‑1 and MAI‑Image‑2 — following their April 2 release. The models target speech recognition, voice generation and image creation, and are being rolled out via Microsoft’s Foundry and MAI Playground platforms.