On September 2, 2026, Indian outlets reported that Meta’s new Muse Voice Transcribe model is now available with native support for five major Indian languages alongside English. The real‑time speech‑to‑text system, developed by Meta Superintelligence Labs, offers streaming transcription, diarisation for over 20 speakers and code‑switching, and is accessible via the Meta Model API and Meta AI products.
This article aggregates reporting from 4 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
Muse Voice Transcribe is Meta’s bid to own real‑time speech as aggressively as it is chasing image and video generation. A single model that can stream transcription, separate 20‑plus speakers, handle long recordings and cleanly code‑switch between English and major Indian languages is a serious step up from today’s patchwork of ASR and diarisation systems. It also plugs directly into Meta AI for Mac, Muse Code and the Meta Model API, making high‑quality voice interfaces much easier to bolt onto apps.
From an AGI perspective, high‑fidelity, low‑latency speech is infrastructure. Agentic systems that can listen continuously, keep track of who is speaking, and respond in near real time start to feel less like chatbots and more like collaborators. Doing that across 70‑plus languages, and particularly in code‑switched Indian contexts, broadens who can realistically participate in the frontier model economy without shifting to English.
Strategically, this keeps Meta on the front foot in India, where language fragmentation has been a moat for local players and for Google’s speech stack. By anchoring Muse Voice Transcribe in Indian use cases and pushing it via widely read local tech media, Meta is signaling that its AGI roadmap is not just about bigger text‑vision models, but about deeply embedding itself in everyday communication flows for hundreds of millions of users.


