New Delhi: Meta is making another push into voice-based artificial intelligence, and this time the move has a clear Indian angle.
The company has introduced a new speech-to-text model that can transcribe conversations in real time, recognise different speakers and work across multiple languages, including five Indian languages.
Meta’s new model, Muse Voice Transcribe, is designed for conversations that do not always follow a neat script. Someone can begin speaking in English, switch to Hindi or another Indian language and continue without stopping to change a setting.
That is particularly relevant in India, where mixing languages in everyday conversation is common rather than unusual.
The model supports Hindi, Tamil, Telugu, Kannada and Malayalam among more than 70 languages. Meta says it can also identify different people speaking in the same conversation, making it possible to produce a transcript that shows not just what was said but who said it.
That could make the technology useful in situations such as meetings, interviews, voice notes and other conversations where waiting for an audio recording to finish before getting a transcript is inconvenient. Muse is built to work while people are still talking, rather than treating transcription as something that happens afterwards.
Meta has also tried to deal with one of the less obvious problems with real-time transcription: speed. If a system waits too long before deciding what someone said, the conversation feels slow. If it responds too quickly, mistakes can creep in.
The company says Muse uses an adaptive approach, giving the model more time when speech is difficult while keeping simpler parts of a conversation moving.
For Indian users, the language support may be the part that gets the most attention. Voice technology has improved rapidly, but understanding the way people actually speak remains difficult. Conversations can include regional accents, English words, local expressions and quick switches between languages.
A system that can handle that naturally has a much better chance of being useful outside controlled demonstrations.
Meta says Muse Voice Transcribe is already being used for dictation in Meta AI for Mac and Muse Code, while developers can access the model through Meta’s Model API. The company has also positioned the technology as part of its broader effort to build AI systems that can understand people through natural conversation rather than relying entirely on typed instructions.
That puts Meta in an increasingly crowded race around voice AI. The competition is no longer just about producing a transcript that is technically correct. The bigger challenge is understanding people when they speak quickly, change languages, interrupt one another and do not speak in perfectly formed sentences.
India could become an important testing ground for that next phase. If AI can understand the way people actually talk at home, at work and on the phone with English mixed into regional languages and several people speaking in the same room it becomes much more than a transcription tool.
For Meta, the latest model is another step towards that goal. For Indian users, the more interesting question is whether the technology will finally make voice AI feel less like a machine listening to a language and more like a system that understands the way people really speak.