AI News Feed
Market watch
Products & Applications

Google launches Gemini Audio models for real-time dialogue and transcription

Google unveiled Gemini Audio, three new AI models for voice recognition and transcription, rolling out across apps like Gemini, Gboard, and Chrome.

Gemini 3.5 Live is an upgraded version of the voice mode that previously ran on Gemini 3.1 Live. According to TechRadar, the new model handles mid-sentence interruptions better, processes live visuals, blends multiple languages on the fly, and can trigger background tools. Android Authority reported that Gemini 3.5 Live Experimental goes further by narrating its reasoning step by step in real time while tackling more complex tasks.

Gemini 3.5 Transcribe is a new model focused on speech-to-text. The Verge quoted Google as saying the model represents a major advancement over its previous transcription model, Chirp 3, and noted it automatically detects more than 85 languages. Engadget added that it can turn unstructured speech into formatted text, remove filler words such as “um” and “ah,” adapt to custom vocabulary and unique spellings, and capture alphanumeric strings like order numbers and postal codes. It can also attribute speech to up to three speakers in pre-recorded audio and provide word-level timestamps.

The rollout has begun for some products but not all. The Verge reported that the Gemini Audio updates are available starting today in English for all macOS Gemini app users, and the Rambler dictation feature on Android in select countries and languages. TechRadar, however, said it was not clear when the models would reach everyone, adding that it expected a broader rollout over the next few days. Engadget said Chrome support is coming soon, allowing users to dictate replies and posts in any web field. The models are also set to arrive in Search Live, Gemini Live, Docs, Keep, and Gmail, according to Android Authority. Developers can access Gemini 3.5 Live Experimental and Gemini 3.5 Transcribe through the Gemini API in AI Studio and Antigravity, as reported by The Verge and Engadget.

Google also detailed new productivity features for Gemini Live, expanding it beyond conversation. Android Authority said Gemini Live can now use an AI agent called Gemini Spark to turn natural voice requests into multi-step tasks. Other additions include Daily Brief, which offers proactive, personalized updates from Google apps, hands-free Gmail management, and Personal Intelligence. Some features require subscriptions: Spark is limited to Google AI Pro and higher, while Daily Brief requires Google AI Plus or higher.

The expansion has drawn criticism about how Google has branded its AI features. TechCrunch wrote that the company’s message — “You shouldn’t have to guess whether a task requires Spark, a Daily Brief, or a quick inbox search” — is undercut by the fact that Gemini users must switch between separate features with their own icons and navigation. TechCrunch also argued that the AI industry as a whole exposes internal architecture to consumers, citing Anthropic’s Claude and OpenAI’s ChatGPT as examples. The publication noted that Apple’s simpler approach to Siri could win consumers, and quoted a16z investment partner Justine Moore as saying, “People don’t want to open an app every time they need help – they want a contact they can text like a friend. And the gold standard is iMessage.”

The Gemini Audio announcement comes several months after Google first introduced the Gemini 3.5 family. The company has since moved on to Gemini 3.6 and 3.7, but it is still expanding 3.5. The Verge noted the new models arrive while users are still waiting for the promised Gemini 3.5 Pro, which Google had said would roll out in June. According to Android Authority, Google said that 63% of Gemini users talk to the assistant out loud, which appears to have driven the company to make Live more than just a conversational mode.