AI News Feed
Market watch
Products & Applications

Meta launches Muse Voice Transcribe for real-time dictation on Mac

Meta has introduced Muse Voice Transcribe, a real-time audio perception model bringing multilingual streaming dictation to Mac and developers.

The model combines streaming automatic speech recognition with speaker diarization and endpointing, the report says. It can transcribe speech as it happens, separate speakers across recordings with more than 20 voices, and determine when someone has finished talking, all without a separate post-processing step. It was trained in over 70 languages, with 25 validated at launch. It supports audio longer than an hour, plus native code-switching within or between sentences. Language, keyword, and context biasing can further improve recognition.

Instead of a fixed tradeoff between speed and accuracy, Muse Voice Transcribe decides how long to listen before committing each word, a feature Meta calls “adaptive delay.” The system can move quickly through easier speech while using more audio context for difficult words. Meta says the model ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1.

Muse Voice Transcribe is available today through the Meta Model API at $3 per 1,000 audio-minutes, equivalent to $0.18 per hour. It already powers dictation in Meta AI for Mac and Muse Code. On Mac, users can hold the Fn key to dictate into any application.