Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Priceline Promo Codes & Coupons: 10% Off September 2026

    September 2, 2026

    Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

    September 2, 2026

    SonicWall Warns of Two SMA1000 Zero-Days Exploited in Attacks

    September 2, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Priceline Promo Codes & Coupons: 10% Off September 2026
    • Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
    • SonicWall Warns of Two SMA1000 Zero-Days Exploited in Attacks
    • Binance Adds Options on 1,000 US Stocks and ETFs
    • High blood sugar may help cancer cells hide from the immune system
    • Southeast Asia’s oak trees highly threatened as report warns of lack of protection
    • US Postal Service whistleblower warns new system could disrupt midterm election voting
    • Russlands Gegenreaktion auf Deutschlands Antwort – POLITICO
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Wednesday, September 2
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKSeptember 2, 2026 Artificial Intelligence No Comments6 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Most production voice stacks are three systems stitched together. One model transcribes, a second separates speakers, and a detector decides when the user stopped talking. Each hand-off adds latency and a new failure mode.

    Muse Voice Transcribe, announced by Meta Superintelligence Labs this week, collapses those three jobs into a single autoregressive model. Meta calls it its first real-time audio perception model. It performs streaming ASR, speaker diarization for 20+ speakers, and endpointing in one pass, with no required post-processing.

    Is it deployable? Yes, but only as a hosted API. It is live on the Meta Model API as muse-voice-transcribe-1.0 at $3.00 per 1,000 audio minutes ($0.18 per hour), and it already powers dictation in Meta AI for Mac and Muse Code. No weights have been released, so there is no self-hosted path.

    Streaming ASR as the foundation

    Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family. Audio arrives in 80ms chunks at 12.5 Hz. Each chunk is transformed into a single soft token.

    After every chunk the model makes one binary choice. It either predicts a <|next_audio|> token and keeps listening, or it emits a text token. When the model predicts <|next_audio|>, that token is replaced by the actual next audio chunk in the input. When the stream ends, an <|empty_audio|> token is inserted, and the model flushes all remaining text without requesting more audio.

    Listening and writing share one decoder loop, so there is no separate alignment stage to drift.

    Adaptive delay, trained with RL

    Because the model controls when it listens, it also controls how much audio context sits behind each word. Meta calls that gap ‘delay.’ Longer delay means a more accurate transcript and higher latency.

    Instead of fixing that trade-off, Meta trains it. Reinforcement learning combines a word error rate reward and a delay reward multiplicatively, producing a policy that varies delay per word by difficulty. Meta reports this puts the model on the Pareto front for speed against accuracy, measured by time to final transcription, ahead of the previous frontier formed by Soniox, Cartesia, and ElevenLabs systems.

    Diarization and endpointing are more tokens

    Meta did not add a second model for speaker attribution. It added special tokens to the same stream.

    For diarization, a <|start_of_turn|> token marks a potential speaker switch, and a <|speaker_{A-Z}|> tag identifies the speaker. The turn token fires as soon as a switch is possible, while the speaker tag is delayed to the end of the chunk. Audio from one speaker can be split across several segments that all resolve to the same tag.

    For endpointing, <|speech_onset|> marks the start of speech and <|speech_endpoint|> marks the point where the user finished. Both tasks are trained jointly with streaming ASR, using extra rewards layered on top of the ASR reward.

    Capabilities

    The model was trained on 70+ languages, of which 25 are extensively verified and recommended at launch. Code-switching is native, both within a sentence and between sentences, which matters for bilingual speakers who mix languages mid-clause. Accuracy can be improved further with language, keyword, and context biasing.

    Long-context handling is a practical differentiator. Meta states the model natively supports audio input exceeding one hour and 20+ speakers, with no required post-processing step.

    Benchmarks

    Meta reports first place on Artificial Analysis for streaming speech-to-text and on public diarization benchmarks, as of September 1, 2026.

    On Artificial Analysis AA-WER Streaming, Muse Voice Transcribe records 3.1% final-transcript WER at 0.16s after end of speech. Cartesia Ink-2 with semantic endpoints is 3.4% at 0.43s. ElevenLabs Scribe v2 Realtime is 3.6% at 0.14s. Cartesia Ink-2 with external endpoints is fastest at 0.07s but least accurate at 4.0%. On first partial transcript, Muse Voice Transcribe records 3.6% WER at 0.13s.

    On diarization, Meta reports a 17.5% average diarization error rate across AMI-IHM, AMI-SDM, and VoxConverse. Five other systems in the same chart range from 21.1% to 28.6%.

    Price is the other axis. At $3.00 per 1,000 minutes, it undercuts Cartesia Ink-2 at $4.00 and is less than half the $6.50 for ElevenLabs Scribe v2 Realtime and Deepgram Flux.

    Interactive explainer

    ASR Diarization Endpointing Labs Meta model Muse RealTime Releases Streaming superintelligence Transcribe voice
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    Reliance’s JioHotstar takes its streaming empire global — without sports

    OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

    OpenAI supports California’s bill to advance youth AI safety

    Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

    Healthcare organizations can now connect EHR and additional industry data to ChatGPT

    Researchers from Princeton, Ant Group and Stanford Introduce AQuA: A Two-Part Agentic Framework for Autonomous Factor Discovery and Model Development in Quantitative Finance

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Priceline Promo Codes & Coupons: 10% Off September 2026

    September 2, 2026

    Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

    September 2, 2026

    SonicWall Warns of Two SMA1000 Zero-Days Exploited in Attacks

    September 2, 2026

    Binance Adds Options on 1,000 US Stocks and ETFs

    September 2, 2026
    Latest Posts

    Bitcoin Only Makes Up 1% Of Legendary Investor Ray Dalio’s Portfolio

    July 30, 2026

    AI Harnesses Burst With Potential Exploit Opps

    July 30, 2026

    LinkedIn actually adds a ‘seems like AI slop’ button

    July 30, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Priceline Promo Codes & Coupons: 10% Off September 2026

    September 2, 2026

    Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

    September 2, 2026

    SonicWall Warns of Two SMA1000 Zero-Days Exploited in Attacks

    September 2, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.