Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Greens just four points behind Labour in urban areas, poll reveals | Green party

    September 16, 2026

    Planning permission for new Scottish AI datacentres suspended for up to a year | Datacentres – UK

    September 16, 2026

    Anthropic merges Claude chat and Cowork into one

    September 16, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Greens just four points behind Labour in urban areas, poll reveals | Green party
    • Planning permission for new Scottish AI datacentres suspended for up to a year | Datacentres – UK
    • Anthropic merges Claude chat and Cowork into one
    • Windows 11 KB5124008 update breaks domain trust for some users
    • Circle (CRCL) debuts Arc blockchain in biggest bet yet beyond $74B USDC stablecoin
    • Scientists find an immune “false alarm” that may drive rapid aging
    • Trump, God, and the Struggle for the Chin State
    • Photos show widespread damage at US sites from Iranian attacks
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Wednesday, September 16
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Artificial Intelligence

    Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 28, 2026 Artificial Intelligence No Comments6 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Google has released Gemini 3.5 Transcribe, a speech-to-text model for real-time voice interfaces and recorded audio. It ships as two endpoints, not one. gemini-3.5-transcribe handles pre-recorded files through the Interactions API. gemini-3.5-transcribe-live handles bidirectional streaming through the Live API. Google reports average word error rates of 4.0% streaming and 2.6% non-streaming, as measured by Artificial Analysis. Time to final transcription improves 70% over Chirp 3, the previous model. Automatic detection covers more than 85 languages, including mid-sentence code-switching. The split between the two endpoints is the part worth planning around. They do not share the same feature set, limits, or price.

    Is it deployable?

    Yes, but API-only. There are no open weights and no self-hosted path. This is a managed-service decision, not an infrastructure one.

    • Company level: Any. Solo developers and startups can start on the Gemini API free tier via Google AI Studio. Mid-market teams move to the paid tier for higher rate limits. The paid tier also guarantees content is not used to improve Google’s products. Regulated enterprises route through the Gemini Enterprise Agent Platform, which adds provisioned throughput, compliance controls, and volume discounts. Both developer and enterprise tracks are in public preview, so treat production commitments accordingly.
    • Industries: Contact centers and CX platforms, clinical documentation, media captioning and localization, legal and insurance intake, meeting tooling, and voice-driven developer tools.
    • Applications: Real-time voice agents, live captioning, post-call analytics pipelines, meeting transcription with speaker attribution, dictation, and voice-controlled interfaces.

    Two API surfaces, two different products

    The Live API delivers sub-second, continuous transcription. It emits interim_input_transcription for speculative partials while someone is still talking, then input_transcription when the turn finalizes. Audio goes in as raw 16-bit PCM at 16kHz mono, in 100ms chunks. It supports automatic, hybrid, and manual voice-activity detection. Ephemeral tokens let mobile and web clients stream without holding an API key.

    The constraints are real. Live sessions cap at 10 minutes of continuous streaming. Speaker diarization is not supported. Word-level timestamps are not supported.

    The Interactions API covers what streaming cannot. It offers speaker diarization, word-level start and end offsets, and custom vocabulary biasing. The vocabulary list takes up to 1,000 terms, with best results below 100. Standard requests accept up to one hour of audio. That drops to 30 minutes once diarization or word timestamps are enabled.

    Verbatim and smart are the real design decision

    Both endpoints expose two modes. verbatim is the default and returns everything, including fillers, repetitions, and false starts. smart removes disfluencies, resolves spoken self-corrections inline, and applies structured formatting.

    Google’s own documented example: “Um, so for the meeting, I think we should, uh, invite Alice and, wait no, Bob and Carol.” Verbatim keeps all of it. Smart returns “For the meeting, I think we should invite Bob and Carol.”

    Smart mode cannot be combined with word timestamps or diarization. That is the tradeoff to plan around. A readable summary and an auditable transcript are now two different API calls.

    Performance

    As measured by Artificial Analysis, Google reports an average word error rate of 4.0% for streaming and 2.6% for non-streaming. On the multilingual FLEURS benchmark, across a set of top languages and locales, the model reports 5.50% streaming and 5.04% non-streaming.

    Against Chirp 3, Google’s previous transcription model, time to final transcription improves by 70%. Language coverage spans over 85 locales with automatic detection and code-switching handled without configuration.

    Ecosystem

    The Live API is already wired into LiveKit, Pipecat, Agora, Fishjam, Vercel, and Vision Agents. On the consumer side, the model powers Rambler on Android, the Gemini app on macOS, and Google Antigravity. Chrome is listed as coming soon.

    Key Takeaways

    • Two endpoints, not one: streaming trades diarization and word timestamps for sub-second latency.
    • Reported WER is 4.0% streaming and 2.6% non-streaming, per Artificial Analysis.
    • Smart mode cannot be combined with timestamps or diarization — pick one per call.
    • Blended cost runs about $0.005/min batch and $0.009/min live; no open weights.
    • Hard limits: 10-minute live sessions, 1-hour files, 30 minutes with diarization on.

    Check out the Technical details here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    Average Gemini Google Languages model Releases Reporting SpeechtoText Transcribe Wer
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    How workers are unlocking new ways of working

    Claude comes for Gemini with its own take on Docs and Slides

    Microsoft AI CEO criticises Anthropic over model ‘rights’

    Nums AI Releases Causilo: A Tabular Foundation Model That Tops TabArena Among Single Models

    ChatGPT pioneer launches Jev model for programmatic logic

    Prior Labs Releases TabPFN-3.5: A Tabular Foundation Model That Beats the Winning Otto Kaggle Solution With Default Settings

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Greens just four points behind Labour in urban areas, poll reveals | Green party

    September 16, 2026

    Planning permission for new Scottish AI datacentres suspended for up to a year | Datacentres – UK

    September 16, 2026

    Anthropic merges Claude chat and Cowork into one

    September 16, 2026

    Windows 11 KB5124008 update breaks domain trust for some users

    September 16, 2026
    Latest Posts

    What is Trump Media’s Truth API and why is it controversial?

    August 4, 2026

    How ProPublica Tested Hundreds of Omaha Homes for Lead — ProPublica

    August 4, 2026

    Golar LNG raises $600 million loan with FLNG business expansion in mind

    August 4, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Greens just four points behind Labour in urban areas, poll reveals | Green party

    September 16, 2026

    Planning permission for new Scottish AI datacentres suspended for up to a year | Datacentres – UK

    September 16, 2026

    Anthropic merges Claude chat and Cowork into one

    September 16, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.