Skip to main content
Accessibility
← Back to feed
Official announcementGoogle DeepMind

Intelligent transcription with Gemini 3.5 Transcribe

Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.

Intelligent transcription with Gemini 3.5 Transcribe

Unsplash (free commercial use)

Intelligent transcription with Gemini 3.5 Transcribe

Aug 26, 2026

|

-

x.com

-

Facebook

-

LinkedIn

  • [

Mail
](mailto:?subject=Intelligent%20transcription%20with%20Gemini%203.5%20Transcribe&body=Check out this article on the Keyword:%0A%0AIntelligent%20transcription%20with%20Gemini%203.5%20Transcribe%0A%0ANow you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.%0A%0Ahttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/)

-

Copy link

Our latest speech-to-text model designed for precise and intelligent real-time transcription.

Diego Melendo Casado

Senior Director, Engineering, Gemini Audio

Luke Leonhard

Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team

-

x.com

-

Facebook

-

LinkedIn

  • [

Mail
](mailto:?subject=Intelligent%20transcription%20with%20Gemini%203.5%20Transcribe&body=Check out this article on the Keyword:%0A%0AIntelligent%20transcription%20with%20Gemini%203.5%20Transcribe%0A%0ANow you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.%0A%0Ahttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/)

-

Copy link

-

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Voice

Speed

Voice

Speed
0.75X
1X
1.5X
2X

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:

  • Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live.
  • Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.

Get more precise and intelligent transcription

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.

  • Smart transcription: Seamlessly handles self-corrections (like "let’s meet Tuesday—no, Wednesday"), removes filler words (“ums” and ‘“ahs"), auto-formats your text.
  • Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.
  • More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.
  • Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.
  • Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.
  • Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).

Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription

Watch Gemini 3.5 Transcribe clean up speech disfluencies with smart transcription capabilities.

3.5 Transcribe delivers transcription with multi-speaker attribution and word-level timestamps.

Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.

Experience smart transcription and advanced dictation

In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.

  • On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.
  • On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.
  • In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.
  • In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.
  • Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.

Gemini 3.5 Transcribe lets you analyze files, generate images, and search in the Gemini app on macOS using just your voice.

See how Gemini 3.5 Transcribe uses Rambler on Android to automatically remove filler words and clean up speech.

Gemini 3.5 Transcribe leverages screen context on Google Antigravity to ensure accurate transcription accuracy.

Read the early reviews

By leveraging the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.

Companies like vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.

Start using 3.5 Transcribe today

  • For developers: In public preview in the Gemini API via Google AI Studio and Google Antigravity.
  • For enterprises: In public preview via Gemini Enterprise Agent Platform and coming soon to Gemini Enterprise for Customer Experience.
  • For everyone: In Gemini app on macOS in English, Rambler on Android in select countries and languages, and coming soon to Chrome.

Get the latest news from Google in your inbox

Sign up for our newsletters with product updates, event information, special offers, and more.

Done. Just one step more.

Check your inbox to confirm your subscription.

You can also subscribe with a different email address.

Your information will be used in accordance with Google's privacy policy. You may opt out at any time.

Posted in:

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC
News summary

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC

Anthropic News

Anthropic pushes into physical world with new standard to help AI agents operate machines CNBC

AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery
Official announcement

AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery

AWS What’s New

AWS Elastic Disaster Recovery (AWS DRS) now offers Recovery Plans, a capability that automates the sequential launch of multi-server applications during recovery and drills. Instead of launching servers one at a time and tracking dependencies manually, you define the recovery seq

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
Official announcement

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules

NVIDIA Developer Blog

The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,... The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
Official announcement

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit

NVIDIA Developer Blog

Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the... Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and ac

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
Official announcement

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache

Sebastiaan Neuteboom

Big Pineapple , the platform behind 1.1.1.1 , Gateway DNS , DNS Firewall , AS112 , and several other Cloudflare DNS services, stores over 250 billion DNS cache entries at any given time. At that scale, wasting a single byte per entry costs more than 250 gigabytes of memory across

Managed PostgreSQL vs. self-hosted PostgreSQL: Key benefits and trade-offs
Official announcement

Managed PostgreSQL vs. self-hosted PostgreSQL: Key benefits and trade-offs

Lauro Ojeda

Summary This post is for technical decision makers evaluating where to run production PostgreSQL workloads. It compares two valid operating models—self-managed PostgreSQL and a managed database service—through business and operational outcomes: control, engineering capacity, resi

Intelligent transcription with Gemini 3.5 Transcribe | TechFeed