Product signal
Google Releases Gemini 3.5 Transcribe
Google is bringing Gemini 3.5 Transcribe to the Gemini API and Gemini Enterprise Agent Platform for real-time and recorded speech transcription.
Google has released Gemini 3.5 Transcribe for the Gemini API and Gemini Enterprise Agent Platform. The change turns a capability previously centered in applications into a standalone model service for enterprise agent workflows. Real-time streaming, custom vocabulary and speaker attribution provide the main mechanism for expanding where the capability can be used.
What changed: transcription becomes a standalone service
Gemini 3.5 Transcribe is more than a user-interface update in the Gemini application. Google is making it available to developers and enterprise agent platforms, allowing speech-to-text, recorded-audio processing and structured timing data to enter customer service, meeting and business-automation workflows.
Mechanism: streaming processing handles complex audio input
Google says the model supports bidirectional streaming, custom vocabulary, speaker attribution and word-level timestamps across a broad language range. It is designed for both live and recorded use, allowing one service to support low-latency interaction as well as later structured processing.
Why it matters: speech becomes an agent infrastructure layer
If reliability and recognition quality meet enterprise requirements, transcription can become an input layer for agents handling meetings, customer interactions and field information rather than merely an endpoint feature. Independent leaderboards do not yet directly list the model, so Google’s accuracy comparison remains a signal to verify.
What to watch next
Watch the API’s regions, pricing and quotas, independent replication of its streaming and non-streaming accuracy, and evidence of deployments through the enterprise platform.