Google has introduced a new audio-focused AI model called Gemini 3.5 Transcribe, whose primary purpose is converting spoken language into structured written text. According to the company, the model uses improved speech-recognition algorithms to better interpret spoken words.
Features and capabilities
- The model can transform "unstructured speech" into formatted, editable text rather than producing only verbatim transcripts.
- It supports voice commands for editing, so users can apply dictation-based instructions to format or correct transcripts.
- During transcription the system removes filler words, yielding cleaner and more concise output.
- Google says the model can learn individual vocabularies and unique spelling conventions, improving accuracy with continued use.
- It records alphanumeric strings (for example, order numbers and ZIP codes) with particular precision, which can be valuable for business scenarios such as order intake or phone-based customer-service confirmations.
Multi-speaker transcription and timestamps
Google claims Gemini 3.5 Transcribe can transcribe up to three speakers simultaneously and attach timestamps to the transcript. This capability is useful for producing transcripts of podcasts or other multi-participant conversations.
Browser integration and developer access
The company promises that the model's text-to-speech features will soon be available across web pages in the Google Chrome browser. Users will also be able to dictate replies directly into web page text fields (for example, comment boxes).
The solution will be made available "soon," and developers will receive access to the model to customize it for their applications.
Announcement and context
Google published the details of the announcement through its own communication channels. The company highlights the improved accuracy of transcripts and the potential business benefits of the model's capabilities.
Why this matters
More accurate, structured transcripts and the ability to edit via voice can speed workflows in media production, customer service, and administrative tasks. The model's precision in capturing alphanumeric data such as order or ZIP codes can reduce errors and improve the reliability of automated processing.



