Google launches Gemini 3.5 Transcribe to enhance real-time voice transcription
Google takes a significant step in the field of voice artificial intelligence with the launch of Gemini 3.5 Transcribe, a new model designed to convert speech to text with greater speed and accuracy. Aimed at both developers and businesses, this technology seeks to meet the growing demand for instant transcription of conversations, meetings, and audio recordings.
The new model stands out for its ability to process speech in complex environments. Background noise, technical vocabulary, hesitations, or corrections made by speakers can be taken into account to improve the quality of the generated text.
According to the provided data, Gemini 3.5 Transcribe has an error rate of 4% for continuous transcription and 2.6% for processing offline audio. The time required to obtain a final transcription has also been reduced by 70% compared to the previous Chirp 3 model.
Real-time or post-recording transcription
Google offers two usage modes tailored to different needs. The first allows for near-instant transcription through a live streaming interface with a latency of less than one second.
The second is intended for pre-recorded audio content, such as business meetings or call logs. This version also allows for the identification and attribution of contributions from different participants.
The model is capable of automatically detecting and transcribing over 85 languages. It can also recognize up to three speakers in the same recording and associate their contributions with timestamps.
Among the announced features are the removal of filler words, automatic text formatting, and the ability to integrate a custom vocabulary to enhance the recognition of sector-specific terms.
Google expands the use of its voice artificial intelligence
On the FLEURS benchmark, Gemini 3.5 Transcribe shows an error rate of 5.50% in streaming mode and 5.04% for processed offline content, according to the provided information.
Google plans to gradually integrate this technology into several of its products and services. The model is notably associated with the Gemini application on macOS, certain features on Android via Gboard, and Google AI Studio.
The launch also comes as part of a broader strategy to enhance the integration of generative artificial intelligence into professional and consumer tools. Developers can access the model in public preview via Google AI Studio, while businesses have integration through the Gemini Enterprise Agent platform.
Ultimately, the Chrome browser is also expected to benefit from new dictation features based on this technology. With Gemini 3.5 Transcribe, Google aims to make voice an increasingly central interface in the daily use of digital tools.
-
16:42
-
16:35
-
16:19
-
16:06
-
15:58
-
15:57
-
15:40
-
15:09
-
14:57
-
14:49
-
14:37
-
14:34
-
14:33
-
14:14
-
14:12
-
14:08
-
14:00
-
13:55
-
13:45
-
13:40
-
13:31
-
13:26
-
13:21
-
13:06
-
13:06
-
13:02
-
11:30
-
11:23
-
11:20
-
11:20
-
11:13
-
11:11
-
11:06
-
11:05
-
10:59
-
10:59
-
10:56
-
10:54
-
10:41
-
10:38
-
10:35
-
10:31
-
10:22
-
10:05
-
09:51
-
09:33
-
09:06
-
09:00
-
08:56
-
08:50
-
08:43
-
08:37
-
08:30
-
08:25
-
08:14
-
08:02
-
07:51
-
07:43
-
07:36
-
07:30
-
07:28
-
17:25
-
17:23
-
17:21
-
17:18
-
17:16
-
17:14
-
17:11
-
17:07
-
17:01
-
16:55
-
16:50