Two transcription paths
Google released `gemini-3.5-transcribe-live` through the Live API for bidirectional streaming and `gemini-3.5-transcribe` through the Interactions API for recorded meetings, calls and other audio. The prerecorded path includes speaker attribution and word-level timestamps.
The company positions the system as intelligent transcription rather than literal dictation. It can remove filler words, resolve spoken self-corrections and format text, which may improve readability but also changes the evidentiary relationship between the recording and transcript.
The performance claims
Google cites Artificial Analysis measurements of 4.0% average word error rate for streaming and 2.6% for non-streaming use, plus a 70% improvement in time to final transcription over Chirp 3. On FLEURS, Google reports 5.50% streaming and 5.04% non-streaming WER across selected languages and locales.
Those numbers are useful starting points, not a guarantee for every microphone, accent or vocabulary. Teams should reproduce accuracy, stability and latency with their own recordings and separately score names, identifiers, negation and speaker attribution.
Availability and practical limits
The developer and enterprise APIs are in public preview. Google says the model automatically detects more than 85 languages, accepts custom vocabulary and identifies up to three speakers in recorded audio, while support beyond three speakers remains experimental.
Visit the AINewsInu homepage and Audio & Voice hub for evaluation guides. Buyers should record the exact endpoint and preview status, verify retention and consent settings, and preserve raw audio whenever a polished transcript will support consequential decisions or quotations.
Sources & further reading
Social-media activity is treated as a signal of attention, not proof. Product claims are attributed to the linked publisher or announcement.