CATEGORY

Audio and voice

Speech, music and audio models evaluated for quality, control and responsible use.

ON THIS PAGE
01Voice generation
02Audio quality
03Rights and consent

EDITOR'S SELECTION

Latest analysis

Audio & Voice

Open ASR Leaderboard adds Hindi and Indian English evaluation sets

Voice Arena and Hugging Face have added speaker-disjoint Hindi and Indian English datasets designed to expose regional, device and demographic performance differences.

· Priya Raman · 10 min read
→
Audio & Voice

Google launches Gemini 3.5 Transcribe in public preview

Google's new speech-to-text model supports streaming and recorded audio, custom vocabulary, word-level timestamps and automatic detection across more than 85 languages.

· Priya Raman · 9 min read
→
Audio & Voice

IBM releases two compact Granite Speech 5 transcription models

IBM's new 470-million-parameter English speech-recognition models target high-throughput transcription, with separate commercial and noncommercial checkpoints.

· Priya Raman · 8 min read
→
Audio & Voice

Hume researchers find signs of benchmark fitting in speech recognition

Tests across 11 open speech-recognition models found cases where systems reproduced benchmark-specific text even when the audio contradicted it. Leaderboard accuracy may overstate real-world transcription quality.

· Priya Raman · 9 min read
→
Audio & Voice

An AI music production checklist for rights and release

Generating a track is only the first step. Producers need records for prompts, source audio, collaborators, likeness, distribution terms and meaningful human authorship.

· Priya Raman · 8 min read
→
Audio & Voice

How speech-to-text AI works—and how to test it

Modern transcription systems are impressive, but word error rate alone cannot tell you whether they will work for meetings, interviews or regulated records.

· Priya Raman · 10 min read
→
Audio & Voice

A safe production workflow for AI voice generation

Synthetic speech can accelerate localization and accessibility. It also creates consent, impersonation and provenance risks that must be addressed before generation begins.

· Priya Raman · 8 min read
→
Audio & Voice

Voice AI finally gets a benchmark that listens to people

Scale’s Voice Showdown uses real spoken prompts across more than 60 languages. It arrives as the industry confronts a basic problem: synthetic tests do not capture the messiness of human speech.

· Priya Raman · 6 min read
→
Research & Data

How to build a held-out evaluation for speech recognition

A useful ASR test set must resemble production without becoming part of the optimization loop. This guide combines temporal separation, error weighting and audio-level review.

· Priya Raman · 10 min read
→
Video Creation

Kling AI 3.0 combines video, image and native audio in one model series

Kuaishou's latest Kling release adds multimodal input, synchronized sound, storyboard control and clips of up to 15 seconds for more complete production workflows.

· Noah Chen · 7 min read
→
Image Generation

FLUX 3 expands Black Forest Labs from images into video and audio

The early-access model jointly learns from images, video and sound, signaling a move from specialist image generation toward a multimodal foundation for visual intelligence.

· Noah Chen · 7 min read
→