CATEGORY
Audio and voice
Speech, music and audio models evaluated for quality, control and responsible use.
ON THIS PAGE
01Voice generation
02Audio quality
03Rights and consent
EDITOR'S SELECTION
Latest analysis
Kling AI 3.0 combines video, image and native audio in one model series
Kuaishou's latest Kling release adds multimodal input, synchronized sound, storyboard control and clips of up to 15 seconds for more complete production workflows.
Noah Chen · 7 min readFLUX 3 expands Black Forest Labs from images into video and audio
The early-access model jointly learns from images, video and sound, signaling a move from specialist image generation toward a multimodal foundation for visual intelligence.
Noah Chen · 7 min readVoice AI finally gets a benchmark that listens to people
Scale’s Voice Showdown uses real spoken prompts across more than 60 languages. It arrives as the industry confronts a basic problem: synthetic tests do not capture the messiness of human speech.
Priya Raman · 6 min read