Tag
6 articles
Alibaba's Qwen Audio 3.0 TTS Plus has topped the Artificial Analysis Speech Arena leaderboard, showcasing advanced multilingual capabilities and expressive controls, though it lags in speed compared to competitors.
Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a hosted text-to-speech model available in Flash and Plus variants across 16 languages.
Learn how to set up OmniVoice Studio, a local, open-source alternative to ElevenLabs, for voice cloning and text-to-speech functionality without cloud dependencies.
Mistral AI's new TTS model, Voxtral, tackles the 'expressivity gap' in voice AI by combining autoregressive and flow-matching techniques for more emotionally expressive, multilingual speech synthesis.
Google introduces Gemini 3.1 Flash TTS, a new text-to-speech model that enhances speech quality, expressive control, and multilingual generation. This release marks a shift toward more controllable and natural AI voice outputs.
French AI startup Mistral has released Voxtral, its first open-weight text-to-speech model that supports nine languages and can clone voices from just three seconds of audio.