Alibaba Cloud has released Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system, bringing major technical advances for users requiring high-quality, customizable text-to-speech. The latest version enhances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, overall efficiency, and robustness. Building on these improvements, the system integrates a 12.5 Hz low-frame-rate speech tokenizer that lowers inference latency, alongside a five-stage progressive training paradigm that tightly coordinates language and frequency modeling optimization. Qwen-Audio-3.0-TTS also introduces free-style natural-language instruction following and fine-grained inline tagging, enabling production-level control for a range of deployment scenarios. For global users, Qwen-Audio-3.0-TTS supports 16 languages, 20 Chinese dialect regions, and offers one-pass long-form synthesis for outputs up to 3 minutes. It maintains high performance ...
Related
YubiKey 5.8 extends passkeys from secure authentication to verifiable authorization
Yubico has officially released YubiKey 5.8, marking an expansion of the role of the passkey from secure authentication to include hardware-backed, verifiable authorization for digi...
Moonshot AI launches Kimi K3, a 2.8T parameter open MoE model with vision support
Last week, Beijing based AI startup Moonshot AI launched Kimi K3, a 2.8 trillion parameter Mixture of Experts model described as the first open model in its class and the largest o...
Alibaba launches Qwen-Audio-3.0-TTS for efficient, robust, and consistent text-to-speech
Alibaba Cloud has released Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system, bringing major technical advances for users requiring high-quality, customizable text-...