Alibaba launches Qwen-Audio-3.0-TTS for efficient, robust, and consistent text-to-speech

Alibaba Cloud has released Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system, bringing major technical advances for users requiring high-quality, customizable text-to-speech. The latest version enhances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, overall efficiency, and robustness. Building on these improvements, the system integrates a 12.5 Hz low-frame-rate speech tokenizer that lowers inference latency, alongside a five-stage progressive training paradigm that tightly coordinates language and frequency modeling optimization. Qwen-Audio-3.0-TTS also introduces free-style natural-language instruction following and fine-grained inline tagging, enabling production-level control for a range of deployment scenarios. For global users, Qwen-Audio-3.0-TTS supports 16 languages, 20 Chinese dialect regions, and offers one-pass long-form synthesis for outputs up to 3 minutes. It maintains high performance ...

Read Original

Related