Qwen-Image-3.0 released with 4.5k token input support for complex AI image generation

Alibaba Cloud has introduced Qwen-Image-3.0, marking the third generation of its foundational image generation model in the Qwen-Image series. The new model supports up to 4,5k token inputs, making it possible to generate visually dense outputs, including newspapers, storyboards, and exam papers, that house complex layouts and large volumes of structured information. Alongside layout complexity, Qwen-Image-3.0 achieves precise rendering of text as small as 10 pixels and captures micro-level visual details such as pores and hair strands, enabling lifelike reproduction of texture and typography. The model offers native support for 12 languages and can simulate a range of mainstream interfaces, including web pages, games, and livestreams. It also generates realistic infographics and draws on extensive world knowledge. Building on these core capabilities, the model demonstrates advanced spatial control with horizontal expansion, allowing orderly placement of multiple concepts within a sing...

Read Original

Related