World's Fair 2025
The Future of Qwen: A Generalist Agent Model — Junyang Lin, Alibaba Qwen
Overview
This talk introduces the Qwen series of large language and multimodal models from Alibaba, focusing on their development towards generalist agent models. It highlights recent advancements in Qwen 3, including its hybrid thinking mode, extensive language support, and enhanced capabilities for agents and coding. The presentation also touches upon the future direction of AI development, emphasizing training agents through reinforcement learning and scaling multimodal capabilities.
Who should watch
- AI Engineers
- Product Managers
- Developers building AI applications
- Researchers interested in LLM advancements
- Those working with open-source models
Key takeaways
- Qwen 3 is the latest iteration, featuring multiple model sizes (dense and Mixture-of-Experts) with competitive performance against state-of-the-art models.
- A key innovation is the hybrid thinking mode, allowing a single model to switch between reflective thinking and direct response, controllable via prompts or hyperparameters.
- The models now support over 119 languages and dialects, a significant expansion from previous versions, enabling broader global application.
- Enhanced support for agentic capabilities and tool usage (MCP) is a focus, demonstrated through examples of models thinking while using tools and interacting with the environment.
- Qwen is also developing multimodal models, including an omni model capable of processing text, vision (images/videos), and audio, and generating text and audio.
- The Qwen team emphasizes open-sourcing their models and checkpoints, encouraging community feedback and development.
- Future directions include advancements in training methods, scaling laws, long-horizon reasoning with environment feedback, and expanding multimodal input/output capabilities.
- The presentation showcases practical applications like WebDev for generating websites and Deep Research for generating comprehensive reports.
Notable quotes
*We believe that there are more potential for a large language model not just becoming an instruction tune model but it can be smarter and smarter with reinforcement learning.*
*The most important features this time it is hybrid thinking mode. So what is hybrid thinking mode? Hybrid thinking mode means that you can use thinking and non-thinking in a single model.*
*We are now moving from the era of training models to training agents.*
Unofficial community note. Prefer the recording for nuance.