World's Fair 2025
2025 in LLMs so far, illustrated by Pelicans on Bicycles — Simon Willison
Overview
The AI landscape has accelerated dramatically over the last six months, with numerous significant model releases making it challenging to assess their quality. Traditional benchmarks and leaderboards are losing credibility, prompting a shift towards more practical, self-devised evaluation methods. The speaker uses a unique benchmark involving generating SVGs of pelicans riding bicycles to assess model capabilities, highlighting the rapid progress in model performance and the increasing accessibility of powerful AI on consumer hardware.
Who should watch
- AI Engineers
- Product Managers
- Builders evaluating LLM capabilities
- Those interested in the rapid evolution of AI models
- Individuals struggling to keep up with new LLM releases
- Developers exploring local model execution
Key takeaways
- The pace of LLM development has accelerated, with dozens of significant model releases in the past six months, making evaluation difficult.
- Traditional benchmarks are becoming less reliable, leading to the use of custom, practical tests like generating SVGs of pelicans on bicycles to assess model performance.
- Models like Meta's Llama 3 70B and Mistral's 7B Small 3 demonstrate that GPT-4 class capabilities are becoming runnable on consumer hardware.
- Open-weight models, such as DeepSeek V3, are achieving state-of-the-art performance at significantly lower training costs than previously estimated.
- The cost of using powerful LLM APIs has drastically decreased, with models like GPT-4.1 Nano being significantly cheaper than older models like GPT-3 Da Vinci.
- Multimodal capabilities are advancing rapidly, but features like memory in conversational AI can lead to loss of user control over context.
- The combination of tools and reasoning is emerging as a powerful technique in AI engineering, enabling models to perform complex tasks like iterative searching and analysis.
- Risks associated with AI systems include prompt injection and the "lethal trifecta" of private data access, malicious instructions, and exfiltration mechanisms.
Notable quotes
*The problem that we have is I counted 30 significant model releases in the past six months.*
*The most exciting trend in the past six months is that the local models are good now.*
*This is the most powerful technique in all of AI engineering right now.*
Unofficial community note. Prefer the recording for nuance.