World's Fair 2026
Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth
Overview
This talk explores the current state and future trajectory of AI models, focusing on advancements in language models and benchmarks. It highlights the rapid progress in AI capabilities, evidenced by performance improvements across various benchmarks, while also addressing limitations such as long context handling and the challenges of reliable evaluation. The discussion delves into the open-source versus closed-source landscape, the complexities of model optimization, and the critical role of software and algorithmic improvements over hardware advancements.
Who should watch
- AI engineers and researchers interested in model performance trends.
- Product managers evaluating AI capabilities for new applications.
- Builders and developers seeking to understand the latest AI tools and techniques.
- Anyone curious about the challenges and future directions of AI development, including benchmarking, optimization, and regulation.
Key takeaways
- AI models are showing exponential progress, with doubling times for capabilities shrinking significantly, though the long-term scaling trend remains a question.
- Long context remains a challenge for current models, with accuracy degrading significantly as context length increases.
- The open-source AI community is rapidly closing the gap with closed-source models, though a lag still exists.
- Benchmarking AI models is complex, with issues like data contamination, reliance on LLMs for verification, and inconsistent results across different benchmarks.
- Software and algorithmic optimizations, such as dynamic quantization and improved training techniques, are becoming more critical than hardware advancements for scaling AI.
- Reinforcement learning is a powerful tool for improving model performance, but it requires careful implementation and can be prone to issues like reward hacking.
- The increasing power of AI models raises concerns about cybersecurity, regulation, and the potential for misuse, prompting discussions on licensing and control.
Notable quotes
*AI progress can continue over time with new ideas and innovation.*
*The main point is the harness, the implementation, the tool is now the most important. It's not the model.*
*Hardware is probably overblown. Hardware is actually not that important. The software was the trick.*
Unofficial community note. Prefer the recording for nuance.