Speaker
Will Brown
4 sessions in this library.
- Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect
This talk delves into modern post-training techniques for AI models, focusing on the open-source Verifiers and Primed RL libraries. It emphasizes the evolution of these tools to handle increasingly complex agent use cases and the need for robust infrastructure to support large-scale, customized model training. The core thesis is that by providing flexible and powerful open-source toolkits, it becomes more accessible for engineers and builders to enhance existing models for their specific applications and workflows, fostering continuous improvement through iterative refinement.
- RL Environments at Scale – Will Brown, Prime Intellect
This talk explores scaling AI research not just through increased data and compute, but by making the practice of AI research more accessible. It introduces the concept of "environments" as a key abstraction for experimentation, akin to web applications for AI research, enabling broader participation beyond large labs. The discussion highlights how these environments facilitate model customization, evaluation, and the development of more effective AI systems.
- Training Agentic Reasoners — Will Brown, Prime Intellect
This talk argues that reasoning and agents are fundamentally the same concept, with reinforcement learning (RL) being the key to developing powerful agentic systems. The speaker posits that traditional approaches to building agents often involve manual, iterative tuning that mirrors RL processes. By framing agent development through the lens of RL, developers can leverage established algorithms and techniques to create more robust and capable agents, especially for complex, multi-turn tasks.
- Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley
This talk explores the potential of reinforcement learning (RL) to advance AI agents beyond current chatbot and reasoner capabilities. It posits that RL offers a path to developing more autonomous systems that can learn and improve through interaction with their environment, moving beyond static prompt engineering and tool-calling approaches. The discussion highlights emerging trends and open-source efforts in this domain, suggesting a future where RL is integral to agent engineering.