Session brief
Creating Agents that Co-Create — Karina Nguyen, OpenAI
Overview
This talk explores the evolution of AI research scaling paradigms, moving from next-token prediction to scaling reinforcement learning on Chain of Thought. These advancements have unlocked new frontiers in product research, enabling rapid iteration cycles and the creation of novel AI capabilities. The future vision is one of AI agents as co-innovators, collaborating with humans to create new knowledge and experiences.
Who should watch
- AI Engineers interested in the latest scaling paradigms and their impact on product development.
- Product Managers seeking to understand how advanced AI capabilities can be integrated into new products.
- Builders looking for insights into the future of human-AI collaboration and co-creation.
- Researchers exploring the challenges and opportunities in creative AI and complex reasoning.
Key takeaways
- Two primary scaling paradigms in AI research are next-token prediction (pre-training) and scaling reinforcement learning on Chain of Thought (complex reasoning).
- Pre-training enables models to learn about the world by predicting the next token, while Chain of Thought scaling allows models to learn to think through complex problems.
- Post-training techniques like Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF) are crucial for refining model usefulness, particularly in areas like code completion.
- The future of AI agents is envisioned as co-innovators, built upon reasoning, tool use, and long context, enhanced by creativity through human-AI collaboration.
- Rapid product development cycles are now possible by using highly reasoning models to distill capabilities into smaller models or to synthetically generate data for new training environments.
- New interaction paradigms are emerging, such as streaming model thoughts to users and creating familiar product factor interfaces for unfamiliar capabilities like 100K context windows via file uploads.
- Bridging real-time interaction with asynchronous task completion requires building trust through human verification, editing, and real-time feedback mechanisms for model self-improvement.
- The concept of a blank canvas interface that adapts to user intent, whether for coding, writing, or data analysis, represents a significant shift in how users interact with AI.
Notable quotes
*The future of Agents is co-innovators, built upon reasoning, tool use, and long context plus creativity.*
*The way you interact with AI fundamentally changes in a way that the way you access the internet will also change.*
Unofficial community note. Prefer the recording for nuance.