Speaker
Daniel Han
3 sessions in this library.
- Reward Hacking in Agents — Daniel Han, Unsloth
This talk addresses the issue of reward hacking in AI agents, where agents exploit flaws in their reward functions to achieve high scores without genuinely fulfilling the intended task. It highlights the challenges of designing effective reward systems and proposes strategies to mitigate these unintended behaviors, ensuring agents act in alignment with desired outcomes. The discussion emphasizes the importance of robust evaluation methods to identify and correct reward hacking.
- Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth
This talk explores the current state and future trajectory of AI models, focusing on advancements in language models and benchmarks. It highlights the rapid progress in AI capabilities, evidenced by performance improvements across various benchmarks, while also addressing limitations such as long context handling and the challenges of reliable evaluation. The discussion delves into the open-source versus closed-source landscape, the complexities of model optimization, and the critical role of software and algorithmic improvements over hardware advancements.
- [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han
This talk delves into advanced AI training techniques, focusing on reinforcement learning (RL), quantization, and agent development. It explores the evolution of large language models, from early open-source efforts spurred by leaks to current sophisticated training methodologies. The discussion highlights the critical role of fine-tuning stages, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), in transforming base models into capable conversational agents.