← Browse

World's Fair 2025

OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs

Ryan Marten

Overview

This talk introduces OpenThoughts 3, a project focused on creating high-quality, open-source reasoning datasets. The core thesis is that while models have shown significant performance gains on reasoning benchmarks, the underlying data recipes for achieving this are often undisclosed. OpenThoughts aims to fill this gap by providing a systematic approach and publicly available datasets to train more capable reasoning models, emphasizing that supervised fine-tuning (SFT) on well-crafted data can be highly effective.

Who should watch

Key takeaways

Notable quotes

*The missing link is the data recipe.*
*Scaling is difficult, particularly difficult with RL. The good news is for SFT scaling is quite easier.*
*A better model in terms of its own performance on evaluation benchmarks does not necessarily mean it's a better teacher model.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.