← Browse

World's Fair 2024

The Hierarchy of Needs for Training Dataset Development: Chang She and Noah Shpak

Overview

This talk addresses the critical importance of training data development for large language models, emphasizing that the quality and format of data directly impact model performance. It proposes a hierarchy of needs for dataset development, starting with clean data and progressing through evaluations, dataset management, and advanced techniques like synthetic data generation and quality scoring. The discussion highlights the challenges posed by the increasing scale and multimodality of AI data and introduces the Lance format as a solution for efficient data handling.

Who should watch

Key takeaways

Notable quotes

*You should really care about what you're training on and you should care for it by giving it a nice format that does a lot of nice things for it.*
*Existing data formats and data infrastructure is good for at most two but but often just one of the three.*
*Speed is probably our our best bet in terms of strategy and a lot of the tools that we've worked with really slow down under load under new multimodal needs.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.