← Browse

World's Fair 2025

Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith

Taylor Jordan Smith

Overview

This talk addresses the critical need for robust evaluation and benchmarking strategies when deploying large language models (LLMs) into production. It highlights the inherent complexities and potential pitfalls of generative AI, emphasizing that scalability, reliability, and safety are paramount. The presentation introduces practical tools and methods to assess LLM performance, ensuring that models meet enterprise-level requirements before and during deployment.

Who should watch

Key takeaways

Notable quotes

*No matter how good your model is, if it's not fast, if it's not reliable, if it's not affordable, you're screwed a little bit from the get-go.*
*Evaluation is a comprehensive process to assess a model end to end and it could include a lot of different kinds of evaluations about a lot of different components.*
*Benchmarking is very specifically controlled specific data sets and specific tasks typically used to compare models against one another.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.