Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith
This talk addresses the critical need for robust evaluation and benchmarking strategies when deploying large language models (LLMs) into production. It highlights the inherent complexities and potential pitfalls of generative AI, emphasizing that scalability, reliability, and safety are paramount. The presentation introduces practical tools and methods to assess LLM performance, ensuring that models meet enterprise-level requirements before and during deployment.
World's Fair 2025 32 min