← Browse

World's Fair 2024

How to Construct Domain Specific LLM Evaluation Systems: Hamel Husain and Emil Sedgh

Overview

This talk outlines a systematic approach to constructing domain-specific LLM evaluation systems. It emphasizes the importance of moving beyond initial "vibe checks" and prompt engineering to establish a robust framework for consistent AI improvement. The core thesis is that a well-defined evaluation system, built on foundational principles, is crucial for developing production-ready AI applications and unlocking advanced capabilities like fine-tuning.

Who should watch

Key takeaways

Notable quotes

*If you don't have a way of measuring progress you can't really build.*
*It's really critical to your overall evaluation system if you can have them and how do you run the assertions one very reasonable way is to use CI.*
*The most important part if you remember anything from this talk it is you need to look at your data and you need to fight as hard as you can to remove all friction in looking at your data.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.