← Browse

Session brief

Mission-Critical Evals at Scale (Learnings from 100k medical decisions)

Overview

This talk outlines the challenges and solutions for building scalable evaluation systems for AI, particularly in mission-critical domains like healthcare where errors have significant consequences. It emphasizes the limitations of traditional human review and offline datasets, advocating for real-time, reference-free evaluation methods to ensure customer trust and enable rapid response to issues.

Who should watch

Key takeaways

Notable quotes

*We've all seen that it's pretty easy to create an MVP product powered by LLMs and it's getting even easier as models get more and more powerful but what about going from MVP to serving customers at scale.*
*Relying only on offline evals is playing with fire.*
*The principles that we followed for building our system and what we would recommend is firstly make sure you build a system you know think big don't just use review data to to audit your performance use it to build audit and improve your auditing system your evaluation system.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.