← Browse

World's Fair 2024

Judging LLMs: Alex Volkov

Overview

This talk uses a courtroom drama format to highlight common pitfalls in AI engineering, particularly concerning Large Language Model (LLM) development and deployment. The core thesis is that rigorous evaluation, logging, and prompt iteration are crucial for successful LLM projects, and neglecting these can lead to significant problems. The presentation emphasizes the importance of a human-in-the-loop approach for effective LLM judging and evaluation.

Who should watch

Key takeaways

Notable quotes

*If you build um non-production stuff in hackathons that's fine but if you put anything of value in production you have to trace and log everything.*
*Folks, it's very important to remember that you have to iterate on prompts before you find tune.*
*The very simpleton lolms of your year are not yet escapable and so don't expect Perfection.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.