← Browse

Session brief

Best Practices for Evaluating Large Language Model Applications with llmeval: Niklas Nielsen

Overview

This talk introduces llmeval, a command-line tool designed to help teams ship reliable large language model (LLM) applications. It addresses the challenge of defining and measuring "good" in generative AI, which is crucial for evolving applications, changing prompts, or considering different models. llmeval provides a structured approach to testing and evaluating LLM outputs, enabling more confident development and deployment.

Who should watch

Key takeaways

Notable quotes

*Without knowing what good means in a generative setting it's really really hard and risky to evolve your applications.*
*llmeval enables teams to ship reliable LLM products.*
*Model based evaluation is a setting where you have typically a larger model discriminate or kind of grade or be a judge over the output from another LLM.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.