Speaker
Phil Hetzel
4 sessions in this library.
- How agent o11y differs from traditional o11y — Phil Hetzel, Braintrust
This talk differentiates traditional observability from agent observability, arguing that agents' non-deterministic nature and complex data structures necessitate a distinct approach. Traditional observability focuses on uptime and technical performance metrics like latency and error rates, using established tools. Agent observability, however, must also account for qualitative aspects of agent performance, such as grounding, tool usage, and adherence to brand standards, which require handling vast amounts of semi-structured and unstructured data.
- The maturity phases of running evals — Phil Hetzel, Braintrust
This talk outlines the maturity phases of running evaluations for AI agents, emphasizing that evals are crucial for ensuring agent quality, mitigating risks, and understanding performance improvements. The speaker suggests a progression from initial human-based assessments to more automated and complex evaluation strategies as agent complexity increases. The core idea is to systematically build confidence in agent behavior before and after deployment.
- Does GenAI \"belong\" to data scientists? — Phil Hetzel, Braintrust
This talk challenges the notion that generative AI agents exclusively belong to data scientists or machine learning engineers. It argues that while these roles bring valuable expertise in model understanding and rigorous testing, the nature of modern AI agents—built upon pre-trained models and adaptable through natural language inputs—opens the door for broader team involvement. The core thesis suggests that successful agent development benefits from diverse skill sets, including product engineers and subject matter experts, to effectively bridge the gap between complex technology and real-world problem-solving.
- Why building eval platforms is hard — Phil Hetzel, Braintrust
Building effective evaluation (eval) platforms for AI agents is a complex systems problem, not just a UI challenge. While starting with simple tools like spreadsheets is a valid first step, maturing requires moving towards more robust solutions that facilitate experimentation and integrate with production data. The core difficulty lies in managing the unique characteristics of AI agent traces, which are often large, unstructured, and high-velocity, demanding specialized data infrastructure.