Europe 2026
Does GenAI \"belong\" to data scientists? — Phil Hetzel, Braintrust
Overview
This talk challenges the notion that generative AI agents exclusively belong to data scientists or machine learning engineers. It argues that while these roles bring valuable expertise in model understanding and rigorous testing, the nature of modern AI agents—built upon pre-trained models and adaptable through natural language inputs—opens the door for broader team involvement. The core thesis suggests that successful agent development benefits from diverse skill sets, including product engineers and subject matter experts, to effectively bridge the gap between complex technology and real-world problem-solving.
Who should watch
- Data scientists and ML engineers who are increasingly tasked with building generative AI agents.
- Product managers and product engineers looking to understand how to leverage and contribute to agent development.
- Builders and engineers exploring new paradigms for AI application development beyond traditional ML.
- Teams struggling to move AI proofs-of-concept into production.
- Organizations seeking to foster cross-functional collaboration in AI initiatives.
Key takeaways
- Traditional ML development involves extensive data pipelines for training and testing models, a phase largely completed by LLM providers.
- Generative AI agents can be influenced and improved primarily by changing inputs like prompts and context, rather than solely through model retraining or feature engineering.
- Data scientists offer crucial skills in understanding model risks, establishing rigorous testing processes, and applying traditional ML metrics, but these may not fully capture agent performance.
- Product engineers are well-suited to work with LLMs as APIs and manage complex, distributed agent systems.
- Subject matter experts and product managers possess vital proximity to the problem agents aim to solve, making their input on prompts and human annotation invaluable.
- Building effective agents requires a diverse team that includes technical and non-technical experts, fostering a holistic approach to development and evaluation.
- Data scientists can provide essential "guardrails" by understanding LLM limitations and can contribute to evaluation processes, including LLM-as-judge scenarios and fine-tuning open-source models.
- Continuous feedback loops from production usage are critical for refining agent performance and ensuring evaluation metrics align with real-world outcomes.
Notable quotes
*Models already built, so we don't necessarily need to do any training and testing.*
*The way that you can change that behavior is just by changing the inputs, the prompts, the context that you're giving that that model.*
*Data scientists can be the adult in the room when they say, you know, the LLM, this is how it's trained. It's just predicting token after token. It doesn't actually know anything really.*
Unofficial community note. Prefer the recording for nuance.