Europe 2026
Shipping complex AI applications — Braintrust & Trainline
Overview
This talk focuses on the practical challenges of shipping complex AI applications, moving beyond initial prototypes to robust, production-ready systems. It emphasizes the need for operational rigor, structured workflows, and continuous evaluation, drawing on experiences from Braintrust and Trainline. The core thesis is that while AI models are increasingly sophisticated, the operational practices for deploying and managing them at scale have lagged, creating a significant hurdle for delivering real customer value.
Who should watch
- AI Engineers
- Product Managers
- Builders working on AI systems
- Those struggling to move AI prototypes into production
- Teams looking to improve the quality and reliability of their AI applications
- Engineers interested in observability and evaluation for AI
Key takeaways
- The gap between AI prototypes and production-ready systems is often due to a lack of operational rigor, not model capability.
- Traditional software engineering principles, like breaking down complex systems into smaller, manageable stages, are crucial for building scalable AI applications.
- Observability, through tracing and logging, is essential for understanding AI system behavior, identifying failure modes, and debugging in production.
- Implementing a continuous feedback loop, or flywheel, involving evaluation, remediation, and monitoring, is key to improving AI application performance over time.
- Tools like Braintrust can help instrument AI systems, manage prompts and tools, and facilitate rigorous evaluation and deployment at scale.
- Moving from local development to managed environments in platforms like Braintrust allows for better collaboration, versioning, and reproducibility.
- Both deterministic and LLM-as-a-judge evaluation methods are valuable for assessing AI application quality, especially for nuanced or non-deterministic aspects.
- Production data is invaluable for evaluating AI systems; starting with a golden dataset and iterating based on real-world signals is a practical approach.
Notable quotes
*The type of operational rigor when it comes to delivering these systems at scale has not kind of kept up.*
*Logs will tell you what has happened, but sometimes you need to go deep into the system and understand its behavior.*
*If you've got a production in application and you're not tracing it, you need to go back to the drawing board and get that done.*
Unofficial community note. Prefer the recording for nuance.