Speaker
Aman Khan
2 sessions in this library.
- Shipping AI That Works: An Evaluation Framework for PMs – Aman Khan, Arize
This talk introduces an evaluation framework for AI product managers (PMs) focused on shipping reliable AI applications. It emphasizes the critical need for robust evaluation, analogous to software testing but adapted for the non-deterministic nature of AI models. The framework aims to provide PMs with tools and methodologies to build confidence in AI product performance, moving beyond subjective "vibe coding" to data-driven "thrive coding."
- Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize
This talk focuses on building scalable LLM evaluation pipelines. It emphasizes that effective evaluation is crucial for developing high-quality AI products and goes beyond simple "LLM as a judge" approaches. The core thesis is that a comprehensive evaluation strategy involves multiple methods, continuous tuning, and integration into the development lifecycle to accelerate iteration and improve AI system performance.