Speaker
Dat Ngo
2 sessions in this library.
- LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize
This talk introduces a platform for LLM observability, evaluation, and experimentation, emphasizing that building AI systems is fundamentally an engineering discipline. It highlights the importance of understanding what an AI system is doing (observability), how to measure its performance (evaluation), and how to systematically improve it (experimentation). The core thesis is that these processes, while complex in the non-deterministic world of AI, can be managed and eventually automated through robust engineering practices and tooling.
- Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize
This talk focuses on building scalable LLM evaluation pipelines. It emphasizes that effective evaluation is crucial for developing high-quality AI products and goes beyond simple "LLM as a judge" approaches. The core thesis is that a comprehensive evaluation strategy involves multiple methods, continuous tuning, and integration into the development lifecycle to accelerate iteration and improve AI system performance.