← Browse

World's Fair 2025

Hacking the Inference Pareto Frontier - Kyle Kranen, NVIDIA

Kyle Kranen

Overview

This talk explores techniques for optimizing AI model inference to break the Pareto frontier, focusing on balancing quality, latency, and cost. The core thesis is that a well-designed system, tailored to specific application constraints, is crucial for successful deployment and application performance. By understanding and manipulating factors like scale, structure, and dynamism, engineers can achieve better service level agreements or reduce costs for existing ones.

Who should watch

Key takeaways

Notable quotes

*A good model and a good system that takes into account the actual constraints for what you need from your deployment is actually key to the success of both your deployment and the application that is backed by it.*
*The three things that we or I like to think about when I'm thinking about whether or not something can actually be deployed and used is really simple. It's quality. Latency. And cost.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.