Hacking the Inference Pareto Frontier - Kyle Kranen, NVIDIA
This talk explores techniques for optimizing AI model inference to break the Pareto frontier, focusing on balancing quality, latency, and cost. The core thesis is that a well-designed system, tailored to specific application constraints, is crucial for successful deployment and application performance. By understanding and manipulating factors like scale, structure, and dynamism, engineers can achieve better service level agreements or reduce costs for existing ones.
World's Fair 2025 20 min