← Browse

Session brief

[Workshop] AI Engineering 201: Inference

103 min

Overview

This workshop focuses on the engineering challenges and considerations for deploying AI inference workloads, moving beyond the initial development of AI-powered applications. It breaks down the process into two main parts: understanding and optimizing inference, and then exploring broader architectural patterns, monitoring, and evaluation for AI applications. The core thesis is that while the capabilities of AI models are rapidly advancing, the engineering behind robust, efficient, and scalable inference is crucial for successful product deployment.

Who should watch

Key takeaways

Notable quotes

*The step is the bottleneck. This is where the vast majority of the engineering time is spent.*
*Proprietary models are fundamentally disincentivized from giving you that level of control despite the fact that it's very critical for actually effectively operating the system.*
*The secret to like the success of Open Source software in general is the ability to do this kind of like highly parallelized development where lots and lots of people are adding tiny little features.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.