← Browse

World's Fair 2025

How fast are LLM inference engines anyway? — Charles Frye, Modal

Charles Frye

Overview

This talk explores the performance of open-source LLM inference engines, highlighting how recent advancements in model quality and inference software have made self-hosting viable. It presents benchmarking data to help engineers understand and optimize LLM performance for various use cases, emphasizing the trade-offs between different configurations and workloads.

Who should watch

Key takeaways

Notable quotes

*The situation has changed. It's really exciting. It sort of finally makes sense to self-host.*
*The requests per second that we're seeing here is about four requests per second for VLM on the same workload but with like context instead of generation.*
*The latency is like almost identical in time in time to first token even though we're doing 10 times as many tokens. Basically a free lunch.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.