← Browse

World's Fair 2025

Introduction to LLM serving with SGLang - Philip Kiely and Yineng Zhang, Baseten

Philip Kiely , Yineng Zhang , Baseten

Overview

This talk introduces SGLang, an open-source framework designed for high-performance serving of large language models (LLMs) and large vision models (LVMs). It emphasizes SGLang's production readiness, speed, and strong community support, highlighting its ability to provide day-zero support for new model releases and allow users to contribute to its development. The framework is presented as a valuable tool for optimizing LLM inference.

Who should watch

Key takeaways

Notable quotes

*SGLang offers excellent performance on a wide variety of GPUs. It's production ready out of the box.*
*If something is broken in SGLang, if you don't like something, you can fix it.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.