← Browse

World's Fair 2024

From model weights to API endpoint with TensorRT LLM: Philip Kiely and Pankaj Gupta

Overview

This talk introduces TensorRT-LLM, an NVIDIA SDK designed for high-performance deep learning inference on NVIDIA GPUs. It focuses on optimizing large language models (LLMs) to achieve higher throughput and lower latency, crucial for production environments. The presentation covers building, configuring, benchmarking, and deploying TensorRT-LLM engines, emphasizing practical application through live coding and detailed explanations.

Who should watch

Key takeaways

Notable quotes

*TensorRT is a SDK for high-performance deep learning inference on Nvidia GPUs.*
*TensorRT LLM is a mechanism on top of that that's going to give us a ton of plugins and a ton of optimization specifically for large language models.*
*The performance gains are worth it.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.