← Browse

Europe 2026

Running LLMs locally: Practical LLM Performance on DGX Spark — Mozhgan Kabiri chimeh, NVIDIA

Overview

This talk explores the practical performance of running large language models (LLMs) locally on the NVIDIA Jetson Spark, a system designed for AI development. It addresses challenges like memory limitations and software stack access that often push AI workloads to the cloud. The presentation emphasizes how local solutions can enhance developer productivity by bringing AI development closer to the user, enabling faster iteration and addressing concerns like cost predictability and data residency.

Who should watch

Key takeaways

Notable quotes

*The key idea here is not replacing the cloud, but bringing powerful AI development closer to the developer.*
*By leveraging Nvidia's NVFB4 4-bit floating-point quantization, we were able to maintain a sophisticated, high-intelligent model at a throughput that is still faster than the average human reading speed.*
*A key takeaway here from this data is that memory capacity is not the same as the memory bandwidth.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.