← Browse

Europe 2026

Running LLMs on your iPhone: 40 tok/s Gemma 4 with MLX — Adrien Grondin, Locally AI

Adrien Grondin , Locally AI

Overview

This talk demonstrates how to run large language models, specifically Gemma 4, on an iPhone using the MLX framework. It highlights the efficiency and speed achievable with on-device AI, enabling applications to run locally without an internet connection. The presentation covers the tools and resources needed to integrate these models into iOS and macOS applications, emphasizing the growing ecosystem and ease of implementation.

Who should watch

Key takeaways

Notable quotes

*MLX is a framework made by Apple that is optimized for Apple Silicon.*
*In less than 10 minutes, you can have an iOS app with a model that is running on your device.*
*4-bit is the lower I would go. And 8-bit is the higher I would go if you're using really small models.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.