Running LLMs on your iPhone: 40 tok/s Gemma 4 with MLX — Adrien Grondin, Locally AI
This talk demonstrates how to run large language models, specifically Gemma 4, on an iPhone using the MLX framework. It highlights the efficiency and speed achievable with on-device AI, enabling applications to run locally without an internet connection. The presentation covers the tools and resources needed to integrate these models into iOS and macOS applications, emphasizing the growing ecosystem and ease of implementation.
Europe 2026 11 min