← Browse

World's Fair 2025

AI Engineering with the Google Gemini 2.5 Model Family - Philipp Schmid, Google DeepMind

Philipp Schmid

Overview

This talk introduces the Google Gemini 2.5 model family, focusing on practical applications for AI engineers and builders. It highlights the multimodal capabilities of Gemini 2.5 Pro and Flash, demonstrating how to leverage these models for text generation, image and audio understanding, function calling, and integrating with external tools via MCP servers. The session emphasizes hands-on learning through a workshop format, encouraging attendees to experiment with the models and SDK.

Who should watch

Key takeaways

Notable quotes

*Gemini 2.5 models are multimodal by default meaning they can understand text, images, audio, videos, documents and can generate text.*
*The best thing always is to test and to explore and evaluate and even if you need to run like a thousand PDFs it's not very cost or like expensive anymore.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.