Any-to-Any: Building Native Multimodal Agents - Patrick Löber, Google DeepMind
This talk introduces the concept of "any-to-any" multimodal agents, enabled by Google DeepMind's Gemini API. The core idea is to build agents that can understand and generate across various modalities including text, code, images, audio, and video. The presentation outlines the architecture for such agents, focusing on multimodal understanding and native generation capabilities, and demonstrates how to build a notebook LM clone as an example.
Europe 2026 16 min