← Browse

Session brief

See, Hear, Speak, Draw: Logan Kilpatrick & Simón Fishman

19 min

Overview

This talk explores the burgeoning field of multimodal AI, moving beyond text-based interactions to incorporate vision and audio. While current applications often treat different modalities as separate "islands" connected by text, the future points towards unified models capable of processing and generating across various inputs and outputs simultaneously. The presentation highlights practical patterns and demos for building with existing multimodal capabilities, anticipating future advancements.

Who should watch

Key takeaways

Notable quotes

*2023 has really been the year of chatbots and I think it's been incredible to see how much people have actually been able to do.*
*I'm excited for 2024 which I think is is really going to be the year of multimodal models.*
*The majority of the work of making multimodal systems today is like how do you hook everything up together and connect the different modalities.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.