← Browse

World's Fair 2025

Vision AI in 2025 — Peter Robicheaux, Roboflow

Peter Robicheaux

Overview

This talk addresses the current state of AI vision, arguing that computer vision models lag significantly behind language models in terms of intelligence and pre-training leverage. The core thesis is that vision models are not yet "smart" due to limitations in evaluation metrics, a lack of effective large-scale pre-training utilization, and challenges in aligning visual and linguistic features. The presentation introduces new benchmarks and models aimed at improving vision AI's capabilities.

Who should watch

Key takeaways

Notable quotes

*Vision models aren't smart.*
*Large language models... cannot see.*
*Coco is too easily solvable.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.