World's Fair 2025
Vision AI in 2025 — Peter Robicheaux, Roboflow
Overview
This talk addresses the current state of AI vision, arguing that computer vision models lag significantly behind language models in terms of intelligence and pre-training leverage. The core thesis is that vision models are not yet "smart" due to limitations in evaluation metrics, a lack of effective large-scale pre-training utilization, and challenges in aligning visual and linguistic features. The presentation introduces new benchmarks and models aimed at improving vision AI's capabilities.
Who should watch
- AI Engineers
- Product Managers
- Builders working with real-world data
- Those interested in the limitations of current vision models
- Researchers focused on computer vision and multimodal AI
Key takeaways
- Current vision evaluation datasets like ImageNet and COCO are saturated and primarily test pattern matching rather than true visual intelligence.
- Unlike LLMs, vision models do not effectively leverage large-scale pre-training, leading to less capable models and embeddings.
- Large language models, despite their advancements, struggle with basic visual perception tasks, as demonstrated by their inability to accurately read clocks or determine object orientation in images.
- Vision-only pre-training methods, such as Dinov2, show promise by discovering meaningful features from vast datasets, but aligning these with language remains a challenge.
- Transformer-based vision models, unlike their convolutional counterparts, show significant performance gains when utilizing large pre-training datasets, mirroring trends seen in LLMs.
- Roboflow has developed RF-DTECTOR, a model using the Dinov2 backbone, which shows improved performance and domain adaptability, particularly on the new RF100VL benchmark.
- The new RF100VL dataset, comprising 100 diverse object detection datasets, is proposed as a more robust measure of visual model intelligence than COCO, focusing on domain adaptability and contextual understanding.
- Current vision-language models excel at linguistic generalization but falter in visual generalization, highlighting a critical area for future research and development.
Notable quotes
*Vision models aren't smart.*
*Large language models... cannot see.*
*Coco is too easily solvable.*
Unofficial community note. Prefer the recording for nuance.