World's Fair 2025
Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3 — Greg Kamradt, ARC Prize Foundation
Greg Kamradt , ARC Prize Foundation
Overview
This talk introduces ARC-AGI-3, a new benchmark designed to measure artificial general intelligence (AGI) by focusing on interactive reasoning and skill acquisition efficiency. The benchmark aims to create problems that are solvable by humans but challenging for current AI, thereby guiding AI research and development towards human-level intelligence. It moves beyond single-turn, static benchmarks to simulate more realistic, open-world exploration and learning scenarios.
Who should watch
- AI Engineers
- Product Managers
- Researchers
- Builders focused on AI evaluation and benchmarking
- Those interested in the practical measurement of AGI capabilities
Key takeaways
- Current AI agents, while capable of impressive feats like playing games, often get stuck, require intervention, or rely on training data similar to the task, indicating a lack of true generalization.
- Human intelligence is proposed as the benchmark for AGI, focusing on the ability to learn new skills efficiently and perform tasks not seen before.
- ARC-AGI-3 is an interactive reasoning benchmark featuring over 1,000 novel tasks, where each task requires a unique skill, preventing simple memorization and testing true generalization.
- The benchmark will utilize games as a medium for interactive reasoning, overcoming limitations of previous game-based evaluations like Atari, such as dense rewards and developer bias.
- ARC-AGI-3 emphasizes exploration and understanding through interaction within controlled environments with defined rules and sparse rewards, rather than providing explicit instructions.
- Core knowledge priors such as basic math, geometry, agentness, and objectness are the only assumed knowledge, stripping away language, text, and trivia to focus on abstract reasoning.
- Evaluation will be based on skill acquisition efficiency, measuring how quickly an AI can explore, intuit, set goals, and complete objectives compared to a human baseline.
- A public sandbox preview with five games and a mini agent competition is planned for next month, with a goal of launching approximately 120 games by Q1 2026.
Notable quotes
*Intelligence is skill acquisition efficiency.*
*As long as we can come up with problems that humans can still do but machines cannot, I would again assert that we do not have AGI.*
Unofficial community note. Prefer the recording for nuance.