← Browse

Code 2025

Coding Evals: From Code Snippets to Codebases – Naman Jain, Cursor

Naman Jain

Overview

This talk explores the evolution of evaluating AI models for coding tasks, from simple code snippets to complex codebases. It highlights challenges like data contamination and brittle test suites, proposing dynamic evaluation sets and LLM-based judges to ensure reliable and relevant assessments as AI capabilities advance. The discussion covers various stages of coding evaluation, emphasizing the need for benchmarks that reflect real-world performance and adapt to the rapid progress in AI.

Who should watch

Key takeaways

Notable quotes

*The first challenge in evaluating language models these days is like data contamination.*
*As we go to more and more real world tasks, this is going to get more challenging and we need to figure ways to combat these kind of reward hacking patterns.*
*Latency is a big concern for acceptance rates.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.