Coding agents
84 sessions
- Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times
This talk explores the application of agentic AI principles to the development of mobile games, focusing on how local agents can enhance game mechanics and player experiences. The presenters discuss the potential for AI to create more dynamic and responsive game environments by processing information and making decisions directly on the device. This approach aims to reduce latency and enable richer, more complex gameplay interactions.
- Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
This talk argues that agentic systems, particularly those involved in coding, require ontologies to manage complexity and ensure reliable operation. Frank Coyle from UC Berkeley posits that as these systems become more sophisticated and capable of independent action, a structured understanding of their environment, tools, and goals becomes crucial for effective development and deployment. Without such a framework, the emergent behaviors of complex agentic systems can become unpredictable and difficult to control.
- Every Harness Will Become A Claw — Sam Bhagwat, Mastra
This talk proposes that AI harnesses are evolving beyond their initial function of managing AI agent workflows. They are increasingly integrating with collaboration tools and becoming more proactive, effectively transforming into "Claws." This evolution is driven by the capabilities of modern coding agents and the desire for AI assistants to be more present and available within team workflows.
- Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk
This talk explores the critical role of architectural decisions in agentic security, emphasizing how these foundational choices impact the overall security posture of AI systems. It delves into the complexities of building secure AI by design, highlighting the need for robust frameworks and thoughtful consideration of security implications from the outset. The core thesis revolves around the idea that effective agentic security is not an add-on but a fundamental architectural requirement.
- Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town
This talk addresses the critical security challenges arising from the increasing use of agentic systems, particularly in software development. It emphasizes the need for robust permission models, verifiable provenance, and a secure supply chain for AI agents to mitigate risks associated with their autonomous capabilities. The core thesis is that without these security measures, agentic systems pose significant threats to code integrity and system security.
- Agentic Development Security — Ezra Tanzer, Snyk
This talk addresses the security implications of agentic development, a paradigm shift where AI agents assist in the software development lifecycle. It highlights the potential risks introduced by these agents, such as new attack vectors and vulnerabilities, and emphasizes the need for robust security practices tailored to this evolving landscape. The core thesis is that securing agentic development requires a proactive and specialized approach to mitigate emerging threats.
- When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain
This talk explores the intersection of AI agents and physical data, moving beyond purely digital interactions. It introduces the concept of agent harnesses as a framework for managing and integrating AI agents with real-world data sources and physical systems. The core thesis is that understanding the "physics" of how agents interact with physical data is crucial for building more robust and capable AI applications.
- From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud
This talk addresses the challenge of silently accumulating performance issues in mature codebases, where pausing feature development for investigation is difficult. It presents a case study on integrating runtime intelligence into coding agents to enable continuous performance optimization in production. The approach involves agents analyzing real production data to identify high-return-on-investment fixes, prioritized by complexity and impact.
- Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer
This talk argues that current approaches to AI agent development, often focused on "harnessing" models or increasing scale, are insufficient for building maintainable software. The core thesis is that the failure of agent software factories stems from a fundamental problem in how coding models are trained, which prioritizes passing tests over architectural quality. This leads to codebases that degrade over time, becoming difficult to manage and debug.
- Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua
This talk introduces advancements in AI agents for computer use, moving beyond earlier, screen-dominating models to more integrated and efficient background operations. It highlights the development of tools like Quad Driver for seamless OS interaction across platforms and Kuabench for robust agent evaluation. The discussion also touches upon optimizing infrastructure for agent training to reduce costs and improve GPU utilization.
- \"The model trains the next model\" — Lee Robinson, Cursor, SpaceXAI
This talk explores the concept of AI models training future AI models, focusing on how this iterative process can lead to significant advancements. It delves into the practical implications of this approach for developers and the potential for AI to accelerate its own development cycle. The core idea is that by leveraging AI to improve AI, we can unlock new levels of performance and capability.
- More Compute In ➤ Better Model Out — Lee Robinson, Cursor
This talk explores the relationship between the amount of compute used to train AI models and the quality of their output, specifically focusing on coding agents. The core thesis suggests that increased computational resources directly correlate with improved model performance, leading to more capable AI tools for developers.
- Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI
This talk explores the process of training AI models, focusing on recursive model improvement and the intricate loops involved. It details how feedback from model usage, combined with increased compute power, drives iterative advancements. The discussion highlights the distinction between outer loops (user feedback, A/B testing) and inner loops (high-quality evaluations, complex training tasks) and emphasizes strategies to accelerate the latter for more efficient model development.
- Forward Deployed Engineering at Cursor — Pauline Brunet
Pauline Brunet discusses the Forward Deployed Engineering (FDE) function, emphasizing its critical role in enterprise AI adoption. She outlines a framework for determining when and how to implement an FDE team, based on customer digital maturity and product customization levels. The core thesis is that FDEs are highly technical individuals who act as agents of change, embedding within customer organizations to drive digital transformation and ensure successful adoption of AI technologies.
- RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI
This talk introduces Recursive Language Models (RLMs) as a method to address the context window limitations of AI coding agents when working with large codebases. The core thesis is to externalize context management into a programmable execution environment, allowing models to operate on the repository as data, curate relevant chunks, and feed them into the main context window. This approach can also serve as a memory layer for coding agents.
- The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab
This talk introduces Project Nanda, an open research effort building the infrastructure for an internet of AI agents. It addresses the current limitations of agents being confined to proprietary platforms by proposing an open, decentralized architecture analogous to the early open web. The core idea is to enable agents from different entities to discover, communicate, and transact with each other permissionlessly.
- Claws Out: Securing and Building with OpenClaw - Nick Taylor, Pomerium
This talk introduces OpenClaw, an open-source project for building and securing AI agent control planes. It highlights a new feature, trusted proxy authentication mode, which enhances security and user experience by eliminating the need for manual token entry and device pairing for web socket connections. The presentation also demonstrates how OpenClaw can be integrated into development workflows, enabling live coding and rapid prototyping of AI-powered applications.
- Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs
This talk details the process of training AI coding agents to perform tasks within spreadsheets, aiming to match their proficiency in programming languages. The project began with a 50% accuracy rate on financial analysis benchmarks and ultimately achieved 92% accuracy. The presentation covers the challenges of representing spreadsheet data to AI, various approaches explored, and the breakthroughs that led to significant improvements.
- Your coding agent doesn't always follow your rules — Talha Sheikh, Checkout.com
This talk addresses the common issue where AI coding agents, despite appearing to complete tasks, often produce outputs that fail upon execution or do not meet specifications. The core argument is that the value is shifting from the agent's ability to generate code to the developer's ability to design and implement robust verification systems, or harnesses, that ensure the agent's output is reliable and deterministic.
- Beyond the Harness: A Journey Towards Adaptative Engineering - Rajiv Chandegra, Annicha Labs
This talk introduces adaptive engineering as a new philosophy for AI development, moving beyond fixed, pre-defined harnesses. It argues that as AI models become more powerful and interact with the dynamic, messy real world, static harnesses will become brittle. Adaptive engineering allows the harness to emerge and adapt during the engineering process, with the engineer's role shifting to designing constraints rather than dictating every step.
- What if the harness mattered more than the model? - Aditya Bhargava, Etsy
This talk argues that the "harness" surrounding an AI model, which includes its tools, safety mechanisms, and reasoning capabilities, is often more critical to performance than the model itself. The speaker proposes shifting focus from solely improving large, proprietary models to developing sophisticated harnesses that can enhance the performance of smaller, open-source, locally runnable models. This approach aims to reduce reliance on closed-source AI and empower developers to build powerful agents with greater control and flexibility.
- SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI
This talk introduces SWE Marathon, a benchmark designed to evaluate the capabilities of coding agents on large-scale, project-level tasks. It addresses the growing need to assess if agents can maintain coherence and successfully complete complex engineering projects over extended operational periods, such as a billion token budget, moving beyond simple bug fixes to end-to-end project ownership.
- The Agentic AI Engineer - Benedikt Sanftl, Mutagent
The Agentic AI Engineer talk introduces a framework for building and iterating on AI agents by applying agentic principles to the development lifecycle itself. This approach aims to automate and optimize the process of agent creation, testing, and deployment, moving beyond manual iteration to a more scalable, agent-driven workflow. The core idea is to create an "agent of agents" that manages the development loop, from initial specification to production monitoring and refinement.
- OpenClaw in Your Hand: Building a Physical AI Terminal - Lech Kalinowski, Callstack
This talk details the creation of OpenClaw in Your Hand, a physical AI terminal designed for interacting with AI agents like OpenClaw. The project explores building an AI-native device using microcontrollers and a dual-display system (OLED and e-paper) for energy efficiency and a unique user experience. It highlights the challenges and potential of creating dedicated hardware for LLM interaction, moving beyond traditional screen-based interfaces.
- Building an Autonomous Engineering Org - Angie Jones, Agentic AI Foundation
This talk details the journey of transforming a large engineering organization into an autonomous one by leveraging AI agents. The core thesis is that true impact from AI enablement comes not just from individual engineers using tools, but from integrating agents deeply into the entire development workflow, enabling them to produce shippable results with minimal human oversight. This transformation involves moving engineers through maturity stages, from basic code generation to complex task delegation and multi-agent collaboration.
- Recursive Coding Agents - Raymond Weitekamp, OpenProse
This talk introduces recursive coding agents, applying the principles of Recursive Language Models (RLMs) to enhance the reliability and capability of AI agents for coding tasks. The core thesis is that current AI agents, while intelligent, are "mismanaged geniuses" lacking a robust system for specifying, managing, reusing, and verifying their work. RLMs offer a new paradigm for inference-time computation by unifying tool calling and reasoning, enabling agents to recursively break down and solve complex problems.
- Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
This talk argues that traditional benchmarks are insufficient for evaluating agentic AI systems in production. Agentic systems, which plan, call tools, and execute workflows, require a shift in evaluation focus from model capability to system behavior. The core thesis is that production telemetry and reliability metrics are paramount for dependable outcomes, moving evaluation from a pre-deployment phase to a continuous operational capability integrated into the system's control plane.
- Build Systems, Not Code - Angie Jones, Agentic AI Foundation
This talk argues that building agentic AI systems requires the same core engineering discipline as traditional software development, just with different primitives. Instead of focusing solely on using agents to write code, the emphasis shifts to architecting complex agentic systems. This approach allows engineers to leverage their existing skills in systems thinking, workflow design, and modularity, recapturing the thrill of building by operating at a higher level of abstraction.
- Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, OpenClaw
This talk explores the rapid development and deployment practices within the OpenClaw project, likening the process to a "dark factory" that ships code at an unprecedented velocity. It argues that the sheer speed of development necessitates a shift from traditional code review to a more managerial approach, where engineers oversee swarms of autonomous agents. The core thesis is that effective engineering in this new paradigm relies on process, intuition, and managing agent workflows rather than solely on individual coding prowess.
- SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius
This talk details the practical lessons learned from evaluating coding agents on real-world software engineering tasks using the rebench leaderboard. It emphasizes the critical need for robust evaluation beyond gut feelings or limited testing, especially as models are deployed to production. The presentation highlights the challenges and methodologies involved in creating and maintaining a "fresh" and "real-world" benchmark, focusing on the complexities of software engineering tasks that require understanding repository structures, writing and running tests, and handling multi-turn interactions and tool use.
- Benchmarking semantic code retrieval on Claude Code — Kuba Rogut, Turbopuffer
This talk explores the effectiveness of semantic code retrieval for AI coding assistants, specifically benchmarking it against traditional grep-based methods on Claude Code. The core thesis is that while grep is simpler and often sufficient, semantic search, powered by vector embeddings, can significantly improve precision and the agent's ability to locate relevant code, especially for complex tasks. This approach offers a form of cached compute, reducing redundant processing and potentially leading to better accuracy and user satisfaction.
- Reverse engineering a Viking VOIP phone protocol with Claude Code — Boris Starkov, Eleven Labs
This talk details the process of reverse-engineering a legacy Viking VOIP phone's protocol using Claude Code to enable connection with a modern conversational AI agent. The primary challenge was the phone's outdated Windows XP-compatible software, which posed setup difficulties for engineers using Mac operating systems. The solution involved leveraging Claude Code to discover communication ports, identify and brute-force two-letter command codes, and ultimately decipher a proprietary protocol, including a one-byte checksum encryption.
- Most Enterprise Agentic Projects Are Doomed, Here's Why — Jess Grogan-Avignon & Jack Wang, Accenture
Most enterprise agentic AI projects are destined to fail due to fundamental tensions between traditional enterprise structures and the demands of machine-speed AI development. Enterprises, built for human pace with layers of control and process, struggle to adapt to the rapid iteration and emergent behaviors characteristic of AI. This talk outlines five key tensions—speed, value, delivery, trust, and moat—and offers a framework for navigating them to achieve successful AI adoption.
- Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
This talk addresses the challenges of AI evaluations, which are currently scattered, quickly become outdated, and often lack transparency and verifiability. The speakers propose solutions to democratize the evaluation process, enabling a broader community to contribute to and benefit from robust AI benchmarking. Their work aims to foster more equitable AI development by allowing diverse expertise to shape evaluation standards.
- Lobster Trap: OpenClaw in Containers from Local to K8s and Back — Sally Ann O'Malley, Red Hat
This talk demonstrates how to run OpenClaw, an open-source coding agent, within containers, from local development environments to Kubernetes clusters. The core thesis is that containerization provides a reproducible, secure, and portable way to deploy and manage AI workloads like OpenClaw, simplifying development, onboarding, and scaling.
- Scaling Agents on Kubernetes with acpx and ACP — Onur Solmaz, OpenClaw
This talk explores the challenges and solutions for scaling open-source AI agents, particularly within the Kubernetes ecosystem. It introduces ACP (Agent Client Protocol) as a standardized way for human users to interact with agents, aiming to reduce duplicated effort in building client integrations. The discussion also covers developing automated workflows for tasks like code review and error reporting, emphasizing the potential for agents to handle complex, multi-step processes.
- Your Coding Agent Should Do AI System Engineering — Ben Burtenshaw, Hugging Face
This talk proposes that AI coding agents should be leveraged for complex AI systems engineering tasks, moving beyond simpler applications. The core idea is to use agents to tackle challenging problems in machine learning engineering and systems design. To enable this, the speaker emphasizes the need for standardized repositories, particularly on platforms like the Hugging Face Hub, which already provide many necessary components.
- Skill issue: Lessons from skilling up coding agents to use Langfuse - Marc Klingen, Clickhouse
This talk explores the challenges and lessons learned in skilling up coding agents to effectively use tools like Langfuse. It highlights the evolution from complex workflows to more autonomous agents and emphasizes the role of formalized skills as shortcuts for reliability. The core thesis is that by leveraging tracing and structured skills, developers can significantly improve agent capabilities and streamline the integration of complex tools into AI applications.
- Harnesses in AI: A Deep Dive — Tejas Kumar, IBM
This talk explores the concept of AI harnesses, defining them as the components surrounding an AI model that provide grounding in reality and ensure reliability. The speaker argues that harnesses are crucial for making AI agents dependable, especially when using rented, black-box models. By building a harness, developers can create stable environments that control agent behavior, regardless of the underlying model's non-deterministic nature.
- Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize
This talk focuses on the practical aspects of evaluating and improving AI agents, moving beyond simple "vibe checks" to implement robust testing frameworks. It emphasizes that effective evaluation is crucial for shipping reliable AI applications, especially agents, which are prone to cascading failures. The session introduces methods for capturing agent behavior through tracing, categorizing failures, and implementing various types of evaluations (code-based, LLM-as-judge, human) to ensure agents perform as expected and to drive iterative improvements.
- Make your own event-sourced agent harness using stream processors — Jonas Templestein, Iterate
This talk introduces a novel approach to building agent harnesses using an event-sourced architecture powered by stream processors. The core thesis is that by treating all agent interactions and state changes as events in a log, debugging becomes significantly easier, and extensibility is greatly enhanced. This event-driven model aims to simplify experimentation with different agent configurations and facilitate composable agent systems.
- Malleable Evals: Why Are We Evaluating Adaptive Systems with Static Tests? — Vincent Koc, OpenClaw
This talk argues that traditional static testing methods are insufficient for evaluating adaptive AI systems. As AI applications become more dynamic and intent-driven, evaluation strategies must evolve to become equally malleable. The core thesis is that static benchmarks fail to capture the emergent behaviors and changing user interactions characteristic of modern AI, necessitating a shift towards more adaptive and continuous evaluation approaches.
- A Piece of Pi: Embedding The OpenClaw Coding Agent In Your Product — Matthias Luebken, Tavon
This talk explores the integration of coding agents, specifically the OpenClaw agent, into products. It emphasizes that coding agents are becoming a fundamental building block for software systems and highlights the Pi framework as an excellent tool for experimentation and development. The presentation demonstrates how agents can be leveraged to automate complex workflows, such as processing sales proposals, by interacting with various tools and systems.
- Agentic Search for Context Engineering — Leonie Monigatti, Elastic
This talk explores agentic search as a critical component of context engineering for Large Language Models (LLMs). The core thesis is that effective context curation for LLMs relies heavily on sophisticated search mechanisms, moving beyond simple retrieval to agent-driven exploration of diverse information sources. The presentation highlights the evolution from basic RAG to agentic RAG and emphasizes the challenges and strategies for building robust search tools that agents can effectively utilize.
- Replacing 12K LoC with a 200 LoC Skill — David Gomes, Cursor
This talk details how Cursor refactored a complex feature, originally spanning 12,000 lines of code, into a concise 200-line skill. This transformation leveraged existing primitives like agent skills and sub-agents, primarily using markdown to redefine functionality. The refactoring aimed to reduce maintenance overhead and improve user experience for advanced features.
- Building your own software factory — Eric Zakariasson, Cursor
This talk explores the concept of building a "software factory" using AI agents, moving beyond simple code completion to autonomous systems that can handle complex development tasks. The core idea is to borrow principles from physical factories, such as assembly lines and robust infrastructure, and apply them to software development to increase throughput and consistency. While fully autonomous factories are still aspirational, significant progress can be made by implementing structured primitives, guardrails, and enablers for AI agents.
- Harness Engineering: How to Build Software When Humans Steer, Agents Execute — Ryan Lopopolo, OpenAI
This talk introduces Harness Engineering, a paradigm shift in software development where AI agents execute tasks guided by human direction. The core thesis is that implementation is no longer the bottleneck; code is abundant and cheap to produce. The focus shifts to systems thinking, design, and delegation, empowering engineers to leverage AI agents for complex, long-horizon work. This approach aims to maximize productivity by automating implementation and freeing human engineers for higher-leverage activities.
- Agentic Engineering: Working With AI, Not Just Using It — Brendan O'Leary
This talk introduces the concept of agentic engineering, shifting the paradigm from merely using AI tools to actively collaborating with them. It highlights the evolution of AI in software development from simple line completion to sophisticated agents capable of executing complex tasks. The core thesis is that understanding and managing AI agents as collaborators, akin to junior developers, is crucial for maximizing their potential and navigating the complexities of modern software development.
- Spec-Driven Development: Agentic Coding at FAANG Scale and Quality — Al Harris, Amazon Kiro
This talk introduces Spec-Driven Development as a method to enhance AI agent development, focusing on improving control, code quality, and reliability. It proposes a structured workflow that compresses the software development lifecycle, moving from prompt to requirements, design, and execution, with an emphasis on creating reproducible results. The approach aims to integrate established software engineering practices with AI capabilities to tackle more complex problems.
- How Claude Code Works - Jared Zoneraich, PromptLayer
This talk explores the advancements in coding agents, particularly focusing on how Claude Code works and the underlying innovations that have made these tools effective. The core thesis is that simplicity in architecture, coupled with better underlying models and a focus on tool calling, has been the key breakthrough. The presentation emphasizes a philosophy of "give it tools and get out of the way," advocating for leaning into the model's capabilities rather than over-engineering complex systems.
- Developer Experience in the Age of AI Coding Agents – Max Kanat-Alexander, Capital One
This talk explores how to optimize developer experience in the era of AI coding agents. It argues that investments in foundational aspects of software development, such as standardized environments, robust validation, and clear documentation, will benefit both human developers and AI agents. The core thesis is that practices good for human developers are also good for AI, ensuring long-term value regardless of AI advancements.
- Coding Evals: From Code Snippets to Codebases – Naman Jain, Cursor
This talk explores the evolution of evaluating AI models for coding tasks, from simple code snippets to complex codebases. It highlights challenges like data contamination and brittle test suites, proposing dynamic evaluation sets and LLM-based judges to ensure reliable and relevant assessments as AI capabilities advance. The discussion covers various stages of coding evaluation, emphasizing the need for benchmarks that reflect real-world performance and adapt to the rapid progress in AI.
- Hard Won Lessons from Building Effective AI Coding Agents – Nik Pash, Cline
This talk argues that the effectiveness of AI coding agents is primarily determined by the underlying model's capability, not by complex engineering scaffolds. Frontier models, when unhindered, can outperform many agent combinations. The core message is to simplify agent engineering and focus on improving model training through robust benchmarks and reinforcement learning environments derived from real-world coding tasks.
- Future-Proof Coding Agents – Bill Chen & Brian Fioca, OpenAI
This talk explores the architecture and development of coding agents, emphasizing the crucial role of the "harness" – the interface layer that connects models to users and tools. It highlights the rapid evolution of AI models and the challenges in adapting agents, proposing that building models and harnesses together, as exemplified by OpenAI's Codeex, offers a more robust and future-proof approach. The discussion also touches upon emerging patterns for integrating these agents into various products and workflows.
- Building Cursor Composer – Lee Robinson, Cursor
Cursor Composer is an AI model designed for real-world software engineering, aiming to balance speed and intelligence. It performs better than top open-source models and rivals frontier models in intelligence, while being significantly more efficient in token generation. The development focused on creating a model that could serve as a daily driver for coding tasks, integrating features like parallel tool calling and semantic search to enhance its capabilities.
- Developing Taste in Coding Agents: Applied Meta Neuro-Symbolic RL — Ahmad Awais, CommandCode
This talk introduces Command Code, a coding agent designed to learn and adapt to an individual developer's unique coding style and preferences, referred to as "taste." Unlike standard AI coding assistants that provide generic code, Command Code aims to internalize a developer's decision-making process, preferences for tools, and architectural choices. This is achieved through a meta-neuro-symbolic reinforcement learning approach, creating a more personalized and efficient coding experience.
- AI Copilots for Tech Architecture: The Highest-ROI Use Case You’re Not Building — Boris B., Catio
This talk argues that AI copilots for tech architecture represent the highest ROI use case currently underutilized by organizations. While coding copilots have become standard, an architecture co-pilot can prevent costly mistakes by ensuring development efforts are directed correctly from the outset. This approach aims to move beyond tribal knowledge and gut instinct, providing data-backed guidance for architectural decisions.
- Infra that fixes itself, thanks to coding agents — Mahmoud Abdelwahab, Railway
This talk introduces a system where an AI coding agent monitors application infrastructure for issues and automatically generates pull requests to fix them. Instead of just alerting developers to problems like memory leaks or high error rates, the system aims to detect, diagnose, and propose solutions, significantly reducing manual debugging and speeding up the resolution process. The core idea is to shift from reactive alerting to proactive, automated self-healing infrastructure.
- Building an Agentic Platform — Ben Kus, CTO Box
This talk explores Box's journey in building an agentic platform, focusing on how agentic approaches can solve complex data extraction challenges that traditional methods and basic LLM calls struggle with. The core thesis is that an agentic abstraction layer provides a clean, flexible, and evolvable architecture for tackling sophisticated AI tasks, moving beyond simple chatbot interactions to orchestrate complex workflows.
- Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith
This talk addresses the critical need for robust evaluation and benchmarking strategies when deploying large language models (LLMs) into production. It highlights the inherent complexities and potential pitfalls of generative AI, emphasizing that scalability, reliability, and safety are paramount. The presentation introduces practical tools and methods to assess LLM performance, ensuring that models meet enterprise-level requirements before and during deployment.
- Piloting agents in GitHub Copilot - Christopher Harrison, Microsoft
This talk explores GitHub Copilot's capabilities beyond basic code completion, focusing on its evolution into an agentic tool. It highlights how providing context, clear instructions, and leveraging features like Copilot Coding Agent can significantly enhance developer productivity. The discussion emphasizes that while AI tools are powerful, fundamental DevOps practices and human oversight remain crucial for secure and effective software development.
- Your Coding Agent Just Got Cloned And Your Brain Isn't Ready - Rustin Banks, Google Jules
This talk introduces Jules, an asynchronous AI coding agent designed to run in the background and handle tasks in parallel, freeing developers to focus on creative coding. The core idea is to shift from serial task execution to a parallel workflow, leveraging AI for both the beginning and end of the software development lifecycle, from task creation to code merging and testing.
- From Copilot to Colleague: Trustworthy Agents for High-Stakes - Joel Hron, CTO Thomson Reuters
This talk explores the evolution of AI assistants from being merely helpful to becoming productive agents capable of making judgments and decisions. It emphasizes that agency in AI is not a binary state but a spectrum that can be tuned based on use case, risk tolerance, and user expectations. The presentation highlights the challenges and lessons learned in building trustworthy AI systems for high-stakes professional environments, particularly within legal, tax, and global trade sectors.
- Agentic GraphRAG: AI’s Logical Edge — Stephen Chin, Neo4j
This talk introduces Agentic GraphRAG as a method to improve the accuracy and reduce hallucinations in AI agent systems. It highlights the limitations of current LLMs in complex reasoning and decision-making, proposing that knowledge graphs, when integrated with LLMs, can provide a more robust and logical framework for AI operations. The approach leverages graph databases to store, manage, and retrieve information, enhancing the AI's ability to understand context and provide reliable outputs.
- Real world MCPs in GitHub Copilot Agent Mode — Jon Peck, Microsoft
This talk introduces GitHub Copilot's Agent Mode, an advanced feature designed for completing moderately complex tasks autonomously. It moves beyond simple code completion and chat interactions to enable deep, iterative engagement with an AI agent. The presentation highlights how Agent Mode can build entire applications from a readme file or perform significant refactoring, with the developer providing permission for actions like terminal interactions.
- The rise of the agentic economy on the shoulders of MCP — Jan Curn, Apify
This talk explores how general intelligence in computing systems may emerge from the interaction of multiple agents, similar to how intelligence arises in biological systems or markets. The speaker posits that the Multi-agent Conversation Protocol (MCP) is a crucial development enabling agents to communicate and form an agentic mesh, facilitating the rise of an agentic economy where agents can discover and utilize services.
- Training Agentic Reasoners — Will Brown, Prime Intellect
This talk argues that reasoning and agents are fundamentally the same concept, with reinforcement learning (RL) being the key to developing powerful agentic systems. The speaker posits that traditional approaches to building agents often involve manual, iterative tuning that mirrors RL processes. By framing agent development through the lens of RL, developers can leverage established algorithms and techniques to create more robust and capable agents, especially for complex, multi-turn tasks.
- Claude Code & the evolution of agentic coding — Boris Cherny, Anthropic
The talk by Boris Cherny from Anthropic discusses the rapid evolution of AI models in coding and the challenges in developing user experiences (UX) that keep pace. It highlights how programming itself has evolved through layers of abstraction and interface changes, from punch cards to modern IDEs. The core thesis is that while models are advancing exponentially, the product development for these AI coding tools is still in its early stages, with a focus on building unopinionated, minimal-viable products to learn and adapt to the unknown optimal UX.
- The emerging skillset of wielding coding agents — Beyang Liu, Sourcegraph / Amp
This talk explores the evolving skillset required to effectively utilize coding agents, moving beyond early chatbot paradigms. It argues that the current era demands a new approach to agent interaction, focusing on empowering agents to perform tasks autonomously rather than micromanaging them. The core thesis is that mastering coding agents is a high-ceiling skill that will significantly enhance developer productivity, akin to learning a new programming language or editor.
- Agentic Excellence: Mastering AI Agent Evals w/ Azure AI Evaluation SDK — Cedric Vidal, Microsoft
This talk focuses on the critical process of evaluating AI agents, emphasizing a shift from ad-hoc testing to a more methodical approach. It highlights that effective evaluation should begin at the earliest stages of AI development, not as an afterthought. The presentation introduces tools and frameworks, particularly the Azure AI Evaluation SDK, to ensure AI agents behave correctly and safely as they gain more independence.
- Agentic GraphRAG: Simplifying Retrieval Across Structured & Unstructured Data — Zach Blumenfeld
This talk introduces Agentic GraphRAG, a method for simplifying data retrieval across both structured and unstructured sources. It proposes using a knowledge graph as an intermediary layer to enhance agentic workflows. By modeling data within a graph, agents can more accurately decompose complex questions, pull relevant information, and perform analytical tasks that go beyond simple semantic search.
- \"Data readiness\" is a Myth: Reliable AI with an Agentic Semantic Layer — Anushrut Gupta, PromptQL
The talk argues that the concept of "data readiness" for AI is a myth, as data is rarely perfect and constantly changing. Instead of striving for pristine data, the focus should be on building AI systems that can reliably work with messy, evolving data. This is achieved by creating an agentic semantic layer that learns and adapts to the specific business domain and its nuances over time, much like an experienced human analyst.
- Building Agentic Applications w/ Heroku Managed Inference and Agents — Julián Duque & Anush Dsouza
This talk introduces Heroku's Managed Inference and Agents service, designed to simplify the development and deployment of agentic AI applications. The service allows developers to integrate AI models directly within their application's infrastructure, enhancing security and control. It provides primitives for inference, model context protocol (MCP) integration, and a managed PostgreSQL database with PGVector for embeddings, enabling the creation of sophisticated AI-powered features.
- Agentic Enterprise - What your CEO must know about AI - Hubert Misztela
This talk explores the concept of agentic enterprise, positing that AI agents could fundamentally reshape organizations within three years. It defines AI agents as autonomous, LLM-based applications capable of planning, tool usage, and dynamic adaptation. The presentation emphasizes that digital assets should evolve to serve as tools or agents themselves, enabling capabilities like multi-step reasoning, adaptability, and computer vision for automation.
- Self Coding Agents — Colin Flaherty, Augment Code
This talk explores the development and capabilities of AI coding agents, focusing on how these agents can contribute to their own creation and improvement. The core thesis is that AI agents are rapidly evolving and will significantly change software engineering, with a notable statistic indicating that over 90% of a 20,000-line codebase for an agent was written by the agent itself under human supervision. The presentation highlights the practical applications, lessons learned, and future implications of this technology.
- How Windsurf writes 90% of your code with an Agentic IDE - Kevin Hou, Windsurf
Windsurf is an AI agent-powered IDE designed to significantly increase developer productivity by automating a large portion of code generation. The core thesis is that agents are the future of software development, capable of handling routine tasks and complex problem-solving, allowing developers to focus on higher-level product building and feature creation. Windsurf aims to keep developers in a state of flow by minimizing manual input and maximizing the agent's contribution.
- Agentic Workflows on Vertex AI: Rukma Sen
This talk explores the concept of AI agents as the primary interface for human interaction with generative AI. It defines an agent as a system designed to achieve specific goals through environmental interaction, comprising a reasoning model, tools for action, and orchestration for memory and state management. The discussion emphasizes the responsibility that comes with building these agents, focusing on ethical considerations, safety, cybersecurity, and data privacy.
- [Full Workshop from Microsoft] Github Copilot - The World's Most Widely Adopted AI Developer Tool
This workshop provides a comprehensive guide to GitHub Copilot, positioning it as the world's most adopted AI developer tool. It covers Copilot's functionality as an AI pair programmer, its integration within Integrated Development Environments (IDEs), and best practices for maximizing its utility. The session emphasizes how Copilot synthesizes code based on context from open files, comments, and direct chat queries, aiming to streamline the software development lifecycle.
- GitHub Copilot: The World's Most Widely Adopted AI Developer Tool
GitHub Copilot has evolved from an AI pair programmer generating code snippets to a comprehensive tool integrated across development workflows. It now offers chat functionalities within IDEs and on GitHub.com, enabling code explanation, debugging, test generation, and interaction with enterprise knowledge bases.
- Build, Evaluate and Deploy a RAG-Based Retail Copilot with Azure AI: Cedric Vidal and David Smith
This talk demonstrates how to build, evaluate, and deploy a Retrieval-Augmented Generation (RAG) based copilot for a retail environment using Azure AI services. The core concept involves creating a chatbot that can answer customer questions by accessing product information from a vector database and customer data from a relational database, all orchestrated through Azure AI Studio's Prompt Flow.
- Creating and scaling your own custom copilots with Azure AI Studio: Hanchi Wang
This talk introduces Azure AI Studio and Prompt Flow, a suite of tools designed to streamline the development, evaluation, and deployment of AI applications. The core thesis is that while large language models are powerful, they require integration with domain-specific knowledge and tools, along with careful filtering, evaluation, and continuous monitoring to be effective in enterprise settings. These tools aim to make AI application building efficient and trustworthy.
- Building Reliable Agentic Systems: Eno Reyes
This talk explores practical lessons learned from building autonomous software engineering systems, referred to as droids. It defines agentic systems by three core characteristics: planning, decision-making, and environmental grounding. The discussion emphasizes strategies for enhancing reliability and effectiveness in these systems, drawing inspiration from control systems, robotics, and symbolic AI.
- Copilots Everywhere: Thomas Dohmke and Eugene Yan
This talk explores the evolution and integration of AI-powered coding assistants, focusing on GitHub Copilot. It posits that AI tools are shifting from simple autocompletion to comprehensive development partners, aiming to enhance developer productivity and democratize access to coding. The core thesis is that AI should augment human capabilities, not replace them, by handling tedious tasks and facilitating exploration within the codebase.
- Harnessing the Power of LLMs Locally: Mithun Hunsur
This talk introduces lm.rs, a Rust library designed for local inference of large language models (LLMs). It highlights the advantages of running LLMs locally, such as greater control, privacy, and potential cost savings compared to cloud-based solutions. The library aims to provide a flexible, Rust-native experience for developers to integrate various LLM architectures into their applications.