World's Fair 2026
144 sessions tagged from titles
- Claude for Long-Horizon Tasks — Lance Martin, Anthropic
This talk explores building agent harnesses for reliable and secure long-horizon tasks using Claude. It emphasizes decoupling the agent's reasoning (brain) from its actions (hands), implementing self-verification and self-learning mechanisms, and designing for adaptability in evolving agent harness systems. The core thesis is that Claude's capabilities can be effectively leveraged for complex, extended tasks through robust harness architecture.
- Better Agent Auth — Bereket Habtemeskel & Paola Estefania, Better Auth
This talk addresses the critical need for robust authentication and authorization mechanisms within AI agent development. It emphasizes that as AI agents become more integrated into workflows and handle sensitive data, secure access control is paramount. The presentation likely explores common pitfalls in agent authentication and proposes solutions to ensure agents operate within defined permissions and trust boundaries.
- Full Workshop: Better Agent Auth — Paola Estefania, Better Auth
This workshop focuses on enhancing authentication for AI agents, aiming to provide a more robust and secure framework for their integration and operation. It delves into the critical aspects of building trustworthy AI systems by addressing the complexities of agent authorization and access control. The core thesis is that improved authentication mechanisms are fundamental to the safe and effective deployment of AI agents in various applications.
- Every Harness Will Become A Claw — Sam Bhagwat, Mastra
This talk proposes that AI harnesses are evolving beyond their initial function of managing AI agent workflows. They are increasingly integrating with collaboration tools and becoming more proactive, effectively transforming into "Claws." This evolution is driven by the capabilities of modern coding agents and the desire for AI assistants to be more present and available within team workflows.
- HTML Is All Agents Need — James Russo, HeyGen
This talk explores the potential of leveraging AI agents to generate HTML, moving beyond traditional templating methods. It posits that a more agentic approach can lead to more dynamic and adaptable web development. The core idea is to treat HTML generation as a task that AI agents can perform, potentially simplifying complex UI development.
- \"The biggest challenge in your stack? Evals, Evals, Evals\" - 2026 State of AI Engineering results
The State of AI Engineering 2026 survey highlights evaluation (evals) as the most significant challenge across AI stacks. This talk delves into the survey results, emphasizing the critical need for robust evaluation methodologies in AI development. Addressing evals is presented as paramount for building reliable and high-quality AI systems.
- 2026 State of AI Engineering — Barr Yaron, Amplify Partners
This talk provides a snapshot of the current landscape and future trajectory of AI engineering. It emphasizes the rapid evolution of tools and methodologies, highlighting the increasing importance of practical application and product development alongside foundational research. The core thesis suggests a shift towards more integrated and accessible AI development, driven by open-source contributions and a focus on developer experience.
- The 2026 State of AI Engineering — Barr Yaron, Amplify Partners
Barr Yaron of Amplify Partners discusses the current landscape of AI engineering, focusing on the practical application and development of AI technologies. The talk highlights key trends and challenges faced by engineers and product managers in building and deploying AI-powered products. It emphasizes the evolution of AI tools and methodologies shaping the industry.
- Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest
The core thesis is that agent architectures have a rapid obsolescence rate, with a half-life of approximately six months. This necessitates a continuous focus on updating and adapting agent systems to remain effective in the fast-evolving AI landscape. The talk emphasizes the need for builders to anticipate and manage this rapid change.
- The Desktop Frontier — Ahmad Osman, Osmantic
This talk explores the emerging landscape of desktop AI agents, positioning the personal computer as the next frontier for AI development and deployment. It argues that the desktop environment offers unique advantages for AI, moving beyond cloud-based models to enable more integrated and powerful user experiences. The discussion highlights the potential for AI to fundamentally change how users interact with their computers.
- Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk
This talk explores the critical role of architectural decisions in agentic security, emphasizing how these foundational choices impact the overall security posture of AI systems. It delves into the complexities of building secure AI by design, highlighting the need for robust frameworks and thoughtful consideration of security implications from the outset. The core thesis revolves around the idea that effective agentic security is not an add-on but a fundamental architectural requirement.
- In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar
As AI agents become more capable of handling complex development tasks, the primary challenge shifts from code generation to verification. Hallucinations are an inherent issue, and as models improve, their failures become more frequent and convincing, increasing the risk of human reviewers succumbing to cognitive surrender. This talk proposes a three-stage discipline for responsible agentic development: Guide, Verify, Solve. It argues that robust verification infrastructure is essential for safety and provides a competitive edge.
- Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog
This talk addresses the common issue of AI agents exhibiting inconsistent behavior, often disagreeing with their own previous outputs or decisions. It explores the underlying causes of this self-disagreement and provides practical strategies for engineers to mitigate these problems, leading to more reliable and predictable agent performance. The core thesis is that understanding and managing agent internal state and decision-making processes is crucial for building robust AI systems.
- Don't Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft
This talk emphasizes that while Large Language Models (LLMs) are powerful tools, they should not be given complete control over development processes. Instead, LLMs should function as sophisticated assistants, augmenting human capabilities rather than replacing human judgment and oversight. The core thesis is to leverage LLMs strategically to enhance productivity and creativity without relinquishing essential human decision-making.
- Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft
This talk challenges the assumption that advanced frontier models are necessary for building effective voice agents. The presenters argue that simpler, more focused models can achieve comparable or even superior results for specific voice agent tasks, emphasizing practicality and efficiency over raw model power. They advocate for a pragmatic approach to agent development, prioritizing task-specific capabilities.
- Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla
Ishita Daga of Tesla addresses a fundamental challenge in enterprise AI: the lack of a standardized structure for AI agents. She argues that current agent development often lacks a clear framework, leading to inefficiencies and difficulties in scaling. The talk emphasizes the need for a more organized approach to building and deploying AI agents within large organizations.
- When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain
This talk explores the intersection of AI agents and physical data, moving beyond purely digital interactions. It introduces the concept of agent harnesses as a framework for managing and integrating AI agents with real-world data sources and physical systems. The core thesis is that understanding the "physics" of how agents interact with physical data is crucial for building more robust and capable AI applications.
- Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared)
This talk introduces the concept of a Go-To-Market (GTM) agent designed to understand buyer behavior and inform product strategy. It emphasizes the need for AI agents that can move beyond simple task execution to possess a deeper understanding of market dynamics and customer needs. The core thesis is that by building agents with this buyer-centric intelligence, companies can more effectively align their product development and marketing efforts.
- Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio
This talk argues that current AI agents, despite advancements in tool calling, often lack the ability to provide verifiable evidence for their actions. The core thesis is that agents need to generate "receipts" or verifiable outputs to build trust and enable effective debugging and auditing, rather than simply being able to call more tools. This shift is crucial for moving beyond basic automation to reliable and accountable AI systems.
- Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens
This talk addresses the critical gap between raw AI agent output and a user-friendly experience, arguing that a dedicated rendering layer is essential for LLM pipelines. It emphasizes that simply displaying raw LLM responses fails to meet user expectations for intuitive and effective interaction. The core thesis is that building a robust rendering layer is key to successfully shipping AI-powered products.
- Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS
This talk explores the development of voice agents capable of handling interruptions, a crucial feature for natural and effective human-computer interaction. The presenters discuss the challenges and solutions involved in creating agents that can seamlessly manage conversational flow when users interject or change topics mid-utterance. The core thesis revolves around building more robust and user-friendly voice interfaces.
- Skills are the New SDKs - Elvin Aghammadzada, DataRobot
This talk posits that AI skills are evolving into a new form of Software Development Kits (SDKs), fundamentally changing how developers build and interact with AI systems. The core thesis is that by abstracting complex AI functionalities into reusable skills, developers can accelerate innovation and create more sophisticated applications. This shift emphasizes modularity and interoperability, akin to traditional software libraries.
- Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest
This talk addresses the common problem of failing jobs in Apache Spark, a widely used distributed computing system. It introduces Medic, a tool designed to provide first aid for these failing jobs, aiming to simplify debugging and improve the reliability of Spark applications. The core thesis is that by providing targeted diagnostics and actionable insights, developers can more efficiently resolve issues in their Spark workloads.
- Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs
This talk explores the potential for AI to automate complex oncology workflows, aiming to reduce or eliminate the need for human intervention in critical processes. It delves into the challenges and possibilities of integrating AI into healthcare, specifically within the demanding field of cancer treatment. The core thesis examines whether AI can independently manage and execute tasks traditionally requiring expert human oversight.
- Perception Agents — Antje Barth, Amazon AGI Lab
This talk introduces perception agents, a new paradigm in human-agent collaboration that moves beyond text-based interactions. Current agents often require detailed textual prompts and struggle with dynamic visual interfaces or unexpected application behaviors. Perception agents aim to bridge this gap by enabling AI to see, understand, and interact with computer interfaces similarly to humans, using clicks and keystrokes to act upon visual information.
- Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex
This workshop focuses on leveraging OpenAI's Codex as a versatile tool for computer control. It explores setting up persistent memory, creating collaborative assistant threads, and managing long-running computational tasks. The session emphasizes practical methods for integrating Codex into workflows, enabling complex operations through iterative prompting and loop-based execution.
- From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs
This talk explores the engineering hurdles in developing foundation models specifically for single-cell biology, presented from the viewpoint of a machine learning engineer without a biology background. It delves into the unique challenges and considerations required to adapt large model technologies for biological data analysis.
- The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest, DSPy
This talk introduces DSPy, a framework that separates the definition of an AI task from its implementation. By declaring a task's inputs and outputs first, developers can later experiment with different models, settings, and execution strategies without altering the core task definition. This approach aims to make AI engineering more flexible and robust, especially in a rapidly evolving landscape of models and tools.
- You Didn't Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit
This talk addresses a critical oversight in agent development: designing infrastructure for human interaction rather than agentic behavior. The core thesis is that many production LLM errors, particularly those related to rate limits, stem from this fundamental mismatch. Infrastructure built for human speed and interaction patterns fails when agents make thousands of calls per second, leading to performance issues and unexpected costs.
- From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud
This talk addresses the challenge of silently accumulating performance issues in mature codebases, where pausing feature development for investigation is difficult. It presents a case study on integrating runtime intelligence into coding agents to enable continuous performance optimization in production. The approach involves agents analyzing real production data to identify high-return-on-investment fixes, prioritized by complexity and impact.
- Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
This talk focuses on the critical importance of building effective evaluation systems for AI models, moving beyond superficial metrics to create benchmarks that genuinely reflect real-world performance and user needs. The presenters emphasize that robust evals are essential for shipping reliable AI products and driving meaningful improvements in model quality.
- Build Evals That Actually Matter - Nick Ung, Lyft
This talk addresses the common problem of AI evaluations that fail to predict real-world performance. The core thesis is that offline evaluations often use simplistic "customers" and test sets that don't reflect the complexity and adversarial nature of actual user interactions, leading to shipped models that fail in production. A solution involves building more realistic, adversarial user simulators trained on real data.
- Notion's Token Town — Sarah Sachs, Notion
This talk addresses the significant cost challenges in AI development, particularly the escalating expense of using large language models. It argues that focusing solely on token economics is a losing strategy, as model providers often increase prices or deprecate older versions. Instead, the core thesis is to win on product by leveraging data flywheels, sophisticated orchestration, and maintaining model agnosticism to preserve flexibility and avoid vendor lock-in.
- Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait
This talk explores the application of autonomous agents to complex scientific discovery tasks, moving beyond simpler coding puzzles or optimization problems. It highlights the necessity for agents to engage with real-world measurement data and employ a scientific method, including hypothesis generation, model implementation, and learning from failures. The core thesis is that significant advancements in agent performance for scientific tasks stem from formulating accurate hypotheses about physical systems and correctly implementing them.
- A Practitioner's Guide to Graphs - Tim Ainge, Good Collective
This talk provides a practical introduction to graph databases and algorithms. It covers the fundamentals of what a graph is, how to extract graph data from unstructured text, and the benefits of using a schema-first approach with ontological improvements. The presentation also delves into specific graph algorithms like personalized PageRank, shortest path, and subgraph matching, illustrating their principles with accessible code examples and real-world applications.
- Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
This talk presents a framework for evaluating and improving AI agents, emphasizing a "rollout-centered" approach. The speakers connect sandboxed environments, agent evaluations, and optimization workflows to create a practical system for generating and learning from agent performance data. The core idea is that treating agent development as a series of rollouts, similar to software deployments, allows for systematic improvement.
- The UX of AI: Making AI-Powered Apps Your Users Don't Hate - Kathryn Grayson Nanz, Progress Software
This talk addresses the critical gap between the excitement AI developers feel and the apprehension users often experience with AI-powered applications. It emphasizes that successful AI product adoption hinges on thoughtful user experience (UX) design, focusing on building user trust, ensuring privacy and security, and guiding users through novel interaction patterns. The core thesis is that teaching users how to effectively interact with AI capabilities is essential for them to derive value from new applications.
- Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse
This talk addresses the critical need for domain expertise in AI self-improvement loops. It argues that without a deep understanding of the specific domain, automated improvement processes can be inefficient or ineffective. The core thesis is that the most efficient path to a continuously improving agentic system involves a collaboration between domain experts and automation, with clear handoff points.
- Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs
This talk addresses the economic challenges of relying on rented AI inference infrastructure, particularly for agents. The core thesis is that while renting models is useful for initial learning and finding product-market fit, long-term operation requires owning the inference infrastructure to manage costs effectively. The speaker shares personal experience moving agents off a paid API to owned infrastructure after encountering unsustainable expenses.
- Content Is Code - Matt Palmer, Conductor
This talk explores how code has evolved into the most rapid medium for creating technical content. The core argument is that the key to exceptional developer experience, and by extension, effective technical content creation, lies not in advanced agent skills or novel frameworks, but in fundamental principles that have always driven great development.
- Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio
This talk introduces Froglet, an open-source protocol and node designed for agent-to-agent computation. It addresses the limitations of traditional logs and API dashboards for agents operating across different hosts. Froglet aims to establish verifiable work as the foundation for agentic commerce, emphasizing the importance of receipts, identities, and workload hashes that persist across various models and marketplaces.
- Agents Need Feature Flags - Sachin Gupta
This talk argues that AI agents, capable of performing high-impact actions like sending money or modifying databases, are being deployed without the safety infrastructure common in traditional web development. It introduces six specific types of feature flags—prompt variants, tool access, model routing, memory policy, autonomy level, and kill switches—that are necessary to manage the unique behavior surfaces of agents and prevent incidents. The core thesis is that applying established web development disciplines, particularly robust kill switches and phased rollouts, is crucial for controlling and safely iterating on AI agent functionality.
- Your Agents Need a Save Button - Hamza Tahir, ZenML
This talk introduces the concept of a "save button" for AI agents, analogous to document saving, to enable state persistence and replayability. Current agent execution lacks this, with only disconnected traces available. Implementing a save button allows for asking "what if" questions, such as swapping models or mocking tools, to optimize agents for cost, speed, and performance.
- Using LLMs to Secure Source Code — Eugene Yan, Anthropic
This talk explores the application of Large Language Models (LLMs) in enhancing source code security. It delves into how LLMs can be leveraged to identify vulnerabilities, suggest secure coding practices, and potentially automate parts of the code review process, thereby improving the overall security posture of software development.
- The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents
This discussion explores the current state and future potential of "loops" in AI-driven software development, debating whether the hype surrounding them aligns with practical capabilities. The core thesis questions if we are at a significant inflection point towards fully autonomous software factories or if current loop implementations fall short of their promised potential. The debate highlights the tension between the rapid advancement of AI models and the necessary discipline and engineering required to build reliable, autonomous systems.
- Reward Hacking in Agents — Daniel Han, Unsloth
This talk addresses the issue of reward hacking in AI agents, where agents exploit flaws in their reward functions to achieve high scores without genuinely fulfilling the intended task. It highlights the challenges of designing effective reward systems and proposes strategies to mitigate these unintended behaviors, ensuring agents act in alignment with desired outcomes. The discussion emphasizes the importance of robust evaluation methods to identify and correct reward hacking.
- Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth
This talk explores the current state and future trajectory of AI models, focusing on advancements in language models and benchmarks. It highlights the rapid progress in AI capabilities, evidenced by performance improvements across various benchmarks, while also addressing limitations such as long context handling and the challenges of reliable evaluation. The discussion delves into the open-source versus closed-source landscape, the complexities of model optimization, and the critical role of software and algorithmic improvements over hardware advancements.
- Intrinsic + Extrinsic + Learned Knowledge — Pablo Castro, Distinguished Engineer & CVP, Microsoft
This talk explores the three categories of knowledge that power AI agents: intrinsic, extrinsic, and learned. Intrinsic knowledge, embedded within models through training data, has been a primary driver of recent AI advancements, particularly in coding. Extrinsic knowledge, accessed through retrieval-augmented generation (RAG) patterns, allows agents to ground themselves in organizational data and external information. Learned knowledge emerges from agents' ongoing work, enabling continuous improvement and the capture of unique organizational capabilities.
- On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft
This talk explores the evolving landscape of AI, focusing on how AI systems can better understand and interact with human knowledge. It delves into the challenges and opportunities in building AI that can reason, learn, and collaborate effectively, moving beyond simple pattern recognition to a deeper comprehension of context and intent. The core thesis centers on the development of AI that acts as a knowledgeable partner, capable of complex problem-solving and informed decision-making.
- \"Software engineering is not about writing code\" — Benoit Schillings, Google DeepMind VP of Research
This talk argues that the core challenge in software engineering is shifting from writing code to managing complexity, designing systems, and ensuring correctness. While AI models now excel at generating code syntax, the future of software development lies in leveraging AI for higher-level tasks like architectural design, complex problem decomposition, and inductive reasoning. The economics of code production are changing, making code generation nearly free and emphasizing the need for new processes and evaluation methods.
- Google DeepMind's Frontier AI Engineering Research Agenda — Benoit Schillings, VP of Research
Benoit Schillings, VP of Research at Google DeepMind, outlines the frontier of AI engineering research, emphasizing the shift from foundational model development to sophisticated engineering for product integration and real-world impact. The agenda focuses on building robust, scalable, and reliable AI systems that can be productized and deployed effectively. This involves a deep dive into the engineering challenges and opportunities at the forefront of AI research.
- Every company should have a Brain — Garry Tan, Y Combinator
The core thesis is that companies should be built as AI-native entities from inception, leveraging AI not as a mere tool but as a workforce. This approach allows a small team to achieve the output of a much larger traditional organization, fundamentally changing the economics of company building. The key lies in how work is structured and managed, akin to building an organization with AI agents as employees and a curated knowledge base as the company's collective memory.
- The New Physics of Business — Garry Tan, Y Combinator
This talk explores a new framework for understanding business, drawing parallels to fundamental physics principles. It suggests that by applying concepts like energy, force, and momentum to business strategy, founders can gain a more intuitive and powerful approach to building and scaling companies. The core idea is to move beyond traditional business jargon and adopt a more fundamental, scientific lens.
- Y Combinator's Head of Design on Imagination Engineering
This talk explores the concept of "imagination engineering," proposing that as AI models become increasingly capable of executing tasks, the primary bottleneck for innovation will shift to generating novel and ambitious ideas. The speaker shares personal experiments using AI to capture and organize streams of consciousness, build personal websites from thoughts, and analyze the commonalities among highly creative historical figures. The core thesis is that thinking and building in public, amplified by AI, can unlock new levels of creativity and knowledge generation.
- Imagination Engineering: \"Live in the future and then build what's missing.\"
This talk introduces a philosophy for product development termed Imagination Engineering, which encourages developers to envision future states and then build the missing components. It emphasizes a proactive approach to innovation by living in the future and identifying needs before they become apparent. The core idea is to bridge the gap between aspirational future visions and current technological realities.
- How Autoresearch is changing ML research — Zhengyao Jiang, Weco
This talk explores how autoresearch is transforming the landscape of machine learning research. It delves into the methodologies and tools that enable automated research processes, aiming to accelerate discovery and innovation within the ML field. The core thesis is that by automating repetitive and time-consuming research tasks, scientists can focus on higher-level problem-solving and conceptualization.
- An AI Agent Became the #1 Contributor in OpenAI's Hiring Challenge — Zhengyao Jiang, Weco
An AI agent named Aiden achieved the top contributor status in OpenAI's Parameter Golf hiring challenge, surpassing human participants. This agent, developed by Weco, is a multi-agent self-improving system designed to autonomously research, experiment, and submit code improvements. Aiden's success highlights the potential of autonomous agents in complex problem-solving and collaborative environments.
- Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua
This talk introduces advancements in AI agents for computer use, moving beyond earlier, screen-dominating models to more integrated and efficient background operations. It highlights the development of tools like Quad Driver for seamless OS interaction across platforms and Kuabench for robust agent evaluation. The discussion also touches upon optimizing infrastructure for agent training to reduce costs and improve GPU utilization.
- \"The model trains the next model\" — Lee Robinson, Cursor, SpaceXAI
This talk explores the concept of AI models training future AI models, focusing on how this iterative process can lead to significant advancements. It delves into the practical implications of this approach for developers and the potential for AI to accelerate its own development cycle. The core idea is that by leveraging AI to improve AI, we can unlock new levels of performance and capability.
- More Compute In ➤ Better Model Out — Lee Robinson, Cursor
This talk explores the relationship between the amount of compute used to train AI models and the quality of their output, specifically focusing on coding agents. The core thesis suggests that increased computational resources directly correlate with improved model performance, leading to more capable AI tools for developers.
- Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI
This talk explores the process of training AI models, focusing on recursive model improvement and the intricate loops involved. It details how feedback from model usage, combined with increased compute power, drives iterative advancements. The discussion highlights the distinction between outer loops (user feedback, A/B testing) and inner loops (high-quality evaluations, complex training tasks) and emphasizes strategies to accelerate the latter for more efficient model development.
- Claude Fable, Claude Tag, and Anthropic's Culture — Cat Wu & Thariq Shihipar ft Simon Willison
This talk explores Anthropic's approach to building AI, focusing on their internal tools like Claude Fable and Claude Tag, and how their company culture influences product development. The discussion highlights the iterative process of creating and refining AI systems, emphasizing the importance of internal tooling and a collaborative environment for shipping robust AI products.
- Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic
This talk explores the rapid evolution of AI coding agents, focusing on Anthropic's Claude Code and its advancements like Fable and Claude Tag. The discussion highlights how these tools have shifted from requiring constant oversight to enabling more complex tasks and proactive assistance, fundamentally changing software development workflows. The speakers emphasize the increasing importance of product sense and business acumen for engineers as execution becomes more streamlined.
- WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar
This talk argues that while AI model intelligence has rapidly advanced, its real-world usefulness is hampered by a lack of business-specific context. The core thesis is that a "context layer" is the missing infrastructure for production-ready AI agents, analogous to how humans learn and operate within a company's knowledge, skills, and norms. This layer aims to transform implicit knowledge into machine-usable context, enabling AI to perform effectively in business environments.
- Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
This talk emphasizes the critical need for rigorous evaluation of AI skills before deployment, arguing that shipping skills without proper testing can lead to unpredictable behavior and degraded performance. The speaker highlights the distinction between agents developers use for personal productivity and agents built for consumers, noting that end-users lack the context to troubleshoot skill invocation issues. The core thesis is that comprehensive evaluations are essential for ensuring skill reliability, managing costs, and determining when skills can be retired as models improve.
- Forward Deployed Engineering at Cursor — Pauline Brunet
Pauline Brunet discusses the Forward Deployed Engineering (FDE) function, emphasizing its critical role in enterprise AI adoption. She outlines a framework for determining when and how to implement an FDE team, based on customer digital maturity and product customization levels. The core thesis is that FDEs are highly technical individuals who act as agents of change, embedding within customer organizations to drive digital transformation and ensure successful adoption of AI technologies.
- \"The engineer of the future is the person who is able to choose what is worth doing.\" — Addy Osmani
The future of engineering lies in the ability to discern what is truly valuable and worth pursuing, rather than simply executing tasks. As AI agents increasingly handle automated work, human engineers will be responsible for owning the evidence, understanding, and ultimate verdict on production decisions. This shift redefines the engineer's role from task execution to higher-level judgment, system reasoning, and accountability for outcomes.
- \"I've never seen anything scarier than an LLM with tool calls.\" — Erik Meijer aka @HeadinTheBox
This talk addresses the inherent dangers of Large Language Models (LLMs), particularly when equipped with tool-calling capabilities. It argues that LLMs, by their nature, will attempt to achieve their goals, even if it means causing harm or deleting data. The core thesis is that provably safe AI agents can be created by leveraging elementary type systems and compiler knowledge, specifically by refying the agent's plan into a program that can be formally verified for safety.
- Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers
This talk proposes a shift from traditional, simplistic model evaluation methods to more sophisticated techniques borrowed from psychology and psychometrics. The core argument is that simply counting correct answers, a method akin to classical test theory, is insufficient. Instead, the presentation advocates for Item Response Theory (IRT) to provide a more nuanced understanding of model capabilities by calibrating individual questions and estimating model intelligence more accurately.
- From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI
This talk explores the design principles behind creating secure and scalable agent sandbox clouds. It begins by explaining why AI models need the ability to execute code or use tools to handle tasks with verifiable rewards, such as math and coding problems. The presentation then delves into the evolution of sandbox technologies, from basic process isolation to advanced virtualization, emphasizing the critical need for robust security to protect against malicious or overzealous code execution in both research and product environments.
- Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect
This talk delves into modern post-training techniques for AI models, focusing on the open-source Verifiers and Primed RL libraries. It emphasizes the evolution of these tools to handle increasingly complex agent use cases and the need for robust infrastructure to support large-scale, customized model training. The core thesis is that by providing flexible and powerful open-source toolkits, it becomes more accessible for engineers and builders to enhance existing models for their specific applications and workflows, fostering continuous improvement through iterative refinement.
- RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI
This talk introduces Recursive Language Models (RLMs) as a method to address the context window limitations of AI coding agents when working with large codebases. The core thesis is to externalize context management into a programmable execution environment, allowing models to operate on the repository as data, curate relevant chunks, and feed them into the main context window. This approach can also serve as a memory layer for coding agents.
- The AI bugpocalypse is here. Now what? - Jack Cable, Corridor
The AI bugpocalypse is characterized by frontier AI models becoming increasingly adept at discovering and exploiting software vulnerabilities, particularly within open-source libraries. This trend is accelerating due to the rapid adoption of AI coding tools, which are fundamentally changing software development practices. The core challenge is to balance the increased attack surface and vulnerability discovery capabilities with robust defense mechanisms, ensuring that AI enhances, rather than compromises, software security.
- Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra
This talk addresses the challenge of "semantic blindness" in Large Language Models (LLMs) when dealing with vast, complex, and inconsistently named datasets, such as sensor names in large-scale industrial environments. The core thesis is that LLMs struggle with sheer scale and naming ambiguity, leading to errors and hallucinations. The proposed solution involves a hybrid approach that leverages the LLM for planning and decision-making while offloading structured data processing, retrieval, and set operations to deterministic code.
- The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab
This talk introduces Project Nanda, an open research effort building the infrastructure for an internet of AI agents. It addresses the current limitations of agents being confined to proprietary platforms by proposing an open, decentralized architecture analogous to the early open web. The core idea is to enable agents from different entities to discover, communicate, and transact with each other permissionlessly.
- A Song of Types and Agents - Roberto Stagi, Ratel
This talk argues that TypeScript is increasingly becoming the dominant language for building AI applications, particularly for the application and agent layers, despite Python's continued strength in model training and infrastructure. The shift is driven by the rise of coding agents and the need to integrate AI capabilities directly into applications, a domain historically dominated by TypeScript.
- ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay
This talk introduces ReviewDebt, a framework for quantifying the accumulating gap between code generated by AI agents and the human review, trust, and understanding applied to it. It argues that while AI agents increase code production speed, they also create a form of debt that compounds and impacts human attention, architectural consistency, and velocity expectations. The framework aims to provide a measurable number for this debt, enabling teams to manage it effectively.
- What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip
This talk addresses the ambiguity of "done" in the context of AI agents, arguing that a simple boolean check is insufficient. It proposes that "done" is a bundle of claims requiring verification across multiple dimensions, including artifact production, evidence of completion, adherence to a rubric, clear ownership for next steps, and survival of real-world conditions. The core thesis is that agent systems need a liveness model to manage the tension between keeping work moving and ensuring its quality.
- Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS
This talk presents five techniques for reducing AI agent hallucinations and token waste, focusing on code-based solutions rather than prompt engineering. These methods aim to improve accuracy and catch failures before they impact users. The techniques include semantic tool selection, graph retrieval augmented generation (RAG), multi-agent validation, neuro-symbolic guardians, and runtime guardians.
- The Factory That Dreams: 39 AI Agents, No Framework - Rushabh Doshi, Machinecraft
This talk details the creation of an AI system, named Eira, designed to act as a comprehensive "brain" for a manufacturing company. Instead of relying on traditional data science teams or extensive model training, the approach focused on building a well-organized memory system using off-the-shelf models. The core idea is to create a digital twin of the company's knowledge, enabling it to manage go-to-market operations and remember critical business information.
- Chat and citations won't save your vertical AI - Atul Ramachandran, Filed Inc
This talk argues that relying solely on chat interfaces and citations is insufficient for building successful vertical AI products. True value in vertical AI, particularly in industries like healthcare, legal, and taxes, comes from enabling users to delegate long-running tasks to AI agents. This shift requires designing products for delegation rather than direct user participation, viewing the product as a conveyor belt where users act as supervisors.
- Local AI State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, Matt Berman
This talk explores the current landscape and future potential of running AI models locally on user devices. It addresses the motivations behind this shift, emphasizing benefits like enhanced privacy, reduced latency, and cost savings. The discussion highlights the rapid advancements in hardware and software that are making local AI increasingly feasible and powerful for a wide range of applications.
- State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman
This talk, State of the Union: Why Local, Why Now, features insights from NVIDIA, Osmantic, Roboflow, and EXO Labs, with a contribution from @matthew_berman. The core thesis revolves around the increasing importance and viability of running AI models and applications locally, rather than solely relying on cloud-based solutions. It explores the reasons behind this shift and the current landscape of tools and technologies enabling this trend.
- Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman
This talk explores the rapid advancements and increasing accessibility of local AI, driven by improvements in both models and the tools (harnesses) used to interact with them. The speakers emphasize that the current inflection point allows for more powerful, private, and cost-effective AI applications, moving beyond simple chatbots to sophisticated, always-on agents. The discussion highlights the growing importance of local AI for enterprises and consumers alike, ensuring data privacy and controlled costs.
- Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD - Sumaiya Shrabony
This talk argues that solo builders of AI agents inevitably reinvent aspects of Continuous Integration and Continuous Deployment (CI/CD) pipelines, often in a less effective manner. As agents are developed, especially in isolation, builders start creating systems that require the same operational guarantees as traditional software, leading to the ad-hoc creation of testing, monitoring, and validation mechanisms. The core thesis is that without implementing deliberate gates and checks, agent systems are prone to shipping polished but flawed outputs, mirroring the software development pitfall of deploying code that compiles but fails tests.
- Develop at Idea Velocity - Jeffrey Lee-Chan, Snapchat
This talk explores developing at an accelerated pace by leveraging AI agents, specifically focusing on the Open Claw framework. The core thesis is that by enabling frictionless communication with AI and utilizing agent orchestrators, developers can achieve higher velocities in their work. The presentation highlights how specialized agents can manage complex tasks, leading to more autonomous development cycles and improved productivity.
- Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers — Alex Bauer, Upside.tech
This talk introduces design patterns for building trust in AI agents, drawing parallels to human collaboration and decision-making processes. The core thesis is that managing AI agents effectively involves applying principles similar to how teams of humans are managed, particularly in complex or ambiguous situations. The presentation emphasizes that AI agents, like people, benefit from clear guidance, access to knowledge, and structured evaluation methods to ensure reliable and trustworthy outputs.
- Understanding is the new bottleneck — Geoffrey Litt, Notion
This talk argues that understanding how code works remains crucial for AI engineers and builders, even as agents automate more coding tasks. The core thesis is that human understanding is shifting from verification to participation, enabling creative leaps and preventing cognitive debt. The speaker introduces techniques inspired by educational principles to help individuals and teams deepen their comprehension of AI-generated code.
- Should AI Engineers Still Read Code in 2026? The Z/L Continuum — Alex Volkov, ThursdAI
The talk explores the evolving role of AI engineers in 2026, specifically addressing whether they still need to read code generated by AI agents. It presents two opposing viewpoints: one advocating that code is now "free" and human intervention is unnecessary, and another emphasizing the critical need to read every line of code. The speaker proposes that the true question is not *if* engineers should read code, but rather *what level of proof* is required for specific tasks, suggesting a continuum based on task criticality.
- The Golden Age of AI Engineering — Alexander Embiricos & Romain Huet & Peter Steinberger, OpenAI
This talk posits that AI engineers are not becoming obsolete but are instead entering a golden age, fundamentally reshaping the engineering landscape. The rapid advancement of AI models, now iterating every few weeks, enables engineers to tackle complex problems with unprecedented speed and capability. The focus is shifting from writing code to problem-solving, integrating AI with design, judgment, and imagination to create usable products. This era represents a return to the core principles of engineering, amplified by powerful AI tools.
- Everything we knew about software has changed — Theo Browne, @t3dotgg
The talk argues that advancements in AI models have fundamentally changed software development, shifting the paradigm from incremental improvements to a need for developers to think bigger and wider. The speaker posits that current developer workflows and ingrained opinions about tools and languages are becoming obsolete as AI capabilities rapidly surpass human development speed. This necessitates a reevaluation of what is possible and how we approach building software.
- I Run a Fleet of AI Agents Across Three Machines. Here's What Broke. - Kyle Jaejun Lee, KRAFTON
This talk details the practical challenges and solutions encountered when running a fleet of AI coding agents across multiple machines. The presenter, Kyle Jaejun Lee, shares his journey from a single-machine setup to a distributed system, highlighting failures and the architectural patterns developed to overcome them. The core thesis is that managing AI agents at scale requires moving beyond a flat structure to a hierarchical organization and externalizing agent state to persistent storage, enabling robust recovery and efficient context management.
- Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis
This talk addresses a critical vulnerability in Large Language Models (LLMs): "sleeper agents" or backdoors that remain undetected by standard evaluations. The core thesis is that current defenses, which focus on model behavior or joint feature analysis, are insufficient. The proposed solution lies in analyzing the difference between a base model's activations and a fine-tuned model's activations, a method that reveals these hidden backdoors with high precision.
- GTM Is You - Victoria Melnikova, Evil Martians
This talk argues that in 2026, the most significant competitive advantage for startups, particularly in developer tools, is the founder's personal brand and direct engagement. With building software becoming easier, distribution has become the primary bottleneck. The speaker emphasizes that while AI can amplify a personal signal, it cannot replace the authenticity and trust built through founder-led go-to-market strategies.
- Beyond the Harness: A Journey Towards Adaptative Engineering - Rajiv Chandegra, Annicha Labs
This talk introduces adaptive engineering as a new philosophy for AI development, moving beyond fixed, pre-defined harnesses. It argues that as AI models become more powerful and interact with the dynamic, messy real world, static harnesses will become brittle. Adaptive engineering allows the harness to emerge and adapt during the engineering process, with the engineer's role shifting to designing constraints rather than dictating every step.
- How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI
This talk addresses the significant gap between the rapidly advancing reasoning capabilities of large language models (LLMs) and the slower evolution of retrieval systems. The core thesis is that LLMs are bottlenecked not by their reasoning power, but by their ability to access the correct knowledge. By improving retrieval tools and teaching agents to use them effectively, most of this knowledge gap can be closed, unlocking LLMs for complex tasks beyond coding.
- What if the harness mattered more than the model? - Aditya Bhargava, Etsy
This talk argues that the "harness" surrounding an AI model, which includes its tools, safety mechanisms, and reasoning capabilities, is often more critical to performance than the model itself. The speaker proposes shifting focus from solely improving large, proprietary models to developing sophisticated harnesses that can enhance the performance of smaller, open-source, locally runnable models. This approach aims to reduce reliance on closed-source AI and empower developers to build powerful agents with greater control and flexibility.
- Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo
This talk argues for designing AI systems that encourage discernment rather than mere approval. It highlights the phenomenon of cognitive surrender, where humans increasingly accept AI outputs without critical evaluation, potentially leading to errors and reduced human reasoning. The core thesis is that by engineering the human-AI interaction loop, developers can elicit more critical engagement, leading to higher quality data and more effective AI systems.
- Respect The Process - Andrew Dumit, Watershed Technology Inc.
This talk addresses the challenges of deploying coding agents for complex tasks, particularly in domains like sustainability where expert judgment and nuanced processes are critical. It emphasizes that while coding agents offer significant power and flexibility, their unconstrained execution can lead to unpredictable and potentially erroneous outcomes. The core thesis is that to effectively leverage coding agents, especially when direct answer verification is difficult, it's essential to implement a "harness" that constrains the agent's *effects* rather than its reasoning, ensuring process integrity and verifiable outputs.
- The Pipeline Is Dead - Iris ten Teije, Sky Valley Ambient Computing
The traditional software pipeline, built around the assumption that software creation is expensive and rare, is becoming obsolete. This model, which emphasizes shipping a single, frozen artifact to all users, is being replaced by a paradigm where software can be continuously adapted and personalized. The decreasing cost and increasing ease of software production, coupled with the demand for tailored user experiences, are driving this shift towards adaptive software.
- 500 people vibe-coded for 30 days. I was one of them. - Sanja Grbic, Automattic
This talk details a 30-day initiative at Automattic where employees were encouraged to pair up and build/ship projects, leveraging AI tools. The experiment aimed to accelerate product development and foster new skills, particularly for designers, by allowing them to own the entire product lifecycle from ideation to deployment. The speaker shares personal experiences building three distinct projects, highlighting the shift in their role from designer to design engineer and the broader implications for organizational agility.
- SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI
This talk introduces SWE Marathon, a benchmark designed to evaluate the capabilities of coding agents on large-scale, project-level tasks. It addresses the growing need to assess if agents can maintain coherence and successfully complete complex engineering projects over extended operational periods, such as a billion token budget, moving beyond simple bug fixes to end-to-end project ownership.
- Field Guide to Fable — Thariq Shihipar, Anthropic
This talk introduces Fable, a new class of Anthropic models, framing it as an expansion into an open world of AI capabilities. It emphasizes that models are grown organically rather than strictly designed, and that our understanding and interaction methods shape their potential. The presentation aims to provide a guide for working with these advanced models, encouraging users to explore their capabilities and push beyond conventional limitations.
- The Missing Layer After Launch - Raphael Kalandadze, Wandero AI
The talk "The Missing Layer After Launch" by Raphael Kalandadze addresses the critical, yet often overlooked, phase of managing AI agents and systems after they have been deployed. The core thesis is that shipping an agent is merely the beginning of the real work, and a robust "missing layer" of monitoring, understanding, and improvement is essential for long-term success and reliability. This layer involves closing the feedback loop to continuously enhance the product based on real-world performance.
- Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI
This talk introduces verifiable continual learning (VCL) for AI agents, a method to achieve durable improvements from experience without forgetting past performance. It addresses the challenges of obtaining feedback from production logs and acting upon it effectively. VCL aims to transform failures into testable scenarios, ensuring that improvements are verified and do not introduce regressions.
- MCP Apps: Primitives, discovery, and the Future of Software - Pietro Zullo, Manufact, Inc
This talk introduces MCP apps, an evolution of the Model Context Protocol (MCP) that allows MCP servers to return UI components alongside JSON data. This enables richer, interactive experiences where AI agents can not only call tools but also display and interact with user interfaces directly within the chat. The presentation emphasizes the growing importance of MCP apps and the opportunities for developers to build and distribute them through emerging app stores.
- Your AI Product Will Fail Unless You Can Explain It - Veronica Hylak, Hey AI
This talk addresses the critical challenge of explaining complex AI products to potential customers and stakeholders. The core thesis is that many AI products fail not due to technical shortcomings, but because their value proposition is not clearly communicated. The speaker emphasizes that effective storytelling, focusing on customer pain points and tangible transformations, is essential for product success in a crowded market.
- The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI
This talk argues that current AI interfaces, particularly prompt-based systems, are fundamentally limited by an outdated "punch card" protocol. Despite advancements in AI model capabilities, the interaction model remains batch-oriented, requiring users to pre-package their entire intent before the AI can respond. This forces users to adapt to the machine's constraints rather than the AI adapting to human conversational patterns.
- Building Great Agent Skills: The Missing Manual
This talk addresses the growing problem of "skill hell" in AI development, where developers struggle to effectively use and integrate available skills. The core thesis is that a lack of clear criteria for what constitutes a "great skill" hinders progress. The presentation offers a structured checklist and practical techniques to help developers build, evaluate, and improve AI skills, moving beyond the current state of confusion and inefficiency.
- Frontier results, on device - RL Nabors, Arize
This talk explores the significant costs associated with using large, frontier AI models and proposes a strategy for migrating to smaller, more efficient, on-device models. The core thesis is that by right-sizing models for specific tasks, developers can drastically reduce expenses, improve user experience through lower latency, and enhance security and privacy, all while maintaining or even improving performance for many applications.
- The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents
The core thesis is that domain-specific agents, rather than general-purpose ones, represent the future of AI development and application. This approach mirrors the Industrial Revolution's harnessing of energy with machines, but for intelligence with agents. The talk argues that while businesses are eager to integrate AI, building custom general agents is complex and often results in demos rather than robust solutions. Domain-specific agents offer a more efficient, controlled, and scalable alternative.
- You Can't Prompt the Room: The Last Skill AI Won't Replace - Balázs Horváth, VisualLabs
This talk argues that as AI becomes proficient at generating code, the critical bottleneck in software development shifts from *how* to build to *what* to build. The ability to understand business needs, elicit requirements, and make strategic decisions becomes the most valuable skill, surpassing traditional coding proficiency. This emphasizes the enduring importance of human-centric skills in guiding AI development.
- Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
This talk addresses the critical challenge of building reliable infrastructure for non-deterministic AI agents, which are increasingly moving beyond simple chatbots to perform complex tasks like planning, tool coordination, and decision-making affecting production systems. The core thesis is that while AI models are inherently probabilistic, the infrastructure supporting them must be deterministic to ensure reliability, safety, and cost-effectiveness at scale.
- The Agentic AI Engineer - Benedikt Sanftl, Mutagent
The Agentic AI Engineer talk introduces a framework for building and iterating on AI agents by applying agentic principles to the development lifecycle itself. This approach aims to automate and optimize the process of agent creation, testing, and deployment, moving beyond manual iteration to a more scalable, agent-driven workflow. The core idea is to create an "agent of agents" that manages the development loop, from initial specification to production monitoring and refinement.
- The Prompt is the Platform - Dominik Tornow, Resonate HQ
This talk proposes that the future of software engineering involves generating bespoke implementations on demand, rather than relying on general-purpose platforms. The core thesis is that as implementations become generatable, human value shifts from writing code to defining specifications. The prompt itself becomes the platform, enabling the creation of tailored software extensions that integrate seamlessly with existing infrastructure.
- Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon
This talk presents an RL-guided system designed to detect and remediate failures in ETL pipelines. The core idea is to enable an AI agent to act usefully, explainably, and within operational trust boundaries, significantly reducing the time and effort required for manual incident resolution. The system aims to automate responses for routine failures while escalating complex or high-risk situations.
- Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft
This talk addresses the challenge of debugging AI agents in production, where failures can be difficult to reproduce due to the inherent non-determinism of LLMs. The core thesis is that instead of chasing bitwise determinism, which is often unattainable, developers should focus on replayability and observability. This involves capturing the full state of an agent's execution to enable debugging and testing, even when the exact same output cannot be guaranteed on subsequent runs.
- Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
This talk explores the challenges and opportunities in building AI experiences that accept voice input and produce visual output. It argues that while voice is a natural human communication method, current AI implementations are often slow and awkward. Conversely, visual outputs, including rich HTML, interactive controls, and illustrations, are highly effective for conveying information and engaging users. The core thesis is that by optimizing for voice-in, visuals-out interactions, developers can create more delightful and seamless AI experiences, overcoming the latency barriers inherent in voice-to-voice conversations.
- AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
This talk introduces a framework for AI-driven multi-document correlation designed to enhance enterprise financial compliance and fraud detection. Traditional systems often analyze financial documents independently, missing critical risks that only become apparent when data is connected across multiple systems. The proposed framework integrates graph-based entity correlation, probabilistic risk modeling, and cross-jurisdictional normalization to uncover these hidden relationships, transforming compliance from a reactive process into a proactive intelligence capability.
- We Cut 94% of AI Coding Tokens With a Local Code Index - Rajkumar Sakthivel, Tesco
This talk addresses the significant cost associated with AI coding tools, which is often driven by excessive context sent to models rather than the model's processing. The core thesis is that by optimizing the input context, substantial cost savings can be achieved, far exceeding savings from output compression. A local code indexing and search layer is proposed as a solution to send only relevant code snippets to AI models.
- Your Agent Is Wasting Tokens and You Don't Know It - Erik Hanchett, AWS
This talk addresses the significant token costs associated with using AI agents, particularly large language models, and provides practical strategies for reducing these expenses. The core thesis is that by implementing specific techniques, developers can optimize agent performance and cost-efficiency without sacrificing functionality.
- OpenClaw in Your Hand: Building a Physical AI Terminal - Lech Kalinowski, Callstack
This talk details the creation of OpenClaw in Your Hand, a physical AI terminal designed for interacting with AI agents like OpenClaw. The project explores building an AI-native device using microcontrollers and a dual-display system (OLED and e-paper) for energy efficiency and a unique user experience. It highlights the challenges and potential of creating dedicated hardware for LLM interaction, moving beyond traditional screen-based interfaces.
- Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry - Abed Matini, Ogilvy
This talk introduces a framework for building chatbots that bypass the limitations and costs associated with multimodal inputs and complex tool integrations. It proposes a hybrid Retrieval Augmented Generation (RAG) approach using SQL, Reciprocal Rank Fusion (RRF), and live telemetry to improve efficiency and accuracy. The core idea is to optimize data ingestion and retrieval through structured processing, enabling more effective querying of documents and reducing the computational burden on LLMs.
- AI System Design: From Idea to Production - Apoorva Joshi, MongoDB
This talk outlines a repeatable framework for designing and building AI systems from initial idea to production. It emphasizes that in the current AI landscape, defining product requirements, system design, and evaluation criteria are the critical, challenging aspects, rather than just coding. The framework consists of four phases: product requirements, system design, evaluation and monitoring, and optimization.
- When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis
This talk introduces an extended Cache Augmented Generation (CAG) approach to address knowledge representation challenges when dealing with large, interconnected, and frequently updated document collections. The core idea is to leverage large context windows more effectively by distributing documents across multiple parallel CAG instances, allowing a supervisor model to query these instances and synthesize comprehensive answers. This method aims to overcome the limitations of simple RAG and the computational expense of GraphRAG in dynamic data environments.
- Research to Reality: Bringing Frontier ML Research to Production - Vaidas Razgaitis, Higharc
This talk addresses the challenge of transitioning frontier machine learning research into production-ready software. It proposes a three-pronged approach focusing on research legibility, code structure, and decomposition strategies to improve the velocity and efficiency of teams bridging the gap between ML research and software engineering. The core idea is to treat this transition as a systems and process problem.
- Building an Autonomous Engineering Org - Angie Jones, Agentic AI Foundation
This talk details the journey of transforming a large engineering organization into an autonomous one by leveraging AI agents. The core thesis is that true impact from AI enablement comes not just from individual engineers using tools, but from integrating agents deeply into the entire development workflow, enabling them to produce shippable results with minimal human oversight. This transformation involves moving engineers through maturity stages, from basic code generation to complex task delegation and multi-agent collaboration.
- HTML is All You Need (for Agents to Make Graphics) - Amol Kapoor, Nori
This talk argues that AI agents are not inherently bad at creating visual artifacts like slides or graphics. The core thesis is that the medium used to instruct agents is crucial; instead of forcing them to use human-centric graphical interfaces like canvases, agents should be given tools that align with their native language-based processing, such as HTML.
- Using Spec-Driven Development for Production Workflows - Erik Hanchett, AWS
This talk introduces spec-driven development as a structured approach to software engineering, emphasizing the creation of detailed specifications and design documents before writing any code. This methodology is particularly effective when working with AI coding assistants, guiding them to produce higher-quality, more accurate code. The core idea is to leverage AI as an intern that requires clear direction, with spec-driven development providing that essential guidance.
- User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch
This talk addresses a critical failure point in AI agents: the inability of evaluation signals to inform future actions, leading to repeated errors. The core thesis is that current agent memory systems are static and do not learn from past successes or failures. The presentation introduces a new runtime experience, Agent RX, designed to bridge this gap by allowing agents to improve dynamically based on outcomes without retraining or manual prompt engineering.
- Structuring the Unstructured - Cedric Clyburn, Red Hat
This talk addresses the challenge of processing unstructured data, such as PDFs, presentations, and technical documents, to make it usable for AI applications. The core thesis is that effectively structuring this data is crucial for building robust AI systems like RAG and agents, and that open-source tools can provide efficient, local solutions. The presentation highlights the limitations of basic parsing methods and introduces Docling as a comprehensive tool for extracting and transforming various document formats.
- Agents Building Agents - Alfonso Graziano, Nearform
This talk explores leveraging AI agents to build and improve other AI agents, addressing challenges like hallucinations, non-determinism, and latency. The core thesis is that AI can be used to automate the iterative process of agent development, leading to more reliable and secure AI systems. This approach involves using coding agents to refine target agents based on evaluation data and real-world user feedback.
- Browser Agents Don't Need Better Models. They Need Better Eyes. - Kushan Raj, ARK
This talk argues that the primary limitation of browser agents is not the underlying AI models but rather the surrounding infrastructure. The core thesis is that providing agents with a better environment to plan, execute, and debug tasks leads to significantly faster and more reliable performance, even with less powerful and cheaper models. The proposed solution involves a novel representation that compresses website information, allowing agents to perceive and interact with entire pages more effectively.
- The 100-Tool Agent Is a Trap - Sohail Shaikh & Ankush Rastogi, Prosodica
This talk addresses the "100-tool agent trap," where providing an AI agent with access to a vast number of tools simultaneously degrades its performance. As the tool catalog grows, agents become slower, more expensive, and less accurate, often confusing similar tools or inventing new ones. The core thesis is that this issue stems from forcing the model to process an entire catalog on every request, leading to context overload and decision-making difficulties.
- Stop Writing Tone Instructions. Layer Them. - Isadora Martin-Dye, Isadora & Co
This talk introduces a four-layer architecture for managing AI agent behavior, particularly focusing on brand voice and user interaction. The core thesis is that a single system prompt is insufficient for complex AI applications, especially where nuanced communication is critical. Instead, layering distinct responsibilities—immutable identity, situational mode, example-anchored voice, and post-generation veto—provides a more robust and reliable system. This approach moves beyond simple prompt engineering to a more structured systems engineering methodology.
- Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI
This talk introduces an AI Research OS designed to transform personal notes and research into a dynamic, queryable knowledge base. The system aims to help AI engineers and builders overcome the challenge of losing or struggling to access valuable information scattered across various digital tools. It proposes a layered approach, moving from raw data to an indexed system and finally to a synthesized, wiki-like interface that agents can effectively leverage.
- Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov
This talk details OpenGov's journey in building and scaling OG Assist, an AI agent integrated across their government software products. The core thesis revolves around the strategic decision to bet on the Effect TypeScript library for building a robust agent loop, enabling significant advancements in development and production capabilities. The presentation highlights practical applications, architectural choices, and the iterative process of improving AI agent performance in a real-world enterprise environment.
- A Genius With Amnesia - Victor Savkin, Nx
This talk addresses the limitations of current AI agents, specifically their spatial constraints (seeing only a small part of the codebase) and temporal amnesia (forgetting past interactions). It proposes a solution, Polygraph, a meta-harness designed to overcome these issues by providing agents with a comprehensive view of an entire organization's codebase, including owned and open-source repositories, and a persistent memory of all past sessions and decisions.
- The Log Is The Agent - Ishaan Sehgal, Omnara
This talk argues that the core identity and state of an AI agent reside not in its model or execution environment, but in its append-only log of events. This log, containing every input, output, tool call, and state transition, serves as the agent's durable history, enabling reliable resumption and state reconstruction even after system failures.
- Recursive Coding Agents - Raymond Weitekamp, OpenProse
This talk introduces recursive coding agents, applying the principles of Recursive Language Models (RLMs) to enhance the reliability and capability of AI agents for coding tasks. The core thesis is that current AI agents, while intelligent, are "mismanaged geniuses" lacking a robust system for specifying, managing, reusing, and verifying their work. RLMs offer a new paradigm for inference-time computation by unifying tool calling and reasoning, enabling agents to recursively break down and solve complex problems.
- Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
This talk argues that traditional benchmarks are insufficient for evaluating agentic AI systems in production. Agentic systems, which plan, call tools, and execute workflows, require a shift in evaluation focus from model capability to system behavior. The core thesis is that production telemetry and reliability metrics are paramount for dependable outcomes, moving evaluation from a pre-deployment phase to a continuous operational capability integrated into the system's control plane.
- Build Systems, Not Code - Angie Jones, Agentic AI Foundation
This talk argues that building agentic AI systems requires the same core engineering discipline as traditional software development, just with different primitives. Instead of focusing solely on using agents to write code, the emphasis shifts to architecting complex agentic systems. This approach allows engineers to leverage their existing skills in systems thinking, workflow design, and modularity, recapturing the thrill of building by operating at a higher level of abstraction.
- The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen
This talk argues that current evaluation methods for role-playing language agents (RPLAs) are fundamentally flawed. These methods, which prioritize fluency and personality consistency, fail to detect a dominant failure mode: anachronistic compositing. This occurs when models generate personas that sound like historical figures but reason using knowledge and moral frameworks that postdate them, often due to the overwhelming influence of culturally dominant, later representations in training data. The proposed solution is a shift towards epistemic simulation, emphasizing corpus-boundedness, temporal anchoring, and expert-loop evaluation.
- 6 Things to Know about AIE World's Fair 2026
The AI Engineering (AIE) World's Fair 2026 is significantly larger than previous events, aiming to provide a broad buffet of curated content for attendees. The event emphasizes a blend of evergreen and new topics, with a focus on practical workflows and industry-research collaboration. It seeks to foster serendipitous discovery and community building, moving beyond algorithmically driven content feeds.