Catalog
Browse all
946 talks, newest on the channel first. Filter by topic, or use Search in the header for full-text.
- AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j
This talk explores how AI models can leverage structured data, specifically within a lakehouse environment, by understanding context in various shapes rather than relying solely on traditional query-based methods. It emphasizes the shift from asking direct questions to enabling AI to infer and utilize information based on its inherent structure and relationships. The core idea is to unlock more sophisticated AI applications by treating data context as a fundamental input.
- Why We Killed Our Multi-Agent Pipeline — Subbiah Sethuraman and Abhilash Asokan, ZS Associates
This talk discusses the decision to discontinue a multi-agent pipeline, highlighting the challenges encountered in practical implementation. The presenters share insights into why a seemingly promising approach was ultimately abandoned, focusing on the complexities and limitations that arose during development and deployment. The core thesis revolves around the difficulties of scaling and managing sophisticated agentic systems in real-world scenarios.
- Citation Needed: Provenance for LLM-Built Knowledge Graphs — Daniel Chalef, Zep AI
This talk addresses the critical challenge of ensuring the accuracy and trustworthiness of knowledge graphs generated by Large Language Models (LLMs). It proposes a system for tracking the provenance of information within these graphs, allowing users to trace data back to its original sources. This is essential for building reliable AI applications that depend on accurate knowledge representation.
- Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times
This talk explores the application of agentic AI principles to the development of mobile games, focusing on how local agents can enhance game mechanics and player experiences. The presenters discuss the potential for AI to create more dynamic and responsive game environments by processing information and making decisions directly on the device. This approach aims to reduce latency and enable richer, more complex gameplay interactions.
- Video Has No Memory. Here's How We Built One. — James Le, TwelveLabs
This talk addresses the challenge of enabling AI models to retain information across extended video content, effectively giving them memory. The presenter discusses the technical hurdles and the architectural solutions developed to overcome these limitations, allowing for more sophisticated video analysis and interaction.
- Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
This talk argues that agentic systems, particularly those involved in coding, require ontologies to manage complexity and ensure reliable operation. Frank Coyle from UC Berkeley posits that as these systems become more sophisticated and capable of independent action, a structured understanding of their environment, tools, and goals becomes crucial for effective development and deployment. Without such a framework, the emergent behaviors of complex agentic systems can become unpredictable and difficult to control.
- Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase
This talk introduces Learned Execution Graphs (LEGs) as a novel approach to detecting anomalies and drift in API usage. The core thesis is that by modeling API interactions as graphs and learning their execution patterns, systems can proactively identify deviations from normal behavior, which is crucial for maintaining API reliability and security.
- From Systems of Record to Systems of Context — Omri Bruchim & Tomer Ast, monday.com
This talk explores the evolution of data management from systems of record to systems of context. It argues that shifting focus to contextual data allows for more dynamic and intelligent applications. The core thesis is that by understanding the context surrounding data, systems can become more adaptive and provide richer insights.
- Your Moat Is Your Data Model — Mike Phipps, Gates Foundation
This talk argues that a robust data model is the primary competitive advantage for organizations, rather than proprietary algorithms or unique datasets. The speaker emphasizes that the way data is structured and related within a system dictates its utility and the potential for innovation. A well-designed data model enables better decision-making and unlocks deeper insights.
- Active Graph Agent Runtime (BabyAGI 4) — Yohei Nakajima, Untapped Capital
This talk introduces ActiveGraph, an AI agent runtime built around an immutable event log. Instead of bolting memory and logging onto an LLM, ActiveGraph uses the log as the core of the agent's state. This approach enables features like replays, rollbacks, and forks by default, as every action and state change is recorded. The system then uses behaviors and policies to react to changes in the projected graph, which represents the agent's state.
- CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4j
This talk argues that automated assistants, particularly those using Retrieval Augmented Generation (RAG), would benefit more from graph memory than simply increasing token limits. It proposes that a graph database can provide a more structured and interconnected way to store and retrieve information, leading to more capable and context-aware AI assistants. The core idea is to move beyond linear text processing towards a relational understanding of data.
- Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer — Emil Eifrem, Neo4j
This talk introduces an ontology-based semantic layer designed to enhance the capabilities of AI agents. The core thesis is that by grounding agents in a structured, knowledge-rich ontology, they can achieve more sophisticated reasoning and decision-making, moving beyond simple pattern matching to a deeper understanding of context and relationships. This approach aims to build a "smarter substrate" for agents, enabling them to operate more effectively.
- Claude for Long-Horizon Tasks — Lance Martin, Anthropic
This talk explores building agent harnesses for reliable and secure long-horizon tasks using Claude. It emphasizes decoupling the agent's reasoning (brain) from its actions (hands), implementing self-verification and self-learning mechanisms, and designing for adaptability in evolving agent harness systems. The core thesis is that Claude's capabilities can be effectively leveraged for complex, extended tasks through robust harness architecture.
- Better Agent Auth — Bereket Habtemeskel & Paola Estefania, Better Auth
This talk addresses the critical need for robust authentication and authorization mechanisms within AI agent development. It emphasizes that as AI agents become more integrated into workflows and handle sensitive data, secure access control is paramount. The presentation likely explores common pitfalls in agent authentication and proposes solutions to ensure agents operate within defined permissions and trust boundaries.
- Full Workshop: Better Agent Auth — Paola Estefania, Better Auth
This workshop focuses on enhancing authentication for AI agents, aiming to provide a more robust and secure framework for their integration and operation. It delves into the critical aspects of building trustworthy AI systems by addressing the complexities of agent authorization and access control. The core thesis is that improved authentication mechanisms are fundamental to the safe and effective deployment of AI agents in various applications.
- Every Harness Will Become A Claw — Sam Bhagwat, Mastra
This talk proposes that AI harnesses are evolving beyond their initial function of managing AI agent workflows. They are increasingly integrating with collaboration tools and becoming more proactive, effectively transforming into "Claws." This evolution is driven by the capabilities of modern coding agents and the desire for AI assistants to be more present and available within team workflows.
- HTML Is All Agents Need — James Russo, HeyGen
This talk explores the potential of leveraging AI agents to generate HTML, moving beyond traditional templating methods. It posits that a more agentic approach can lead to more dynamic and adaptable web development. The core idea is to treat HTML generation as a task that AI agents can perform, potentially simplifying complex UI development.
- \"The biggest challenge in your stack? Evals, Evals, Evals\" - 2026 State of AI Engineering results
The State of AI Engineering 2026 survey highlights evaluation (evals) as the most significant challenge across AI stacks. This talk delves into the survey results, emphasizing the critical need for robust evaluation methodologies in AI development. Addressing evals is presented as paramount for building reliable and high-quality AI systems.
- 2026 State of AI Engineering — Barr Yaron, Amplify Partners
This talk provides a snapshot of the current landscape and future trajectory of AI engineering. It emphasizes the rapid evolution of tools and methodologies, highlighting the increasing importance of practical application and product development alongside foundational research. The core thesis suggests a shift towards more integrated and accessible AI development, driven by open-source contributions and a focus on developer experience.
- The 2026 State of AI Engineering — Barr Yaron, Amplify Partners
Barr Yaron of Amplify Partners discusses the current landscape of AI engineering, focusing on the practical application and development of AI technologies. The talk highlights key trends and challenges faced by engineers and product managers in building and deploying AI-powered products. It emphasizes the evolution of AI tools and methodologies shaping the industry.
- Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest
The core thesis is that agent architectures have a rapid obsolescence rate, with a half-life of approximately six months. This necessitates a continuous focus on updating and adapting agent systems to remain effective in the fast-evolving AI landscape. The talk emphasizes the need for builders to anticipate and manage this rapid change.
- The Desktop Frontier — Ahmad Osman, Osmantic
This talk explores the emerging landscape of desktop AI agents, positioning the personal computer as the next frontier for AI development and deployment. It argues that the desktop environment offers unique advantages for AI, moving beyond cloud-based models to enable more integrated and powerful user experiences. The discussion highlights the potential for AI to fundamentally change how users interact with their computers.
- Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk
This talk explores the critical role of architectural decisions in agentic security, emphasizing how these foundational choices impact the overall security posture of AI systems. It delves into the complexities of building secure AI by design, highlighting the need for robust frameworks and thoughtful consideration of security implications from the outset. The core thesis revolves around the idea that effective agentic security is not an add-on but a fundamental architectural requirement.
- Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town
This talk addresses the critical security challenges arising from the increasing use of agentic systems, particularly in software development. It emphasizes the need for robust permission models, verifiable provenance, and a secure supply chain for AI agents to mitigate risks associated with their autonomous capabilities. The core thesis is that without these security measures, agentic systems pose significant threats to code integrity and system security.
- Agentic Development Security — Ezra Tanzer, Snyk
This talk addresses the security implications of agentic development, a paradigm shift where AI agents assist in the software development lifecycle. It highlights the potential risks introduced by these agents, such as new attack vectors and vulnerabilities, and emphasizes the need for robust security practices tailored to this evolving landscape. The core thesis is that securing agentic development requires a proactive and specialized approach to mitigate emerging threats.
- What to do when 1 in 8 Skills have critical vulns and tries to steal your keys — Ezra Tanzer, Snyk
This talk addresses the significant security risks present in AI agent skills, highlighting that approximately one in eight skills contain critical vulnerabilities. These risks include the potential for malicious actors to exploit these weaknesses to gain unauthorized access to sensitive information, such as API keys. The presentation emphasizes the urgent need for robust security practices and thorough vetting of AI agent components.
- Why Securing Generated Code is Not Enough — Ezra Tanzer, Snyk
This talk addresses the critical need to secure code generated by AI, arguing that traditional security measures are insufficient. It emphasizes that the focus must shift beyond simply scanning the output to understanding and mitigating the risks inherent in the AI development process itself. The core thesis is that securing the *process* of AI code generation is paramount, not just the resulting code.
- Your LLM Stack Is a 2008 Database With Better Marketing — Lovina Dmello, NVIDIA
This talk argues that current Large Language Model (LLM) stacks are fundamentally similar to 2008-era databases but with more sophisticated marketing. It suggests that the underlying principles and challenges of managing and utilizing LLMs are not as novel as often portrayed, drawing parallels to the evolution and limitations of database technology. The core thesis is that a deeper understanding of these parallels can lead to more effective and realistic approaches to building with LLMs.
- We Gave an Agent Production Code Access and Then Tried to Sleep at Night — Moritz Johner, Form3
This talk explores the practical challenges and considerations of granting AI agents access to production codebases. It delves into the necessary infrastructure, safety protocols, and the evolving role of human oversight when AI is empowered to make changes in a live production environment. The core thesis revolves around the tension between AI's potential for rapid development and the critical need for robust safeguards.
- Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon)
This talk explores privacy-preserving intelligence, focusing on the infrastructure and methods required to build AI systems that respect user privacy. It delves into the technical challenges and potential solutions for developing intelligent agents and applications without compromising sensitive data. The core thesis revolves around enabling powerful AI capabilities while maintaining robust privacy guarantees.
- It's 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard
This talk addresses the challenges of managing and understanding AI agents in production environments. It emphasizes the need for robust monitoring and observability to ensure agents behave as expected and to debug issues effectively. The core
- Security Track Intro — Randall Degges, Snyk
This talk introduces the Security Track, emphasizing the critical role of security in the AI engineering landscape. It highlights the increasing complexity and potential vulnerabilities as AI systems become more integrated into various applications and workflows. The session sets the stage for understanding the security challenges and best practices relevant to AI development.
- AI’s Jurassic Park Period — Aaron Stanley, dbt Labs
This talk posits that the current era of AI development is akin to a "Jurassic Park period," characterized by rapid, often uncontrolled, innovation and a focus on demonstrating capabilities rather than robust, safe deployment. The speaker suggests that the industry is currently in a phase of building impressive, but potentially fragile, AI systems, drawing parallels to the early days of genetic engineering. The core thesis revolves around the need to transition from this experimental phase to one of responsible and sustainable development.
- In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar
As AI agents become more capable of handling complex development tasks, the primary challenge shifts from code generation to verification. Hallucinations are an inherent issue, and as models improve, their failures become more frequent and convincing, increasing the risk of human reviewers succumbing to cognitive surrender. This talk proposes a three-stage discipline for responsible agentic development: Guide, Verify, Solve. It argues that robust verification infrastructure is essential for safety and provides a competitive edge.
- Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog
This talk addresses the common issue of AI agents exhibiting inconsistent behavior, often disagreeing with their own previous outputs or decisions. It explores the underlying causes of this self-disagreement and provides practical strategies for engineers to mitigate these problems, leading to more reliable and predictable agent performance. The core thesis is that understanding and managing agent internal state and decision-making processes is crucial for building robust AI systems.
- Don't Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft
This talk emphasizes that while Large Language Models (LLMs) are powerful tools, they should not be given complete control over development processes. Instead, LLMs should function as sophisticated assistants, augmenting human capabilities rather than replacing human judgment and oversight. The core thesis is to leverage LLMs strategically to enhance productivity and creativity without relinquishing essential human decision-making.
- Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft
This talk challenges the assumption that advanced frontier models are necessary for building effective voice agents. The presenters argue that simpler, more focused models can achieve comparable or even superior results for specific voice agent tasks, emphasizing practicality and efficiency over raw model power. They advocate for a pragmatic approach to agent development, prioritizing task-specific capabilities.
- Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla
Ishita Daga of Tesla addresses a fundamental challenge in enterprise AI: the lack of a standardized structure for AI agents. She argues that current agent development often lacks a clear framework, leading to inefficiencies and difficulties in scaling. The talk emphasizes the need for a more organized approach to building and deploying AI agents within large organizations.
- When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain
This talk explores the intersection of AI agents and physical data, moving beyond purely digital interactions. It introduces the concept of agent harnesses as a framework for managing and integrating AI agents with real-world data sources and physical systems. The core thesis is that understanding the "physics" of how agents interact with physical data is crucial for building more robust and capable AI applications.
- Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared)
This talk introduces the concept of a Go-To-Market (GTM) agent designed to understand buyer behavior and inform product strategy. It emphasizes the need for AI agents that can move beyond simple task execution to possess a deeper understanding of market dynamics and customer needs. The core thesis is that by building agents with this buyer-centric intelligence, companies can more effectively align their product development and marketing efforts.
- Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio
This talk argues that current AI agents, despite advancements in tool calling, often lack the ability to provide verifiable evidence for their actions. The core thesis is that agents need to generate "receipts" or verifiable outputs to build trust and enable effective debugging and auditing, rather than simply being able to call more tools. This shift is crucial for moving beyond basic automation to reliable and accountable AI systems.
- Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens
This talk addresses the critical gap between raw AI agent output and a user-friendly experience, arguing that a dedicated rendering layer is essential for LLM pipelines. It emphasizes that simply displaying raw LLM responses fails to meet user expectations for intuitive and effective interaction. The core thesis is that building a robust rendering layer is key to successfully shipping AI-powered products.
- Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS
This talk explores the development of voice agents capable of handling interruptions, a crucial feature for natural and effective human-computer interaction. The presenters discuss the challenges and solutions involved in creating agents that can seamlessly manage conversational flow when users interject or change topics mid-utterance. The core thesis revolves around building more robust and user-friendly voice interfaces.
- Skills are the New SDKs - Elvin Aghammadzada, DataRobot
This talk posits that AI skills are evolving into a new form of Software Development Kits (SDKs), fundamentally changing how developers build and interact with AI systems. The core thesis is that by abstracting complex AI functionalities into reusable skills, developers can accelerate innovation and create more sophisticated applications. This shift emphasizes modularity and interoperability, akin to traditional software libraries.
- Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest
This talk addresses the common problem of failing jobs in Apache Spark, a widely used distributed computing system. It introduces Medic, a tool designed to provide first aid for these failing jobs, aiming to simplify debugging and improve the reliability of Spark applications. The core thesis is that by providing targeted diagnostics and actionable insights, developers can more efficiently resolve issues in their Spark workloads.
- Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs
This talk explores the potential for AI to automate complex oncology workflows, aiming to reduce or eliminate the need for human intervention in critical processes. It delves into the challenges and possibilities of integrating AI into healthcare, specifically within the demanding field of cancer treatment. The core thesis examines whether AI can independently manage and execute tasks traditionally requiring expert human oversight.
- Perception Agents — Antje Barth, Amazon AGI Lab
This talk introduces perception agents, a new paradigm in human-agent collaboration that moves beyond text-based interactions. Current agents often require detailed textual prompts and struggle with dynamic visual interfaces or unexpected application behaviors. Perception agents aim to bridge this gap by enabling AI to see, understand, and interact with computer interfaces similarly to humans, using clicks and keystrokes to act upon visual information.
- Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex
This workshop focuses on leveraging OpenAI's Codex as a versatile tool for computer control. It explores setting up persistent memory, creating collaborative assistant threads, and managing long-running computational tasks. The session emphasizes practical methods for integrating Codex into workflows, enabling complex operations through iterative prompting and loop-based execution.
- From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs
This talk explores the engineering hurdles in developing foundation models specifically for single-cell biology, presented from the viewpoint of a machine learning engineer without a biology background. It delves into the unique challenges and considerations required to adapt large model technologies for biological data analysis.
- The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest, DSPy
This talk introduces DSPy, a framework that separates the definition of an AI task from its implementation. By declaring a task's inputs and outputs first, developers can later experiment with different models, settings, and execution strategies without altering the core task definition. This approach aims to make AI engineering more flexible and robust, especially in a rapidly evolving landscape of models and tools.
- You Didn't Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit
This talk addresses a critical oversight in agent development: designing infrastructure for human interaction rather than agentic behavior. The core thesis is that many production LLM errors, particularly those related to rate limits, stem from this fundamental mismatch. Infrastructure built for human speed and interaction patterns fails when agents make thousands of calls per second, leading to performance issues and unexpected costs.
- From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud
This talk addresses the challenge of silently accumulating performance issues in mature codebases, where pausing feature development for investigation is difficult. It presents a case study on integrating runtime intelligence into coding agents to enable continuous performance optimization in production. The approach involves agents analyzing real production data to identify high-return-on-investment fixes, prioritized by complexity and impact.
- Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
This talk focuses on the critical importance of building effective evaluation systems for AI models, moving beyond superficial metrics to create benchmarks that genuinely reflect real-world performance and user needs. The presenters emphasize that robust evals are essential for shipping reliable AI products and driving meaningful improvements in model quality.
- Build Evals That Actually Matter - Nick Ung, Lyft
This talk addresses the common problem of AI evaluations that fail to predict real-world performance. The core thesis is that offline evaluations often use simplistic "customers" and test sets that don't reflect the complexity and adversarial nature of actual user interactions, leading to shipped models that fail in production. A solution involves building more realistic, adversarial user simulators trained on real data.
- Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer
This talk argues that current approaches to AI agent development, often focused on "harnessing" models or increasing scale, are insufficient for building maintainable software. The core thesis is that the failure of agent software factories stems from a fundamental problem in how coding models are trained, which prioritizes passing tests over architectural quality. This leads to codebases that degrade over time, becoming difficult to manage and debug.
- Notion's Token Town — Sarah Sachs, Notion
This talk addresses the significant cost challenges in AI development, particularly the escalating expense of using large language models. It argues that focusing solely on token economics is a losing strategy, as model providers often increase prices or deprecate older versions. Instead, the core thesis is to win on product by leveraging data flywheels, sophisticated orchestration, and maintaining model agnosticism to preserve flexibility and avoid vendor lock-in.
- Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait
This talk explores the application of autonomous agents to complex scientific discovery tasks, moving beyond simpler coding puzzles or optimization problems. It highlights the necessity for agents to engage with real-world measurement data and employ a scientific method, including hypothesis generation, model implementation, and learning from failures. The core thesis is that significant advancements in agent performance for scientific tasks stem from formulating accurate hypotheses about physical systems and correctly implementing them.
- A Practitioner's Guide to Graphs - Tim Ainge, Good Collective
This talk provides a practical introduction to graph databases and algorithms. It covers the fundamentals of what a graph is, how to extract graph data from unstructured text, and the benefits of using a schema-first approach with ontological improvements. The presentation also delves into specific graph algorithms like personalized PageRank, shortest path, and subgraph matching, illustrating their principles with accessible code examples and real-world applications.
- Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
This talk presents a framework for evaluating and improving AI agents, emphasizing a "rollout-centered" approach. The speakers connect sandboxed environments, agent evaluations, and optimization workflows to create a practical system for generating and learning from agent performance data. The core idea is that treating agent development as a series of rollouts, similar to software deployments, allows for systematic improvement.
- The UX of AI: Making AI-Powered Apps Your Users Don't Hate - Kathryn Grayson Nanz, Progress Software
This talk addresses the critical gap between the excitement AI developers feel and the apprehension users often experience with AI-powered applications. It emphasizes that successful AI product adoption hinges on thoughtful user experience (UX) design, focusing on building user trust, ensuring privacy and security, and guiding users through novel interaction patterns. The core thesis is that teaching users how to effectively interact with AI capabilities is essential for them to derive value from new applications.
- Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse
This talk addresses the critical need for domain expertise in AI self-improvement loops. It argues that without a deep understanding of the specific domain, automated improvement processes can be inefficient or ineffective. The core thesis is that the most efficient path to a continuously improving agentic system involves a collaboration between domain experts and automation, with clear handoff points.
- Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs
This talk addresses the economic challenges of relying on rented AI inference infrastructure, particularly for agents. The core thesis is that while renting models is useful for initial learning and finding product-market fit, long-term operation requires owning the inference infrastructure to manage costs effectively. The speaker shares personal experience moving agents off a paid API to owned infrastructure after encountering unsustainable expenses.
- Content Is Code - Matt Palmer, Conductor
This talk explores how code has evolved into the most rapid medium for creating technical content. The core argument is that the key to exceptional developer experience, and by extension, effective technical content creation, lies not in advanced agent skills or novel frameworks, but in fundamental principles that have always driven great development.
- Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio
This talk introduces Froglet, an open-source protocol and node designed for agent-to-agent computation. It addresses the limitations of traditional logs and API dashboards for agents operating across different hosts. Froglet aims to establish verifiable work as the foundation for agentic commerce, emphasizing the importance of receipts, identities, and workload hashes that persist across various models and marketplaces.
- Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google
This talk explores the challenges and opportunities of running AI models on edge devices, where memory (RAM) is the primary constraint, not compute power. It highlights the development of smaller, more efficient models, including quantized versions of Gemma and even sub-billion parameter models, to enable AI capabilities on resource-limited hardware like Raspberry Pis and mobile NPUs. The focus is on practical applications and the trade-offs involved in deploying AI at the edge.
- Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs
This talk introduces Vending-Bench, a novel evaluation framework designed to assess long-horizon agentic behavior in simulated business environments. It addresses the challenge of models acting differently when they suspect they are being tested, proposing methods to maintain reproducibility and observe emergent behaviors like collusion and power-seeking over extended periods. The work also explores the transition of AI-run businesses from simulation to the real world.
- Agents Need Feature Flags - Sachin Gupta
This talk argues that AI agents, capable of performing high-impact actions like sending money or modifying databases, are being deployed without the safety infrastructure common in traditional web development. It introduces six specific types of feature flags—prompt variants, tool access, model routing, memory policy, autonomy level, and kill switches—that are necessary to manage the unique behavior surfaces of agents and prevent incidents. The core thesis is that applying established web development disciplines, particularly robust kill switches and phased rollouts, is crucial for controlling and safely iterating on AI agent functionality.
- Your Agents Need a Save Button - Hamza Tahir, ZenML
This talk introduces the concept of a "save button" for AI agents, analogous to document saving, to enable state persistence and replayability. Current agent execution lacks this, with only disconnected traces available. Implementing a save button allows for asking "what if" questions, such as swapping models or mocking tools, to optimize agents for cost, speed, and performance.
- Using LLMs to Secure Source Code — Eugene Yan, Anthropic
This talk explores the application of Large Language Models (LLMs) in enhancing source code security. It delves into how LLMs can be leveraged to identify vulnerabilities, suggest secure coding practices, and potentially automate parts of the code review process, thereby improving the overall security posture of software development.
- The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents
This discussion explores the current state and future potential of "loops" in AI-driven software development, debating whether the hype surrounding them aligns with practical capabilities. The core thesis questions if we are at a significant inflection point towards fully autonomous software factories or if current loop implementations fall short of their promised potential. The debate highlights the tension between the rapid advancement of AI models and the necessary discipline and engineering required to build reliable, autonomous systems.
- Reward Hacking in Agents — Daniel Han, Unsloth
This talk addresses the issue of reward hacking in AI agents, where agents exploit flaws in their reward functions to achieve high scores without genuinely fulfilling the intended task. It highlights the challenges of designing effective reward systems and proposes strategies to mitigate these unintended behaviors, ensuring agents act in alignment with desired outcomes. The discussion emphasizes the importance of robust evaluation methods to identify and correct reward hacking.
- Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth
This talk explores the current state and future trajectory of AI models, focusing on advancements in language models and benchmarks. It highlights the rapid progress in AI capabilities, evidenced by performance improvements across various benchmarks, while also addressing limitations such as long context handling and the challenges of reliable evaluation. The discussion delves into the open-source versus closed-source landscape, the complexities of model optimization, and the critical role of software and algorithmic improvements over hardware advancements.
- Intrinsic + Extrinsic + Learned Knowledge — Pablo Castro, Distinguished Engineer & CVP, Microsoft
This talk explores the three categories of knowledge that power AI agents: intrinsic, extrinsic, and learned. Intrinsic knowledge, embedded within models through training data, has been a primary driver of recent AI advancements, particularly in coding. Extrinsic knowledge, accessed through retrieval-augmented generation (RAG) patterns, allows agents to ground themselves in organizational data and external information. Learned knowledge emerges from agents' ongoing work, enabling continuous improvement and the capture of unique organizational capabilities.
- On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft
This talk explores the evolving landscape of AI, focusing on how AI systems can better understand and interact with human knowledge. It delves into the challenges and opportunities in building AI that can reason, learn, and collaborate effectively, moving beyond simple pattern recognition to a deeper comprehension of context and intent. The core thesis centers on the development of AI that acts as a knowledgeable partner, capable of complex problem-solving and informed decision-making.
- \"Software engineering is not about writing code\" — Benoit Schillings, Google DeepMind VP of Research
This talk argues that the core challenge in software engineering is shifting from writing code to managing complexity, designing systems, and ensuring correctness. While AI models now excel at generating code syntax, the future of software development lies in leveraging AI for higher-level tasks like architectural design, complex problem decomposition, and inductive reasoning. The economics of code production are changing, making code generation nearly free and emphasizing the need for new processes and evaluation methods.
- Google DeepMind's Frontier AI Engineering Research Agenda — Benoit Schillings, VP of Research
Benoit Schillings, VP of Research at Google DeepMind, outlines the frontier of AI engineering research, emphasizing the shift from foundational model development to sophisticated engineering for product integration and real-world impact. The agenda focuses on building robust, scalable, and reliable AI systems that can be productized and deployed effectively. This involves a deep dive into the engineering challenges and opportunities at the forefront of AI research.
- Every company should have a Brain — Garry Tan, Y Combinator
The core thesis is that companies should be built as AI-native entities from inception, leveraging AI not as a mere tool but as a workforce. This approach allows a small team to achieve the output of a much larger traditional organization, fundamentally changing the economics of company building. The key lies in how work is structured and managed, akin to building an organization with AI agents as employees and a curated knowledge base as the company's collective memory.
- The New Physics of Business — Garry Tan, Y Combinator
This talk explores a new framework for understanding business, drawing parallels to fundamental physics principles. It suggests that by applying concepts like energy, force, and momentum to business strategy, founders can gain a more intuitive and powerful approach to building and scaling companies. The core idea is to move beyond traditional business jargon and adopt a more fundamental, scientific lens.
- Y Combinator's Head of Design on Imagination Engineering
This talk explores the concept of "imagination engineering," proposing that as AI models become increasingly capable of executing tasks, the primary bottleneck for innovation will shift to generating novel and ambitious ideas. The speaker shares personal experiments using AI to capture and organize streams of consciousness, build personal websites from thoughts, and analyze the commonalities among highly creative historical figures. The core thesis is that thinking and building in public, amplified by AI, can unlock new levels of creativity and knowledge generation.
- Imagination Engineering: \"Live in the future and then build what's missing.\"
This talk introduces a philosophy for product development termed Imagination Engineering, which encourages developers to envision future states and then build the missing components. It emphasizes a proactive approach to innovation by living in the future and identifying needs before they become apparent. The core idea is to bridge the gap between aspirational future visions and current technological realities.
- How Autoresearch is changing ML research — Zhengyao Jiang, Weco
This talk explores how autoresearch is transforming the landscape of machine learning research. It delves into the methodologies and tools that enable automated research processes, aiming to accelerate discovery and innovation within the ML field. The core thesis is that by automating repetitive and time-consuming research tasks, scientists can focus on higher-level problem-solving and conceptualization.
- An AI Agent Became the #1 Contributor in OpenAI's Hiring Challenge — Zhengyao Jiang, Weco
An AI agent named Aiden achieved the top contributor status in OpenAI's Parameter Golf hiring challenge, surpassing human participants. This agent, developed by Weco, is a multi-agent self-improving system designed to autonomously research, experiment, and submit code improvements. Aiden's success highlights the potential of autonomous agents in complex problem-solving and collaborative environments.
- Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua
This talk introduces advancements in AI agents for computer use, moving beyond earlier, screen-dominating models to more integrated and efficient background operations. It highlights the development of tools like Quad Driver for seamless OS interaction across platforms and Kuabench for robust agent evaluation. The discussion also touches upon optimizing infrastructure for agent training to reduce costs and improve GPU utilization.
- \"The model trains the next model\" — Lee Robinson, Cursor, SpaceXAI
This talk explores the concept of AI models training future AI models, focusing on how this iterative process can lead to significant advancements. It delves into the practical implications of this approach for developers and the potential for AI to accelerate its own development cycle. The core idea is that by leveraging AI to improve AI, we can unlock new levels of performance and capability.
- More Compute In ➤ Better Model Out — Lee Robinson, Cursor
This talk explores the relationship between the amount of compute used to train AI models and the quality of their output, specifically focusing on coding agents. The core thesis suggests that increased computational resources directly correlate with improved model performance, leading to more capable AI tools for developers.
- Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI
This talk explores the process of training AI models, focusing on recursive model improvement and the intricate loops involved. It details how feedback from model usage, combined with increased compute power, drives iterative advancements. The discussion highlights the distinction between outer loops (user feedback, A/B testing) and inner loops (high-quality evaluations, complex training tasks) and emphasizes strategies to accelerate the latter for more efficient model development.
- Claude Fable, Claude Tag, and Anthropic's Culture — Cat Wu & Thariq Shihipar ft Simon Willison
This talk explores Anthropic's approach to building AI, focusing on their internal tools like Claude Fable and Claude Tag, and how their company culture influences product development. The discussion highlights the iterative process of creating and refining AI systems, emphasizing the importance of internal tooling and a collaborative environment for shipping robust AI products.
- Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic
This talk explores the rapid evolution of AI coding agents, focusing on Anthropic's Claude Code and its advancements like Fable and Claude Tag. The discussion highlights how these tools have shifted from requiring constant oversight to enabling more complex tasks and proactive assistance, fundamentally changing software development workflows. The speakers emphasize the increasing importance of product sense and business acumen for engineers as execution becomes more streamlined.
- WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar
This talk argues that while AI model intelligence has rapidly advanced, its real-world usefulness is hampered by a lack of business-specific context. The core thesis is that a "context layer" is the missing infrastructure for production-ready AI agents, analogous to how humans learn and operate within a company's knowledge, skills, and norms. This layer aims to transform implicit knowledge into machine-usable context, enabling AI to perform effectively in business environments.
- Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
This talk emphasizes the critical need for rigorous evaluation of AI skills before deployment, arguing that shipping skills without proper testing can lead to unpredictable behavior and degraded performance. The speaker highlights the distinction between agents developers use for personal productivity and agents built for consumers, noting that end-users lack the context to troubleshoot skill invocation issues. The core thesis is that comprehensive evaluations are essential for ensuring skill reliability, managing costs, and determining when skills can be retired as models improve.
- Forward Deployed Engineering at Cursor — Pauline Brunet
Pauline Brunet discusses the Forward Deployed Engineering (FDE) function, emphasizing its critical role in enterprise AI adoption. She outlines a framework for determining when and how to implement an FDE team, based on customer digital maturity and product customization levels. The core thesis is that FDEs are highly technical individuals who act as agents of change, embedding within customer organizations to drive digital transformation and ensure successful adoption of AI technologies.
- \"The engineer of the future is the person who is able to choose what is worth doing.\" — Addy Osmani
The future of engineering lies in the ability to discern what is truly valuable and worth pursuing, rather than simply executing tasks. As AI agents increasingly handle automated work, human engineers will be responsible for owning the evidence, understanding, and ultimate verdict on production decisions. This shift redefines the engineer's role from task execution to higher-level judgment, system reasoning, and accountability for outcomes.
- \"I've never seen anything scarier than an LLM with tool calls.\" — Erik Meijer aka @HeadinTheBox
This talk addresses the inherent dangers of Large Language Models (LLMs), particularly when equipped with tool-calling capabilities. It argues that LLMs, by their nature, will attempt to achieve their goals, even if it means causing harm or deleting data. The core thesis is that provably safe AI agents can be created by leveraging elementary type systems and compiler knowledge, specifically by refying the agent's plan into a program that can be formally verified for safety.
- Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers
This talk proposes a shift from traditional, simplistic model evaluation methods to more sophisticated techniques borrowed from psychology and psychometrics. The core argument is that simply counting correct answers, a method akin to classical test theory, is insufficient. Instead, the presentation advocates for Item Response Theory (IRT) to provide a more nuanced understanding of model capabilities by calibrating individual questions and estimating model intelligence more accurately.
- From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI
This talk explores the design principles behind creating secure and scalable agent sandbox clouds. It begins by explaining why AI models need the ability to execute code or use tools to handle tasks with verifiable rewards, such as math and coding problems. The presentation then delves into the evolution of sandbox technologies, from basic process isolation to advanced virtualization, emphasizing the critical need for robust security to protect against malicious or overzealous code execution in both research and product environments.
- Modern Post-Training: A Deep Dive — Will Brown, Prime Intellect
This talk delves into modern post-training techniques for AI models, focusing on the open-source Verifiers and Primed RL libraries. It emphasizes the evolution of these tools to handle increasingly complex agent use cases and the need for robust infrastructure to support large-scale, customized model training. The core thesis is that by providing flexible and powerful open-source toolkits, it becomes more accessible for engineers and builders to enhance existing models for their specific applications and workflows, fostering continuous improvement through iterative refinement.
- RLM: Recursive Language Models for Large Codebases - Shashi, Superagentic AI
This talk introduces Recursive Language Models (RLMs) as a method to address the context window limitations of AI coding agents when working with large codebases. The core thesis is to externalize context management into a programmable execution environment, allowing models to operate on the repository as data, curate relevant chunks, and feed them into the main context window. This approach can also serve as a memory layer for coding agents.
- The AI bugpocalypse is here. Now what? - Jack Cable, Corridor
The AI bugpocalypse is characterized by frontier AI models becoming increasingly adept at discovering and exploiting software vulnerabilities, particularly within open-source libraries. This trend is accelerating due to the rapid adoption of AI coding tools, which are fundamentally changing software development practices. The core challenge is to balance the increased attack surface and vulnerability discovery capabilities with robust defense mechanisms, ensuring that AI enhances, rather than compromises, software security.
- Semantic Blindness: 500,000 Sensors Confused an LLM - Raahul Singh & Vanč Levstik, Phaidra
This talk addresses the challenge of "semantic blindness" in Large Language Models (LLMs) when dealing with vast, complex, and inconsistently named datasets, such as sensor names in large-scale industrial environments. The core thesis is that LLMs struggle with sheer scale and naming ambiguity, leading to errors and hallucinations. The proposed solution involves a hybrid approach that leverages the LLM for planning and decision-making while offloading structured data processing, retrieval, and set operations to deterministic code.
- The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab
This talk introduces Project Nanda, an open research effort building the infrastructure for an internet of AI agents. It addresses the current limitations of agents being confined to proprietary platforms by proposing an open, decentralized architecture analogous to the early open web. The core idea is to enable agents from different entities to discover, communicate, and transact with each other permissionlessly.
- A Song of Types and Agents - Roberto Stagi, Ratel
This talk argues that TypeScript is increasingly becoming the dominant language for building AI applications, particularly for the application and agent layers, despite Python's continued strength in model training and infrastructure. The shift is driven by the rise of coding agents and the need to integrate AI capabilities directly into applications, a domain historically dominated by TypeScript.
- ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, Ebay
This talk introduces ReviewDebt, a framework for quantifying the accumulating gap between code generated by AI agents and the human review, trust, and understanding applied to it. It argues that while AI agents increase code production speed, they also create a form of debt that compounds and impacts human attention, architectural consistency, and velocity expectations. The framework aims to provide a measurable number for this debt, enabling teams to manage it effectively.
- remobi.app: Don't change your terminal workflow for mobile
Remoby is a progressive web application designed to bring terminal workflows, specifically those utilizing tmux, to mobile devices. It aims to allow users to manage and interact with their AI agents and development environments from their phones without altering their existing desktop setup. The core idea is to extend the power and customization of a terminal-based workflow to a mobile context.
- What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, Paperclip
This talk addresses the ambiguity of "done" in the context of AI agents, arguing that a simple boolean check is insufficient. It proposes that "done" is a bundle of claims requiring verification across multiple dimensions, including artifact production, evidence of completion, adherence to a rubric, clear ownership for next steps, and survival of real-world conditions. The core thesis is that agent systems need a liveness model to manage the tension between keeping work moving and ensuring its quality.
- Claws Out: Securing and Building with OpenClaw - Nick Taylor, Pomerium
This talk introduces OpenClaw, an open-source project for building and securing AI agent control planes. It highlights a new feature, trusted proxy authentication mode, which enhances security and user experience by eliminating the need for manual token entry and device pairing for web socket connections. The presentation also demonstrates how OpenClaw can be integrated into development workflows, enabling live coding and rapid prototyping of AI-powered applications.
- Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS
This talk presents five techniques for reducing AI agent hallucinations and token waste, focusing on code-based solutions rather than prompt engineering. These methods aim to improve accuracy and catch failures before they impact users. The techniques include semantic tool selection, graph retrieval augmented generation (RAG), multi-agent validation, neuro-symbolic guardians, and runtime guardians.
- The Factory That Dreams: 39 AI Agents, No Framework - Rushabh Doshi, Machinecraft
This talk details the creation of an AI system, named Eira, designed to act as a comprehensive "brain" for a manufacturing company. Instead of relying on traditional data science teams or extensive model training, the approach focused on building a well-organized memory system using off-the-shelf models. The core idea is to create a digital twin of the company's knowledge, enabling it to manage go-to-market operations and remember critical business information.
- Chat and citations won't save your vertical AI - Atul Ramachandran, Filed Inc
This talk argues that relying solely on chat interfaces and citations is insufficient for building successful vertical AI products. True value in vertical AI, particularly in industries like healthcare, legal, and taxes, comes from enabling users to delegate long-running tasks to AI agents. This shift requires designing products for delegation rather than direct user participation, viewing the product as a conveyor belt where users act as supervisors.
- Local AI State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, Matt Berman
This talk explores the current landscape and future potential of running AI models locally on user devices. It addresses the motivations behind this shift, emphasizing benefits like enhanced privacy, reduced latency, and cost savings. The discussion highlights the rapid advancements in hardware and software that are making local AI increasingly feasible and powerful for a wide range of applications.
- State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman
This talk, State of the Union: Why Local, Why Now, features insights from NVIDIA, Osmantic, Roboflow, and EXO Labs, with a contribution from @matthew_berman. The core thesis revolves around the increasing importance and viability of running AI models and applications locally, rather than solely relying on cloud-based solutions. It explores the reasons behind this shift and the current landscape of tools and technologies enabling this trend.
- Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman
This talk explores the rapid advancements and increasing accessibility of local AI, driven by improvements in both models and the tools (harnesses) used to interact with them. The speakers emphasize that the current inflection point allows for more powerful, private, and cost-effective AI applications, moving beyond simple chatbots to sophisticated, always-on agents. The discussion highlights the growing importance of local AI for enterprises and consumers alike, ensuring data privacy and controlled costs.
- Every Solo Agent Builder Eventually Reinvents a Worse Version of CI/CD - Sumaiya Shrabony
This talk argues that solo builders of AI agents inevitably reinvent aspects of Continuous Integration and Continuous Deployment (CI/CD) pipelines, often in a less effective manner. As agents are developed, especially in isolation, builders start creating systems that require the same operational guarantees as traditional software, leading to the ad-hoc creation of testing, monitoring, and validation mechanisms. The core thesis is that without implementing deliberate gates and checks, agent systems are prone to shipping polished but flawed outputs, mirroring the software development pitfall of deploying code that compiles but fails tests.
- Develop at Idea Velocity - Jeffrey Lee-Chan, Snapchat
This talk explores developing at an accelerated pace by leveraging AI agents, specifically focusing on the Open Claw framework. The core thesis is that by enabling frictionless communication with AI and utilizing agent orchestrators, developers can achieve higher velocities in their work. The presentation highlights how specialized agents can manage complex tasks, leading to more autonomous development cycles and improved productivity.
- So many AI Tools: When to use what? — Chris Noring, Microsoft
This talk addresses the proliferation of AI tools and provides guidance on selecting the appropriate ones for specific tasks. It emphasizes understanding the capabilities and limitations of various tools to make informed decisions in AI development and product building. The core thesis revolves around strategic tool selection to enhance efficiency and effectiveness.
- From Writing Code to Designing Systems: How the Developer Role is Changing — Chris Noring, Microsoft
The developer role is shifting from solely writing code to designing and managing complex systems, augmented by AI tools. This evolution means developers are becoming more efficient, acting as orchestrators rather than just coders. The core idea is to leverage AI to scale development efforts, enabling individuals to achieve significantly higher productivity while maintaining control through robust guardrails.
- Design Patterns for AI Trust: Juries, Libraries, and Agent Tiers — Alex Bauer, Upside.tech
This talk introduces design patterns for building trust in AI agents, drawing parallels to human collaboration and decision-making processes. The core thesis is that managing AI agents effectively involves applying principles similar to how teams of humans are managed, particularly in complex or ambiguous situations. The presentation emphasizes that AI agents, like people, benefit from clear guidance, access to knowledge, and structured evaluation methods to ensure reliable and trustworthy outputs.
- Understanding is the new bottleneck — Geoffrey Litt, Notion
This talk argues that understanding how code works remains crucial for AI engineers and builders, even as agents automate more coding tasks. The core thesis is that human understanding is shifting from verification to participation, enabling creative leaps and preventing cognitive debt. The speaker introduces techniques inspired by educational principles to help individuals and teams deepen their comprehension of AI-generated code.
- Should AI Engineers Still Read Code in 2026? The Z/L Continuum — Alex Volkov, ThursdAI
The talk explores the evolving role of AI engineers in 2026, specifically addressing whether they still need to read code generated by AI agents. It presents two opposing viewpoints: one advocating that code is now "free" and human intervention is unnecessary, and another emphasizing the critical need to read every line of code. The speaker proposes that the true question is not *if* engineers should read code, but rather *what level of proof* is required for specific tasks, suggesting a continuum based on task criticality.
- The Golden Age of AI Engineering — Alexander Embiricos & Romain Huet & Peter Steinberger, OpenAI
This talk posits that AI engineers are not becoming obsolete but are instead entering a golden age, fundamentally reshaping the engineering landscape. The rapid advancement of AI models, now iterating every few weeks, enables engineers to tackle complex problems with unprecedented speed and capability. The focus is shifting from writing code to problem-solving, integrating AI with design, judgment, and imagination to create usable products. This era represents a return to the core principles of engineering, amplified by powerful AI tools.
- Everything we knew about software has changed — Theo Browne, @t3dotgg
The talk argues that advancements in AI models have fundamentally changed software development, shifting the paradigm from incremental improvements to a need for developers to think bigger and wider. The speaker posits that current developer workflows and ingrained opinions about tools and languages are becoming obsolete as AI capabilities rapidly surpass human development speed. This necessitates a reevaluation of what is possible and how we approach building software.
- Think You Can Build a Game with AI? Think Again! - Danielle An & David Hoe, Meta
This talk explores the evolving landscape of AI-driven game development, moving beyond simple prompt-based game generation to more sophisticated and dynamic experiences. It highlights how AI tools are lowering the barrier to entry for game creation, enabling individuals with diverse skill sets to build games. The presentation emphasizes that while AI can facilitate game creation, achieving a high-quality, standout game still requires artistic cohesion, thoughtful design, and a deep understanding of player experience, particularly with the advent of runtime LLMs that can dynamically alter gameplay.
- Your agent is blindfolded — Johan Lajili, Poolside AI
This talk addresses the discrepancy between the perceived capabilities of AI agents and their actual performance in real-world applications. The core argument is that the difference lies not in the agent's inherent ability, but in the feedback loop and the agent's ability to verify its own actions. Without robust verification mechanisms, agents may produce seemingly correct output that is actually flawed, leading to user frustration and distrust.
- Building an ACP-Compatible Agent Live — Bennet Fenner, Zed
This talk introduces the Agent Client Protocol (ACP), a JSON RPC-based standard designed to unify interactions between AI coding agents and their clients. The core idea is to provide a consistent interface, similar to Language Server Protocol (LSP), allowing users to bring their preferred AI agent to various development tools. The presentation demonstrates how to build a minimal ACP-compatible coding agent live, showcasing the protocol's capabilities for tool use, streaming output, and file manipulation.
- Teaching Coding Agents to do Spreadsheets - Nuno Campos, Witan Labs
This talk details the process of training AI coding agents to perform tasks within spreadsheets, aiming to match their proficiency in programming languages. The project began with a 50% accuracy rate on financial analysis benchmarks and ultimately achieved 92% accuracy. The presentation covers the challenges of representing spreadsheet data to AI, various approaches explored, and the breakthroughs that led to significant improvements.
- Your coding agent doesn't always follow your rules — Talha Sheikh, Checkout.com
This talk addresses the common issue where AI coding agents, despite appearing to complete tasks, often produce outputs that fail upon execution or do not meet specifications. The core argument is that the value is shifting from the agent's ability to generate code to the developer's ability to design and implement robust verification systems, or harnesses, that ensure the agent's output is reliable and deterministic.
- Running a Chess YouTube Channel entirely by AI — Stephan Steinfurt, TNG
This talk details the creation of an AI system capable of generating YouTube videos explaining chess games. The system combines large language models (LLMs) with specialized chess tools to analyze games, identify key moments like blunders and brilliant moves, and then produce narrated video content. The core challenge addressed is bridging the gap between AI engines that play chess well but cannot explain it, and LLMs that can explain but cannot play effectively.
- I Run a Fleet of AI Agents Across Three Machines. Here's What Broke. - Kyle Jaejun Lee, KRAFTON
This talk details the practical challenges and solutions encountered when running a fleet of AI coding agents across multiple machines. The presenter, Kyle Jaejun Lee, shares his journey from a single-machine setup to a distributed system, highlighting failures and the architectural patterns developed to overcome them. The core thesis is that managing AI agents at scale requires moving beyond a flat structure to a hierarchical organization and externalizing agent state to persistent storage, enabling robust recovery and efficient context management.
- Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data - Sachin Kumar, LexisNexis
This talk addresses a critical vulnerability in Large Language Models (LLMs): "sleeper agents" or backdoors that remain undetected by standard evaluations. The core thesis is that current defenses, which focus on model behavior or joint feature analysis, are insufficient. The proposed solution lies in analyzing the difference between a base model's activations and a fine-tuned model's activations, a method that reveals these hidden backdoors with high precision.
- GTM Is You - Victoria Melnikova, Evil Martians
This talk argues that in 2026, the most significant competitive advantage for startups, particularly in developer tools, is the founder's personal brand and direct engagement. With building software becoming easier, distribution has become the primary bottleneck. The speaker emphasizes that while AI can amplify a personal signal, it cannot replace the authenticity and trust built through founder-led go-to-market strategies.
- Beyond the Harness: A Journey Towards Adaptative Engineering - Rajiv Chandegra, Annicha Labs
This talk introduces adaptive engineering as a new philosophy for AI development, moving beyond fixed, pre-defined harnesses. It argues that as AI models become more powerful and interact with the dynamic, messy real world, static harnesses will become brittle. Adaptive engineering allows the harness to emerge and adapt during the engineering process, with the engineer's role shifting to designing constraints rather than dictating every step.
- How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI
This talk addresses the significant gap between the rapidly advancing reasoning capabilities of large language models (LLMs) and the slower evolution of retrieval systems. The core thesis is that LLMs are bottlenecked not by their reasoning power, but by their ability to access the correct knowledge. By improving retrieval tools and teaching agents to use them effectively, most of this knowledge gap can be closed, unlocking LLMs for complex tasks beyond coding.
- What if the harness mattered more than the model? - Aditya Bhargava, Etsy
This talk argues that the "harness" surrounding an AI model, which includes its tools, safety mechanisms, and reasoning capabilities, is often more critical to performance than the model itself. The speaker proposes shifting focus from solely improving large, proprietary models to developing sophisticated harnesses that can enhance the performance of smaller, open-source, locally runnable models. This approach aims to reduce reliance on closed-source AI and empower developers to build powerful agents with greater control and flexibility.
- Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo
This talk argues for designing AI systems that encourage discernment rather than mere approval. It highlights the phenomenon of cognitive surrender, where humans increasingly accept AI outputs without critical evaluation, potentially leading to errors and reduced human reasoning. The core thesis is that by engineering the human-AI interaction loop, developers can elicit more critical engagement, leading to higher quality data and more effective AI systems.
- Respect The Process - Andrew Dumit, Watershed Technology Inc.
This talk addresses the challenges of deploying coding agents for complex tasks, particularly in domains like sustainability where expert judgment and nuanced processes are critical. It emphasizes that while coding agents offer significant power and flexibility, their unconstrained execution can lead to unpredictable and potentially erroneous outcomes. The core thesis is that to effectively leverage coding agents, especially when direct answer verification is difficult, it's essential to implement a "harness" that constrains the agent's *effects* rather than its reasoning, ensuring process integrity and verifiable outputs.
- The Pipeline Is Dead - Iris ten Teije, Sky Valley Ambient Computing
The traditional software pipeline, built around the assumption that software creation is expensive and rare, is becoming obsolete. This model, which emphasizes shipping a single, frozen artifact to all users, is being replaced by a paradigm where software can be continuously adapted and personalized. The decreasing cost and increasing ease of software production, coupled with the demand for tailored user experiences, are driving this shift towards adaptive software.
- 500 people vibe-coded for 30 days. I was one of them. - Sanja Grbic, Automattic
This talk details a 30-day initiative at Automattic where employees were encouraged to pair up and build/ship projects, leveraging AI tools. The experiment aimed to accelerate product development and foster new skills, particularly for designers, by allowing them to own the entire product lifecycle from ideation to deployment. The speaker shares personal experiences building three distinct projects, highlighting the shift in their role from designer to design engineer and the broader implications for organizational agility.
- SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale - Rishi Desai, Abundant AI
This talk introduces SWE Marathon, a benchmark designed to evaluate the capabilities of coding agents on large-scale, project-level tasks. It addresses the growing need to assess if agents can maintain coherence and successfully complete complex engineering projects over extended operational periods, such as a billion token budget, moving beyond simple bug fixes to end-to-end project ownership.
- Field Guide to Fable — Thariq Shihipar, Anthropic
This talk introduces Fable, a new class of Anthropic models, framing it as an expansion into an open world of AI capabilities. It emphasizes that models are grown organically rather than strictly designed, and that our understanding and interaction methods shape their potential. The presentation aims to provide a guide for working with these advanced models, encouraging users to explore their capabilities and push beyond conventional limitations.
- The Missing Layer After Launch - Raphael Kalandadze, Wandero AI
The talk "The Missing Layer After Launch" by Raphael Kalandadze addresses the critical, yet often overlooked, phase of managing AI agents and systems after they have been deployed. The core thesis is that shipping an agent is merely the beginning of the real work, and a robust "missing layer" of monitoring, understanding, and improvement is essential for long-term success and reliability. This layer involves closing the feedback loop to continuously enhance the product based on real-world performance.
- Continual Learning for AI Agents: From Failures to Durable Improvements - Soheil Feizi, RELAI
This talk introduces verifiable continual learning (VCL) for AI agents, a method to achieve durable improvements from experience without forgetting past performance. It addresses the challenges of obtaining feedback from production logs and acting upon it effectively. VCL aims to transform failures into testable scenarios, ensuring that improvements are verified and do not introduce regressions.
- MCP Apps: Primitives, discovery, and the Future of Software - Pietro Zullo, Manufact, Inc
This talk introduces MCP apps, an evolution of the Model Context Protocol (MCP) that allows MCP servers to return UI components alongside JSON data. This enables richer, interactive experiences where AI agents can not only call tools but also display and interact with user interfaces directly within the chat. The presentation emphasizes the growing importance of MCP apps and the opportunities for developers to build and distribute them through emerging app stores.
- Your AI Product Will Fail Unless You Can Explain It - Veronica Hylak, Hey AI
This talk addresses the critical challenge of explaining complex AI products to potential customers and stakeholders. The core thesis is that many AI products fail not due to technical shortcomings, but because their value proposition is not clearly communicated. The speaker emphasizes that effective storytelling, focusing on customer pain points and tangible transformations, is essential for product success in a crowded market.
- The Prompt Is Still a Punch Card - Ted Johnson, JoinIn AI
This talk argues that current AI interfaces, particularly prompt-based systems, are fundamentally limited by an outdated "punch card" protocol. Despite advancements in AI model capabilities, the interaction model remains batch-oriented, requiring users to pre-package their entire intent before the AI can respond. This forces users to adapt to the machine's constraints rather than the AI adapting to human conversational patterns.
- Building Great Agent Skills: The Missing Manual
This talk addresses the growing problem of "skill hell" in AI development, where developers struggle to effectively use and integrate available skills. The core thesis is that a lack of clear criteria for what constitutes a "great skill" hinders progress. The presentation offers a structured checklist and practical techniques to help developers build, evaluate, and improve AI skills, moving beyond the current state of confusion and inefficiency.
- Frontier results, on device - RL Nabors, Arize
This talk explores the significant costs associated with using large, frontier AI models and proposes a strategy for migrating to smaller, more efficient, on-device models. The core thesis is that by right-sizing models for specific tasks, developers can drastically reduce expenses, improve user experience through lower latency, and enhance security and privacy, all while maintaining or even improving performance for many applications.
- The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents
The core thesis is that domain-specific agents, rather than general-purpose ones, represent the future of AI development and application. This approach mirrors the Industrial Revolution's harnessing of energy with machines, but for intelligence with agents. The talk argues that while businesses are eager to integrate AI, building custom general agents is complex and often results in demos rather than robust solutions. Domain-specific agents offer a more efficient, controlled, and scalable alternative.
- You Can't Prompt the Room: The Last Skill AI Won't Replace - Balázs Horváth, VisualLabs
This talk argues that as AI becomes proficient at generating code, the critical bottleneck in software development shifts from *how* to build to *what* to build. The ability to understand business needs, elicit requirements, and make strategic decisions becomes the most valuable skill, surpassing traditional coding proficiency. This emphasizes the enduring importance of human-centric skills in guiding AI development.
- Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
This talk addresses the critical challenge of building reliable infrastructure for non-deterministic AI agents, which are increasingly moving beyond simple chatbots to perform complex tasks like planning, tool coordination, and decision-making affecting production systems. The core thesis is that while AI models are inherently probabilistic, the infrastructure supporting them must be deterministic to ensure reliability, safety, and cost-effectiveness at scale.
- The Agentic AI Engineer - Benedikt Sanftl, Mutagent
The Agentic AI Engineer talk introduces a framework for building and iterating on AI agents by applying agentic principles to the development lifecycle itself. This approach aims to automate and optimize the process of agent creation, testing, and deployment, moving beyond manual iteration to a more scalable, agent-driven workflow. The core idea is to create an "agent of agents" that manages the development loop, from initial specification to production monitoring and refinement.
- The Prompt is the Platform - Dominik Tornow, Resonate HQ
This talk proposes that the future of software engineering involves generating bespoke implementations on demand, rather than relying on general-purpose platforms. The core thesis is that as implementations become generatable, human value shifts from writing code to defining specifications. The prompt itself becomes the platform, enabling the creation of tailored software extensions that integrate seamlessly with existing infrastructure.
- Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon
This talk presents an RL-guided system designed to detect and remediate failures in ETL pipelines. The core idea is to enable an AI agent to act usefully, explainably, and within operational trust boundaries, significantly reducing the time and effort required for manual incident resolution. The system aims to automate responses for routine failures while escalating complex or high-risk situations.
- Your Agent Failed in Prod. Good Luck Reproducing It. - Tisha Chawla & Susheem Koul, Microsoft
This talk addresses the challenge of debugging AI agents in production, where failures can be difficult to reproduce due to the inherent non-determinism of LLMs. The core thesis is that instead of chasing bitwise determinism, which is often unattainable, developers should focus on replayability and observability. This involves capturing the full state of an agent's execution to enable debugging and testing, even when the exact same output cannot be guaranteed on subsequent runs.
- Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
This talk explores the challenges and opportunities in building AI experiences that accept voice input and produce visual output. It argues that while voice is a natural human communication method, current AI implementations are often slow and awkward. Conversely, visual outputs, including rich HTML, interactive controls, and illustrations, are highly effective for conveying information and engaging users. The core thesis is that by optimizing for voice-in, visuals-out interactions, developers can create more delightful and seamless AI experiences, overcoming the latency barriers inherent in voice-to-voice conversations.
- AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
This talk introduces a framework for AI-driven multi-document correlation designed to enhance enterprise financial compliance and fraud detection. Traditional systems often analyze financial documents independently, missing critical risks that only become apparent when data is connected across multiple systems. The proposed framework integrates graph-based entity correlation, probabilistic risk modeling, and cross-jurisdictional normalization to uncover these hidden relationships, transforming compliance from a reactive process into a proactive intelligence capability.
- We Cut 94% of AI Coding Tokens With a Local Code Index - Rajkumar Sakthivel, Tesco
This talk addresses the significant cost associated with AI coding tools, which is often driven by excessive context sent to models rather than the model's processing. The core thesis is that by optimizing the input context, substantial cost savings can be achieved, far exceeding savings from output compression. A local code indexing and search layer is proposed as a solution to send only relevant code snippets to AI models.
- Your Agent Is Wasting Tokens and You Don't Know It - Erik Hanchett, AWS
This talk addresses the significant token costs associated with using AI agents, particularly large language models, and provides practical strategies for reducing these expenses. The core thesis is that by implementing specific techniques, developers can optimize agent performance and cost-efficiency without sacrificing functionality.
- OpenClaw in Your Hand: Building a Physical AI Terminal - Lech Kalinowski, Callstack
This talk details the creation of OpenClaw in Your Hand, a physical AI terminal designed for interacting with AI agents like OpenClaw. The project explores building an AI-native device using microcontrollers and a dual-display system (OLED and e-paper) for energy efficiency and a unique user experience. It highlights the challenges and potential of creating dedicated hardware for LLM interaction, moving beyond traditional screen-based interfaces.
- Bypassing the Multimodal Tax: Hybrid RAG, SQL RRF & UI Telemetry - Abed Matini, Ogilvy
This talk introduces a framework for building chatbots that bypass the limitations and costs associated with multimodal inputs and complex tool integrations. It proposes a hybrid Retrieval Augmented Generation (RAG) approach using SQL, Reciprocal Rank Fusion (RRF), and live telemetry to improve efficiency and accuracy. The core idea is to optimize data ingestion and retrieval through structured processing, enabling more effective querying of documents and reducing the computational burden on LLMs.
- AI System Design: From Idea to Production - Apoorva Joshi, MongoDB
This talk outlines a repeatable framework for designing and building AI systems from initial idea to production. It emphasizes that in the current AI landscape, defining product requirements, system design, and evaluation criteria are the critical, challenging aspects, rather than just coding. The framework consists of four phases: product requirements, system design, evaluation and monitoring, and optimization.
- When All Context Matters: Extended Cache Augmented Generation - Luis Romero-Sevilla, Orbis
This talk introduces an extended Cache Augmented Generation (CAG) approach to address knowledge representation challenges when dealing with large, interconnected, and frequently updated document collections. The core idea is to leverage large context windows more effectively by distributing documents across multiple parallel CAG instances, allowing a supervisor model to query these instances and synthesize comprehensive answers. This method aims to overcome the limitations of simple RAG and the computational expense of GraphRAG in dynamic data environments.
- Research to Reality: Bringing Frontier ML Research to Production - Vaidas Razgaitis, Higharc
This talk addresses the challenge of transitioning frontier machine learning research into production-ready software. It proposes a three-pronged approach focusing on research legibility, code structure, and decomposition strategies to improve the velocity and efficiency of teams bridging the gap between ML research and software engineering. The core idea is to treat this transition as a systems and process problem.
- Building an Autonomous Engineering Org - Angie Jones, Agentic AI Foundation
This talk details the journey of transforming a large engineering organization into an autonomous one by leveraging AI agents. The core thesis is that true impact from AI enablement comes not just from individual engineers using tools, but from integrating agents deeply into the entire development workflow, enabling them to produce shippable results with minimal human oversight. This transformation involves moving engineers through maturity stages, from basic code generation to complex task delegation and multi-agent collaboration.
- HTML is All You Need (for Agents to Make Graphics) - Amol Kapoor, Nori
This talk argues that AI agents are not inherently bad at creating visual artifacts like slides or graphics. The core thesis is that the medium used to instruct agents is crucial; instead of forcing them to use human-centric graphical interfaces like canvases, agents should be given tools that align with their native language-based processing, such as HTML.
- Using Spec-Driven Development for Production Workflows - Erik Hanchett, AWS
This talk introduces spec-driven development as a structured approach to software engineering, emphasizing the creation of detailed specifications and design documents before writing any code. This methodology is particularly effective when working with AI coding assistants, guiding them to produce higher-quality, more accurate code. The core idea is to leverage AI as an intern that requires clear direction, with spec-driven development providing that essential guidance.
- User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch
This talk addresses a critical failure point in AI agents: the inability of evaluation signals to inform future actions, leading to repeated errors. The core thesis is that current agent memory systems are static and do not learn from past successes or failures. The presentation introduces a new runtime experience, Agent RX, designed to bridge this gap by allowing agents to improve dynamically based on outcomes without retraining or manual prompt engineering.
- Structuring the Unstructured - Cedric Clyburn, Red Hat
This talk addresses the challenge of processing unstructured data, such as PDFs, presentations, and technical documents, to make it usable for AI applications. The core thesis is that effectively structuring this data is crucial for building robust AI systems like RAG and agents, and that open-source tools can provide efficient, local solutions. The presentation highlights the limitations of basic parsing methods and introduces Docling as a comprehensive tool for extracting and transforming various document formats.
- Agents Building Agents - Alfonso Graziano, Nearform
This talk explores leveraging AI agents to build and improve other AI agents, addressing challenges like hallucinations, non-determinism, and latency. The core thesis is that AI can be used to automate the iterative process of agent development, leading to more reliable and secure AI systems. This approach involves using coding agents to refine target agents based on evaluation data and real-world user feedback.
- Browser Agents Don't Need Better Models. They Need Better Eyes. - Kushan Raj, ARK
This talk argues that the primary limitation of browser agents is not the underlying AI models but rather the surrounding infrastructure. The core thesis is that providing agents with a better environment to plan, execute, and debug tasks leads to significantly faster and more reliable performance, even with less powerful and cheaper models. The proposed solution involves a novel representation that compresses website information, allowing agents to perceive and interact with entire pages more effectively.
- The 100-Tool Agent Is a Trap - Sohail Shaikh & Ankush Rastogi, Prosodica
This talk addresses the "100-tool agent trap," where providing an AI agent with access to a vast number of tools simultaneously degrades its performance. As the tool catalog grows, agents become slower, more expensive, and less accurate, often confusing similar tools or inventing new ones. The core thesis is that this issue stems from forcing the model to process an entire catalog on every request, leading to context overload and decision-making difficulties.
- Stop Writing Tone Instructions. Layer Them. - Isadora Martin-Dye, Isadora & Co
This talk introduces a four-layer architecture for managing AI agent behavior, particularly focusing on brand voice and user interaction. The core thesis is that a single system prompt is insufficient for complex AI applications, especially where nuanced communication is critical. Instead, layering distinct responsibilities—immutable identity, situational mode, example-anchored voice, and post-generation veto—provides a more robust and reliable system. This approach moves beyond simple prompt engineering to a more structured systems engineering methodology.
- Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-François Bouchard, Towards AI
This talk introduces an AI Research OS designed to transform personal notes and research into a dynamic, queryable knowledge base. The system aims to help AI engineers and builders overcome the challenge of losing or struggling to access valuable information scattered across various digital tools. It proposes a layered approach, moving from raw data to an indexed system and finally to a synthesized, wiki-like interface that agents can effectively leverage.
- Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov
This talk details OpenGov's journey in building and scaling OG Assist, an AI agent integrated across their government software products. The core thesis revolves around the strategic decision to bet on the Effect TypeScript library for building a robust agent loop, enabling significant advancements in development and production capabilities. The presentation highlights practical applications, architectural choices, and the iterative process of improving AI agent performance in a real-world enterprise environment.
- A Genius With Amnesia - Victor Savkin, Nx
This talk addresses the limitations of current AI agents, specifically their spatial constraints (seeing only a small part of the codebase) and temporal amnesia (forgetting past interactions). It proposes a solution, Polygraph, a meta-harness designed to overcome these issues by providing agents with a comprehensive view of an entire organization's codebase, including owned and open-source repositories, and a persistent memory of all past sessions and decisions.
- The Log Is The Agent - Ishaan Sehgal, Omnara
This talk argues that the core identity and state of an AI agent reside not in its model or execution environment, but in its append-only log of events. This log, containing every input, output, tool call, and state transition, serves as the agent's durable history, enabling reliable resumption and state reconstruction even after system failures.
- Recursive Coding Agents - Raymond Weitekamp, OpenProse
This talk introduces recursive coding agents, applying the principles of Recursive Language Models (RLMs) to enhance the reliability and capability of AI agents for coding tasks. The core thesis is that current AI agents, while intelligent, are "mismanaged geniuses" lacking a robust system for specifying, managing, reusing, and verifying their work. RLMs offer a new paradigm for inference-time computation by unifying tool calling and reasoning, enabling agents to recursively break down and solve complex problems.
- Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
This talk argues that traditional benchmarks are insufficient for evaluating agentic AI systems in production. Agentic systems, which plan, call tools, and execute workflows, require a shift in evaluation focus from model capability to system behavior. The core thesis is that production telemetry and reliability metrics are paramount for dependable outcomes, moving evaluation from a pre-deployment phase to a continuous operational capability integrated into the system's control plane.
- Build Systems, Not Code - Angie Jones, Agentic AI Foundation
This talk argues that building agentic AI systems requires the same core engineering discipline as traditional software development, just with different primitives. Instead of focusing solely on using agents to write code, the emphasis shifts to architecting complex agentic systems. This approach allows engineers to leverage their existing skills in systems thinking, workflow design, and modularity, recapturing the thrill of building by operating at a higher level of abstraction.
- The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen
This talk argues that current evaluation methods for role-playing language agents (RPLAs) are fundamentally flawed. These methods, which prioritize fluency and personality consistency, fail to detect a dominant failure mode: anachronistic compositing. This occurs when models generate personas that sound like historical figures but reason using knowledge and moral frameworks that postdate them, often due to the overwhelming influence of culturally dominant, later representations in training data. The proposed solution is a shift towards epistemic simulation, emphasizing corpus-boundedness, temporal anchoring, and expert-loop evaluation.
- 6 Things to Know about AIE World's Fair 2026
The AI Engineering (AIE) World's Fair 2026 is significantly larger than previous events, aiming to provide a broad buffet of curated content for attendees. The event emphasizes a blend of evergreen and new topics, with a focus on practical workflows and industry-research collaboration. It seeks to foster serendipitous discovery and community building, moving beyond algorithmically driven content feeds.
- The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks
This talk outlines a production AI playbook, a framework designed to guide the deployment of AI systems at an enterprise scale. It addresses common pitfalls encountered when moving AI from demos to production, emphasizing the need for structured approaches to evaluation, observability, data foundation, orchestration, and governance. The framework aims to ensure AI systems are measurable, accountable, and reliable in real-world applications.
- Your Agent's Biggest Lie: \"I Searched the Web\" — Rafael Levi, Bright Data
This talk addresses the common issue of Large Language Models (LLMs) fabricating information, particularly when attempting to access real-time web data. The core problem stems from LLMs being programmed to please users, leading them to generate plausible but incorrect answers rather than admitting they cannot fulfill a request. This is exacerbated by increasing web-based anti-bot measures, such as CAPTCHAs and AI-detection systems, which prevent LLMs from accessing accurate, up-to-date information.
- You Might Not Need 50 Diffusion Steps — Ziv Ilan, Nvidia
This talk explores techniques to accelerate diffusion models for image and video generation, addressing the challenge of high latency in practical applications. The core thesis is that by employing methods like quantization, caching, and distillation, developers can significantly reduce the number of diffusion steps required, thereby improving performance and enabling real-time generation without sacrificing quality.
- Why MCP and ChatGPT Apps Use Double Iframes — Frédéric Barthelet, Alpic
This talk explores the technical reasons behind the double iFrame implementation used by ChatGPT and other MCP (Meta-Conversational Platform) applications to render third-party UIs. It delves into the security implications of Content Security Policy (CSP) and sandboxing, explaining why direct iFrame embedding or using source-doc attributes are problematic. The presentation highlights the double iFrame approach as a solution to isolate third-party code while enabling necessary functionalities.
- Why Can't Anyone Answer Questions About the Business? — Garrett Galow, WorkOS
This talk addresses the challenge of accessing and answering business-related questions within an organization, particularly for non-technical teams. It introduces "Studio," an internal workspace developed at WorkOS, designed to empower employees to build tools and dashboards for self-service data analysis. Studio leverages AI agents to interact with various data sources, enabling users to ask questions and generate reusable data visualizations and reports without direct engineering intervention.
- The agent-ready web: Simplify user actions with WebMCP — Tara Agyemang, Google
This talk introduces WebMCP, a proposed web standard designed to simplify how AI agents interact with websites. Instead of agents relying on brittle methods like screen scraping, WebMCP allows websites to expose their functionalities as structured tools. This approach aims to significantly improve the performance and reliability of AI agents performing actions on behalf of users, ultimately creating better user experiences.
- Your Attention Is the Bottleneck, Not Your Agents — Zack Proser, WorkOS
This talk argues that AI agents are no longer the primary bottleneck in developer productivity; human attention is. While agents can scale infinitely and handle complex tasks with verification, human cognitive limits remain the hard constraint. The presentation suggests strategies for developers to manage their attention, leverage agents effectively, and maintain balance in an increasingly powerful AI-assisted workflow.
- Stop Making Models Bigger, Make Them Behave — Kobie Crawford, Snorkel
This talk challenges the assumption that larger AI models are always superior, particularly for enterprise use cases requiring reliability and security. It presents research demonstrating that a smaller, 4-billion parameter model, when fine-tuned with Reinforcement Learning (RL) and high-quality data, can outperform a significantly larger 235-billion parameter model on a tool-use task for financial analysis. The core thesis is that focusing on improving model behavior and tool discipline through targeted data and training can yield substantial performance gains, often more effectively and efficiently than simply increasing model size.
- Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind
This talk introduces Google DeepMind's Gemma family of open models, highlighting their capabilities and the benefits of using open-source AI for greater ownership and customization. The presenters emphasize that while proprietary models like Gemini offer cutting-edge performance, open models like Gemma provide crucial advantages for specific use cases, such as running models on local hardware, handling sensitive data, and adapting models to unique requirements.
- Self Driving Products: Product Signals to Pull Requests — Joshua Snyder, PostHog
This talk introduces a pipeline designed to automate the process of turning product observability data into actionable code changes. Instead of engineers manually reviewing dashboards and creating pull requests (PRs) for issues, the system aims to automatically detect product signals, diagnose problems, and generate PRs for review or even direct merging. The ultimate vision is a product that can largely build and maintain itself by learning from user interactions and data.
- GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod
This talk introduces RunPod's Flash, a Python SDK designed to streamline the development and deployment of AI models on GPU infrastructure. The core thesis is that developers can significantly reduce iteration time by deploying functions directly to the cloud from their local IDE, eliminating the need for manual commits, Docker builds, and server configuration. This allows for rapid testing and iteration on GPU-accelerated workloads.
- RAG is dead, right?? — Kuba Rogut, Turbopuffer
This talk challenges the notion that Retrieval Augmented Generation (RAG) is obsolete, arguing instead that hybrid, tool-rich retrieval is becoming the standard for sophisticated agentic search. While simple RAG, often equated with basic vector search, was effective in early AI development, more advanced applications now leverage agentic search. This approach involves agents iteratively reasoning over context using a suite of tools, including vector search, full-text search, and other filters, to progressively refine their understanding and complete tasks.
- From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind
This talk explores advancements in AI audio processing, focusing on Google DeepMind's Gemini models. The core thesis is that Gemini's sophisticated audio understanding capabilities enable richer transcription, robust reasoning, and more nuanced speech generation, moving beyond simple speech-to-text to a more comprehensive audio comprehension and synthesis system.
- 2026 AI Engineer Vibe Reel
This talk, titled 2026 AI Engineer Vibe Reel, explores the evolving landscape of AI engineering, focusing on prototyping and shipping AI-powered products. It offers insights into the current trends and future directions that AI engineers, product managers, and builders should be aware of as they develop and deploy AI solutions. The reel aims to capture the essence of what it means to be an AI engineer in the near future.
- Road to 5 Million Tokens: Breaking Barriers in Long Context Training — Max Ryabinin, Together AI
This talk details Together AI's research into training large context models, specifically aiming to break the 5 million token barrier. The core challenge lies in overcoming memory and computational bottlenecks inherent in standard transformer architectures when dealing with extremely long sequences. The research explores and combines various techniques to enable efficient training at unprecedented context lengths.
- Why More Context Makes Your Agent Dumber and What to Do About It — Nupur Sharma, Qodo
This talk explores the counterintuitive finding that providing more context to AI agents can sometimes lead to dumber or less effective results. The core thesis is that current LLMs tend to focus on the initial and final parts of a given context, often ignoring or "purging" information in the middle. This phenomenon, termed the "U-curve" effect, necessitates a shift from simply dumping more data to optimizing context strategically.
- Why Eval++ Is the Next Great Compute Primitive — Sunil Pai & Matt Carey, Cloudflare
This talk introduces Eval++, a new compute primitive designed to enhance the development and deployment of AI agents. The core idea is to leverage Cloudflare's stateful serverless technology, specifically Durable Objects, as an ideal execution environment for agents. This approach enables long-running processes, persistence, and efficient scaling, fundamentally changing how developers can build and interact with AI-powered applications.
- LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize
This talk introduces a platform for LLM observability, evaluation, and experimentation, emphasizing that building AI systems is fundamentally an engineering discipline. It highlights the importance of understanding what an AI system is doing (observability), how to measure its performance (evaluation), and how to systematically improve it (experimentation). The core thesis is that these processes, while complex in the non-deterministic world of AI, can be managed and eventually automated through robust engineering practices and tooling.
- Under 5 minutes to a deployed LLM endpoint — Audry Hsu, RunPod
This talk introduces RunPod as a cloud AI infrastructure platform designed to simplify GPU access and model deployment for developers. The core thesis is that managing complex infrastructure, especially GPU hardware, is a significant hurdle for builders, and RunPod aims to abstract this away. The platform allows users to bring their own code and models, whether private or open-source, and deploy them quickly, focusing on enabling developers to build applications rather than manage hardware.
- From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data
This talk explores building self-maintaining data collection pipelines using Large Language Models (LLMs) and specialized tools. It addresses the challenge of collecting large-scale data from websites, particularly those with anti-bot measures, by shifting from direct LLM parsing of every page to an agent-driven approach that builds and maintains scrapers. The core idea is to empower LLMs with the ability to create and manage the tools needed for data extraction, thereby saving significant computational resources and development time.
- Building Interactive UIs in VS Code with MCP Apps — Marlene Mhangami & Liam Hampton, GitHub
This talk introduces Model Context Protocol (MCP) apps, which enable server tools to return rich, interactive components directly within chat interfaces. MCP apps enhance user experience by allowing direct interaction with generated UIs, moving beyond simple text or emoji responses. This technology facilitates complex tasks like data exploration and e-commerce within a chat environment, improving efficiency and engagement.
- Evals Are Broken, Use Them Anyway — Ara Khan, Cline
This talk critiques the current state of AI model evaluations, arguing that while flawed, they remain essential tools for development. The speaker proposes a three-stage approach to using evals effectively: leveraging external evals, using them to improve internal agents, and building custom evals for specific use cases. The core message is to use evals pragmatically, understanding their limitations while still employing them for iterative improvement.
- Building safe Payment Infrastructure for the autonomous economy — Steve Kaliski, Stripe
This talk addresses the critical need for secure payment infrastructure in an increasingly autonomous economy, where AI agents will act as economic actors. It highlights the inherent risks of agents transacting online, such as purchasing from incorrect vendors, selecting the wrong items, or spending unintended amounts. The core thesis is that while discovery and exploration benefit from non-determinism, financial transactions require strict determinism to ensure safety and prevent fraud.
- Building Agent Interfaces: Lessons from Chrome DevTools (MCP) for Agents — Michael Hablich, Google
This talk introduces Chrome DevTools for Agents, a purpose-built debugging tool for AI agents, drawing parallels to the widely used Chrome DevTools for human web developers. The core thesis is that agents, as a distinct user class, require specialized interfaces and debugging capabilities to overcome their limitations, such as processing vast amounts of data or understanding complex tool descriptions. The presentation emphasizes engineering lessons learned from building and deploying this tool, focusing on effectiveness, efficiency, discoverability, and trust.
- Dark Factory: OpenClaw Ships Faster Than You Can Read the Diff — Vincent Koc, OpenClaw
This talk explores the rapid development and deployment practices within the OpenClaw project, likening the process to a "dark factory" that ships code at an unprecedented velocity. It argues that the sheer speed of development necessitates a shift from traditional code review to a more managerial approach, where engineers oversee swarms of autonomous agents. The core thesis is that effective engineering in this new paradigm relies on process, intuition, and managing agent workflows rather than solely on individual coding prowess.
- Beyond Transcription: Building Voice AI That Understands Conversations — Hervé Bredin, pyannoteAI
This talk explores building AI systems that understand conversations beyond simple transcription. It highlights the importance of identifying who spoke, when they spoke, and even how they spoke to gain deeper insights. The presentation introduces speaker diarization as a core technology for this purpose and discusses the challenges and advancements in achieving accurate speaker-attributed transcription.
- Text Diffusion — Brendan O’Donoghue, Google DeepMind
This talk introduces text diffusion, a forward-looking research area in AI that offers an alternative to traditional autoregressive text generation. Unlike models that generate text one token at a time, text diffusion models start with a noisy sequence and iteratively refine it to produce clean output. This approach promises significant advantages in inference speed and enables unique capabilities such as bidirectional reasoning and in-place editing.
- The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI
This talk addresses the growing gap between the capabilities of AI agents and the methods used to evaluate them. It argues that while progress is evident in areas like coding agents, current evaluation benchmarks are falling behind, creating hesitation in deploying these agents in high-stakes environments. The presentation outlines key principles for building effective benchmarks, distinguishing between the scientific rigor required for measurement and the artistic vision needed to shape the future of AI development.
- SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius
This talk details the practical lessons learned from evaluating coding agents on real-world software engineering tasks using the rebench leaderboard. It emphasizes the critical need for robust evaluation beyond gut feelings or limited testing, especially as models are deployed to production. The presentation highlights the challenges and methodologies involved in creating and maintaining a "fresh" and "real-world" benchmark, focusing on the complexities of software engineering tasks that require understanding repository structures, writing and running tests, and handling multi-turn interactions and tool use.
- Beyond Components: Designing Generative UI for MCP Apps — Ruben Casas, Postman
This talk explores the evolution of generative UI for Multi-Component Platform (MCP) applications, moving beyond static components to more dynamic and collaborative interfaces. It posits that current AI models are capable of generating sophisticated UI code, prompting a re-evaluation of how users interact with AI-powered applications. The core thesis is that the future of generative UI lies in collaborative experiences rather than just component-based outputs.
- Benchmarking semantic code retrieval on Claude Code — Kuba Rogut, Turbopuffer
This talk explores the effectiveness of semantic code retrieval for AI coding assistants, specifically benchmarking it against traditional grep-based methods on Claude Code. The core thesis is that while grep is simpler and often sufficient, semantic search, powered by vector embeddings, can significantly improve precision and the agent's ability to locate relevant code, especially for complex tasks. This approach offers a form of cached compute, reducing redundant processing and potentially leading to better accuracy and user satisfaction.
- BDD, ADR, PRD, WTF: Capturing Decisions for Humans and AI Alike — Michal Cichra, Safe Intelligence
This talk addresses the challenge of maintaining consistency and understanding in software development, particularly with the advent of AI agents. It proposes a system for capturing and enforcing decisions using established documentation practices like Architecture Decision Records (ADRs), Product Requirements Documents (PRDs), and Behavior-Driven Development (BDD). The core thesis is that by formalizing why and how decisions are made, developers and AI agents can operate more effectively and consistently.
- What Lies Beneath the API — Benjamin Cowen, Modal
As AI applications mature, companies increasingly turn to fine-tuning models for improved performance and cost efficiency. This talk explores the transition from using frontier APIs to custom model development, highlighting the emerging middle ground that simplifies fine-tuning and deployment. It suggests that differentiated products naturally evolve towards domain specificity, prompting a reevaluation of when custom models become necessary and feasible.
- Task Fidelity Scaling Laws — Kobie Crawdord, Snorkel
This talk explores the critical role of task quality in the performance of AI models, particularly in agentic tasks. The core thesis is that data quality and task quality are fundamentally intertwined, and improving task quality directly leads to better model training outcomes and performance uplifts. The research validates this by comparing model performance on high-quality versus low-quality tasks, demonstrating a significant difference in learning efficiency.
- How Lovable self-improves every hour — Benjamin Verbeek, Lovable
This talk details Lovable's approach to continuous learning and self-improvement for AI agents, aiming to prevent repeated mistakes. The core thesis is that by systematically identifying and learning from user friction points, AI systems can become more robust and user-friendly, particularly for non-technical users. The presentation outlines two primary methods Lovable employs to achieve this: a "Lovable Stack Overflow" for learning from solvable issues and a "venting tool" for agents to report unsolvable problems.
- 20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
This talk challenges the conventional understanding of "state-of-the-art" AI models, arguing that relying solely on public leaderboards or internal manual evaluations can lead to suboptimal choices. It emphasizes that true state-of-the-art is context-dependent and that efficiency, not just raw quality, is a critical factor in model selection for practical applications. The core thesis is that a more nuanced approach to benchmarking, considering specific use cases and efficiency metrics, reveals a landscape of multiple specialized, high-performing models rather than a single dominant one.
- What if the network was the sandbox? — Remy Guercio, Tailscale
This talk proposes a shift in how we conceptualize sandboxes for AI agents, suggesting the network itself can serve as the sandbox. Instead of relying on traditional methods like API keys or OAuth within a VM or container, the approach leverages network-level identity and permissions. This allows for granular control over agent access and behavior directly at the network layer, enabling more secure and observable AI agent deployments.
- How to talk to statues — Joe Reeve, ElevenLabs
This talk explores the rapid prototyping and viral success of an AI-powered app that allows users to converse with statues. The speaker, Joe Reeve from ElevenLabs, details how the app was built in just two hours using existing APIs, highlighting the power of "vibe coding" and the importance of storytelling in product development. The discussion also touches upon the broader implications of AI in creative fields, user interaction patterns, and the future of voice interfaces.
- Can LLMs generate Enterprise Quality Code? — Prasenjit Sarkar, Sonar
This talk addresses the quality of code generated by Large Language Models (LLMs) for enterprise use. While LLMs show high functional correctness on standard benchmarks, they often produce code that is verbose, complex, insecure, and difficult to maintain. The presentation introduces an evaluation framework and a proposed agent-centric development cycle to address these shortcomings and ensure LLM-generated code meets engineering standards.
- Engineering voice agents: Latency, quality, and scale — Rishabh Bhargava, Together AI
This talk explores the engineering challenges and architectural patterns for building high-quality, low-latency voice agents at scale. It highlights that voice is a natural interface for human-computer interaction, moving beyond customer service to more complex applications. The core of the discussion revolves around the dominant pipeline architecture, its components, and the critical trade-offs involved in achieving real-time, intelligent, and natural-sounding voice interactions.
- Spec-Driven Testing for Agents With A Brain the Size of A Planet — Steven Willmott, SafeIntelligence
This talk introduces the concept of spec-driven testing for AI agents, arguing that simply using larger models or extensive datasets is insufficient for ensuring agent safety and reliability. It proposes a more comprehensive approach to defining and testing agent behavior by specifying not just expected outputs but also the context, rules, domain knowledge, and robustness requirements relevant to the agent's intended task.
- How I deleted 95% of my agent skills and got better results — Nick Nisi, WorkOS
This talk explores how to improve AI agent performance by reducing complexity and focusing on verifiable outcomes. The speaker advocates for enforcing actions through structured pipelines rather than relying on agent instructions alone. By minimizing unnecessary agent skills and implementing robust verification mechanisms, developers can achieve better results and reduce errors. The core idea is to shift from trusting agents to demanding proof of their work.
- How We Built Zeta2: Training an Edit Prediction Model in Production — Ben Kunkle, Zed
This talk details the process of training Zeta2, an edit prediction model for the Zed code editor. The core thesis revolves around using distillation from frontier models and leveraging opt-in production data to create a specialized, fast model for predicting user edits. The training pipeline emphasizes data processing, prompt engineering for teacher models, and a novel approach using "settled data" to generate high-quality training examples.
- Why (Senior) Engineers Struggle to Build AI Agents — Philipp Schmid, Google DeepMind
Building AI agents presents unique challenges compared to traditional software development. Instead of acting as traffic controllers with predefined rules, engineers now function more like dispatchers, defining goals for agents without dictating every step. This shift requires embracing the non-deterministic nature of AI, treating errors as inputs, and moving from rigid unit tests to broader evaluations of agent reliability and success rates.
- Reachy Mini: the $300 open source robot you can actually hack — Andres Marafioti, Hugging Face
This talk introduces Reachy Mini, an open-source robot designed to be affordable and hackable, making advanced robotics accessible to hackers, researchers, and students. The project aims to democratize interaction with robots, moving beyond human-like imitation to foster creative new forms of engagement. Reachy Mini is presented as a platform for exploring expressive, voice-interactive robotic experiences.
- Why your agents need decision traces, not just documents — Zach Blumenfeld, Neo4j
This talk introduces context graphs as a method to enhance AI agent decision-making beyond simple document retrieval. Context graphs store not only factual information but also past decision traces, precedents, and causal chains. This allows agents to make more informed and explainable decisions by understanding the reasoning behind previous choices, acting with a form of subject matter expertise.
- Reverse engineering a Viking VOIP phone protocol with Claude Code — Boris Starkov, Eleven Labs
This talk details the process of reverse-engineering a legacy Viking VOIP phone's protocol using Claude Code to enable connection with a modern conversational AI agent. The primary challenge was the phone's outdated Windows XP-compatible software, which posed setup difficulties for engineers using Mac operating systems. The solution involved leveraging Claude Code to discover communication ports, identify and brute-force two-letter command codes, and ultimately decipher a proprietary protocol, including a one-byte checksum encryption.
- How agent o11y differs from traditional o11y — Phil Hetzel, Braintrust
This talk differentiates traditional observability from agent observability, arguing that agents' non-deterministic nature and complex data structures necessitate a distinct approach. Traditional observability focuses on uptime and technical performance metrics like latency and error rates, using established tools. Agent observability, however, must also account for qualitative aspects of agent performance, such as grounding, tool usage, and adherence to brand standards, which require handling vast amounts of semi-structured and unstructured data.
- Most Enterprise Agentic Projects Are Doomed, Here's Why — Jess Grogan-Avignon & Jack Wang, Accenture
Most enterprise agentic AI projects are destined to fail due to fundamental tensions between traditional enterprise structures and the demands of machine-speed AI development. Enterprises, built for human pace with layers of control and process, struggle to adapt to the rapid iteration and emergent behaviors characteristic of AI. This talk outlines five key tensions—speed, value, delivery, trust, and moat—and offers a framework for navigating them to achieve successful AI adoption.
- Context Graphs for Explainable, Decision-Aware AI Agents — Andreas Kollegger & Zaid Zaim, Neo4j
This talk introduces context graphs as a method to enhance AI agents' decision-making capabilities. By integrating knowledge graphs with AI agents, the goal is to move beyond simple knowledge provision to enabling agents to understand and act upon rules and policies, thereby making more informed decisions. This approach aims to fill the gap in AI agents' reasoning by providing them with the necessary context and decision-making frameworks.
- The AI Skill I Rely On Daily — Priscila Andre de Oliveira, Sentry
This talk emphasizes that the most significant benefit of AI in large codebases is not code generation, but comprehension. The speaker shares personal experience using AI, particularly Claude, to navigate complex codebases at Sentry. The core argument is that AI acts as an invaluable teammate for understanding code, which is crucial for maintaining quality and avoiding technical debt, especially when working on projects that are critical to one's livelihood.
- Why Rust is the Ideal Language for Vibe-Coding — Daniel Szoke, Sentry
This talk argues that Rust is a superior language for agentic coding, often referred to as vibe coding, despite conventional wisdom favoring languages like Python and TypeScript. The speaker contends that while dynamic languages are easier for LLMs to generate initially, this ease comes at the cost of increased error potential. Rust's strict compiler and numerous constraints, though posing a steeper learning curve for LLMs, provide deterministic guardrails that significantly reduce bugs and improve code reliability.
- The maturity phases of running evals — Phil Hetzel, Braintrust
This talk outlines the maturity phases of running evaluations for AI agents, emphasizing that evals are crucial for ensuring agent quality, mitigating risks, and understanding performance improvements. The speaker suggests a progression from initial human-based assessments to more automated and complex evaluation strategies as agent complexity increases. The core idea is to systematically build confidence in agent behavior before and after deployment.
- Run Frontier AI at Home — Alex Cheema, EXO Labs
This talk explores the challenges and opportunities of running advanced AI models locally on consumer hardware. The core thesis is that significant advancements in performance and efficiency are achievable by optimizing the entire AI stack, from hardware and software to model architecture, rather than solely relying on cloud-based solutions. The presentation advocates for a future where powerful AI can be accessed privately and affordably at home.
- What the Best Agents Share — Mardu Swanepoel, Flinn AI
This talk explores four key patterns observed in effective AI agents, drawing inspiration from existing tools like Cursor, Claude, and Harvey. The core thesis is that by studying and adapting these successful patterns, developers can build more capable and user-friendly agents. The presentation emphasizes learning from successful designs to improve agent development.
- Stop babysitting your agents... — Brandon Waselnuk, Unblocked
This talk addresses the challenge of agents lacking necessary context, forcing engineers to act as the "context engine." It proposes that instead of babysitting agents, a dedicated context engine is needed to provide agents with comprehensive, runtime-relevant information. This engine should understand organizational specifics, resolve conflicts, and deliver token-optimized responses, ultimately enabling agents to produce code and actions that reflect deep team integration.
- Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind
This talk addresses the challenges of AI evaluations, which are currently scattered, quickly become outdated, and often lack transparency and verifiability. The speakers propose solutions to democratize the evaluation process, enabling a broader community to contribute to and benefit from robust AI benchmarking. Their work aims to foster more equitable AI development by allowing diverse expertise to shape evaluation standards.
- Does GenAI \"belong\" to data scientists? — Phil Hetzel, Braintrust
This talk challenges the notion that generative AI agents exclusively belong to data scientists or machine learning engineers. It argues that while these roles bring valuable expertise in model understanding and rigorous testing, the nature of modern AI agents—built upon pre-trained models and adaptable through natural language inputs—opens the door for broader team involvement. The core thesis suggests that successful agent development benefits from diverse skill sets, including product engineers and subject matter experts, to effectively bridge the gap between complex technology and real-world problem-solving.
- Bounded Autonomy: Between Free Will and Determinism — Angus J. McLean, Oliver
This talk explores the evolving relationship between humans and large language models (LLMs), particularly in the context of agentic systems. It argues against the hype surrounding rapid AI advancements, emphasizing that core LLM capabilities have not fundamentally changed. Instead, many current tools act as temporary fixes for inherent model limitations like data inefficiency and a lack of continuous learning. The core thesis suggests that understanding and leveraging these limitations through self-imposed constraints and diverse representation structures can lead to more effective and creative AI applications.
- How Google DeepMind Runs Agents at Scale — KP Sawhney & Ian Ballantyne, Google DeepMind
This talk discusses Google DeepMind's approach to building and scaling agentic software. It highlights the Antigravity framework as an integrated IDE and agent management platform, showcasing its capabilities in code generation, task execution, and human-in-the-loop feedback. The discussion also touches upon the challenges and strategies for managing agentic systems at scale, including resource allocation, cost efficiency, and robust evaluation methods.
- Let's Talk About FOMAT: Fear of Missing Agent Time — Michael Richman, Cmd+Ctrl
This talk addresses the challenge of managing and interacting with coding agents, a problem termed FOMAT (Fear of Missing Agent Time). The core thesis is that current agentic workflows often require constant oversight and can lead to developers being blocked or unable to effectively manage multiple agent sessions, especially when away from their development machines. The presented solution, Command and Control, aims to provide a unified interface for monitoring, interacting with, and launching agent sessions from any device.
- Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum
This talk proposes a shift from homogeneous AI systems, which rely on scaling single models across identical hardware, to a new paradigm of heterogeneous intelligence. This approach leverages the inherent complexity of real-world problems by decomposing them into sub-problems that can be addressed by diverse models, workflows, and hardware working in concert. The core thesis is that this heterogeneity, when properly orchestrated, leads to more efficient, faster, and cheaper AI systems.
- Introducing WebMCP: Agents in the Browser — RL Nabors
This talk introduces WebMCP, a system that enables agents to interact with web content directly within the browser. It explores how existing browser primitives can be leveraged to transform websites into interactive canvases for AI agents, moving beyond simple chat interfaces to richer, embedded experiences. The core idea is to make web content accessible and actionable for agents, similar to how humans navigate and interact with websites.
- The Missing Primitive for Agent Swarms — Lou Bichard, Ona
This talk explores the concept of a software factory, defined as the incremental removal of humans from the Software Development Life Cycle (SDLC) to achieve automated workflows. It highlights that while runtimes, orchestration, and triggers for agents are largely solved, agent coordination remains a significant challenge. The presentation suggests that breaking down the SDLC into micro-steps and developing robust coordination mechanisms are crucial for realizing truly autonomous software development.
- Prompt to Pipeline: Building with Google's Gen Media Stack — Paige & Guillaume, Google DeepMind
This talk introduces Google DeepMind's latest advancements in generative AI, focusing on the Gemini family of models and their applications. The presentation highlights the multimodal capabilities of Gemini, its cost-effectiveness, and its integration into tools like AI Studio for building applications. It also touches upon the development of specialized models for generative media, including image, video, and music creation, and the increasing accessibility of powerful AI models for on-device and local execution.
- Fast Models Need Slow Developers — Sarah Chieng, Cerebras
The talk "Fast Models Need Slow Developers" by Sarah Chieng of Cerebras argues that the advent of significantly faster AI coding models necessitates a fundamental shift in developer habits. While models like Codex Spark can generate code at 1,200 tokens per second, up to 20 times faster than previous generations, developers must adapt their workflows. Without this adaptation, the increased speed will simply lead to the rapid generation of low-quality code and technical debt, rather than unlocking new capabilities.
- Lobster Trap: OpenClaw in Containers from Local to K8s and Back — Sally Ann O'Malley, Red Hat
This talk demonstrates how to run OpenClaw, an open-source coding agent, within containers, from local development environments to Kubernetes clusters. The core thesis is that containerization provides a reproducible, secure, and portable way to deploy and manage AI workloads like OpenClaw, simplifying development, onboarding, and scaling.
- Gemini Nano on device — Florina Muntenescu & Oli Gaymond, Google DeepMind
This talk introduces Gemini Nano, Google's on-device AI model for Android, and the ML Kit GenAI APIs that provide access to it. The presentation emphasizes the benefits of on-device processing, including enhanced privacy, offline capabilities, and cost savings, while also outlining hybrid approaches that leverage cloud models when necessary. The goal is to provide developers with a comprehensive offering for building intelligent experiences across a range of Android devices.
- Cooking with Agents in VS Code — Liam Hampton, Microsoft
This talk explores how AI agents can be integrated into the VS Code environment to enhance developer productivity. It categorizes agents into local, background, and cloud types, each suited for different tasks. The core thesis is that VS Code can serve as a unified interface for managing and utilizing these diverse AI agents, streamlining workflows and reducing cognitive load for developers.
- Scaling Agents on Kubernetes with acpx and ACP — Onur Solmaz, OpenClaw
This talk explores the challenges and solutions for scaling open-source AI agents, particularly within the Kubernetes ecosystem. It introduces ACP (Agent Client Protocol) as a standardized way for human users to interact with agents, aiming to reduce duplicated effort in building client integrations. The discussion also covers developing automated workflows for tasks like code review and error reporting, emphasizing the potential for agents to handle complex, multi-step processes.
- Your Coding Agent Should Do AI System Engineering — Ben Burtenshaw, Hugging Face
This talk proposes that AI coding agents should be leveraged for complex AI systems engineering tasks, moving beyond simpler applications. The core idea is to use agents to tackle challenging problems in machine learning engineering and systems design. To enable this, the speaker emphasizes the need for standardized repositories, particularly on platforms like the Hugging Face Hub, which already provide many necessary components.
- Any-to-Any: Building Native Multimodal Agents - Patrick Löber, Google DeepMind
This talk introduces the concept of "any-to-any" multimodal agents, enabled by Google DeepMind's Gemini API. The core idea is to build agents that can understand and generate across various modalities including text, code, images, audio, and video. The presentation outlines the architecture for such agents, focusing on multimodal understanding and native generation capabilities, and demonstrates how to build a notebook LM clone as an example.
- Skill issue: Lessons from skilling up coding agents to use Langfuse - Marc Klingen, Clickhouse
This talk explores the challenges and lessons learned in skilling up coding agents to effectively use tools like Langfuse. It highlights the evolution from complex workflows to more autonomous agents and emphasizes the role of formalized skills as shortcuts for reliability. The core thesis is that by leveraging tracing and structured skills, developers can significantly improve agent capabilities and streamline the integration of complex tools into AI applications.
- From 46% to 90%: Fine-Tuning Tiny LLMs for On-Device Agents — Cormac Brick, Google
This talk explores the development and application of small and tiny Large Language Models (LLMs) for on-device AI agents. It highlights the benefits of on-device processing, such as improved latency, privacy, and offline functionality. The presentation introduces the Google AI Edge stack, including MediaPipe and the lighter TLM runtime, and discusses two primary approaches: leveraging system-level GenAI like Gemini Nano and developing App GenAI with customizable tiny LLMs.
- What Breaks When You Build AI Under Sovereignty Constraints - Bilge Yücel, deepset GmbH
This talk explores the concept of AI sovereignty, defined as an organization's ability to design, deploy, and operate AI systems on its own terms. It breaks down sovereignty into four pillars: data, model, infrastructure, and operational control. The presentation highlights the challenges and trade-offs encountered when implementing AI systems under these constraints, particularly when moving from convenience-based SaaS solutions to more controlled, self-hosted environments.
- Don't Build Slop (4 Levels of AI Agent Maturity) - Ara Khan, Cline
This talk proposes a structured approach to building AI agents, moving beyond the hype to focus on practical maturity. It outlines four distinct levels of agent development, from initial framework experimentation to sophisticated cloud deployment. The core thesis is to slow down, identify solvable problems, and build useful agents by understanding their lifecycle and potential pitfalls.
- Personalization in the Era of LLMs - Shivam Verma, Spotify
This talk explores Spotify's approach to personalization in the age of Large Language Models (LLMs), moving beyond traditional recommendation systems. It details how Spotify builds foundational user and content models, integrating them into a steerable, personalized generative system. The core thesis is that by leveraging LLMs and advanced modeling techniques, Spotify can create more relevant and user-controlled recommendations.
- Rewiring the State — Eoin Mulgrew, No. 10 (Downing Street)
This talk outlines an initiative by the No. 10 data science team to radically scale AI engineering and development capabilities within the UK government. The core thesis is that by establishing small, agile, and well-resourced "insurgent units" operating with high political backing and autonomy, the government can overcome traditional bureaucratic hurdles to drive AI adoption and address public service delivery crises. This approach aims to attract top technical talent by offering challenging work and competitive compensation, thereby accelerating innovation and improving efficiency across government functions.
- Let's go Bananas with GenMedia — Guillaume Vernade, Google DeepMind
This talk demonstrates how to use Google DeepMind's GenMedia models to create rich multimedia content, such as images, videos, and music, by illustrating a book. The presentation walks through practical examples of generating character images, scene illustrations, video clips, and musical scores, all while emphasizing the integration and capabilities of various AI models. The core idea is to leverage AI to bring creative works to life through diverse media.
- Anthropic Workshop: Build Agents That Run for Hours — Ash Prabaker & Andrew Wilson
This talk explores the evolution of AI agents, focusing on enabling them to run for extended periods, from minutes to hours or even days. It details the challenges of long-running agents, such as context limitations and model judgment, and presents Anthropic's approach to overcoming these through model improvements and harness development. The discussion highlights how co-evolving models and scaffolding have led to more capable and autonomous AI systems.
- Harnesses in AI: A Deep Dive — Tejas Kumar, IBM
This talk explores the concept of AI harnesses, defining them as the components surrounding an AI model that provide grounding in reality and ensure reliability. The speaker argues that harnesses are crucial for making AI agents dependable, especially when using rented, black-box models. By building a harness, developers can create stable environments that control agent behavior, regardless of the underlying model's non-deterministic nature.
- Fighting AI with AI — Lawrence Jones, Incident
This talk explores how AI engineers can leverage AI itself to manage the complexity of AI products. It details strategies for using internal tools to understand, debug, and improve AI systems, particularly when dealing with intricate prompt chains and complex decision-making processes. The core thesis is that effective AI development requires applying AI-powered solutions to the engineering and debugging workflows.
- Why Your AI UX Is Broken (and It's Not the Model's Fault) — Mike Christensen, Ably
This talk argues that the common direct HTTP streaming approach for AI chat applications fundamentally limits user experience quality. The default pattern, often using Server-Sent Events (SSE), creates a single, point-to-point connection between a client and an agent, which breaks down when users switch devices, experience network interruptions, or need to interact with a working agent. The presentation proposes a shift towards a decoupled architecture using durable sessions to enable more resilient, multi-surface, and interactive AI experiences.
- Beyond Code Coverage: Functionality Testing with Playwright MCP — Marlene Mhangami, Microsoft
This talk explores how to move beyond traditional code coverage metrics to ensure true functionality testing, especially in the context of AI-assisted development. It argues that while AI can increase code output, its true productivity gains are amplified by clean codebases and robust testing. The presentation introduces Playwright as a tool for end-to-end browser testing, demonstrating how it can be integrated with AI agents to accelerate development workflows, particularly within a test-driven development (TDD) framework.
- How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs
This talk argues that building successful vertical AI products is fundamentally an organizational challenge, not solely a technical one. The core thesis is that effectively integrating domain expertise into AI development processes is more critical than the sophistication of the AI models themselves. The speaker proposes a framework for structuring organizations around domain experts, enabling them to contribute at different levels, from direct input to system design, ultimately leading to more differentiated and effective AI products.
- Connecting the Dots with Context Graphs — Stephen Chin, Neo4j
This talk introduces context graphs as a solution to the problem of AI agents and systems lacking comprehensive understanding due to siloed data. It argues that by connecting disparate data sources, previous decision traces, and tool-use reasoning into a knowledge graph, AI can move from controlling engineers to being a controllable tool. This approach aims to provide grounded, complete information for more reliable AI-driven decision-making and application building.
- Agents Don't Do Standups: Building the Post-Engineer Engineering Org — Mike Spitz, PFF
This talk explores a case study at PFF, a sports data company, where they transformed their engineering organization by integrating AI agents. The core thesis is that by shifting focus from optimizing individual engineer output to enhancing agent capabilities, significant gains in deployment frequency and product quality can be achieved. This approach led to a reimagining of traditional engineering processes, moving away from rigid structures like Scrum towards more agile, feedback-driven workflows.
- Combine Skills and MCP to Close the Context Gap — Pedro Rodrigues, Supabase
This talk explores how to effectively guide AI agents when they interact with complex products like Superbase. It argues that while agents are capable of many tasks, they require specific guidance to handle novel situations, security protocols, and optimized workflows. The presentation introduces the concept of "skills" as a method to provide this guidance, demonstrating how they improve agent performance and safety, particularly when integrated with tool-calling capabilities.
- How Building with AI Can Double the Throughput of Your Engineering Team — Brian Scanlan, Intercom
This talk explores how Intercom leveraged AI, specifically focusing on integrating AI agents into their software development lifecycle, to significantly increase engineering throughput. The core thesis is that by treating AI as a platform and empowering engineers with the right tools and guidance, organizations can achieve substantial productivity gains, potentially doubling output without increasing headcount.
- Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize
This talk focuses on the practical aspects of evaluating and improving AI agents, moving beyond simple "vibe checks" to implement robust testing frameworks. It emphasizes that effective evaluation is crucial for shipping reliable AI applications, especially agents, which are prone to cascading failures. The session introduces methods for capturing agent behavior through tracing, categorizing failures, and implementing various types of evaluations (code-based, LLM-as-judge, human) to ensure agents perform as expected and to drive iterative improvements.
- Mind the Gap (In your Agent Observability) — Amy Boyd & Nitya Narasimhan, Microsoft
This talk addresses the critical need for robust observability in AI agents, highlighting the gap between an agent's intended functionality and its actual behavior in production. The presenters introduce the "Mind the Gap" analogy to emphasize the importance of continuous evaluation and monitoring throughout an agent's lifecycle. They advocate for a proactive approach to identify and address discrepancies, ensuring agent reliability, quality, and safety.
- Make your own event-sourced agent harness using stream processors — Jonas Templestein, Iterate
This talk introduces a novel approach to building agent harnesses using an event-sourced architecture powered by stream processors. The core thesis is that by treating all agent interactions and state changes as events in a log, debugging becomes significantly easier, and extensibility is greatly enhanced. This event-driven model aims to simplify experimentation with different agent configurations and facilitate composable agent systems.
- Self-Training Agents: Hermes Agent, HF Traces, Skills, MCP & Finetuning — Merve Noyan, Hugging Face
This talk explores the burgeoning open-source AI agent ecosystem, emphasizing the accessibility and power of open models and tools. It highlights how advancements in Hugging Face's platform, including the Hub, Traces, Skills, and MCP tools, are democratizing the development and deployment of sophisticated AI agents. The core thesis is that open-source AI is rapidly catching up to and even surpassing closed-source alternatives, offering greater control, privacy, and customization for developers.
- Building a Chess Coach — Anant Dole and Asbjorn Steinskog, Take Take Take
This talk details the development of an AI-powered chess coach, focusing on how to bridge the gap between traditional chess engines and Large Language Models (LLMs). The core challenge addressed is LLMs' inherent weakness in calculation and reasoning for games like chess, despite their ability to explain concepts. The solution involves a pipeline that leverages strong chess engines for analysis and LLMs for generating human-readable explanations, with an autonomous agent loop for continuous improvement.
- CI/CD Is Dead, Agents Need Continuous Compute and Computers — Hugo Santos and Madison Faulkner
This talk argues that traditional Continuous Integration and Continuous Deployment (CI/CD) pipelines are becoming obsolete due to the rise of agentic software development. The core thesis is that the increasing volume and complexity of changes generated by autonomous agents necessitate a shift towards a model of Continuous Compute. This new paradigm focuses on accelerating the entire development lifecycle, from code generation to deployment, by integrating compute and caching more deeply into the workflow.
- Give Your Agent a Computer — Nico Albanese, Vercel
This talk introduces a framework for building AI agents that can interact with a computer, leveraging the AI SDK. The core thesis is that effective agents require an agent runtime, tools for interaction, and a computational environment like a sandbox file system for state persistence and code execution. The presentation demonstrates how to integrate these components to create more capable and autonomous AI agents.
- Lessons from Trillion Token Deployments at Fortune 500s — Alessandro Cappelli, Adaptive ML
This talk argues that reinforcement learning (RL) is crucial for moving AI models, particularly large language models (LLMs), from pilot stages to production. The speaker posits that the common failure of GenAI pilots stems from the "myth of the last mile," where initial MVPs built on proprietary models or instruction fine-tuning lack systematic improvement pathways. RL, by its nature, allows for the mathematical integration of feedback, leading to more effective model steering and enabling scaled, cost-efficient, and faster deployments.
- Malleable Evals: Why Are We Evaluating Adaptive Systems with Static Tests? — Vincent Koc, OpenClaw
This talk argues that traditional static testing methods are insufficient for evaluating adaptive AI systems. As AI applications become more dynamic and intent-driven, evaluation strategies must evolve to become equally malleable. The core thesis is that static benchmarks fail to capture the emergent behaviors and changing user interactions characteristic of modern AI, necessitating a shift towards more adaptive and continuous evaluation approaches.
- A Piece of Pi: Embedding The OpenClaw Coding Agent In Your Product — Matthias Luebken, Tavon
This talk explores the integration of coding agents, specifically the OpenClaw agent, into products. It emphasizes that coding agents are becoming a fundamental building block for software systems and highlights the Pi framework as an excellent tool for experimentation and development. The presentation demonstrates how agents can be leveraged to automate complex workflows, such as processing sales proposals, by interacting with various tools and systems.
- Viktor: AI Coworker That Lives in Slack — Fryderyk Wiatrowski
Victor is presented as an AI coworker that integrates directly into Slack, aiming to function like a human employee by participating in discussions and utilizing company tools. Unlike personal agents, Victor is designed as a company-wide asset, inheriting access to a vast array of integrations and possessing a broad, horizontal context across the organization. The system addresses challenges such as memory management for multiple users and the complexity of context across different Slack interaction modes.
- Why MLX — Prince Canuma, Neywa Labs
This talk introduces MLX, an array framework for Apple silicon, enabling developers to run AI models, including large language and multimodal models, entirely on-device. The core thesis is that on-device AI offers a viable and often superior alternative to cloud-based solutions, providing enhanced privacy, reduced costs, and greater accessibility, especially in areas with limited internet connectivity. MLX aims to democratize AI by making powerful models runnable on everyday Apple devices.
- Two Roads to Durable Agents: Replay vs. Snapshot — Eric Allam, CEO, Trigger.dev
This talk explores the fundamental shift required for AI agents in backend infrastructure, moving from stateless to stateful compute. It proposes two primary methods for achieving durable agents: replay and snapshotting. Replay, an evolution of traditional workflow engines, logs every interaction but can become unwieldy with complex agents. Snapshotting, a newer approach, preserves the entire execution state of a machine, offering a more robust solution for long-running, complex agent sessions.
- How we solved Context Management in Agents — Sally-Ann Delucia
This talk addresses the critical challenge of context management in AI agents, arguing that context engineering, rather than prompt engineering, is the key to agent success. The core thesis is that effective context management allows agents to retain necessary information while discarding irrelevant details, preventing failures caused by overloaded or insufficient context windows.
- Feedback Loops are All You Need — Mehedi Hassan, Granola
This talk argues that effectively integrating AI into products requires more than just one-shot prompting. The core thesis is that building robust feedback loops, enabling detailed tracing, and creating flexible development environments are crucial for iterating on AI features and ensuring they meet user needs. This approach allows for a more magical user experience rather than relying on hope and guesswork.
- Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral
This talk explores the recent architectural trends in text-to-speech (TTS) models, highlighting their convergence with Large Language Model (LLM) approaches. The core thesis is that TTS is increasingly being treated as a sequence modeling problem, similar to how LLMs process text, enabling more natural and lower-latency speech generation, particularly for agentic interfaces.
- Voice AI: when is the \"Her\" moment? — Neil Zeghidour, CEO, Gradium AI
The talk explores the current state of voice AI and the challenges in achieving a "Her moment," where AI voices are indistinguishable from humans and seamlessly integrated into conversations. While significant progress has been made in speech-to-text and text-to-speech, true conversational fluency, low latency, and the ability to handle complex tool calls remain major hurdles. The presentation contrasts cascaded systems with end-to-end speech-to-speech models, highlighting the limitations of current half-duplex speech-to-speech models in replicating human conversational nuances like overlapping speech and backchanneling.
- Give Your Chat Agent a Voice — Luke Harries, Head of Growth, ElevenLabs
This talk explores the evolution of chat agents, arguing that voice is the next natural and accessible interface, surpassing text-based interactions. It introduces Voice Engine, a product designed to easily integrate voice capabilities into existing chat agents, enabling richer, more interactive, and omni-channel experiences. The core thesis is that chat agents will either adopt voice or become obsolete, highlighting voice as a crucial upgrade for conversational AI.
- How Transformers Finally Ate Vision – Isaac Robinson, Roboflow
This talk explores the evolution of computer vision models, detailing how transformers, despite lacking inherent inductive biases for vision, ultimately surpassed traditional convolutional neural networks (CNNs). This shift was driven by massive, specialized pre-training techniques and leveraged infrastructure advancements from large language models, enabling transformers to learn visual patterns effectively.
- FLUX, Open Research, and the Future of Visual AI — Stephen Batifol, Black Forest Labs
This talk introduces FLUX, an open-source visual AI model developed by Black Forest Labs, and discusses its evolution and the future of visual AI. The presentation highlights FLUX's capabilities in text-to-image generation, image editing, and its progression towards visual intelligence. It also delves into a novel training methodology called Self Flow, designed to improve multimodal generative models by integrating representation learning directly into the generation process, eliminating the need for external encoders.
- Agentic Search for Context Engineering — Leonie Monigatti, Elastic
This talk explores agentic search as a critical component of context engineering for Large Language Models (LLMs). The core thesis is that effective context curation for LLMs relies heavily on sophisticated search mechanisms, moving beyond simple retrieval to agent-driven exploration of diverse information sources. The presentation highlights the evolution from basic RAG to agentic RAG and emphasizes the challenges and strategies for building robust search tools that agents can effectively utilize.
- Agent Optimization with Pydantic AI: GEPA, Evals, Feedback Loops — Samuel Colvin, Pydantic
This talk introduces Pydantic AI's agent optimization capabilities, focusing on the Jeppa library for prompt optimization and managed variables for dynamic configuration. The core idea is to systematically improve agent performance by iteratively refining prompts and configurations, demonstrated through a practical example of analyzing political family ties in MP data. The presentation highlights how these tools can lead to more accurate and efficient AI agents.
- Vibe Engineering Effect Apps — Michael Arnaldi, Effectful
This talk explores the "Vibe Engineering Effect" by demonstrating how to build applications from scratch using AI agents, focusing on practical development workflows rather than hype. The core thesis is that by structuring projects to leverage AI effectively, developers can significantly reduce manual coding, even for complex library-level tasks. The session emphasizes practical techniques for integrating AI into the development process, particularly when working with new or unfamiliar codebases.
- Everything You Need To Know About Agent Observability — Danny Gollapalli & Zubin Koticha, Raindrop
This talk addresses the critical need for agent observability in AI systems, highlighting how traditional testing and evaluation methods are insufficient for non-deterministic and complex agents. It proposes a shift towards a monitoring paradigm, emphasizing the importance of collecting both explicit signals (like error rates and latency) and implicit signals (like user frustration and refusals) to understand and improve agent behavior in production. The discussion also touches upon self-diagnostics as a low-effort method for agents to report their own issues.
- Full Walkthrough: Writing & Using Skills — Nick Nisi and Zack Proser
This workshop explores the concept and practical application of "skills" in AI agents, presenting them as discrete, composable units of work designed to enhance agentic workflows. The core thesis is that skills provide a structured way to encode essential context, instructions, and behaviors, preventing repetitive information input and ensuring consistent, predictable agent performance across various tasks and projects.
- The Multi-Agent Architecture That Actually Ships — Luke Alvoeiro, Factory
This talk introduces Missions, a multi-agent system designed to tackle complex software development tasks that extend beyond a single agent's capacity. The core thesis is that human attention, not intelligence, is the current bottleneck in software engineering. Missions aims to overcome this by enabling systems to execute tasks autonomously for extended periods, allowing human engineers to focus on higher-level strategic decisions.
- MCP UI: Extending the frontier — Liad Yosef and Ido Salomon, MCP Apps
This talk introduces MCP apps, a standardized way for tools and companies to send their user interfaces directly into chat applications. Instead of relying on text-based responses, which can be suboptimal and strip away brand identity, MCP apps allow for interactive UI components to be embedded within platforms like ChatGPT and Claude. This approach aims to preserve existing UI/UX knowledge and branding while enabling a more seamless integration of external services into conversational AI experiences.
- The Small Model Infrastructure Nobody Built (So We Did) — Filip Makraduli, Superlinked
This talk addresses the gap in infrastructure for small model inference, particularly for AI search and document processing. The speaker, Filip Makraduli, shares his journey from overlooking inference to co-founding the open-source Superlinked Inference Engine (SIE). The core thesis is that effective inference for agentic workflows requires a holistic approach combining robust model support with scalable infrastructure, enabling developers to efficiently deploy and manage diverse small models.
- Accelerating AI on Edge — Chintan Parikh and Weiyi Wang, Google DeepMind
This talk focuses on accelerating AI deployment on edge devices, highlighting Google DeepMind's Gemma models and the Lite RT framework. The core thesis is that by optimizing models for edge hardware and leveraging a unified cross-platform architecture, developers can achieve significant performance gains, enhanced privacy, and reduced latency for a wide range of AI applications.
- Demand-Driven Context: A Methodology for Coherent Knowledge Bases Through Agent Failure
This talk introduces a "demand-driven context" methodology to address the challenge of integrating institutional knowledge into AI agents. The core thesis is that instead of a "push" strategy where all knowledge is pre-loaded, a "pull" approach is more effective. This involves agents actively seeking and documenting information as they encounter problems, gradually building a coherent and useful knowledge base.
- Training an LLM from Scratch, Locally — Angelos Perivolaropoulos, ElevenLabs
This talk provides a hands-on guide to training a transformer-based Large Language Model (LLM) from scratch using PyTorch. It covers the fundamental building blocks of LLMs, including tokenization, model architecture, and the training loop, demonstrating how to implement these components with minimal code. The session emphasizes practical application, enabling participants to train a small model locally or on cloud platforms like Google Colab.
- Skill Issue: How We Used AI to Make Agents Actually Good at Supabase — Pedro Rodrigues, Supabase
This talk explores how to improve AI agent performance through the use of skills, which are essentially structured folders containing instructions and files for running workflows or providing custom information. The speaker, Pedro Rodrigues from Supabase, details the process of creating, testing, and automating the evaluation of these skills, emphasizing their role in enhancing agent capabilities within products like Supabase. The core idea is to leverage skills for progressive disclosure, allowing agents to access information only when needed, thereby optimizing context window usage.
- Ralph Loops: Build Dumb AI Loops That Ship — Chris Parsons, Cherrypick
This talk introduces "Ralph loops," a method for building AI agents that can iteratively improve their own work by repeatedly attempting a task. The core idea is to leverage the ability of modern AI models to self-correct and refine outputs through repeated execution of a prompt or instruction. This approach moves away from complex, brittle orchestration workflows towards simpler, more robust agentic loops that can be applied to various tasks, including coding and content generation.
- TLMs: Tiny LLMs and Agents on Edge Devices with LiteRT-LM — Cormac Brick, Google
This talk explores the advancements and applications of tiny Large Language Models (LLMs) and agent skills on edge devices. It highlights how these technologies enable powerful on-device AI experiences, focusing on reduced latency, enhanced privacy, and offline functionality. The presentation delves into the capabilities of Google's Gemma models and the LiteRT-LM runtime, showcasing their potential for building sophisticated AI-powered applications directly on user devices.
- Mergeable by default: Building the context engine to save time and tokens — Peter Werry, Unblocked
This talk introduces the concept of a "context engine" for AI agents, which aims to provide agents with the necessary information to perform tasks effectively without overwhelming them with irrelevant data. The core idea is to move beyond simple access to information and achieve true understanding, enabling agents to operate with the efficiency and insight of a seasoned team member. The presentation debunks common myths about context engines and shares lessons learned from building and implementing such systems.
- Context Is the New Code — Patrick Debois, Tessl
This talk proposes that context is the new code in software development, shifting focus from writing explicit code to generating and managing contextual information for AI agents. It outlines a "context development life cycle" analogous to the DevOps infinity loop, encompassing generation, testing, distribution, and observation of context. The core idea is that by engineering and refining context, developers can significantly improve the performance and reliability of AI-driven development processes.
- Human-in-the-Loop Automation with n8n — Liam McGarrigle
This talk demonstrates how to build a human-in-the-loop automation agent using n8n, focusing on managing Gmail and Google Calendar. The core idea is to create an agent that can perform tasks but requires human approval for sensitive or critical actions, ensuring control and preventing errors. The session covers setting up n8n, integrating AI models, defining tools for the agent, and implementing a human review step before execution.
- I Gave an AI Agent the Keys to My Life (Here's What Happened) — Radek Sienkiewicz (@velvetshark-com)
This talk details how an AI agent, specifically Open Claw, was incrementally integrated into an individual's daily life, granting it access to emails, notes, files, calendars, and system automations. The process was gradual, starting with simple chat functionalities and evolving to encompass a comprehensive knowledge base and automated daily operations, ultimately aiming to assist a future self.
- Software Engineering Is Becoming Plan and Review — Louis Knight-Webb, Vibe Kanban
This talk posits that as AI capabilities advance, the role of software engineers will increasingly shift from direct code writing to planning and reviewing AI-generated code. This evolution necessitates new approaches to managing workflows, particularly as AI agents begin to run for longer durations, requiring engineers to adopt a more parallel and managerial style of work.
- Mastering AI Pricing — Mayank Pant, Stripe
AI companies are experiencing hypergrowth, growing three times faster than traditional SaaS, but this rapid expansion presents significant pricing challenges. Traditional subscription or pure usage-based models are insufficient due to unpredictable external costs and the risk of margin erosion from power users. Pricing must evolve rapidly alongside product development to maintain a competitive advantage.
- Agents on the Canvas in tldraw — Steve Ruiz, tldraw
This talk explores the evolution of AI agents within the tldraw collaborative canvas environment. It begins with early experiments like Make Real, which translated drawings into functional prototypes, and progresses to integrating AI as a direct collaborator on the canvas. The core thesis is that by bringing agents directly onto the canvas, rather than confining them to sidebars, we can unlock more intuitive and powerful forms of human-AI collaboration, especially for creative and development tasks.
- Shipping complex AI applications — Braintrust & Trainline
This talk focuses on the practical challenges of shipping complex AI applications, moving beyond initial prototypes to robust, production-ready systems. It emphasizes the need for operational rigor, structured workflows, and continuous evaluation, drawing on experiences from Braintrust and Trainline. The core thesis is that while AI models are increasingly sophisticated, the operational practices for deploying and managing them at scale have lagged, creating a significant hurdle for delivering real customer value.
- Agents for Everything Else — swyx
This talk explores the expanding role of AI agents beyond traditional coding tasks, emphasizing their potential to enhance productivity across various business functions. The core thesis is that AI agents are becoming indispensable tools for "everything else," enabling individuals and small teams to achieve significant output by automating complex workflows and reducing manual effort. The speaker highlights personal experiences demonstrating how agents can streamline tasks from design implementation to data management and even purchasing.
- Building Conversational Agents — Thor Schaeff and Philipp Schmid, Google DeepMind
This talk introduces the Gemini Interactions API, a new unified API designed to simplify building with large language models and agents. It emphasizes a more developer-friendly interface, akin to industry standards, and introduces server-side state management for improved agent development. The session also showcases the Gemini Live API for real-time conversational AI applications, including audio and video processing, and demonstrates how to build a coding agent with file I/O and bash command execution capabilities.
- LLM codegen fails and how to stop 'em — Danilo Campos, PostHog
This talk addresses common failure modes in Large Language Model (LLM) code generation and offers strategies to mitigate them. The core thesis is that while LLMs can automate complex tasks like software integration, their outputs are prone to issues like outdated knowledge, architectural inconsistencies, and unpredictable behavior. By understanding these failure points and implementing specific techniques, developers can improve the reliability and effectiveness of autonomous coding agents.
- Replacing 12K LoC with a 200 LoC Skill — David Gomes, Cursor
This talk details how Cursor refactored a complex feature, originally spanning 12,000 lines of code, into a concise 200-line skill. This transformation leveraged existing primitives like agent skills and sub-agents, primarily using markdown to redefine functionality. The refactoring aimed to reduce maintenance overhead and improve user experience for advanced features.
- OpenAI Codex Masterclass — Vaibhav Srivastav & Katia Gil Guzman
This talk introduces OpenAI's Codex, an AI software engineering agent capable of performing a wide range of tasks beyond just writing code, such as running commands, executing tests, and exploring codebases. It highlights the underlying models, the unified agent harness for managing agent behavior, and the various interfaces through which users can interact with Codex, including a dedicated app, IDE extensions, and CLI. The presentation emphasizes the continuous improvement of Codex through model advancements and the development of features like plugins and automations to enhance developer productivity.
- Build & deploy AI-powered apps — Paige Bailey, Google DeepMind
This talk explores building and deploying AI-powered applications, focusing on the capabilities and accessibility of Google DeepMind's latest models and tools. It highlights the multimodal nature of Gemini models, their integration into AI Studio for rapid prototyping, and the potential for creating complex applications with features like code execution, search grounding, and dynamic world generation. The presentation emphasizes practical demonstrations over slides, showcasing how developers can leverage these advancements for diverse projects.
- Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI
This talk explores the unique challenges and opportunities in training small, frontier AI models, emphasizing that they are not simply scaled-down versions of larger models. The core thesis is that small models, designed for edge deployment, possess distinct characteristics like being memory-bound, having lower knowledge capacity, and being latency-sensitive. Addressing these specific traits requires tailored architectural choices, training methodologies, and problem-solving approaches, particularly for issues like doom-looping.
- Building your own software factory — Eric Zakariasson, Cursor
This talk explores the concept of building a "software factory" using AI agents, moving beyond simple code completion to autonomous systems that can handle complex development tasks. The core idea is to borrow principles from physical factories, such as assembly lines and robust infrastructure, and apply them to software development to increase throughput and consistency. While fully autonomous factories are still aspirational, significant progress can be made by implementing structured primitives, guardrails, and enablers for AI agents.
- Why building eval platforms is hard — Phil Hetzel, Braintrust
Building effective evaluation (eval) platforms for AI agents is a complex systems problem, not just a UI challenge. While starting with simple tools like spreadsheets is a valid first step, maturing requires moving towards more robust solutions that facilitate experimentation and integrate with production data. The core difficulty lies in managing the unique characteristics of AI agent traces, which are often large, unstructured, and high-velocity, demanding specialized data infrastructure.
- One Login to Rule Them All: Cross-App Access for MCP — Garrett Galow, WorkOS
This talk addresses the friction of repeated authentication for Multi-Cloud Platform (MCP) tools, where users must consent to access for each tool individually. It introduces Cross-App Access (XAA), a solution that leverages an Identity Provider (IDP) to act as a trust intermediary. XAA enables MCP clients to obtain credentials for MCP servers without manual user intervention after an initial single sign-on, streamlining the user experience and improving security management for IT.
- Gemma 4 Deep Dive — Cassidy Hardin, Researcher, Google DeepMind
This talk provides a deep dive into Gemma 4, Google DeepMind's latest family of open-source AI models. The presentation highlights significant architectural improvements and new capabilities, including enhanced multimodal support and optimized performance for both on-device and complex reasoning tasks. Gemma 4 aims to set a new standard for small, open-source models, making advanced AI more accessible to developers.
- Scaling GitHub for your Agents — Sam Morrow, GitHub
This talk discusses the challenges and solutions encountered by GitHub in building and scaling their MCP (Multi-modal Communication Protocol) server for AI agents. The core thesis revolves around managing the complexity of tool integration, optimizing context window usage, and enhancing security in agentic systems. The presentation highlights how GitHub has iterated on its server architecture and features to improve agent performance and user experience.
- Gateways are All You Need — Karan Sampath, Anthropic
This talk argues that gateways are the essential infrastructure for enterprises adopting Multi-Model Communication Protocol (MCP) tools. While MCPs offer significant potential, enterprises face challenges with observability, access control, and security. Gateways act as a crucial middle layer, abstracting these complexities away from individual MCP server developers and establishing a root of trust for secure and scalable agent deployments.
- Collaborative AI Engineering: One Dev, Two Dozen Agents, Zero Alignment — Maggie Appleton, GitHub
The talk argues that the current trend of scaling individual developer productivity with AI agents overlooks the inherently collaborative nature of software development. While AI can rapidly implement code, the critical bottleneck has shifted to aligning on *what* to build. Existing tools are ill-equipped for this new paradigm, leading to wasted effort and misaligned projects. The proposed solution involves creating shared, intelligent environments where teams can collaborate on planning and development alongside AI agents.
- MCP = Mega Context Problem - Matt Carey
This talk addresses the challenge of providing AI agents with access to a vast number of APIs, a problem termed the Mega Context Problem (MCP). Traditional methods like bundling all tools into an agent's context window lead to an explosion of tokens, rendering them unusable. The speaker proposes that instead of dumping tools into context, agents should be empowered to write code against APIs, leveraging typed SDKs and secure execution environments.
- AgentCraft: Putting the Orc in Orchestration — Ido Salomon
Agent Craft is an orchestrator designed to enhance human-agent collaboration by drawing parallels from gaming mechanics. The core thesis is that while individual agents are powerful, human orchestration is the current bottleneck. Agent Craft aims to raise this ceiling by providing better visibility, enabling greater agent autonomy, and facilitating seamless human-to-agent collaboration, ultimately transforming how developers work with AI agents.
- Full Walkthrough: Workflow for AI Coding — Matt Pocock
This talk explores a structured workflow for AI coding, emphasizing that traditional software engineering fundamentals remain crucial when working with AI. The core thesis is that by understanding LLM constraints—specifically their "smart zone" and tendency to forget—developers can build effective AI-powered workflows. The approach focuses on breaking down large tasks into manageable chunks that fit within the LLM's optimal performance window, ensuring better decision-making and code quality.
- What Do Models Still Suck At? - Peter Gostev, Arena.ai, BullshitBench
This talk challenges the perception that AI models are rapidly approaching general intelligence, as suggested by steadily increasing benchmark scores. It argues that despite impressive progress, models still exhibit significant weaknesses, particularly in handling nonsensical or complex, real-world tasks. The presentation introduces a benchmark focused on nonsensical questions and analyzes user dissatisfaction data to reveal areas where models continue to struggle.
- \"Software Fundamentals Matter More Than Ever\" — Matt Pocock
This talk argues that fundamental software engineering principles are more critical than ever in the age of AI. The speaker critiques the "specs-to-code" movement, which suggests generating code from specifications using AI, asserting that this approach often leads to deteriorating code quality. Instead, the talk emphasizes that well-structured, understandable, and maintainable codebases are essential for effectively leveraging AI's capabilities and that AI should be viewed as a tactical tool requiring strategic human oversight.
- The End of Apps — Kitze, Sizzy.co
This talk explores the evolution of personal productivity tools and agents, moving from early to-do lists and text files to sophisticated AI assistants. The core thesis suggests that traditional applications are becoming obsolete, replaced by integrated "life OS" systems and AI agents that proactively manage tasks and information. The future likely involves a shift where AI prompts the user, rather than the other way around, fundamentally changing how we interact with computers.
- Agents need more than a chat - Jacob Lauritzen, CTO Legora
This talk argues that complex AI agents require more than simple chat interfaces for effective collaboration. The core thesis is that as AI agents become more capable of performing complex, end-to-end tasks, the bottlenecks shift from the execution of the work itself to the planning and review stages. To address this, a more sophisticated approach to agent-human interaction is needed, focusing on control and trust within specialized workflows.
- Building Generative Image & Video models at Scale - Sander Dieleman, Google DeepMind
This talk provides a behind-the-scenes look at training generative image and video diffusion models at scale. It covers essential aspects from data curation and representation to modeling, architecture, training, sampling, and control mechanisms. The core thesis emphasizes that while diffusion models are powerful, their effectiveness relies heavily on careful data handling, efficient latent representations, and sophisticated sampling techniques like guidance.
- How AI is changing Software Engineering: A Conversation with Gergely Orosz, @The Pragmatic Engineer
This talk explores the evolving landscape of software engineering driven by AI, focusing on the phenomenon of "token maxing" and its implications for developer productivity and company culture. It questions whether current AI tools are truly making engineers faster and discusses how the role of a software engineer is expanding to encompass broader responsibilities.
- Taste & Craft: A Conversation with Tuomas Artman, CTO Linear & Gergely Orosz, @The Pragmatic Engineer
This talk explores the potential pitfalls of rapid software development fueled by AI, emphasizing the importance of deliberate decision-making and high-quality craft over sheer speed. It argues that while AI tools can accelerate development, they may lead to a proliferation of mediocre or confusing products if not guided by a strong sense of taste and user focus. The conversation highlights the need for engineers to maintain a critical perspective, prioritizing user needs and product quality even as development cycles shorten.
- Running LLMs on your iPhone: 40 tok/s Gemma 4 with MLX — Adrien Grondin, Locally AI
This talk demonstrates how to run large language models, specifically Gemma 4, on an iPhone using the MLX framework. It highlights the efficiency and speed achievable with on-device AI, enabling applications to run locally without an internet connection. The presentation covers the tools and resources needed to integrate these models into iOS and macOS applications, emphasizing the growing ecosystem and ease of implementation.
- Full Workshop: Build Your Own Deep Research Agents - Louis-François Bouchard, Paul Iusztin, Samridhi
This workshop introduces a system for automating the creation of technical articles, moving beyond generic AI-generated content. It details the architecture and implementation of a deep research agent and a writing agent, designed to work together to produce high-quality, human-like content. The system aims to augment human writers by handling the research and initial drafting, allowing for more efficient and scalable content production.
- Gemma, DeepMind's Family of Open Models — Omar Sanseviero, Google DeepMind
Gemma is Google DeepMind's family of open models, designed to be downloadable, runnable on personal infrastructure, and fine-tunable for specific use cases. The latest release, Gemma 4, offers models ranging from 2 to 32 billion parameters, with capabilities extending to on-device agentic tasks, multimodal understanding, and multilingual support. These models are developed with a focus on developer-friendly sizes and an Apache 2.0 license for greater flexibility.
- The New Application Layer - Malte Ubl, CTO Vercel
AI engineering represents the successor to web development, poised to shape the next decade of software. Agents are emerging as a new form of software, making previously economically unviable automation tasks feasible. This shift is expanding the scope of what can be automated and increasing the demand for software engineers.
- Code Mode: Let the Code do the Talking - Sunil Pai, Cloudflare
This talk introduces "code mode," an alternative approach to AI agent interaction that moves beyond traditional JSON-based tool calling. Instead of complex back-and-forth with models to select and execute tools, code mode allows AI agents to generate and execute code directly within a controlled environment. This method leverages the inherent capabilities of code, such as looping and state management, and significantly reduces token usage and execution time, especially when dealing with extensive API surfaces.
- The Future of MCP — David Soria Parra, Anthropic
This talk explores the evolution and future of the Message Communication Protocol (MCP), emphasizing its role in enabling agents to achieve full connectivity. The core thesis is that MCP is a crucial connective tissue, allowing agents to interact with various applications and services, moving beyond simple coding tasks to handle complex knowledge work. The future of agent development in 2026 hinges on seamless integration of MCP with other tools like CLIs and skills to achieve robust functionality.
- How Google DeepMind is researching the next Frontier of AI for Gemini — Raia Hadsell, VP of Research
This talk explores Google DeepMind's research into the next frontiers of AI, focusing on advancements beyond traditional language models. It highlights progress in creating more robust and versatile AI systems, including multimodal embedding models, advanced weather prediction systems, and interactive 3D world models. The core thesis is that significant breakthroughs lie in developing AI that can understand and interact with the world in more comprehensive and nuanced ways, moving towards more general intelligence.
- The Friction is Your Judgment — Armin Ronacher & Cristina Poncela Cubeiro, Earendil
This talk explores the challenges and potential pitfalls of using AI agents in software development, arguing that while these tools offer significant productivity gains, they can also introduce subtle forms of friction that undermine code quality and engineering judgment. The core thesis is that intentionally reintroducing certain types of friction is crucial for maintaining control, ensuring quality, and leveraging human experience in an agent-assisted development workflow.
- State of the Claw — Peter Steinberger
Peter Steinberger discusses the rapid growth and current state of Open Claw, an open-source AI project he created. He highlights its significant adoption and contributor base, while also addressing the substantial security challenges the project faces due to its popularity and the evolving threat landscape. Steinberger emphasizes the importance of user education and responsible deployment of AI agents.
- Harness Engineering: How to Build Software When Humans Steer, Agents Execute — Ryan Lopopolo, OpenAI
This talk introduces Harness Engineering, a paradigm shift in software development where AI agents execute tasks guided by human direction. The core thesis is that implementation is no longer the bottleneck; code is abundant and cheap to produce. The focus shifts to systems thinking, design, and delegation, empowering engineers to leverage AI agents for complex, long-horizon work. This approach aims to maximize productivity by automating implementation and freeing human engineers for higher-leverage activities.
- Building pi in a World of Slop — Mario Zechner
This talk critiques the current state of AI coding agents and development tools, arguing that many are overly complex, lack transparency, and introduce "booboos" (errors) that compound over time. The speaker advocates for a shift towards more malleable, self-modifying agents and emphasizes the importance of human oversight, discipline, and understanding in the development process. The core thesis is that developers should regain control of their tools and workflows, prioritizing essential features and human understanding over unchecked agent-driven development.
- $1 AI Guardrails: The Unreasonable Effectiveness of Finetuned ModernBERTs – Diego Carpentero
This talk addresses the escalating sophistication of AI attacks, moving beyond simple prompt injection to complex exploits targeting LLM interfaces, data, and even internal mechanisms. It proposes a practical, low-latency, self-hosted defensive layer built by fine-tuning a modern encoder model, specifically ModernBERT, to act as a guardrail against these threats. The approach emphasizes efficiency and cost-effectiveness, aiming for a defense under a dollar per instance.
- Paperclip: Open Source Human Control Plane for AI Labor — Dotta Bippa
Paperclip is an open-source human control plane for AI labor, designed to orchestrate AI agents for accomplishing real-world tasks. It allows users to set up an organizational chart of agents, manage their interactions, and invoke specific preferences to complete work. The system emphasizes user involvement from high-level design to execution, enabling the creation of businesses that can largely run themselves with human oversight.
- One Registry to Rule them All - Sonny Merla, Mauro Luchetti, & Mattia Redaelli, Quantyca
This talk addresses the chaos arising from multiple teams building AI agents independently, leading to duplicated efforts in security, infrastructure, and deployment. Amplifon's "Amplify" program established a unified approach through an enterprise-grade registry system for MCP (Machine Communication Protocol) and A2A (Agent-to-Agent) agents. This system aims to standardize development, ensure governance, and provide traceability for AI solutions across the organization.
- Judge the Judge: Building LLM Evaluators That Actually Work with GEPA — Mahmoud Mabrouk, Agenta AI
This talk introduces GEPA, a method for building calibrated LLM evaluators that align with human annotations. The core idea is to move beyond generic LLM judges, which often fail in production, by optimizing prompts using algorithms like GEPA. This approach aims to accelerate development cycles by providing reliable signals for both offline evaluations and online monitoring, ultimately contributing to the creation of a data flywheel for continuous AI improvement.
- AI Didn’t Kill the Web, It Moved in! — Olivier Leplus (AWS) & Yohan Lasorsa (Microsoft)
This talk explores how AI is transforming web development, moving beyond simple coding assistance to integrate into every stage of the web application lifecycle. It highlights how AI can aid in debugging, performance tuning, and even how web applications need to adapt to be consumable by AI agents. The presenters emphasize that mastering AI tools and understanding new integration patterns are becoming essential skills for web developers.
- Running LLMs locally: Practical LLM Performance on DGX Spark — Mozhgan Kabiri chimeh, NVIDIA
This talk explores the practical performance of running large language models (LLMs) locally on the NVIDIA Jetson Spark, a system designed for AI development. It addresses challenges like memory limitations and software stack access that often push AI workloads to the cloud. The presentation emphasizes how local solutions can enhance developer productivity by bringing AI development closer to the user, enabling faster iteration and addressing concerns like cost predictability and data residency.
- Contact Center Voice AI: Low-Latency Intelligence Extraction from Messy Audio Streams — Dippu Singh
This talk addresses the significant engineering challenge of extracting actionable business intelligence from messy, low-latency audio streams in contact centers. It proposes a four-stage pipeline to transform raw audio into structured data, aiming to reduce after-call work (ACW) and improve operational efficiency. The core idea is to leverage generative AI to automate summarization and data extraction, thereby reducing operator stress and enhancing customer experience analysis.
- OpenRAG: An open-source stack for RAG — Phil Nash
This talk introduces OpenRAG, an open-source stack designed to simplify the creation of powerful and customizable Retrieval-Augmented Generation (RAG) systems. It addresses the complexity often encountered in RAG pipelines, from document ingestion and chunking to embedding and search, by integrating existing open-source tools into a cohesive framework. OpenRAG aims to provide a robust baseline that developers can easily extend to meet specific project requirements.
- From Chaos to Choreography: Multi-Agent Orchestration Patterns That Actually Work — Sandipan Bhaumik
This talk addresses the challenges of scaling multi-agent AI systems beyond a single agent, emphasizing that complexity grows exponentially, turning simple AI problems into distributed systems challenges. It introduces key patterns for coordination, state management, and failure recovery, arguing that robust system design, not just AI capabilities, is crucial for production-grade multi-agent applications. The core thesis is that adopting established distributed systems principles is essential for building reliable and valuable multi-agent systems.
- Cognitive Exhaust Fumes, or: Read-Only AI Is Underrated — Šimon Podhajský, Head of AI, Waypoint
This talk introduces the concept of "cognitive exhaust fumes" – the byproduct of digital activity that, when analyzed across multiple sources, can reveal insights into an individual's cognition, attention, and relationships. The speaker advocates for read-only AI systems, termed "observers," which analyze this exhaust without writing back to the original data sources. This approach is presented as a safer alternative to action-oriented AI agents, mitigating risks and potentially yielding more authentic self-analysis.
- Platforms for Humans and Machines: Engineering for the Age of Agents — Juan Herreros Elorza
This talk focuses on engineering platforms to effectively support both human developers and AI agents. It argues that the challenges developers face in deploying applications and managing infrastructure become more pronounced and limiting when AI agents are involved. The core thesis is that by adopting specific platform engineering best practices, organizations can unlock greater productivity and enable AI agents to function more autonomously and effectively.
- Why, and how you need to sandbox AI-Generated Code? — Harshil Agrawal, Cloudflare
This talk addresses the critical need to sandbox AI-generated code, emphasizing that such code should be treated as untrusted. The core thesis is that while AI can accelerate development, running its output without proper security measures is akin to executing code from an anonymous internet source, posing significant risks. The presentation advocates for applying established sandboxing principles, particularly capability-based security, to mitigate these threats.
- Your Insecure MCP Server Won't Survive Production — Tun Shwe, Lenses
This talk addresses the critical security and design considerations for Multi-modal Communication Protocol (MCP) servers intended for production environments. It argues that insecure or poorly designed MCP servers are vulnerable to exploitation by agentic AI systems. The presentation emphasizes that robust security and effective design are intertwined, advocating for a product engineering mindset when building interfaces for AI agents.
- Let LLMs Wander: Engineering RL Environments — Stefano Fiorucci
This talk explores engineering reinforcement learning (RL) environments for language models, enabling them to learn through interaction, exploration, and feedback. It highlights how these environments serve as crucial training grounds for LLM agents, allowing them to develop skills in tool use, code execution, and complex task solving. The presentation introduces Verifiers, an open-source library for building such environments, and demonstrates their application through an experiment transforming a basic tic-tac-toe playing model into a master.
- Bending a Public MCP Server Without Breaking It — Nimrod Hauser, Baz
This talk explores strategies for effectively integrating and managing third-party tools within AI agent workflows, specifically focusing on Multi-Call Protocol (MCP) servers. It addresses the common challenges of generic tool descriptions, unexpected agent behavior, and potential performance degradation or security risks when using off-the-shelf tools. The presentation outlines a framework of five best practices to tailor these tools for specific use cases, ensuring agents perform as intended and applications remain stable.
- Agentic Engineering: Working With AI, Not Just Using It — Brendan O'Leary
This talk introduces the concept of agentic engineering, shifting the paradigm from merely using AI tools to actively collaborating with them. It highlights the evolution of AI in software development from simple line completion to sophisticated agents capable of executing complex tasks. The core thesis is that understanding and managing AI agents as collaborators, akin to junior developers, is crucial for maximizing their potential and navigating the complexities of modern software development.
- How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR
This talk explores the challenges and potential of measuring AI capabilities, particularly focusing on developer productivity and long-term task completion. It questions the extrapolation of current AI progress trends, suggesting that physical and economic constraints, alongside potential technological breakthroughs, could alter the trajectory of AI development. The discussion also delves into the complexities of evaluating AI in real-world scenarios beyond controlled benchmarks, highlighting the gap between AI capabilities and practical application in fields like software engineering and data science.
- Identity for AI Agents - Patrick Riley & Carlos Galan, Auth0
This talk addresses the critical need for robust identity and authorization mechanisms for AI agents. It introduces new features and concepts designed to enable agents to securely interact with resources and perform actions on behalf of users. The core thesis is that as agents become more autonomous and capable, establishing clear identity, managing permissions, and ensuring user consent are paramount for safe and effective integration into various workflows.
- OpenAI + @Temporalio : Building Durable, Production Ready Agents - Cornelia Davis, Temporal
This talk explores building durable, production-ready AI agents by integrating the OpenAI Agents SDK with Temporal. It highlights how Temporal's distributed systems capabilities can provide essential durability, visibility, and scalability to agentic applications, which are often inherently stateful and prone to failures. The presentation demonstrates how to leverage Temporal's workflow and activity abstractions to manage agent execution, tool invocation, and error handling, ensuring that agent processes can reliably resume after interruptions.
- Your MCP Server is Bad (and you should feel bad) - Jeremiah Lowin, Prefect
This talk argues that MCP servers, which act as interfaces for AI agents, should be designed with agents' unique strengths and weaknesses in mind, rather than simply mirroring human-facing APIs. The core thesis is that agents have different needs regarding discovery, iteration, and context compared to humans, and MCP server design must account for these differences to create effective and usable interfaces.
- Spec-Driven Development: Agentic Coding at FAANG Scale and Quality — Al Harris, Amazon Kiro
This talk introduces Spec-Driven Development as a method to enhance AI agent development, focusing on improving control, code quality, and reliability. It proposes a structured workflow that compresses the software development lifecycle, moving from prompt to requirements, design, and execution, with an emphasis on creating reproducible results. The approach aims to integrate established software engineering practices with AI capabilities to tackle more complex problems.
- DSPy: The End of Prompt Engineering - Kevin Madura, AlixPartners
This talk introduces DSPy, a framework designed to streamline the development of applications that leverage large language models (LLMs). It proposes a shift from traditional prompt engineering to a more programmatic approach, treating LLMs as first-class citizens within Python programs. DSPy aims to simplify the creation of modular, optimizable, and maintainable AI systems by providing a structured way to define intents, manage logic, and integrate LLM calls.
- Automating Large Scale Refactors with Parallel Agents - Robert Brennan, OpenHands
This talk explores automating large-scale software refactoring and maintenance tasks using parallel AI agents. The core idea is that while individual agents are effective for smaller tasks, complex, multi-faceted operations like dependency updates, code modernization, or vulnerability remediation require orchestrating multiple agents working in parallel. This approach aims to tackle significant technical debt and improve developer productivity by automating laborious processes that are currently too large or complex for a single agent.
- Build a Prompt Learning Loop - SallyAnn DeLucia & Fuad Ali, Arize
This talk introduces prompt learning, a method for optimizing AI agent prompts by creating a feedback loop that incorporates human and LLM-generated evaluations. The core idea is to move beyond static prompts and leverage detailed feedback, including explanations for failures, to continuously improve agent performance. This approach aims to enhance agent reliability and effectiveness without requiring extensive fine-tuning or architectural changes.
- Building durable Agents with Workflow DevKit & AI SDK - Peter Wielander, Vercel
This talk introduces the Workflow DevKit and AI SDK from Vercel, designed to simplify the process of building durable AI agents. The core thesis is that by adopting a workflow pattern, developers can abstract away the complexities of production deployment, error handling, and observability, allowing them to focus on agent capabilities. The DevKit enables agents to be resumable, reliable, and easily integrated with human-in-the-loop processes.
- Claude Agent SDK [Full Workshop] — Thariq Shihipar, Anthropic
This talk introduces the Claude Agent SDK, a framework designed to simplify the creation of autonomous AI agents. The core thesis is that by leveraging Unix primitives like bash and the file system, developers can build more powerful and flexible agents. The SDK aims to encapsulate best practices learned from building agents like Claude Code, offering an opinionated yet effective approach to agent development.
- Welcome to AIE CODE - Jed Borovik, Google DeepMind
This talk introduces the AI Engineering Code Summit, emphasizing AI coding as the most critical problem in technology today. It highlights the event's focus on bringing together experts to advance the AI coding industry across all companies, distinguishing it as a single-track summit dedicated to this theme. The summit aims to explore the patterns, systems, and products enabling AI's transformation of software development.
- Building Intelligent Research Agents with Manus - Ivan Leo, Manus AI (now Meta Superintelligence)
This talk introduces Manus, an AI action engine designed to go beyond providing answers and instead execute tasks, automate workflows, and extend human capabilities. The presentation highlights the platform's evolution, including Manus 1.5, which offers increased speed and quality. Manus aims to be a general AI agent accessible through various interfaces like web applications, Slack, and a newly launched API, meeting users wherever they are.
- Jack Morris: Stuffing Context is not Memory, Updating Weights is
This talk challenges the conventional approaches of stuffing context or using Retrieval Augmented Generation (RAG) for AI knowledge integration. The speaker argues that these methods are fundamentally limited by cost, speed, and reasoning capabilities. Instead, the focus shifts to training information directly into model weights, proposing this as a more effective and scalable solution for building AI systems that truly know and can reason with specific data.
- AGI: The Path Forward – Jason Warner & Eiso Kant, Poolside
Poolside is developing AI models from scratch, focusing on pairing next-token prediction with reinforcement learning to bridge the gap between models and human intelligence. Their second-generation model, Malibu agent, is designed for high-consequence code environments, demonstrating capabilities in code conversion, testing, and feature implementation. The company emphasizes building foundational intelligence and scaling it, with plans for public API access and integration with platforms like Amazon Bedrock.
- Shipping AI That Works: An Evaluation Framework for PMs – Aman Khan, Arize
This talk introduces an evaluation framework for AI product managers (PMs) focused on shipping reliable AI applications. It emphasizes the critical need for robust evaluation, analogous to software testing but adapted for the non-deterministic nature of AI models. The framework aims to provide PMs with tools and methodologies to build confidence in AI product performance, moving beyond subjective "vibe coding" to data-driven "thrive coding."
- How Claude Code Works - Jared Zoneraich, PromptLayer
This talk explores the advancements in coding agents, particularly focusing on how Claude Code works and the underlying innovations that have made these tools effective. The core thesis is that simplicity in architecture, coupled with better underlying models and a focus on tool calling, has been the key breakthrough. The presentation emphasizes a philosophy of "give it tools and get out of the way," advocating for leaning into the model's capabilities rather than over-engineering complex systems.
- Why Agent Hype can fall short of reality – Joel Becker, METR
This talk addresses the discrepancy between AI capabilities suggested by benchmarks and real-world performance, particularly in developer productivity. It introduces two distinct methods of evaluation: benchmark-style assessments measuring AI performance on diverse tasks against human baselines, and field experiments examining AI's impact on experienced developers in complex, real-world coding environments. The core thesis is that while benchmarks show rapid AI advancement, practical application, especially in messy, high-context scenarios, reveals a more nuanced and sometimes even negative impact on productivity.
- Small Bets, Big Impact Building GenBI at a Fortune 100 – Asaf Bord, Northwestern Mutual
This talk discusses the development of GenBI, a Generative AI-powered Business Intelligence agent, at Northwestern Mutual. The core thesis is that by taking small, incremental bets and focusing on building trust through transparency and gradual delivery, large organizations can successfully innovate with GenAI despite inherent risk aversion. The approach emphasizes using real, messy data from the outset to bridge the gap between prototypes and production.
- Developer Experience in the Age of AI Coding Agents – Max Kanat-Alexander, Capital One
This talk explores how to optimize developer experience in the era of AI coding agents. It argues that investments in foundational aspects of software development, such as standardized environments, robust validation, and clear documentation, will benefit both human developers and AI agents. The core thesis is that practices good for human developers are also good for AI, ensuring long-term value regardless of AI advancements.
- The Unreasonable Effectiveness of Prompt Learning – Aparna Dhinakaran, Arize
This talk explores prompt learning as a method to improve coding agents, contrasting it with traditional Reinforcement Learning (RL). Prompt learning leverages English feedback on agent outputs to iteratively refine system prompts, offering a potentially more efficient and data-light approach for building agents compared to RL's reliance on scalar rewards and extensive data. The core idea is to use LLM-based evaluations to generate actionable feedback that directly informs prompt adjustments.
- Amp Code: Next Generation AI Coding – Beyang Liu, Amp Code
AMP Code is an opinionated frontier AI coding agent designed to help developers navigate the rapidly changing landscape of AI-assisted software development. It aims to embrace the awe and absurdity of agents writing significant amounts of code, positioning itself as a research lab exploring the future of AI in coding. AMP offers both terminal and editor interfaces, focusing on providing the right amount of information without overwhelming the user and streamlining the code review process.
- Making Codebases Agent Ready – Eno Reyes, Factory AI
This talk argues that making codebases "agent-ready" is crucial for unlocking the full potential of AI in software engineering. The core thesis is that robust, automated validation and verification systems within a codebase are the primary enablers for successful and scalable AI agent adoption, rather than solely focusing on the capabilities of the AI tools themselves. By enhancing these internal validation mechanisms, organizations can significantly increase development velocity and the effectiveness of AI-powered workflows.
- The 3 Pillars of Autonomy – Michele Catasta, Replit
This talk explores the concept of autonomy in AI coding agents, particularly for non-technical users. It argues that true autonomy for this audience means offloading all technical decision-making, allowing users to focus solely on their desired outcomes. The presentation redefines autonomy beyond just long runtimes, emphasizing the importance of scoped tasks and robust verification to build powerful and stable agents.
- No More Slop – swyx
This talk declares war on "slop," defined as low-quality, inauthentic, or inaccurate content, whether generated by humans or AI. The speaker argues that while AI has advanced rapidly, there's a growing problem of low-taste output that dilutes genuine progress. The core thesis is that the AI engineering community must actively combat this trend by prioritizing quality, taste, and accountability in AI-generated content and code.
- The Infinite Software Crisis – Jake Nations, Netflix
This talk addresses the accelerating trend of developers shipping code they don't fully understand, a phenomenon amplified by AI code generation tools. It posits that this issue stems from a historical pattern of software complexity outpacing human comprehension, exacerbated by a modern confusion between what is simple and what is merely easy. The core argument is that while AI can make coding easier, it doesn't inherently make it simpler, and without intentional effort to maintain understanding, this leads to an unsustainable accumulation of complexity.
- From Arc to Dia: Lessons learned building AI Browsers – Samir Mody, The Browser Company of New York
This talk details the journey of The Browser Company in evolving from their initial browser, Arc, to their AI-native browser, Dia. It emphasizes the critical lessons learned in building AI-powered products, focusing on optimizing iteration speed, treating model behavior as a craft, and integrating AI security as a fundamental aspect of product development. The core thesis is that embracing technological shifts with conviction is essential for company-wide evolution, not just product updates.
- Leadership in AI Assisted Engineering – Justin Reock, DX (acq. Atlassian)
This talk explores the current impact and future potential of AI-assisted engineering, emphasizing that while AI can augment developers, its effective integration requires careful consideration of productivity metrics, organizational culture, and strategic implementation across the software development lifecycle. The core thesis is that AI's true value lies in enhancing, not replacing, engineers, and achieving positive outcomes depends on addressing variability in adoption and impact through education, clear communication, and a focus on psychological safety.
- Paying Engineers like Salespeople – Arman Hezarkhani, Tenex
This talk proposes a shift in how engineers are compensated, moving away from traditional hourly or salary models towards a system that directly incentivizes productivity and the adoption of AI tools. The core idea is to pay engineers based on completed work units, similar to how salespeople are compensated based on sales, to encourage faster and more efficient delivery of high-quality code. This approach aims to unlock greater potential by aligning compensation with the increased capabilities offered by AI.
- Welcome to AIE LEAD - Alex Lieberman, Tenex
This talk, delivered by Alex Lieberman at the AI Engineer Code Summit 2025, serves as an introduction to the event. Lieberman, co-founder of 10x.co, an AI transformation firm, frames the summit as a look back at the year's advancements in AI and a tactical view of future directions. The event aims to bring together diverse perspectives from leading AI labs, startups, academics, consultants, and major brands.
- Dispatch from the Future: building an AI-native Company – Dan Shipper, Every, AI & I
This talk explores the emerging playbook for building AI-native companies, emphasizing that the process is currently being invented collaboratively. The core thesis is that achieving 100% AI adoption among engineers creates a significant, tenfold difference in organizational capability. This shift enables a single engineer to build and maintain complex production products, fundamentally altering how software development is approached.
- AI Consulting in Practice – NLW, Superintelligent, @AIDailyBrief
This talk explores the current state of enterprise AI adoption, moving beyond the narrative of an AI bubble to focus on tangible value and return on investment (ROI). It presents findings from a study collecting self-reported ROI data across various AI use cases, highlighting that organizations are indeed finding value, with a significant portion reporting modest to high ROI. The discussion emphasizes the growing adoption of AI, particularly in software engineering, and the increasing deployment of agents within enterprises, despite challenges in scaling beyond pilot phases.
- Code World Model: Building World Models for Computation – Jacob Kahn, FAIR Meta
The Code World Model (CWM) project aims to build AI models capable of reasoning, planning, and decision-making, using code as a constrained environment for developing these capabilities. Instead of solely focusing on code syntax, CWM models program execution more explicitly by predicting program states and transitions. This approach seeks to bridge the gap between traditional world models and large language models, enabling more efficient agentic reasoning through simulated execution.
- AI Kernel Generation: What's working, what's not, what's next – Natalie Serrino, Gimlet Labs
This talk explores the use of AI for generating and optimizing low-level computational kernels, which are crucial for the performance of complex AI workloads. The core idea is to leverage AI to automatically port and optimize these kernels for diverse hardware, addressing the scarcity of human experts in this specialized field. While AI shows promise in tasks like kernel fusion and translating optimizations, it's not a replacement for human expertise in highly complex algorithmic advancements.
- Your Support Team Should Ship Code – Lisa Orr, Zapier
This talk details Zapier's initiative to empower their support team to ship code, addressing the constant challenge of "app erosion" caused by third-party API changes. By enabling support to fix bugs directly, Zapier aims to improve integration reliability and customer experience. The project, named Scout, evolved from initial experiments in enabling support for bug fixing and exploring AI-assisted code generation to a comprehensive agent-based system.
- What We Learned Deploying AI within Bloomberg’s Engineering Organization – Lei Zhang, Bloomberg
This talk discusses Bloomberg's experience integrating AI into its engineering organization, focusing on practical applications and lessons learned. The core thesis is that AI tooling can significantly alter the cost function of software engineering, enabling a re-evaluation of fundamental principles and a shift towards higher-quality software development. The organization aimed to leverage AI to improve developer productivity and system reliability across its vast codebase and numerous functions.
- Building in the Gemini Era – Kat Kampf & Ammaar Reshi, Google DeepMind
This talk introduces Google DeepMind's latest advancements, Gemini 3 Pro and Nano Banana Pro, highlighting their capabilities in building AI-powered applications. The core thesis is that these models empower anyone to build software by simplifying complex tasks through intuitive interfaces and powerful AI features, enabling a new generation of engineers to create at an unprecedented scale.
- Coding Evals: From Code Snippets to Codebases – Naman Jain, Cursor
This talk explores the evolution of evaluating AI models for coding tasks, from simple code snippets to complex codebases. It highlights challenges like data contamination and brittle test suites, proposing dynamic evaluation sets and LLM-based judges to ensure reliable and relevant assessments as AI capabilities advance. The discussion covers various stages of coding evaluation, emphasizing the need for benchmarks that reflect real-world performance and adapt to the rapid progress in AI.
- From Vibe Coding To Vibe Engineering – Kitze, Sizzy
This talk explores the evolution of front-end development, contrasting traditional coding practices with emerging AI-assisted workflows. It introduces the concept of "vibe engineering" as a more sophisticated approach to using AI agents for coding, emphasizing the need for skilled engineers to guide and refine AI-generated code. The core thesis is that while AI can accelerate development, human expertise remains crucial for quality, abstraction, and complex problem-solving.
- Minimax M2: Building the #1 Open Model – Olive Song, MiniMax
This talk introduces Minimax M2, an open-weight, 10-billion parameter model specifically designed for coding and agentic tasks. It emphasizes the model's cost-efficiency and strong performance on both intelligence and agent benchmarks, aiming to provide developers with a practical and efficient tool. The presentation details the training methodologies and characteristics that contribute to M2's effectiveness in coding, long-horizon tasks, generalization, and multi-agent scalability.
- Proactive Agents – Kath Korevec, Google Labs
This talk introduces the concept of proactive AI agents, moving beyond reactive tools that require explicit user commands. The core thesis is that AI agents should act as collaborators, anticipating developer needs and handling tasks autonomously to reduce mental load and increase productivity. This shift aims to free developers to focus on creative and high-level problem-solving rather than managing agent workflows.
- Moving away from Agile: What's Next – Martin Harrysson & Natasha Maniar, McKinsey & Company
This talk argues that the rapid advancements in AI necessitate a fundamental shift in software development operating models, moving beyond traditional Agile methodologies. While AI tools offer significant individual productivity gains, realizing broader organizational value requires addressing bottlenecks in collaboration, code review, and work allocation. The presenters propose a transition to AI-native workflows and roles, emphasizing smaller teams, continuous planning, and spec-driven development to achieve substantial improvements in delivery speed and quality.
- Hard Won Lessons from Building Effective AI Coding Agents – Nik Pash, Cline
This talk argues that the effectiveness of AI coding agents is primarily determined by the underlying model's capability, not by complex engineering scaffolds. Frontier models, when unhindered, can outperform many agent combinations. The core message is to simplify agent engineering and focus on improving model training through robust benchmarks and reinforcement learning environments derived from real-world coding tasks.
- The State of AI Code Quality: Hype vs Reality — Itamar Friedman, Qodo
This talk addresses the gap between the hype surrounding AI code generation and the reality of its impact on software quality. While AI tools can significantly boost productivity in writing code, they introduce new challenges related to quality assurance, security, and maintainability. The core thesis is that achieving substantial, reliable gains requires moving beyond basic code generation to implement robust, agentic quality workflows throughout the entire software development lifecycle.
- Can you prove AI ROI in Software Eng? (Stanford 120k Devs Study) – Yegor Denisov-Blanch, Stanford
This talk explores the impact of AI tools on software engineering productivity, questioning whether current enterprise adoption is driven by genuine ROI or hype. Research spanning two years and multiple companies uses a machine learning model, trained on expert evaluations of code commits, to measure productivity gains. Initial findings suggest a median productivity increase of around 10% for AI-using teams, but with a widening gap between top and bottom performers, highlighting the need for companies to understand their position in this trend.
- Agent Reinforcement Fine Tuning – Will Hang & Cathy Zhou, OpenAI
This talk introduces Agent Reinforcement Fine-Tuning (Agent RFT), a method for enhancing the performance of AI agents that interact with the outside world through tools. Agent RFT modifies model weights based on a specified learning signal, teaching the agent to distinguish between good and bad behavior. This process allows agents to explore various tool-calling strategies to solve tasks more effectively, leading to improved reasoning, lower latency, and better adaptation to specific business contexts.
- RL Environments at Scale – Will Brown, Prime Intellect
This talk explores scaling AI research not just through increased data and compute, but by making the practice of AI research more accessible. It introduces the concept of "environments" as a key abstraction for experimentation, akin to web applications for AI research, enabling broader participation beyond large labs. The discussion highlights how these environments facilitate model customization, evaluation, and the development of more effective AI systems.
- Efficient Reinforcement Learning – Rhythm Garg & Linden Li, Applied Compute
This talk introduces an efficient reinforcement learning (RL) approach for training large language models, focusing on making the process faster and cheaper for enterprise use cases. The core idea is to break the synchronous dependency between sampling data and training the model, enabling asynchronous operations. This method aims to improve ROI by moving AI beyond simple productivity tasks into real-world automations.
- Don't Build Agents, Build Skills Instead – Barry Zhang & Mahesh Murag, Anthropic
This talk proposes a shift from building standalone AI agents to developing modular "skills" that agents can utilize. The core argument is that while agents possess intelligence, they often lack the specific domain expertise required for real-world tasks. Skills, defined as organized collections of files containing procedural knowledge and tools, offer a more composable and expertise-driven approach to enhancing agent capabilities. This paradigm aims to make AI more accessible and adaptable by packaging expertise into shareable units.
- 2026: The Year The IDE Died — Steve Yegge & Gene Kim, Authors, Vibe Coding
The talk posits that the traditional Integrated Development Environment (IDE) is becoming obsolete, predicting that by 2026, most code will be generated by advanced AI systems overseen by engineers. This shift moves from current tools like code assistants, which are seen as powerful but potentially dangerous like power tools, towards more precise, automated systems akin to CNC machines. The core argument is that AI will fundamentally reshape technology organizations and the economy, enabling unprecedented productivity and new ways of building software.
- VoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWS
This talk introduces VoiceVision RAG, a system that integrates visual document intelligence with voice response capabilities. It explores multimodal RAG techniques, focusing on vision-based retrieval using the Pali model. The presentation demonstrates how to build agentic applications with the Strands framework, enabling voice-enabled retrieval and response generation from visual documents.
- Government Agents: AI Agents Meet Tough Regulations — Mark Myshatyn, Los Alamos National Lab
This talk explores the application of AI agents within a national laboratory setting, emphasizing their potential to accelerate scientific discovery and address complex national security challenges. It highlights the evolution from early computational methods to modern agentic systems, stressing the need for these tools to operate within stringent regulatory and security frameworks. The core thesis is that AI agents can significantly enhance scientific and operational capabilities, provided they are developed with explainability, isolation, governance, and speed in mind, fostering crucial partnerships between government entities and the broader tech community.
- Future-Proof Coding Agents – Bill Chen & Brian Fioca, OpenAI
This talk explores the architecture and development of coding agents, emphasizing the crucial role of the "harness" – the interface layer that connects models to users and tools. It highlights the rapid evolution of AI models and the challenges in adapting agents, proposing that building models and harnesses together, as exemplified by OpenAI's Codeex, offers a more robust and future-proof approach. The discussion also touches upon emerging patterns for integrating these agents into various products and workflows.
- Katelyn Lesse – Evolving Claude APIs for Agents, Anthropic
This talk outlines Anthropic's platform evolution to empower developers in building advanced agentic systems with Claude. The core thesis is that by providing tools to harness Claude's capabilities, manage its context window effectively, and grant it access to computational resources, developers can achieve peak performance in their AI-driven applications. The platform aims to enable agents to operate more autonomously within secure, sandboxed environments.
- No Vibes Allowed: Solving Hard Problems in Complex Codebases – Dex Horthy, HumanLayer
This talk addresses the challenges of using AI for complex software engineering tasks, particularly within large, existing codebases. The core thesis is that current AI models struggle with brownfield projects, often leading to rework and "slop." The presented solution focuses on advanced context engineering and intentional compaction to maximize the effectiveness of AI coding agents, enabling them to tackle difficult problems without introducing technical debt.
- Defying Gravity - Kevin Hou, Google DeepMind
Anti-gravity is a new AI developer platform from Google DeepMind designed with an agent-first approach. It integrates an editor, a browser, and an agent manager to provide a cohesive environment for AI-powered software development. The platform aims to leverage advancements in AI models, particularly in areas like reasoning, multimodal capabilities, and tool use, to enable more ambitious and complex agentic workflows.
- Building Cursor Composer – Lee Robinson, Cursor
Cursor Composer is an AI model designed for real-world software engineering, aiming to balance speed and intelligence. It performs better than top open-source models and rivals frontier models in intelligence, while being significantly more efficient in token generation. The development focused on creating a model that could serve as a daily driver for coding tasks, integrating features like parallel tool calling and semantic search to enhance its capabilities.
- Music from AIE Code Summit - Instrumentals
This content is a placeholder for instrumental music from the AIE Code Summit livestream and venue stage. It is provided as a YouTube link for community members who enjoyed the background music during the event.
- The Unbearable Lightness of Agent Optimization — Alberto Romero, Jointly
This talk introduces Meta AC, a novel framework designed to optimize AI agents by orchestrating multiple adaptation strategies beyond single-dimensional approaches. It addresses limitations in existing context engineering methods by employing a meta-controller that dynamically selects and allocates strategies based on task complexity, uncertainty, verifiability, and resource constraints. This multi-dimensional adaptation aims to improve agent performance, robustness, and efficiency across various tasks.
- Backlog.md: Terminal Kanban Board for Managing Tasks with AI Agents — Alex Gavrilescu, Funstage
BacklogMD is a terminal-based Kanban board designed for managing tasks with AI agents and humans. It addresses the common issue of AI agents going off-track or running out of context by breaking down large features into smaller, manageable Markdown tasks. This approach facilitates clearer communication, better scope definition, and a more robust review process for AI-assisted development.
- Agents are Robots Too: What Self-Driving Taught Me About Building Agents — Jesse Hu, Abundant
This talk draws parallels between building self-driving cars and developing AI agents, emphasizing that agents, like robots, operate in complex environments and require more than just sophisticated models. The core thesis is that the 99% of work beyond the AI model—encompassing infrastructure, tooling, feedback loops, and robust offline stacks—is crucial for reliable agent development and deployment.
- Vision: Zero Bugs — Johann Schleier-Smith, Temporal
This talk presents a vision for achieving zero software bugs by drawing parallels with highly reliable systems in industries like aerospace. It argues that while current AI coding agents introduce new challenges, established techniques for building robust software can be adapted and scaled, potentially making high-assurance code significantly more accessible and affordable. The core thesis is that the path to widespread adoption of AI in software development hinges on solving the quality and reliability problem.
- Compilers in the Age of LLMs — Yusuf Olokoba, Muna
This talk addresses the fundamental challenges AI engineers face in deploying models, moving beyond the hype of voice agents and MCP tools. The core thesis is that building a Python compiler to generate self-contained binaries offers a simpler, more standardized approach to running AI models anywhere, from local devices to the cloud, without extensive infrastructure changes. This method aims to replicate the ease of using models like OpenAI's API but with the flexibility of any open-source model.
- Developing Taste in Coding Agents: Applied Meta Neuro-Symbolic RL — Ahmad Awais, CommandCode
This talk introduces Command Code, a coding agent designed to learn and adapt to an individual developer's unique coding style and preferences, referred to as "taste." Unlike standard AI coding assistants that provide generic code, Command Code aims to internalize a developer's decision-making process, preferences for tools, and architectural choices. This is achieved through a meta-neuro-symbolic reinforcement learning approach, creating a more personalized and efficient coding experience.
- From Stateless Nightmares to Durable Agents — Samuel Colvin, Pydantic
This talk introduces Pydantic AI's integration with durable execution frameworks like Temporal, addressing the challenge of maintaining state and preventing data loss in long-running AI agent workflows. It demonstrates how these frameworks enable agents to recover from failures and resume execution without starting from scratch, significantly improving reliability for complex tasks. The presentation also touches upon Pydantic Evals for model evaluation.
- Enterprise Deep Research: The Next Killer App for Enterprise AI — Ofer Mendelevitch, Vectara
This talk introduces enterprise deep research as a powerful application for AI agents, focusing on its ability to conduct in-depth, multi-step investigations within private enterprise data. It highlights the challenges of factual accuracy in generative AI and presents an agent operating system designed to mitigate hallucinations and ensure robust information retrieval for enterprise-grade deployments.
- What Data from 20m Pull Requests Reveal About AI Transformation — Nick Arcolano, Jellyfish
This talk analyzes data from 20 million pull requests across 200,000 developers to reveal real-world trends in AI adoption and its impact on software engineering. It highlights that while interactive AI coding tools are seeing widespread adoption and significant productivity gains, fully autonomous agents are still in early stages. The data suggests that code architecture plays a crucial role in realizing AI-driven productivity benefits.
- AI Copilots for Tech Architecture: The Highest-ROI Use Case You’re Not Building — Boris B., Catio
This talk argues that AI copilots for tech architecture represent the highest ROI use case currently underutilized by organizations. While coding copilots have become standard, an architecture co-pilot can prevent costly mistakes by ensuring development efforts are directed correctly from the outset. This approach aims to move beyond tribal knowledge and gut instinct, providing data-backed guidance for architectural decisions.
- Infra that fixes itself, thanks to coding agents — Mahmoud Abdelwahab, Railway
This talk introduces a system where an AI coding agent monitors application infrastructure for issues and automatically generates pull requests to fix them. Instead of just alerting developers to problems like memory leaks or high error rates, the system aims to detect, diagnose, and propose solutions, significantly reducing manual debugging and speeding up the resolution process. The core idea is to shift from reactive alerting to proactive, automated self-healing infrastructure.
- Context Platform Engineering to Reduce Token Anxiety — Val Bercovici, WEKA
This talk introduces context platform engineering as a method to reduce token anxiety and improve AI agent performance. The core thesis is that by optimizing how context is managed and cached, developers can significantly increase KV cache hit rates, leading to more efficient and cost-effective AI systems. The presenters announce the open-sourcing of their context platform engineering toolkit, designed to help engineers achieve these optimizations.
- Context Engineering: Connecting the Dots with Graphs — Stephen Chin, Neo4j
This talk explores context engineering as a method to enhance AI applications by moving beyond simple prompt engineering. It emphasizes the importance of providing AI models with dynamic, curated, and structured context to improve signal over noise, leading to more relevant and reliable outputs. The core thesis is that by treating AI development as information architecture, engineers can achieve superior results and gain greater control over AI behavior.
- The Cure for the Vibe Coding Hangover — Corey J. Gallon, Rexmore
This talk introduces a framework designed to combat "vibe coding," an approach to AI-accelerated development that prioritizes speed and intuition over planning and maintainability, leading to brittle, unmanageable software. The framework emphasizes treating AI engineering as a continuous learning experience, with the engineer acting as the architect and the AI as the implementer. It advocates for a deliberate, iterative process that compounds understanding and productivity, ultimately enabling the creation of robust, production-ready applications.
- Hacking Subagents Into Codex CLI — Brian John, Betterup
This talk explores a method for integrating sub-agents into the CodeX CLI, enabling users to leverage existing workflows with alternative tools. The core idea is to create a wrapper script that manages the execution of child CodeX processes as sub-agents, allowing them to perform tasks and return results to the parent session. This approach aims to overcome the limitations of being locked into a single tool or model family while retaining the benefits of sub-agent context management.
- AI changes *Nothing* — Dax Raad, OpenCode
This talk argues that despite the rapid advancements in AI, the fundamental principles of building successful products remain unchanged. The speaker contends that AI does not inherently solve the core challenges of product development, such as capturing user attention, guiding them to a product's value, and retaining them long-term. Instead, these critical aspects still require human creativity, strategic thinking, and relentless execution.
- Z.ai GLM 4.6: What We Learned From 100 Million Open Source Downloads — Yuxuan Zhang, Z.ai
This talk introduces Z.ai's GLM 4.6 model series, highlighting its advancements in both language and multimodal understanding. The series has achieved over 100 million downloads across various platforms, indicating significant community adoption and contribution. GLM 4.6 demonstrates strong performance on public benchmarks, rivaling and sometimes surpassing commercial models, particularly in coding and reasoning tasks.
- Rishabh Garg, Tesla Optimus — Challenges in High Performance Robotics Systems
This talk addresses the complexities of high-performance robotics systems, focusing on the critical interplay between control policies and the underlying software and hardware infrastructure. It highlights how issues that appear to stem from the control policy often originate in the communication protocols, threading, synchronization, logging, and priority management within the system. The presentation uses a simplified toy robot architecture to illustrate common pitfalls and debugging strategies.
- Building an Agentic Platform — Ben Kus, CTO Box
This talk explores Box's journey in building an agentic platform, focusing on how agentic approaches can solve complex data extraction challenges that traditional methods and basic LLM calls struggle with. The core thesis is that an agentic abstraction layer provides a clean, flexible, and evolvable architecture for tackling sophisticated AI tasks, moving beyond simple chatbot interactions to orchestrate complex workflows.
- Five hard earned lessons about Evals — Ankur Goyal, Braintrust
This talk emphasizes the critical role of robust evaluation (evals) in the development and deployment of AI systems. It argues that evals should be actively engineered, not passively accepted, and that they are essential for both playing offense by identifying new use cases and defense by ensuring quality. The core thesis is that by treating evals as a continuous engineering process, organizations can better adapt to rapid model advancements and ship higher-quality AI products.
- Perceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.ai
This talk explores the limitations of current AI evaluation metrics, particularly in generative media, by highlighting how they often fail to account for human perception and aesthetic judgment. The speaker argues that traditional metrics, like FID scores, can be misled by factors such as compression artifacts, leading to inaccurate assessments of AI model quality. The core thesis is that AI evaluation needs to evolve to incorporate human perceptual nuances and subjective qualities, moving beyond easily quantifiable but potentially superficial measures.
- How BlackRock Builds Custom Knowledge Apps at Scale — Vaibhav Page & Infant Vasanth, BlackRock
BlackRock has developed a framework to accelerate the creation of custom AI and knowledge applications at scale within their asset management operations. This system addresses the complexities of data extraction, LLM strategy selection, and deployment challenges, aiming to reduce app development time from months to days by empowering domain experts.
- Form factors for your new AI coworkers — Craig Wattrus, Flatfile
This talk explores new ways to interact with AI, moving beyond traditional interfaces to create more intuitive and collaborative "AI coworkers." It emphasizes understanding the capabilities and limitations of AI models as if they were physical materials, allowing for the development of novel form factors that enable AI to perform tasks more effectively and align with user goals. The core idea is to shift from simply automating tasks to fostering emergent behaviors and deeper collaboration between humans and AI.
- Fuzzing in the GenAI Era — Leonard Tang, Haize Labs
This talk introduces hazing, a form of fuzz testing adapted for Generative AI systems. It addresses the critical challenge of validating and verifying AI outputs, which are inherently subjective and unstructured. Traditional evaluation methods relying on static datasets are insufficient due to the brittle and non-deterministic nature of GenAI, where minor input variations can lead to drastically different outputs. Hazing aims to pressure-test AI systems through large-scale simulation and optimization before deployment to ensure robustness and reliability.
- Multi Agent AI and Network Knowledge Graphs for Change — Ola Mabadeje, Cisco
This talk introduces a system designed to reduce failures in network change management by leveraging multi-agent AI and network knowledge graphs. The core thesis is that by creating a digital twin of the production network, represented by a knowledge graph, and enabling specialized AI agents to interact with it, complex network operations can be made more robust and efficient. The system aims to provide a natural language interface for network operations teams and integrate with existing IT service management tools.
- Wisdom-Driven Knowledge Augmented Generation at Scale - Chin Keong Lam, Patho AI
This talk introduces Knowledge Augmented Generation (KAG) as an enhancement over traditional Retrieval Augmented Generation (RAG). KAG integrates structured knowledge graphs with large language models to enable more accurate, insightful, and decision-driving responses. The core thesis is that by representing wisdom as interconnected knowledge, AI systems can move beyond simple data retrieval to understanding and strategic advisory roles, particularly in complex domains.
- The Next Unicorns: 7 Top AI startups from the HF0 Residency
This talk highlights seven AI startups from the HF0 Residency program, showcasing innovative approaches to AI product development and deployment. The startups address diverse challenges, from creating personalized user experiences and improving data attribution to developing novel AI architectures and enabling widespread AI adoption. The core thesis emphasizes the shift from simply scaling AI models to focusing on reliability, developer ecosystems, and practical applications that solve real-world problems.
- #define AI Engineer - Greg Brockman, OpenAI (ft. Jensen Huang)
This talk features Greg Brockman discussing the evolution of AI engineering, his personal journey into coding and AI, and the future of building AI systems. He emphasizes the importance of practical building, the synergy between research and engineering, and the potential for AI to fundamentally transform various industries by lowering barriers to entry and enabling new forms of economic activity.
- The Future of Evals - Ankur Goyal, Braintrust
This talk discusses the evolution of AI model evaluation (evals), highlighting the shift from manual processes to automated optimization. It introduces Loop, an agent designed to enhance prompts, datasets, and scorers, thereby revolutionizing the eval process. The core thesis is that frontier models, particularly recent advancements, are now capable of significantly improving AI development workflows.
- Designing AI-Intensive Applications - swyx
This talk explores the evolving landscape of AI engineering, emphasizing the need for new standard models to guide the development of AI-intensive applications. The speaker posits that the field is moving beyond simple wrappers and demos towards robust production systems, drawing parallels to foundational periods in physics and other engineering disciplines. The core thesis is that identifying and adopting these new standard models will be crucial for building valuable and intelligent AI products.
- How to look at your data — Jeff Huber (Chroma) + Jason Liu (567)
This talk emphasizes the critical importance of measuring and analyzing both the inputs and outputs of AI systems to drive systematic improvement. The core thesis is that effective measurement, akin to Peter Drucker's adage, is essential for making informed decisions and enabling continuous enhancement. By looking at data, practitioners can move beyond guesswork and build more robust and user-centric AI products.
- On Engineering AI Systems that Endure The Bitter Lesson - Omar Khattab, DSPy & Databricks
This talk addresses the challenge of engineering AI systems in a rapidly evolving landscape, drawing parallels to Rich Sutton's "bitter lesson" in AI research. The core thesis is that while scaling and general methods are crucial for intelligence, AI engineering must focus on building reliable, robust, and understandable systems by abstracting away from rapidly changing low-level components. This involves investing in system design, clear specifications, and modularity rather than premature optimization or tightly coupled, model-specific implementations.
- Evals Are Not Unit Tests — Ido Pesok, Vercel v0
This talk introduces the concept of evals at the application layer, distinguishing them from traditional unit tests. It emphasizes that Large Language Models (LLMs) can be unreliable, leading to unexpected failures in AI applications even when basic functionality appears to work. The core thesis is that robust evals are crucial for building reliable AI products by systematically testing and measuring performance across a spectrum of user-driven scenarios.
- 2025 is the Year of Evals! Just like 2024, and 2023, and … — John Dickerson, CEO Mozilla AI
The core thesis of this talk is that 2025 will finally be the year for robust AI evaluation, driven by the convergence of several key factors. The increasing understanding of AI by business leaders, coupled with budget shifts towards generative AI projects, has set the stage. Crucially, the rise of agentic systems, which make decisions and take actions, introduces significant complexity and risk, making rigorous evaluation a necessity rather than an option.
- Vibe Coding with Confidence — Itamar Friedman, Qodo
This talk explores the evolution of AI in software development, moving beyond simple autocompletion and chat interfaces to more sophisticated, agent-driven workflows. The core thesis is that the Command Line Interface (CLI) will become a central hub for "vibe coding with confidence," enabling developers to manage complex, end-to-end tasks with greater reliability and control. This shift aims to address the limitations of current AI tools, particularly in enterprise settings, by integrating AI across the entire Software Development Life Cycle (SDLC).
- Full Workshop: Realtime Voice AI — Mark Backman, Daily
This workshop introduces Pipcat, an open-source Python framework for building real-time voice and multimodal AI agents. It emphasizes the challenges of creating natural, fast, and conversational voice AI, highlighting the advancements in speech-to-speech models that simplify pipeline complexity. The session provides a hands-on approach to building a voice bot, demonstrating Pipcat's modularity and orchestration capabilities.
- Vision AI in 2025 — Peter Robicheaux, Roboflow
This talk addresses the current state of AI vision, arguing that computer vision models lag significantly behind language models in terms of intelligence and pre-training leverage. The core thesis is that vision models are not yet "smart" due to limitations in evaluation metrics, a lack of effective large-scale pre-training utilization, and challenges in aligning visual and linguistic features. The presentation introduces new benchmarks and models aimed at improving vision AI's capabilities.
- Practical tactics to build reliable AI apps — Dmitry Kuchin, Multinear
This talk emphasizes a practical, user-centric approach to building reliable AI applications, moving beyond generic data science metrics. The core thesis is that AI app reliability stems from defining and testing against real-world scenarios and desired business outcomes, rather than abstract measures like factuality or groundness. This method allows for continuous improvement and confidence in the application's performance.
- How to Improve your Vibe Coding — Ian Butler
This talk addresses the current limitations of AI coding agents in identifying and fixing bugs, highlighting their tendency to generate false positives and struggle with complex codebases. It proposes practical strategies for developers to improve the effectiveness of these agents, focusing on structured rules, better context management, and leveraging advanced "thinking" models. The core thesis is that while agents show promise, their current implementation requires careful guidance and configuration to be truly useful for developers.
- Vibes won't cut it — Chris Kelly, Augment Code
This talk argues that the current hype around AI code generation overlooks the complexities of production software engineering. While AI can generate code, it doesn't inherently understand the nuanced decision-making, maintenance, and safety requirements of large-scale, production-ready systems. The core thesis is that professional software engineers remain essential, and the focus should shift from the quantity of AI-generated code to how AI can augment the existing, rigorous practices of software development.
- Real World Development with GitHub Copilot and VS Code — Harald Kirschner, Christopher Harrison
- Building Agents at Cloud Scale — Antje Barth, AWS
This talk explores building AI agents at cloud scale, emphasizing the reinvention of customer experiences and the development of new AI-powered applications. It highlights the significant integration of agentic capabilities and large language models (LLMs) in reimagining existing services like Alexa, which now operates on over 600 million devices. The presentation introduces a model-driven approach and an open-source Python SDK called Strand Agents, designed to accelerate the development and deployment of production-ready AI agents.
- State of Startups and AI 2025 - Sarah Guo, Conviction
The AI landscape is rapidly evolving, with significant advancements in reasoning capabilities and the emergence of sophisticated AI agents. While the potential for AI is vast, building impactful AI products is more challenging than anticipated, yet the value creation is immense, enabling companies to scale faster than ever before. The market for model capabilities is becoming increasingly competitive, with open-source contributions playing a vital role.
- Useful General Intelligence — Danielle Perszyk, Amazon AGI
This talk proposes a shift in the goal of AI development from creating "thinking machines" to augmenting human intelligence. It argues that general intelligence is not solely contained within an individual machine but emerges from social interaction and co-evolution with technology. The core thesis is that by building AI agents that can reliably interact with digital environments and align with human representations, we can enhance human capabilities and create useful general intelligence.
- The 2025 AI Engineering Report — Barr Yaron, Amplify
The 2025 AI Engineering Report survey reveals key trends in the rapidly evolving field. Despite varied job titles, the community is broad and technical, with a significant influx of experienced software engineers new to AI. The report highlights widespread LLM adoption for diverse use cases, with OpenAI models leading for external products. Customization methods like RAG and fine-tuning are prevalent, and frequent model and prompt updates are common.
- Agents vs Workflows: Why Not Both? — Sam Bhagwat, Mastra.ai
This talk explores the synergy between AI agents and workflows, arguing that they are not mutually exclusive but rather complementary tools for building robust AI applications. The core thesis is that combining agentic capabilities with structured workflows offers a more powerful and flexible approach than relying on either in isolation. The speaker critiques the tendency for some to advocate for one approach exclusively, emphasizing the practical benefits of integrating both.
- Why We Don’t Need More Data Centers - Dr. Jasper Zhang, Hyperbolic
The talk argues that the escalating demand for AI compute, particularly GPUs, does not necessitate solely building more data centers. Instead, it proposes that a GPU marketplace and intelligent resource allocation can significantly address this demand more efficiently and sustainably. The core idea is to leverage underutilized existing GPU capacity rather than exclusively expanding physical infrastructure.
- Infrastructure for the Singularity — Jesse Han, Morph
This talk introduces Morph's Infinibranch technology, a virtualization, storage, and networking solution designed for advanced AI agents. It posits that as AI develops intelligence and personhood, it requires infrastructure that can match its speed and complexity. Infinibranch aims to provide a substrate for a cloud environment where AI agents can operate with zero latency, enabling reversible actions, parallel exploration of possibilities, and a more dynamic interaction with the digital world.
- Hacking the Inference Pareto Frontier - Kyle Kranen, NVIDIA
This talk explores techniques for optimizing AI model inference to break the Pareto frontier, focusing on balancing quality, latency, and cost. The core thesis is that a well-designed system, tailored to specific application constraints, is crucial for successful deployment and application performance. By understanding and manipulating factors like scale, structure, and dynamism, engineers can achieve better service level agreements or reduce costs for existing ones.
- Pipecat Cloud: Enterprise Voice Agents Built On Open Source - Kwindla Hultman Kramer, Daily
This talk introduces Pipecat, an open-source, vendor-neutral framework for building reliable and performant voice AI agents. It emphasizes the challenges in voice AI development, such as achieving fast response times (targeting under 800 milliseconds for voice-to-voice) and accurate turn detection. The presentation also highlights Pipecat Cloud, a new offering designed to simplify the deployment and scaling of these agents by abstracting away complexities like Kubernetes.
- [Full Workshop] Building Conversational AI Agents - Thor Schaeff, ElevenLabs
This workshop focuses on building multilingual conversational AI agents, detailing the pipeline from speech-to-text to text-to-speech. It highlights the integration of large language models as the agent's "brain" and showcases ElevenLabs' tools for creating dynamic, responsive AI interactions across numerous languages. The session emphasizes practical application and developer experience, offering insights into configuring and deploying these agents.
- From Self-driving to Autonomous Voice Agents — Brooke Hopkins, Coval
This talk draws parallels between the development of self-driving car technology and the creation of robust evaluation strategies for autonomous voice agents. The core thesis is that the challenges and solutions in achieving reliability and scalability in self-driving can inform how we build trustworthy and effective voice AI systems. By adopting principles like large-scale simulation and continuous evaluation loops, developers can move beyond the limitations of deterministic approaches and unpredictable autonomous agents.
- Your realtime AI is ngmi — Sean DuBois (OpenAI), Kwindla Kramer (Daily)
This talk emphasizes the critical role of low latency in building effective voice AI experiences, arguing that current real-time AI is often inadequate. The core thesis is that achieving natural, human-like voice interactions requires a fundamental shift in how audio and video are handled over networks, moving beyond traditional methods like WebSockets to leverage technologies like WebRTC. The speakers suggest that voice is poised to become a primary interface for the next generation of AI applications.
- Why ChatGPT Keeps Interrupting You — Dr. Tom Shapland, LiveKit
This talk addresses the persistent issue of voice AI agents, like ChatGPT's advanced voice mode, interrupting users. The core problem lies in current voice AI's simplistic approach to turn-taking, which contrasts sharply with the complex, predictive, and parallel processing humans use in conversation. The presentation explores lessons from human conversation and introduces emerging techniques to improve voice AI's ability to manage conversational flow.
- Serving Voice AI at $1/hr: Open-source, LoRAs, Latency, Load Balancing - Neil Dwyer, Gabber
This talk details the experience of hosting an open-source voice AI model, Orpheus, for real-time consumer applications. The core challenge addressed is achieving low latency and high fidelity voice generation at a cost viable for widespread consumer use, which often requires near-free operation. The presentation highlights the technical hurdles in serving these models efficiently and presents solutions involving model fine-tuning, optimized infrastructure, and load balancing.
- How to defend your sites from AI bots — David Mytton, Arcjet
This talk addresses the increasing problem of automated traffic on websites, exacerbated by the rise of AI. While bots have always existed, AI crawlers and agents are intensifying issues like increased costs, bandwidth consumption, and denial-of-service attacks. The presentation outlines various defense mechanisms, ranging from voluntary standards to sophisticated detection and mitigation techniques, to help site owners manage and control bot traffic.
- The Unofficial Guide to Apple’s Private Cloud Compute - Jmo, CONFSEC
This talk provides an unofficial guide to Apple's Private Cloud Compute (PCC) system, explaining how it enables remote AI computation while maintaining user privacy. The core thesis is that Apple addresses the inherent privacy risks of sending data to remote servers by implementing a system with five key requirements: stateless computation, enforceable guarantees, non-targetability, no privileged runtime access, and verifiable transparency. The presentation outlines Apple's conceptual architecture and technical components designed to meet these requirements, offering insights into how similar privacy-preserving techniques can be adopted by developers outside the Apple ecosystem.
- How to Secure Agents using OAuth — Jared Hanson (Keycard, Passport.js)
This talk addresses the critical need for robust security in AI agents by advocating for the adoption of OAuth 2.0. It highlights the current security risks associated with broadly scoped, long-lived API keys used by agents and proposes OAuth as a solution to transition from static secrets to dynamic, delegated access. The discussion emphasizes that while OAuth is complex, its core principles are straightforward and essential for securing increasingly connected and useful AI agents.
- How we hacked YC Spring 2025 batch’s AI agents — Rene Brandel, Casco
This talk explores common security vulnerabilities found in AI agents by detailing a process of "hacking" Y Combinator batch AI agents. The core thesis is that agent security extends beyond traditional LLM prompt injection to encompass broader system-level risks, emphasizing the need to treat agents with the same security considerations as human users and to avoid building custom code execution environments.
- OpenAI on Securing Code-Executing AI Agents — Fouad Matin (Codex, Agent Robustness)
This talk addresses the critical security and safety considerations for AI agents capable of executing code. As AI models become increasingly proficient at writing and running code, the focus shifts from mere capability to responsible deployment and robust guardrails. The presentation emphasizes that code execution is becoming a standard feature for AI agents, moving beyond traditional software engineering tasks to achieve objectives more efficiently across various applications.
- Evaluating AI Search: A Practical Framework for Augmented AI Systems — Quotient AI + Tavily
This talk introduces a practical framework for evaluating AI search systems, particularly those operating in dynamic, real-world environments. It highlights the limitations of traditional static benchmarks and proposes dynamic datasets and reference-free metrics as crucial components for assessing the performance and reliability of augmented AI systems. The core thesis is that continuous improvement in AI systems requires robust evaluation methods that adapt to the evolving nature of information and user interactions.
- Scaling Enterprise-Grade RAG: Lessons from Legal Frontier - Calvin Qi (Harvey), Chang She (Lance)
This talk addresses the complexities of building enterprise-grade Retrieval Augmented Generation (RAG) systems, particularly within specialized domains like legal documents. It highlights the challenges of handling massive, complex datasets, sophisticated user queries, and stringent security requirements. The discussion emphasizes the critical role of robust evaluation strategies and modern data infrastructure to support these demanding applications.
- Building Alice’s Brain: an AI Sales Rep that Learns Like a Human - Sherwood & Satwik, 11x
This talk details the development of Alice's Brain, an AI sales representative's knowledge base designed to learn and operate more like a human. The system shifts from a manual context-pushing model to an automated one where the AI proactively pulls and utilizes relevant seller information. This approach aims to improve personalization and efficiency in sales outreach.
- Layering every technique in RAG, one query at a time - David Karam, Pi Labs (fmr. Google Search)
This talk provides a practical framework for improving Retrieval Augmented Generation (RAG) systems by systematically layering techniques based on their complexity and impact. It emphasizes a quality engineering approach, starting with defined outcomes and product problems, then analyzing failures to select appropriate RAG techniques. The core thesis is that understanding where and why a system fails is crucial for making informed decisions about which RAG enhancements to implement, rather than getting lost in hype or theoretical debates.
- Building a Smarter AI Agent with Neural RAG - Will Bryk, Exa.ai
This talk introduces Exa, a search engine designed for AI agents, moving beyond traditional keyword-based search. It argues that current search engines, optimized for human users, are insufficient for the complex, data-intensive needs of AI. Exa aims to provide a more powerful and flexible API that allows AI to query and retrieve information from the web with greater precision and comprehensiveness.
- [Full Workshop] Building Metrics that actually work — David Karam, Pi Labs (fmr Google Search)
This workshop focuses on building effective evaluation metrics for AI systems, moving beyond basic testing to create robust scoring systems. It emphasizes that evaluations are not just for testing but are the primary place where domain knowledge resides, enabling significant improvements in AI development. The session introduces a methodology for creating nuanced, calibrated metrics that correlate with desired outcomes, ultimately simplifying the AI development stack.
- Make your LLM app a Domain Expert: How to Build an Expert System — Christopher Lovejoy, Anterior
This talk outlines a playbook for building domain-native LLM applications, emphasizing that the system surrounding the model is more critical than the model's sophistication itself. The core challenge, termed the "last mile problem," involves providing LLMs with specific contextual understanding of industry workflows. By developing an adaptive domain intelligence engine, companies can leverage domain experts to translate industry insights into performance improvements, enabling continuous iteration and refinement of AI applications.
- Shipping Products When You Don't Know What they Can Do — Ben Stein, Teammates
This talk addresses the challenges of building and shipping products in the rapidly evolving AI agent space, where the capabilities of the underlying LLMs are not fully understood. It proposes a shift in product development from defining specific requirements to focusing on affordances and emergent behaviors, emphasizing the need for new tools and practices to navigate this uncertainty. The core thesis is that product management must transform to embrace discovery and adaptation in a probabilistic world.
- Shipping something to someone always wins — Kenneth Auchenberg (ex. Stripe, VSCode)
The core thesis is that successful product development, especially in the age of AI, hinges on rapid, iterative shipping and continuous user feedback, rather than large, infrequent releases. The speaker advocates for building "skateboard" equivalents—minimally viable products that deliver immediate value and allow for quick iteration based on real user input, which is far more effective than a traditional, linear development approach.
- Why your product needs an AI product manager, and why it should be you — James Lowe, i.AI
This talk argues for the critical role of an AI product manager, emphasizing that AI expertise is essential for this position. It builds on the idea that as AI coding agents and product features reduce the cost of software development, the demand for individuals who can effectively decide *what* to build will increase. The core thesis is that a dedicated AI product manager, or at least an AI product management mindset within a team, is crucial for navigating the complexities of AI product development.
- Everything is ugly, so go build something that isn't — Raiza Martin, Huxe (ex NotebookLM)
This talk argues that the current era of AI development is characterized by "ugly" or clunky products, which are merely the precursors to a more refined future. The core thesis is that building truly great AI products requires deep personal clarity, a singular focus on purpose, and a commitment to building user trust, which ultimately enables delight. The speaker emphasizes that restraint and judgment are key innovation multipliers in a landscape often dominated by overwhelming capabilities.
- Building the platform for agent coordination — Tom Moor, Linear
This talk explores Linear's journey in integrating AI into its product development tool, moving from early, pragmatic applications like natural language filters to more sophisticated agent-based features. The core thesis is that AI should be seamlessly integrated to provide practical value, acting as scalable, cloud-based teammates that enhance productivity and quality without being overly intrusive. The platform is evolving to support these agents as first-class citizens, enabling complex workflows and interactions.
- What Is a Humanoid Foundation Model? An Introduction to GR00T N1 - Annika & Aastha
This talk introduces GR00T N1, a humanoid robotics foundation model developed by Nvidia. It addresses the challenge of bridging the gap between the intelligence of large language models and the physical world by enabling robots to operate in human environments. The model's development is framed within a physical AI lifecycle involving data generation, model training, and deployment on edge devices.
- Real-time Experiments with an AI Co-Scientist - Stefania Druga, fmr. Google Deepmind
This talk introduces the concept of an AI co-scientist, analogous to a pair programmer but for real-world scientific experiments. The system integrates various sensors and cameras to collect empirical data in real-time, which an AI then analyzes to provide insights, generate hypotheses, and potentially accelerate scientific discovery. The presented system is built using open-source hardware and software, demonstrating a low-cost, accessible approach to AI-assisted experimentation.
- Scaling AI Agents Without Breaking Reliability — Preeti Somal, Temporal
This talk addresses the challenges of building reliable and scalable AI agent applications. It posits that these agents are complex distributed systems requiring robust orchestration, state management, and error handling, especially given the inherent unreliability of LLMs and external tools. The core thesis is that platforms like Temporal can abstract away the complexities of reliability and scalability, allowing developers to focus on core business logic and agent functionality.
- Ship Agents that Ship: A Hands-On Workshop - Kyle Penfound, Jeremy Adams, Dagger
This talk demonstrates how to build and deploy AI agents that can actively contribute to software development workflows using Dagger. It emphasizes creating agents with specific tools and environments, enabling them to perform tasks like code editing and testing within a sandboxed, containerized system. The core idea is to integrate AI capabilities into existing software engineering pipelines, providing guardrails and structure to manage the output of generative AI.
- The AI Engineer’s Guide to Raising VC — Dani Grant (Jam), Chelcie Taylor (Notable)
This talk, The AI Engineer’s Guide to Raising VC, features Dani Grant and Chelcie Taylor discussing the process and nuances of securing venture capital funding, particularly for early-stage companies. The core thesis is that founders, especially those from technical backgrounds, often overemphasize product and technology details when pitching VCs. Instead, investors primarily bet on the founder's vision, unique market insights, and ability to build a strong team. The discussion highlights that revenue and even a fully developed product are not prerequisites for raising capital, and emphasizes the importance of compelling storytelling and understanding investor psychology.
- Strategies for LLM Evals (GuideLLM, lm-eval-harness, OpenAI Evals Workshop) — Taylor Jordan Smith
This talk addresses the critical need for robust evaluation and benchmarking strategies when deploying large language models (LLMs) into production. It highlights the inherent complexities and potential pitfalls of generative AI, emphasizing that scalability, reliability, and safety are paramount. The presentation introduces practical tools and methods to assess LLM performance, ensuring that models meet enterprise-level requirements before and during deployment.
- Why you should care about AI interpretability - Mark Bissell, Goodfire AI
This talk explores mechanistic interpretability, a field focused on reverse-engineering neural networks to understand their internal workings. It argues that interpretability is moving from research labs into practical applications, offering AI engineers new tools for debugging, enhancing user experiences, and advancing scientific discovery. The core thesis is that understanding how AI models function internally is becoming crucial for building more reliable, controllable, and insightful AI systems.
- Information Retrieval from the Ground Up - Philipp Krenn, Elastic
This talk explores the fundamentals of information retrieval, focusing on the "R" in Retrieval Augmented Generation (RAG). It delves into both traditional keyword search and modern vector search, explaining their underlying mechanisms, strengths, and limitations. The presentation emphasizes that effective retrieval is crucial for accurate and relevant results in AI applications.
- Introduction to LLM serving with SGLang - Philip Kiely and Yineng Zhang, Baseten
This talk introduces SGLang, an open-source framework designed for high-performance serving of large language models (LLMs) and large vision models (LVMs). It emphasizes SGLang's production readiness, speed, and strong community support, highlighting its ability to provide day-zero support for new model releases and allow users to contribute to its development. The framework is presented as a valuable tool for optimizing LLM inference.
- Waymo's EMMA: Teaching Cars to Think - Jyh Jing Hwang, Waymo
This talk explores the evolution of autonomous driving systems, highlighting the transition from early, less capable models to sophisticated L4 systems like Waymo's. It introduces EMMA, an experimental system leveraging multimodal large language models (LLMs) like Gemini to enhance driving capabilities, particularly in handling rare and complex scenarios. The core thesis is that LLMs can significantly improve the generalizability and safety of autonomous driving by understanding and reacting to diverse real-world situations.
- Robotics: why now? - Quan Vuong and Jost Tobias Springberg, Physical Intelligence
This talk explores the advancements in robotics driven by the emergence of vision-language action (VLA) models. It highlights the transition from robots operating in highly constrained environments to performing complex tasks in semistructured real-world settings. The core thesis is that significant progress in robotics is now enabled by general AI developments, particularly VLA models, and that the primary bottleneck is shifting from hardware to software and model intelligence.
- A2A & MCP Workshop: Automating Business Processes with LLMs — Damien Murphy, Bench
This talk explores the automation of business processes using Large Language Models (LLMs) through two key protocols: A2A (Agent-to-Agent) and MCP (Model Context Protocol). A2A facilitates communication between remote agents, enabling specialization and parallel processing, while MCP acts as a standardized interface for agents to access context and tools, akin to a USB-C for AI. The workshop demonstrates how to integrate these protocols to build multi-agent systems triggered by webhooks, automating tasks like bug reporting and information dissemination.
- Piloting agents in GitHub Copilot - Christopher Harrison, Microsoft
This talk explores GitHub Copilot's capabilities beyond basic code completion, focusing on its evolution into an agentic tool. It highlights how providing context, clear instructions, and leveraging features like Copilot Coding Agent can significantly enhance developer productivity. The discussion emphasizes that while AI tools are powerful, fundamental DevOps practices and human oversight remain crucial for secure and effective software development.
- Ship Production Software in Minutes, Not Months — Eno Reyes, Factory
This talk argues that the future of software development is agent-driven, moving beyond human-driven approaches. It posits that AI agents, when properly integrated and provided with sufficient context, can handle a majority of tasks across the software lifecycle, enabling production software to be built in minutes rather than months. The core thesis is that organizations must transition to an agent-native development paradigm to unlock the true power of AI.
- Beyond the Prototype: Using AI to Write High-Quality Code - Josh Albrecht, Imbue
This talk addresses the gap between AI-generated code prototypes and production-ready software. It introduces Sculptor, an experimental coding agent environment designed to build trust in AI-generated code by focusing on identifying and preventing defects. The core thesis is that AI should be leveraged not just for code generation but also for ensuring its quality and reliability.
- Software Development Agents: What Works and What Doesn't - Robert Brennan, OpenHands
This talk explores the practical application and effectiveness of software development agents, emphasizing the shift from manual coding to higher-level problem-solving. It argues that while AI excels at the iterative process of writing and running code, human engineers remain crucial for critical thinking, user empathy, and architectural decisions. The presentation details the core components of these agents, their underlying mechanisms, and best practices for their integration into the development workflow.
- Devin 2.0 and the Future of SWE - Scott Wu, Cognition
The talk discusses the rapid advancement of AI agents in software engineering, highlighting a trend of capability doubling approximately every 70 days for coding tasks. This exponential growth has transformed AI's role from simple tab completion to handling complex, multi-file development tasks, with the interface and most critical capabilities evolving every few months.
- Your Coding Agent Just Got Cloned And Your Brain Isn't Ready - Rustin Banks, Google Jules
This talk introduces Jules, an asynchronous AI coding agent designed to run in the background and handle tasks in parallel, freeing developers to focus on creative coding. The core idea is to shift from serial task execution to a parallel workflow, leveraging AI for both the beginning and end of the software development lifecycle, from task creation to code merging and testing.
- Latent Space Paper Club: AIEWF Special Edition (Test of Time, DeepSeek R1/V3) — VIbhu Sapra
This talk reviews recent advancements in AI research, focusing on the DeepSeek models and the evolution of training methodologies. It introduces a new "Test of Time" paper club initiative aimed at systematically covering foundational AI concepts. The discussion highlights how reinforcement learning and extended inference time are enabling models to develop advanced reasoning capabilities, leading to significant performance improvements.
- Human seeded Evals — Samuel Colvin, Pydantic
This talk focuses on building AI applications more quickly and safely, emphasizing the importance of type safety in refactoring and avoiding bugs. It introduces techniques for developing AI applications, particularly highlighting the use of Pydantic AI for extracting structured data and managing agentic loops. The discussion also touches upon the challenges of determining when an agent loop should terminate and the benefits of using validation errors to guide model retries.
- Building AI Products That Actually Work — Ben Hylak (Raindrop), Sid Bendre (Oleve)
This talk focuses on the practical aspects of building AI products that are reliable and effective, moving beyond theoretical discussions of evaluations. It emphasizes that while AI technology is advancing, the core challenge lies in effective communication and managing the inherent complexity and undefined behaviors of AI systems. The speakers advocate for an iterative approach to product development, where continuous refinement based on real-world user data and signals is crucial for success.
- Rise of the AI Architect — Clay Bavor, Cofounder, Sierra w/ Alessio Fanelli
This talk introduces the concept of the AI Architect, a role emerging in the AI era analogous to the webmaster of the internet's early days. AI Architects are responsible for understanding the technology, shaping the agent's brand and user experience, and driving business outcomes. The discussion highlights the complexity of building AI agents, emphasizing that successful strategies involve a spirit of exploration, a focus on solving real problems, and a willingness to rearchitect existing teams and processes.
- AI That Pays: Lessons from Revenue Cycle — Nathan Wan, Ensemble Health
This talk explores the application of AI in healthcare's revenue cycle management (RCM), an often overlooked but critical financial process. The core thesis is that significant inefficiencies and lost revenue in healthcare stem from manual, complex, and error-prone administrative tasks, rather than clinical costs. AI offers a powerful opportunity to disrupt this area by reducing friction, preventing errors, and shifting resources towards more productive activities like patient care.
- Structuring a modern AI team — Denys Linkov, Wisedocs
This talk emphasizes that building a successful AI team hinges on understanding specific organizational needs and problems rather than solely chasing the latest AI research or hiring specialized roles prematurely. The core thesis is that technology adoption is often slow, and the effectiveness of AI solutions depends more on how they are integrated and utilized within a business context than on the cutting-edge nature of the technology itself. The speaker advocates for a pragmatic approach to team structure, prioritizing domain knowledge and business acumen alongside technical skills.
- The Rise of Open Models in the Enterprise — Amir Haghighat, Baseten
Enterprises are increasingly adopting AI, moving beyond simply purchasing verticalized solutions to building their own AI capabilities. While initial adoption often leverages closed-source models from providers like OpenAI and Anthropic, several factors are driving a shift towards open-source models. These include the need for specialized quality, reduced latency, better unit economics for agentic use cases, and a desire for competitive differentiation.
- Mentoring the Machine — Eric Hou, Augment Code
This talk explores how AI agents can be integrated into software development workflows, shifting the engineer's role from direct implementation to mentorship and orchestration. The core thesis is that by treating AI like junior engineers—providing context, guidance, and structured environments—teams can overcome the limitations of AI and unlock significant productivity gains. This approach transforms chaotic, interruption-filled days into more focused and efficient work, ultimately making software creation more of a science.
- Building Applications with AI Agents — Michael Albada, Microsoft
This talk explores the development of AI agents, highlighting both their promise and the significant obstacles encountered in production. It defines agents as entities capable of reasoning, acting, communicating, and adapting. The presentation emphasizes that agency exists on a spectrum and should be a tool to enhance effectiveness, not a goal in itself, cautioning against systems with low efficacy despite high agency.
- AX is the only Experience that Matters - Ivan Burazin, Daytona
The core thesis is that the future of software development and tooling lies in building for agents, not humans. As agents become the primary users of digital environments, the concept of "agent experience" (AX) will supersede traditional user and developer experiences. Tools that require human intervention will become obsolete, necessitating a shift towards building systems that agents can operate within autonomously.
- How to build Enterprise Aware Agents - Chau Tran, Glean
This talk explores the distinction between AI workflows and agents, highlighting their respective strengths and weaknesses. It proposes that agents can be viewed as tools that generate workflows, suggesting a powerful synergy between the two. The discussion also delves into methods for building enterprise-aware agents, emphasizing the importance of incorporating business-specific knowledge and processes.
- Monetizing AI — Alvaro Morales, Orb
This talk addresses the complexities of monetizing AI products, emphasizing the need for strategic pricing beyond intuition. It highlights the rapid evolution of AI technology, its impact on margins, and customer demand for clear ROI, all of which contribute to unique pricing challenges. The speaker proposes moving away from pricing based on vague notions towards data-driven frameworks and tools for effective monetization.
- Does AI Actually Boost Developer Productivity? (100k Devs Study) - Yegor Denisov-Blanch, Stanford
- How agents will unlock the $500B promise of AI - Donald Hruska, Retool
The talk argues that AI agents are poised to unlock significant economic value, moving beyond simple chatbots and code generation to integrate with real-world production systems. While building basic agents is becoming increasingly accessible, deploying them effectively within enterprises requires careful consideration of security, cost, and integration challenges. The future likely involves a mix of custom-built agents for core business logic and platform-hosted agents for broader workflows.
- How Intuit uses LLMs to explain taxes to millions of taxpayers - Jaspreet Singh, Intuit
Intuit leverages Large Language Models (LLMs) within TurboTax to enhance taxpayer understanding of their tax situations and potential deductions. The company has developed a proprietary generative OS called Geno, designed for large-scale, secure, and reliable use in the highly regulated tax industry. This platform integrates various components, including UI elements and an orchestrator, to manage different LLM solutions and deliver a cohesive user experience.
- 3 ingredients for building reliable enterprise agents - Harrison Chase, LangChain/LangGraph
Building reliable enterprise agents hinges on three core ingredients: maximizing value when the agent is correct, minimizing cost when it is wrong, and increasing the probability of success. This approach provides a foundational framework for developing agents that are more likely to be adopted and effective within an enterprise setting. The future vision involves numerous agents operating autonomously, coordinating tasks, and requiring careful management.
- From Hype to Habit: How We’re Building an AI-First SaaS Company—While Still Shipping the Roadmap
This talk explores the complexities of transitioning a SaaS company to an AI-first model, emphasizing that it's an evolutionary journey rather than a binary switch. It proposes a framework focusing on strategy, ways of working, and people to navigate this transformation. The core thesis is that becoming AI-first requires reimagining product development, embracing ambiguity, and fostering a cultural shift across the entire organization, not just within AI teams.
- Machines of Buying and Selling Grace - Adam Behrens, New Generation
This talk explores the evolution of the store concept in the age of AI, moving from physical locations to digitized online presences, and now towards an agentic future. The core thesis is that AI will digitize not just merchandise and distribution, but also the participants and their interactions, leading to dynamic, real-time, and generative interfaces for both human and agentic consumers. This shift aims to enhance transactions by moving beyond static websites to sophisticated agent-to-agent or agent-to-API interactions.
- How to Build Planning Agents without losing control - Yogendra Miraje, Factset
This talk explores how to build planning agents for AI applications, focusing on agentic workflows that combine the flexibility of agents with the reliability of traditional workflows. The core thesis is that by moving beyond simple reactive agents to proactive ones, and by employing a "planning by sub-goal division" design pattern, developers can create more controlled, scalable, and enterprise-ready AI systems. The approach emphasizes leveraging existing enterprise microservices and robust evaluation frameworks.
- Building Agents (the hard parts!) - Rita Kozlov, Cloudflare
This talk focuses on the practical challenges and components involved in building AI agents. It emphasizes that moving beyond simple AI augmentation to true automation requires a structured approach. The core idea is that agents need a client interface, an AI reasoning engine, workflow management for execution, and access to tools to perform actions.
- POC to PROD: Hard Lessons from 200+ Enterprise GenAI Deployments - Randall Hunt, Caylent
This talk shares hard-earned lessons from over 200 enterprise Generative AI deployments, emphasizing that AI is not a universal solution. It highlights the importance of understanding customer needs, robust evaluation, and efficient architecture over simply adopting the latest AI trends. The core thesis is that successful AI implementation requires a deep understanding of inputs, outputs, and user behavior, rather than relying solely on advanced models or prompt engineering.
- Build Dynamic Products, and Stop the AI Sideshow — Eliza Cabrera (Workday) + Jeremy Silva (Freeplay)
This talk argues for building dynamic, deeply integrated AI products rather than AI "sideshows" that are bolted onto existing systems. The core thesis is that companies should move beyond using AI primarily to demonstrate technological capability and instead focus on solving customer problems by integrating AI strategically into their core product development. This approach requires aligning AI and product strategies, teams, and roadmaps, and embracing a crawl, walk, run methodology for iterative development.
- The Billable Hour is Dead; Long Live the Billable Hour — Kevin Madura + Mo Bhasin, Alix Partners
This talk explores how AI, particularly generative AI, is reshaping knowledge work and professional services. The core thesis is that AI can significantly compress the initial data ingestion and analysis phases of engagements, freeing up experienced professionals for higher-value strategic tasks. This shift moves beyond traditional leverage models to a future where AI replicates and scales the expertise of senior individuals.
- From Copilot to Colleague: Trustworthy Agents for High-Stakes - Joel Hron, CTO Thomson Reuters
This talk explores the evolution of AI assistants from being merely helpful to becoming productive agents capable of making judgments and decisions. It emphasizes that agency in AI is not a binary state but a spectrum that can be tuned based on use case, risk tolerance, and user expectations. The presentation highlights the challenges and lessons learned in building trustworthy AI systems for high-stakes professional environments, particularly within legal, tax, and global trade sectors.
- How to Hire AI Engineers when EVERYONE is cheating with AI — Beth Glenfield, DevDay
The current technical hiring process is fundamentally broken due to the widespread use of AI, making traditional methods like LeetCode puzzles ineffective. Companies are struggling to identify genuine talent as AI assistants significantly boost candidate performance on these outdated assessments. This shift necessitates a reimagining of how to evaluate candidates, focusing on skills relevant to modern AI development and collaboration.
- Stateful environments for vertical agents — Josh Purtell, Synth Labs
This talk introduces the concept of stateful environments for AI agents, particularly for vertical applications. The core idea is to externalize and containerize the logic of a task or application, creating a distinct workspace that the agent can interact with. This separation allows for more robust agent development, easier updates, and advanced capabilities like multi-agent collaboration and state rollback.
- Books reimagined: AI to create new experiences for things you know — Lukasz Gandecki, TheBrain.pro
This talk explores reimagining books through AI to create novel user experiences. The core idea is to augment familiar content with dynamic, AI-generated elements like contextual summaries, character visualizations, and scene-appropriate music, transforming passive reading into an interactive, immersive journey. The approach emphasizes integrating AI subtly to enhance, rather than replace, the human element in content creation.
- AI powered entomology: Lessons from millions of AI code reviews — Tomas Reimers, Graphite
This talk explores the challenges and opportunities of using AI, specifically Large Language Models (LLMs), for code review. It highlights that while LLMs can identify bugs, their effectiveness is limited by the type of feedback they provide and the user's receptiveness to it. The core thesis is that by understanding the different categories of bugs and user preferences, AI can be trained to provide more valuable and actionable code review comments.
- Critical AI Inference your CIO can Trust — Sahil Yadav, Hariharan Ganesan, Telemetrak
This talk addresses the critical need for trustworthy AI, especially in mission-critical applications where AI inferences directly impact business decisions and financial outcomes. It highlights a significant gap between AI adoption and AI governance, leading to potential silent failures with substantial financial and operational consequences. The presentation introduces a framework for building and scaling AI systems that instill confidence through explainability, traceability, and robust guardrails.
- How to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, Adobe
This talk emphasizes the critical role of evaluations (evals) in developing robust AI applications, moving beyond simple accuracy or similarity metrics. It highlights that effective evals are fundamental for aligning applications with system goals, ensuring continuous improvement, and building user trust. The core thesis is that a strategic, data-driven approach to evals, tailored to specific use cases, is essential for scaling AI development.
- Continuous Profiling for GPUs — Matthias Loibl, Polar Signals
This talk introduces continuous profiling for GPUs, a method to monitor and analyze GPU performance. It highlights the importance of profiling for improving application performance and reducing operational costs by enabling more efficient resource utilization. The approach leverages Linux eBPF for low-overhead, always-on profiling in production environments without requiring application instrumentation.
- Top Ten Challenges to Reach AGI — Stephen Chin, Andreas Kollegger
This talk explores ten potential challenges on the path to Artificial General Intelligence (AGI), drawing parallels with science fiction concepts. The speakers suggest that by examining these fictional scenarios, we can better understand and prepare for the real-world implications and ethical considerations of developing advanced AI. The core thesis is that a proactive, science-fiction-informed approach is crucial for responsible AGI development.
- Practical GraphRAG: Making LLMs smarter with Knowledge Graphs — Michael, Jesus, and Stephen, Neo4j
This talk introduces Graph RAG, a method for enhancing Large Language Models (LLMs) by integrating knowledge graphs. It addresses the limitations of standard LLMs, such as a lack of domain-specific knowledge, tendency to hallucinate, and difficulty in explaining answers. Graph RAG aims to provide more accurate, contextual, and explainable responses by leveraging structured data within knowledge graphs, moving beyond the limitations of purely vector-based retrieval.
- Knowledge Graphs in Litigation Agents — Tom Smoker, WhyHow
This talk explores the application of knowledge graphs and multi-agent systems within the legal industry, specifically for identifying and supporting class-action lawsuits. It details how unstructured data from web scraping and legal discovery can be transformed into structured graphs to aid legal professionals in case research and analysis. The core thesis is that by leveraging these technologies, complex legal data can be made more accessible, accurate, and actionable for lawyers.
- When Vectors Break Down: Graph-Based RAG for Dense Enterprise Knowledge - Sam Julien, Writer
This talk explores the limitations of traditional vector-based Retrieval Augmented Generation (RAG) for complex enterprise knowledge and proposes a graph-based RAG approach. It highlights how preserving relationships within data through knowledge graphs, combined with advanced techniques like fusion-decoder, significantly improves accuracy and reduces hallucinations in AI applications, especially for dense, specialized datasets.
- Stop Using RAG as Memory — Daniel Chalef, Zep
This talk argues against using Retrieval-Augmented Generation (RAG) as a generic memory solution for AI agents. Instead, it proposes modeling memory after specific business domains to create more cogent and capable memory systems. The core thesis is that semantic similarity alone is insufficient for accurate memory recall, and domain-aware memory structures are necessary.
- HybridRAG: A Fusion of Graph and Vector Retrieval - Mitesh Patel, NVIDIA
This talk introduces HybridRAG, a system that combines graph and vector retrieval for enhanced information retrieval. It highlights the advantages of knowledge graphs in capturing detailed relationships between entities, offering a more comprehensive view than purely semantic approaches. The system is broken down into data processing, graph and vector database creation, and inferencing, with a focus on optimizing retrieval strategies and evaluating performance.
- tldraw.computer - Steve Ruiz, tldraw
This talk introduces tldraw.computer, a platform that leverages AI for creative and functional applications. It showcases how a hackable canvas, built with standard web technologies, can serve as a foundation for AI-driven tools. The core thesis is that by treating AI models as programmable components within a visual interface, users can build complex applications, iterate rapidly, and achieve novel results, blurring the lines between design, coding, and AI execution.
- Excalidraw: AI and Human Whiteboarding Partnership - Christopher Chedeau
This talk explores the evolution of whiteboarding tools, drawing parallels between the transition from physical to digital whiteboards and the current integration of AI. It emphasizes that successful AI integrations should enhance, not merely replicate, existing functionalities, focusing on AI-native interactions rather than simply adding AI features. The core thesis is that AI and human collaboration in digital whiteboarding is entering a new phase, moving beyond basic AI augmentation to a more symbiotic partnership.
- The Bitter Layout or: How I Learned to Love the Model Picker — Maximillian Piras, Yutori
This talk explores the evolution of AI user interfaces, arguing that the current dominant chat-based layout, while functional, is a "bitter layout" that prioritizes adaptability to new models over optimal user experience. It suggests that as models become less commoditized, the interface itself becomes the commodity, necessitating a shift in design philosophy towards goal-oriented and constraint-based approaches, akin to gardening rather than construction.
- UX Design Principles for Semi Autonomous Multi Agent Systems — Victor Dibia, Microsoft
This talk explores the design principles for creating effective user experiences in semi-autonomous multi-agent systems. It emphasizes that while multi-agent systems offer powerful capabilities, their complexity requires careful consideration of user interaction, observability, and control. The presentation highlights a practical approach to building such systems, starting with defining the goal and tools before agent development, and advocates for an evaluation-driven design process.
- Agentic GraphRAG: AI’s Logical Edge — Stephen Chin, Neo4j
This talk introduces Agentic GraphRAG as a method to improve the accuracy and reduce hallucinations in AI agent systems. It highlights the limitations of current LLMs in complex reasoning and decision-making, proposing that knowledge graphs, when integrated with LLMs, can provide a more robust and logical framework for AI operations. The approach leverages graph databases to store, manage, and retrieve information, enhancing the AI's ability to understand context and provide reliable outputs.
- CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS
This talk addresses the critical need for robust identity, authentication, and authorization solutions for AI agents, a rapidly growing segment in the B2B SaaS landscape. As agents become more integrated into workflows, they require first-class identity support, distinct from traditional bots or human users. The presentation emphasizes the urgency of developing new standards for agent identity to ensure user safety and enable scalable adoption of AI technologies.
- Good design hasn’t changed with AI — John Pham, SF Compute
This talk argues that fundamental principles of good design remain unchanged despite advancements in AI tools. Design is defined not by aesthetics but by the entirety of a user's experience across all touchpoints, influencing how they feel about a product. In the current AI landscape, where feature parity is no longer a differentiator, design becomes the key element for creating unique and memorable user experiences.
- Building Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAI
This talk explores the advancements and practical applications of building effective voice agents, moving beyond traditional text-based AI. It highlights the emergence of speech-to-speech technology as a key component of the multimodal AI era, emphasizing that current models are now fast, expressive, and accurate enough for scalable production use. The discussion covers architectural patterns, key trade-offs in development, and best practices for creating robust and engaging audio experiences.
- What every AI engineer needs to know about GPUs — Charles Frye, Modal
This talk explains why AI engineers need to understand GPUs, shifting focus from API-based development to leveraging hardware capabilities. It draws an analogy to database usage, where developers don't build databases but must understand how to query them effectively. Similarly, AI engineers will increasingly need to understand GPU architecture, particularly tensor cores, to optimize performance for tasks like language model inference.
- Robots as professional Chefs - Nikhil Abraham, CloudChef
CloudChef has developed culinary intelligent robots capable of acting, sensing, and reasoning like human chefs. These robots utilize a combination of foundation models, teleoperation for edge cases, and specialized thermal and visual embeddings for food understanding. The system is designed to adapt to new kitchens, learn recipes from single demonstrations, and handle variations in ingredients and appliances, aiming to make high-quality food affordable through automation.
- [Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han
This talk delves into advanced AI training techniques, focusing on reinforcement learning (RL), quantization, and agent development. It explores the evolution of large language models, from early open-source efforts spurred by leaks to current sophisticated training methodologies. The discussion highlights the critical role of fine-tuning stages, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF), in transforming base models into capable conversational agents.
- A Taxonomy for Next-gen Reasoning — Nathan Lambert, Allen Institute (AI2) & Interconnects.ai
This talk explores the evolution of AI reasoning capabilities, moving beyond benchmark scores to focus on practical applications and future development. It posits that while current models excel at specific skills like math and code, the next frontier lies in planning, strategy, and abstraction. Achieving this requires a shift in how models are trained, emphasizing human effort in developing new algorithmic methods and data acquisition strategies.
- How to Train Your Agent: Building Reliable Agents with RL — Kyle Corbitt, OpenPipe
This talk details the process of building a reliable AI agent using reinforcement learning (RL), focusing on practical lessons learned from the ART E project, an email assistant. The core thesis is that while RL can significantly improve agent performance beyond prompted models, it's crucial to start with a strong prompted baseline, carefully design the training environment and reward functions, and be vigilant against reward hacking. The project demonstrates how RL can lead to substantial gains in accuracy, cost reduction, and latency improvements for specialized tasks.
- OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs
This talk introduces OpenThoughts 3, a project focused on creating high-quality, open-source reasoning datasets. The core thesis is that while models have shown significant performance gains on reasoning benchmarks, the underlying data recipes for achieving this are often undisclosed. OpenThoughts aims to fill this gap by providing a systematic approach and publicly available datasets to train more capable reasoning models, emphasizing that supervised fine-tuning (SFT) on well-crafted data can be highly effective.
- Google Photos Magic Editor: GenAI Under the Hood of a Billion-User App - Kelvin Ma, Google Photos
This talk details the engineering journey behind Google Photos' Magic Editor, highlighting the transition from on-device ML models to server-side generative AI. It emphasizes the challenges and learnings in integrating cutting-edge AI into a widely used application, focusing on user experience, model efficiency, and the inherent unpredictability of AI systems. The core thesis is that while AI offers powerful new capabilities, successful product integration requires rigorous engineering to manage its randomness and deliver reliable, user-friendly features.
- Dream Machine: Scaling to 1m users in 4 days — Keegan McCallum, Luma AI
This talk details Luma AI's experience scaling their Dream Machine video generation model to one million users in four days. It covers the immense infrastructure challenges faced during the launch, including the need to rapidly scale GPU resources from an initial 500 H100s to 5,000. The presentation also discusses the evolution of their serving stack from a brittle, coupled system to a more robust, decoupled architecture built on PyTorch, and the strategies implemented for fair job scheduling and efficient model management.
- ComfyUI Full Workshop — first workshop from ComfyAnonymous himself!
ComfyUI is an open-source, node-based design canvas for generative AI, supporting multimodal creative applications including image, video, audio, 3D, and text. It aims to provide maximal control by allowing users to interact with models at a granular level, going beyond simple prompts to manipulate elements like depth maps and masks. The platform emphasizes extensibility through community-developed custom nodes and a unique workflow sharing mechanism where generated images contain embedded metadata of the original workflow.
- Design like Karpathy is watching — Zeke Sikelianos, Replicate
This talk explores how to design AI products and services with language models as a primary audience, drawing lessons from Andrej Karpathy's experience building the Menuguen app. It emphasizes the shift towards LLMs consuming structured data like markdown and API schemas, rather than just human-readable web pages. The core thesis is that embracing these formats and improving developer experience is crucial for building successful AI-powered applications.
- On Curiosity — Sharif Shameem, Lexica
This talk argues that curiosity is the primary driving force for innovation, enabling the translation of future ideas into present realities. The speaker emphasizes that building and sharing demos is the most effective way to explore the potential of AI models, as these models are not fully understood even by their creators. Demos serve as a crucial tool for AI engineers, akin to excavators uncovering hidden capabilities, with curiosity acting as the guide.
- Real world MCPs in GitHub Copilot Agent Mode — Jon Peck, Microsoft
This talk introduces GitHub Copilot's Agent Mode, an advanced feature designed for completing moderately complex tasks autonomously. It moves beyond simple code completion and chat interactions to enable deep, iterative engagement with an AI agent. The presentation highlights how Agent Mode can build entire applications from a readme file or perform significant refactoring, with the developer providing permission for actions like terminal interactions.
- The rise of the agentic economy on the shoulders of MCP — Jan Curn, Apify
This talk explores how general intelligence in computing systems may emerge from the interaction of multiple agents, similar to how intelligence arises in biological systems or markets. The speaker posits that the Multi-agent Conversation Protocol (MCP) is a crucial development enabling agents to communicate and form an agentic mesh, facilitating the rise of an agentic economy where agents can discover and utilize services.
- MCP is all you need — Samuel Colvin, Pydantic
This talk argues that the Messaging, Communication, and Protocol (MCP) framework, particularly its tool-calling capabilities, can simplify complex agentic workflows. The speaker, Samuel Colvin, creator of Pydantic, suggests that many current approaches to agent-to-agent communication are overcomplicated and that MCP offers a more streamlined solution. The core idea is to leverage MCP's primitives, especially dynamic tool calling, to build more robust and efficient AI systems.
- Full Spec MCP: Hidden Capabilities of the MCP spec — Harald Kirschner, Microsoft/VSCode
This talk explores the underutilized capabilities of the Model Contract Protocol (MCP) beyond basic tool calling. It highlights how embracing the full MCP specification can unlock richer, stateful interactions for AI agents. The presentation emphasizes that while many are shipping products quickly using only tools, the spec offers advanced features like dynamic discovery, resources, and sampling that enable more sophisticated agent behavior and development experiences.
- Shipping an Enterprise Voice AI Agent in 100 Days - Peter Bar, Intercom Fin
This talk details the process of developing and shipping Finn Voice, an AI-powered voice agent for customer support, within a 100-day timeframe. The core thesis is that voice AI represents a significant frontier in customer service, offering substantial benefits in cost savings, availability, and user experience compared to traditional phone support. The presentation emphasizes that successful deployment requires more than just advanced AI models; it necessitates a product-centric approach that considers use cases, conversation design, integration with existing workflows, and building user trust.
- The State of Generative Media - Gorkem Yurtseven, FAL
The generative media landscape is rapidly evolving, with significant advancements in video and image generation models. The marginal cost of creation is approaching zero, leading to transformative impacts across various industries like advertising, e-commerce, and entertainment. While image generation has seen substantial progress, video generation is poised for exponential growth, potentially dwarfing the image market in scale and utility.
- Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+DAU - Devansh Tandon
This talk details the adaptation of large language models (LLMs), specifically Gemini, for YouTube's recommendation system, aiming to improve user engagement for over 2 billion daily active users. The core thesis is that LLMs can revolutionize recommendation systems, potentially surpassing search applications in consumer impact. The approach involves creating a domain-specific language for videos and training a bilingual model capable of understanding both natural language and this new video token language.
- Transforming search and discovery using LLMs — Tejaswi & Vinesh, Instacart
This talk details how Instacart is leveraging Large Language Models (LLMs) to significantly enhance its search and discovery functionalities. The core thesis is that LLMs can overcome limitations of traditional search models, particularly for tail queries and new item discovery, by better understanding user intent and context, ultimately leading to improved user experience and business metrics.
- Netflix's Big Bet: One model to rule recommendations: Yesu Feng, Netflix
Netflix is betting on a single foundation model, based on transformer architecture, to handle all recommendation use cases. This approach aims to improve personalization by scaling semi-supervised learning and create high leverage by integrating the model across all systems, simultaneously enhancing downstream applications. The model learns user representations and incorporates rich event data, addressing challenges like cold-start problems and enabling faster innovation.
- 360Brew: LLM-based Personalized Ranking and Recommendation - Hamed and Maziar, LinkedIn AI
This talk details LinkedIn's journey in developing and deploying 360Brew, a large language model-based system for personalized ranking and recommendations. The core thesis is that a single, holistic LLM can replace disjointed, task-specific models, offering zero-shot capabilities, in-context learning, and instruction following to understand user journeys and serve diverse personalization needs on the LinkedIn platform.
- What We Learned from Using LLMs in Pinterest — Mukuntha Narayanan, Han Wang, Pinterest
This talk details Pinterest's experience integrating Large Language Models (LLMs) into their search relevance system. The core thesis is that LLMs significantly improve relevance prediction, and that techniques like knowledge distillation are crucial for productionizing these models at scale. The presentation also highlights the value of synthetic captions and user engagement data as content annotations.
- Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3 — Greg Kamradt, ARC Prize Foundation
This talk introduces ARC-AGI-3, a new benchmark designed to measure artificial general intelligence (AGI) by focusing on interactive reasoning and skill acquisition efficiency. The benchmark aims to create problems that are solvable by humans but challenging for current AI, thereby guiding AI research and development towards human-level intelligence. It moves beyond single-turn, static benchmarks to simulate more realistic, open-world exploration and learning scenarios.
- RL for Autonomous Coding — Aakanksha Chowdhery, Reflection.ai
This talk explores the progression of large language models (LLMs) from early scaling laws to the current frontier of autonomous coding. It highlights how techniques like chain-of-thought prompting and reinforcement learning with human feedback have improved LLM capabilities. The core thesis is that reinforcement learning, particularly in domains with automated verification like coding, represents the next significant step in scaling LLM performance and building more intelligent systems.
- Recsys Keynote: Improving Recommendation Systems & Search in the Age of LLMs - Eugene Yan, Amazon
This talk explores the integration of Large Language Models (LLMs) with recommendation systems and search functionalities. It proposes three key areas for advancement: semantic IDs to better represent item content and address cold-start problems, data augmentation using LLMs to improve data quality and scale for training, and unified models to streamline complex recommendation infrastructures.
- Benchmarks Are Memes: How What We Measure Shapes AI—and Us - Alex Duffy, Every.to
This talk argues that AI benchmarks function as memes, spreading ideas and shaping the development of AI models. As benchmarks become popular, models are trained and tested on them until saturation occurs. This cycle presents an opportunity for individuals to create new benchmarks that can influence the future direction of AI development, emphasizing the importance of human-defined goals and values in the process.
- Small AI Teams with Huge Impact — Vik Paruchuri, Datalab
This talk challenges the conventional Silicon Valley belief that increasing headcount directly correlates with greater productivity. The speaker argues that smaller, highly capable teams of generalists can achieve significantly more by focusing on core competencies, leveraging AI for low-leverage tasks, and maintaining a high degree of trust and customer focus. This approach prioritizes efficient collaboration and rapid feedback loops over bureaucratic processes often found in larger organizations.
- Rethinking Team Building: how a 30-person Startup serves 50 Million Users — Grant Lee, Gamma
This talk challenges traditional startup team-building and scaling strategies, advocating for a leaner, more agile approach. The speaker, Grant Lee of Gamma, proposes that instead of hierarchical growth, companies can achieve significant user reach with a small, highly capable team by embracing generalists, player-coaches, and a strong, shared culture. This model aims to foster innovation and productivity by prioritizing content and user needs over design complexities.
- Building a 10 person unicorn - Max Brodeur-Urbas, Gumloop
This talk focuses on how to scale a startup team to a "unicorn" status with fewer than ten people, emphasizing culture, hiring, and operational efficiency. The core thesis is that by being extremely selective in hiring, fostering a product-led approach, and automating extensively, small teams can achieve rapid growth and significant impact without ballooning in size. The speaker shares insights from their experience founding Gumloop, a workflow automation tool.
- Using OSS models to build AI apps with millions of users — Hassan El Mghari
This talk explores how to leverage open-source AI models to build applications capable of reaching millions of users. The speaker emphasizes that it's an opportune time for builders due to lowered barriers in development tools and the rapid release of groundbreaking AI models. The core thesis is that by focusing on simplicity, iterating quickly, and embracing open-source principles, developers can successfully create and scale AI-powered applications.
- Bolt.new: How we scaled $0-20m ARR in 60 days, with 15 people — Eric Simons, Bolt
This talk details Bolt.new's rapid growth from $0 to $20 million in Annual Recurring Revenue (ARR) within 60 days with a small team. The core thesis emphasizes the power of a lean, highly aligned team with deep context, enabling rapid iteration and scaling, even with an initial MVP product. It highlights the importance of independent decision-making and fostering user love through direct engagement and community building.
- Prompt Engineering and AI Red Teaming — Sander Schulhoff, HackAPrompt/LearnPrompting
This talk explores the evolving landscape of prompt engineering and AI red teaming, emphasizing that prompt engineering remains a critical skill despite claims of its demise. The speaker highlights the inherent difficulty in securing generative AI systems, drawing parallels and distinctions with classical cybersecurity. The presentation delves into various prompting techniques, the challenges of AI security, and the practical implications of prompt injection and jailbreaking.
- Survive the AI Knife Fight: Building Products That Win — Brian Balfour, Reforge
In a rapidly evolving AI landscape, building successful products requires a strategic approach beyond simply integrating AI features. The core thesis is that competitive advantage stems not from the AI technology itself, but from the unique combination of a product's proprietary data, its specific functionality, and a deep understanding of unmet customer needs. This talk emphasizes treating AI as modular components, or "Lego blocks," that can be assembled to create differentiated offerings.
- Automating Escrow with USDC and AI - Corey Cooper, Circle
This talk explores the integration of AI with programmable money, specifically Circle's USDC, to automate complex financial workflows like escrow. The core idea is to leverage AI for verifying conditions and USDC for instant, borderless settlement, creating a more efficient and programmable financial system. The presentation introduces Circle's developer tooling and demonstrates an open-sourced escrow agent application.
- How LLMs work for Web Devs: GPT in 600 lines of Vanilla JS - Ishan Anand
This talk breaks down how Large Language Models (LLMs) like GPT-2 work, demystifying them for web developers by implementing a core GPT-2 model in approximately 600 lines of vanilla JavaScript. The presentation emphasizes understanding the underlying mechanics rather than requiring deep ML expertise, using analogies and code walkthroughs to make complex concepts accessible. The goal is to transform the perception of LLMs from magic to understandable machinery.
- [Workshop] AI Pipelines and Agents in Pure TypeScript with Mastra.ai — Nick Nisi, Zack Proser
This workshop introduces Mastra.ai, a framework for building AI applications in TypeScript. It demonstrates how to create agentic workflows, focusing on composable pipelines, tools, and agents. The session highlights using TypeScript for production-ready AI applications, emphasizing patterns applicable to deploying AI solutions.
- AI Engineering with the Google Gemini 2.5 Model Family - Philipp Schmid, Google DeepMind
This talk introduces the Google Gemini 2.5 model family, focusing on practical applications for AI engineers and builders. It highlights the multimodal capabilities of Gemini 2.5 Pro and Flash, demonstrating how to leverage these models for text generation, image and audio understanding, function calling, and integrating with external tools via MCP servers. The session emphasizes hands-on learning through a workshop format, encouraging attendees to experiment with the models and SDK.
- The New Code — Sean Grove, OpenAI
This talk argues that structured communication, embodied in specifications, is more valuable than code itself for AI development. Specifications serve as the primary artifact for aligning human intent and values, enabling clearer communication, better testing, and more robust AI systems. As AI models advance, the ability to write effective specifications will become the most critical skill for programmers.
- Production software keeps breaking and it will only get worse — Anish Agarwal, Traversal.ai
The talk argues that as AI tools increasingly automate code development, the complexity of system design and production troubleshooting will grow, potentially leading engineers to spend more time on on-call duties. The core thesis is that current approaches to AI-assisted troubleshooting are insufficient and a new, more integrated method is needed to handle the escalating complexity of production incidents.
- Thinking Deeper in Gemini — Jack Rae, Google DeepMind
This talk explores the concept of "thinking" within Gemini, a new paradigm for AI models that allows for iterative computation and deeper reasoning before generating a final response. The core thesis is that by enabling models to spend more test-time compute on complex problems, we can unlock significant advancements in AI capabilities, moving beyond immediate response generation towards more profound problem-solving. This approach addresses bottlenecks in current AI by allowing for dynamic allocation of computational resources based on task difficulty.
- A year of Gemini progress + what comes next — Logan Kilpatrick, Google DeepMind
Logan Kilpatrick of Google DeepMind discussed a year of progress with Gemini models and outlined future developments. The talk highlighted the release of a new Gemini 2.5 Pro model, emphasizing its improved performance across benchmarks and its role as a turning point for the Gemini family. Kilpatrick also touched upon the organizational shifts within Google that have integrated research and product teams, enabling faster delivery of AI capabilities to both consumers and developers.
- 2025 in LLMs so far, illustrated by Pelicans on Bicycles — Simon Willison
The AI landscape has accelerated dramatically over the last six months, with numerous significant model releases making it challenging to assess their quality. Traditional benchmarks and leaderboards are losing credibility, prompting a shift towards more practical, self-devised evaluation methods. The speaker uses a unique benchmark involving generating SVGs of pelicans riding bicycles to assess model capabilities, highlighting the rapid progress in model performance and the increasing accessibility of powerful AI on consumer hardware.
- Trends Across the AI Frontier — George Cameron, ArtificialAnalysis.ai
This talk explores multiple frontiers in AI beyond just raw intelligence, emphasizing the trade-offs involved in accessing advanced AI capabilities. It highlights that the most intelligent models are not always the most suitable due to implications for cost, latency, and verbosity. The presentation uses benchmarking data to illustrate these trade-offs across reasoning capabilities, open-weight models, cost-effectiveness, and speed.
- Training Agentic Reasoners — Will Brown, Prime Intellect
This talk argues that reasoning and agents are fundamentally the same concept, with reinforcement learning (RL) being the key to developing powerful agentic systems. The speaker posits that traditional approaches to building agents often involve manual, iterative tuning that mirrors RL processes. By framing agent development through the lens of RL, developers can leverage established algorithms and techniques to create more robust and capable agents, especially for complex, multi-turn tasks.
- New York Times' Connections: A Case Study on NLP in Word Games — Shafik Quoraishee, NYT Games
This talk explores the New York Times Connections word game as a case study for evaluating AI's abstract reasoning capabilities. The presenter, a game developer at the NYT Games team, details independent research into how AI can approach solving the game, which involves grouping 16 words into four sets of four related terms. The research investigates the game's potential as a benchmarking tool, highlighting how its intentional decoys and need for explainable reasoning challenge current AI models.
- Claude Code & the evolution of agentic coding — Boris Cherny, Anthropic
The talk by Boris Cherny from Anthropic discusses the rapid evolution of AI models in coding and the challenges in developing user experiences (UX) that keep pace. It highlights how programming itself has evolved through layers of abstraction and interface changes, from punch cards to modern IDEs. The core thesis is that while models are advancing exponentially, the product development for these AI coding tools is still in its early stages, with a focus on building unopinionated, minimal-viable products to learn and adapt to the unknown optimal UX.
- 12-Factor Agents: Patterns of reliable LLM applications — Dex Horthy, HumanLayer
This talk proposes the "12-Factor Agents" pattern, drawing parallels to the 12-Factor App methodology for building reliable software. The core thesis is that building robust LLM applications requires applying established software engineering principles, focusing on modularity, control flow, and state management rather than solely relying on agentic capabilities. The approach emphasizes treating agents as software components that can be engineered for reliability and maintainability.
- MCP Is Not Good Yet — David Cramer, Sentry
David Cramer of Sentry discusses the current state of MCP (Model Communication Protocol), framing it as a pluggable architecture for agents. He emphasizes that while the concept is powerful, current implementations are often rough and require significant design effort. Cramer suggests that MCP's true value lies in its ability to integrate services into agent workflows, particularly for B2B SaaS companies, by providing context that LLMs can reason about.
- Your Personal Open-Source Humanoid Robot for $8,999 — JX Mo, K-Scale Labs
K-Scale Labs is developing open-source humanoid robots, the Kbot and Zbot, aimed at making advanced robotics accessible to developers and researchers. Their core mission is to democratize robotics by open-sourcing the entire hardware and software stack, enabling widespread adoption and innovation. The Kbot is a full-sized humanoid robot designed for affordability and modularity, while the Zbot is a smaller, more accessible option derived from a hackathon project.
- The Build-Operate Divide: Bridging Product Vision and AI Operational Reality
This talk addresses the common challenge of AI product concepts failing to reach their full potential due to operational difficulties. It emphasizes that bridging the gap between product vision and AI operational reality requires a deep understanding of how to deliver quality through rigorous evaluation, human review, and strategic team building. The core thesis is that successful AI product development hinges on a robust operational foundation that supports continuous iteration and quality assurance.
- The New Lean Startup — Sid Bendre, Oleve
This talk introduces a new model for building companies, termed the new lean startup, driven by the advent of AI tooling. It emphasizes the shift towards smaller, more profitable companies that achieve significant ARR with minimal teams. The core thesis is that AI tools enable tiny teams to build and scale successful products rapidly, challenging the traditional startup growth model.
- Conquering Agent Chaos — Rick Blalock, Agentuity
This talk addresses the significant challenges in deploying and running AI agents, often referred to as agent chaos. The speaker highlights common issues like timeouts, state management, and the need for agents to run for extended periods, pause, and resume. The core thesis is that successful agent deployment requires specific infrastructure and capabilities beyond typical stateless web applications, leading to the development of a platform designed to manage these complexities.
- Optimizing inference for voice models in production - Philip Kiely, Baseten
This talk focuses on optimizing inference for voice models in production, emphasizing runtime performance and infrastructure considerations. It highlights how the architectural similarity between Text-to-Speech (TTS) models and Large Language Models (LLMs) allows for the application of LLM optimization techniques. The core thesis is that while runtime optimizations are crucial, non-runtime factors like infrastructure and client code implementation can significantly impact overall latency and cost-efficiency.
- [Evals Workshop] Mastering AI Evaluation: From Playground to Production
This talk focuses on mastering AI evaluation, moving from initial development to production. It emphasizes that even the best large language models (LLMs) require a robust testing framework due to issues like hallucinations and performance degradation with changes. Effective evaluation helps answer critical questions about model selection, cost-effectiveness, brand consistency, and ongoing improvement, ultimately reducing development time, costs, and enabling faster iteration.
- Intro to GraphRAG — Zach Blumenfeld
This talk introduces GraphRAG, an architecture that integrates knowledge graphs with AI agents to enhance data retrieval and reasoning. It demonstrates how to build a knowledge graph using Neo4j and leverage it with tools like LangChain and LangGraph for applications such as talent management and skill analysis. The session covers data modeling, querying with Cypher, incorporating vector search for semantic similarity, and building agents that can interact with the graph.
- Securing Agents with Open Standards — Bobby Tiernay and Kam Sween, Auth0
This talk addresses the critical security challenges arising from increasingly capable AI agents that perform actions in the real world. It emphasizes the need for robust identity and access control mechanisms to prevent issues like secrets in prompts, overly broad scopes, and difficult troubleshooting. The core thesis is that by adopting open standards and implementing proper identity management, developers can build more secure and trustworthy AI agent systems.
- The emerging skillset of wielding coding agents — Beyang Liu, Sourcegraph / Amp
This talk explores the evolving skillset required to effectively utilize coding agents, moving beyond early chatbot paradigms. It argues that the current era demands a new approach to agent interaction, focusing on empowering agents to perform tasks autonomously rather than micromanaging them. The core thesis is that mastering coding agents is a high-ceiling skill that will significantly enhance developer productivity, akin to learning a new programming language or editor.
- Agents, Access, and the Future of Machine Identity — Nick Nisi (WorkOS) + Lizzie Siegle (Cloudflare)
This talk explores the evolving landscape of AI agents, emphasizing the critical need for robust authorization and identity management as these agents act on behalf of users. It highlights how existing human-centric authorization frameworks, like OAuth, are becoming essential for AI agents to securely interact with tools and services. The discussion also touches upon the infrastructure required to support these agents, including persistent storage and edge computing capabilities.
- Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier
Building effective AI agents and the platforms that enable non-technical users to create them is a complex challenge due to the inherent nondeterminism of AI and unpredictable user behavior. The process requires a shift from traditional software development to building a data flywheel that continuously collects feedback, understands usage patterns, and identifies failures to improve the product iteratively. This involves robust instrumentation, strategic feedback collection, and a tiered approach to evaluation.
- Building voice agents with OpenAI — Dominik Kundel, OpenAI
This talk introduces OpenAI's Agents SDK for TypeScript, focusing on building voice agents. It defines agents as systems that accomplish tasks independently using a model, instructions, and tools, all managed by a runtime. The SDK offers features like handoffs, guardrails, streaming, tool calling, and native voice support, aiming to make technology more accessible and information exchange richer through voice.
- Containing Agent Chaos — Solomon Hykes, Dagger
This talk addresses the chaos and complexity arising from the use of coding agents, particularly from the perspective of platform engineers who enable developers. It proposes a new approach to managing agents by leveraging containerization and Git-like versioning to create isolated, customizable, and collaborative development environments. The core idea is to provide agents with dedicated, manageable environments that allow for background work, clear constraints, seamless human intervention, and flexibility in choosing underlying tools and infrastructure.
- Evals 101 — Doug Guthrie, Braintrust
This talk introduces the concept and practical application of evals for AI systems, emphasizing their role in ensuring quality, reliability, and correctness. Evals provide a structured testing framework to move beyond the non-deterministic nature of LLMs, enabling developers to rigorously build and iterate on AI applications. The presentation highlights how evals facilitate a feedback loop between offline development and online production monitoring, ultimately leading to improved AI products.
- Why should anyone care about Evals? — Manu Goyal, Braintrust
This talk argues that evals are crucial for the success and iteration of AI products, moving beyond simple unit tests for AI. Evals provide a simulated laboratory environment, allowing developers to test and refine models extensively before deploying to production. This process significantly speeds up development cycles and increases confidence in shipping AI features.
- Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize
This talk focuses on building scalable LLM evaluation pipelines. It emphasizes that effective evaluation is crucial for developing high-quality AI products and goes beyond simple "LLM as a judge" approaches. The core thesis is that a comprehensive evaluation strategy involves multiple methods, continuous tuning, and integration into the development lifecycle to accelerate iteration and improve AI system performance.
- To the moon! Navigating deep context in legacy code with Augment Agent — Forrest Brazeal, Matt Ball
This talk demonstrates how Augment Agent can be used to navigate and modernize complex, legacy codebases. By leveraging AI to understand historical code, developers can accelerate comprehension, identify potential issues, and even automate modernization tasks. The presentation uses the Apollo 11 guidance computer code as a case study to illustrate these capabilities.
- Serving Voice AI at Scale — Arjun Desai (Cartesia) & Rohit Talluri (AWS)
This talk addresses the challenges and advancements in serving voice AI at scale, focusing on the critical need for low latency and high quality in real-time interactive applications. It introduces state space models (SSMs) as a more efficient alternative to transformers for handling long sequences, enabling faster and more natural voice interactions across various devices. The discussion highlights how these advancements are crucial for enterprise voice AI, customer support, gaming, and content creation.
- Ship it! Building Production Ready Agents — Mike Chambers, AWS
This talk focuses on building production-ready AI agents, moving beyond simple prototypes to scalable, cloud-hosted solutions. It outlines the essential components of an agent, including the model, prompt, loop, history, and tools, and demonstrates how to leverage AWS services like Amazon Bedrock Agents to deploy these agents at cloud scale. The presentation emphasizes practical steps for integrating custom tools via Lambda functions and preparing agents for a production environment.
- Introducing Strands Agents, an Open Source AI Agents SDK — Suman Debnath, AWS
Strands Agents is an open-source SDK designed to simplify the creation of AI agents by minimizing scaffolding. The core philosophy is to leverage the increasing intelligence of AI models, allowing them to handle the reasoning process with minimal explicit prompting. Strands focuses on integrating models and tools, enabling developers to build agentic applications with greater ease.
- Data is Your Differentiator: Building Secure and Tailored AI Systems — Mani Khanuja, AWS
This talk emphasizes that data is the critical differentiator for building secure and tailored AI systems. Generative AI applications require special data treatment beyond standard transformation and loading, focusing on how data interacts with technology and people, and avoiding data silos. The approach to data must be tailored to specific use cases, such as travel agents, employee productivity chatbots, or marketing tools, each with unique data needs and responsibilities.
- How to build world-class AI products — Sarah Sachs (AI lead @ Notion) & Carlos Esteban (Braintrust)
This talk, presented by Sarah Sachs of Notion AI and Carlos Esteban of Braintrust, focuses on the critical role of observability and rigorous evaluation in building world-class AI products. The core thesis is that the quality and scalability of AI products stem directly from robust evaluation processes, which allow teams to iterate effectively and ensure consistent performance beyond simple demos.
- From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva
This talk explores scaling large language models (LLMs) beyond traditional Mixture of Experts (MoE) architectures by introducing the concept of Mixture of Agents (MoA). MoA leverages multiple specialized LLMs, akin to agents, to collectively solve complex problems, aiming for higher intelligence and efficiency compared to monolithic models. The presentation also highlights Cerebras' hardware, designed for extremely fast inference, which enables practical implementation of MoA systems.
- Forget RAG Pipelines—Build Production Ready Agents in 15 Mins: Nina Lopatina, Rajiv Shah, Contextual
This talk introduces Contextual AI's platform for building production-ready Retrieval Augmented Generation (RAG) agents quickly, aiming to simplify the RAG pipeline. The core thesis is that RAG can be treated as a managed service, abstracting away the complexities of building and maintaining individual components like vector databases and LLM training. The platform offers an end-to-end solution for ingesting data, retrieving relevant information, and generating grounded responses, with a focus on enterprise-grade accuracy and modularity.
- Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat
This talk explores the development of real-time voice AI workflows, emphasizing voice as a natural and universal interface for the next generation of AI applications. It highlights the complexities involved in creating seamless voice interactions, from foundational LLMs and real-time APIs like Gemini Live API to orchestration frameworks such as Pipecat and application code. The presentation suggests that while significant progress has been made, many aspects of voice AI are still in early stages of development, with capabilities progressively moving down the technology stack.
- Realtime Conversational Video with Pipecat and Tavus — Chad Bailey and Brian Johnson, Daily & Tavus
This talk introduces Pipecat, an open-source framework for building real-time conversational AI, and Tavus, a platform for creating conversational video replicas. It highlights the three key components needed for real-time AI: models, orchestration, and deployment. The presenters explain how Pipecat acts as an orchestration layer, managing the flow of data (frames) through various processors to enable low-latency audio and video communication.
- Vector Search Benchmark[eting] - Philipp Krenn, Elastic
This talk addresses the common pitfalls and deceptive practices in vector search benchmarking, often referred to as "benchmarketing." The core thesis is that many published benchmarks are misleading due to biased scenario selection, outdated competitor versions, and a lack of focus on crucial factors like data freshness and query relevance. The speaker emphasizes that truly meaningful benchmarks require rigorous, reproducible, and self-conducted evaluations tailored to specific use cases.
- Taming Rogue AI Agents with Observability-Driven Evaluation — Jim Bennett, Galileo
This talk addresses the challenge of ensuring AI agents function reliably by introducing observability-driven evaluation. It highlights that AI's non-deterministic nature makes traditional testing methods insufficient. The core thesis is that by using AI itself to evaluate AI outputs, developers can gain crucial insights into agent performance, identify failures at granular levels, and implement targeted improvements.
- Building agent fleet architectures your CISO doesn't hate — Lou Bichard, Gitpod
This talk discusses the evolution of Gitpod's architecture to support secure development environments, particularly for regulated industries, and how this foundation enables a privacy-first AI agent offering. The core thesis is that simplifying infrastructure, moving away from complex systems like Kubernetes towards native cloud services, leads to better security, lower operational overhead for customers, and a more robust platform for deploying advanced tools like AI agents.
- Don’t get one-shotted: Use AI to test, review, merge, and deploy code — Tomas Reimers, Graphite
The increasing adoption of AI tools in software development, particularly in the inner loop of coding, is creating a bottleneck in the outer loop of review, testing, merging, and deployment. This talk proposes that AI can also be leveraged to streamline and automate these outer loop processes, transforming the entire developer workflow rather than just the IDE.
- Effective agent design patterns in production — Laurie Voss, LlamaIndex
This talk introduces LlamaIndex as a framework for building generative AI applications, with a particular focus on agents. It emphasizes the necessity of Retrieval Augmented Generation (RAG) for agents to effectively process and utilize large, unstructured datasets. The presentation outlines several high-level design patterns that enhance agent performance in production environments.
- Foundry Local: Cutting-Edge AI experiences on device with ONNX Runtime/Olive — Emma Ning, Microsoft
Foundry Local is a Microsoft solution designed to enable developers to build cross-platform AI applications that run directly on user devices. It addresses the need for local AI by providing reasons such as offline access, enhanced privacy and security for sensitive data, cost efficiency for high-volume inference, and real-time latency requirements. The platform leverages ONNX Runtime for accelerated on-device inference across various hardware, integrating with Azure AI Foundry for model management and on-demand downloads.
- [Full Workshop] Vibe Coding at Scale: Customizing AI Assistants for Enterprise Environments
This talk explores "Vibe Coding," a methodology for leveraging AI assistants to accelerate software development by focusing on output rather than the intricacies of code. It outlines a progression from "YOLO vibes" for rapid prototyping to "structured vibes" for maintainability and "spectrum vibes" for scale and reliability, emphasizing trust-building and guardrails for enterprise environments.
- Unlocking AI Powered DevOps Within Your Organization — Jon Peck, GitHub
This talk explores how organizations can effectively integrate AI into their DevOps workflows to enhance efficiency and developer productivity. It emphasizes moving beyond basic AI code generation to leverage AI for planning, security, testing, deployment, and documentation. The presentation also touches on the evolution towards more autonomous AI agents and the importance of governance and safety.
- Vibe Coding at Scale: Customizing AI Assistants for Enterprise Environments - Harald Kirshner,
This talk explores different approaches to AI-assisted coding, moving from rapid, experimental "YOLO vibe coding" to more structured and spec-driven methods suitable for enterprise environments. The core thesis is that by understanding and applying these different "vibes," developers can leverage AI assistants more effectively, leading to increased productivity and maintainability in software development.
- The Agent Awakens: Collaborative Development with Copilot - Christopher Harrison, GitHub
- AI Red Teaming Agent: Azure AI Foundry — Nagkumar Arkalgud & Keiji Kanazawa, Microsoft
This talk introduces the AI Red Teaming Agent, a tool developed within Azure AI Foundry to help AI engineers proactively identify and mitigate risks in their AI systems. It emphasizes that building trustworthy AI is a collaborative effort, akin to building bridges and dams, and that AI engineering requires a systematic approach to testing and iteration. The agent provides a practical way for developers to simulate attacks and evaluate their models' vulnerabilities.
- Collaborating with Agents in your Software Dev Workflow - Jon Peck & Christopher Harrison, Microsoft
This talk explores how developers can effectively collaborate with AI agents, specifically focusing on GitHub Copilot's capabilities within the software development workflow. The core thesis is that understanding and providing proper context is crucial for maximizing the AI pair programmer's utility, moving beyond simple prompt engineering to encompass code readability, project structure, and clear intent. The discussion highlights various Copilot features, from basic code completion to advanced agent modes, and emphasizes best practices for integrating these tools into daily development tasks.
- Agentic Excellence: Mastering AI Agent Evals w/ Azure AI Evaluation SDK — Cedric Vidal, Microsoft
This talk focuses on the critical process of evaluating AI agents, emphasizing a shift from ad-hoc testing to a more methodical approach. It highlights that effective evaluation should begin at the earliest stages of AI development, not as an afterthought. The presentation introduces tools and frameworks, particularly the Azure AI Evaluation SDK, to ensure AI agents behave correctly and safely as they gain more independence.
- Building Code First AI Agents with Azure AI Agent Service — Cedric Vidal, Microsoft
This talk explores building AI agents using Azure AI Agent Service, focusing on a practical approach for developers. It defines an agent as a semi-autonomous software that pursues a goal by reasoning, integrating with data, and acting on the world. The presentation demonstrates how to leverage Azure AI Agent Service to simplify agent development by managing state, context, and tool integration in the cloud, enabling the creation of applications that can dynamically generate queries, user interfaces, and visualizations.
- How fast are LLM inference engines anyway? — Charles Frye, Modal
This talk explores the performance of open-source LLM inference engines, highlighting how recent advancements in model quality and inference software have made self-hosting viable. It presents benchmarking data to help engineers understand and optimize LLM performance for various use cases, emphasizing the trade-offs between different configurations and workloads.
- RAG in 2025: State of the Art and the Road Forward — Tengyu Ma, MongoDB (acq. Voyage AI)
This talk explores Retrieval Augmented Generation (RAG) as a key technology for enabling large language models (LLMs) to access and utilize proprietary enterprise data. The speaker argues that RAG offers a more efficient, cost-effective, and manageable approach compared to fine-tuning or relying solely on long context windows. The presentation also touches upon the evolution of RAG, advancements in embedding models, and future directions for improving retrieval accuracy and simplifying workflows.
- The State of AI Powered Search and Retrieval — Frank Liu, MongoDB (prev Voyage AI)
This talk explores the evolution and future of AI-powered search and retrieval, moving beyond traditional keyword matching to understand user intent and conceptual relationships. It highlights how AI search systems can provide more grounded and relevant responses, particularly in applications like Retrieval-Augmented Generation (RAG), by leveraging embeddings and advanced techniques. The discussion also touches upon the increasing importance of multimodality and agentic capabilities in shaping the next generation of search technologies.
- Architecting Agent Memory: Principles, Patterns, and Best Practices — Richmond Alake, MongoDB
This talk explores the critical role of memory in developing advanced AI agents, moving beyond stateless applications to create more believable, capable, and reliable systems. It posits that memory is fundamental to mimicking human intelligence and enabling agents to be reflective, interactive, proactive, and autonomous. The discussion emphasizes memory management as a core task for AI engineers, focusing on principles and practical patterns for integrating memory into agentic systems.
- Building Multimodal AI Agents From Scratch — Apoorva Joshi, MongoDB
This talk introduces the concept of building multimodal AI agents from scratch, focusing on integrating text and image processing capabilities. It explains the evolution from simple prompting and RAG to AI agents, highlighting agents' suitability for complex, multi-step tasks requiring reasoning and action. The session details the core components of an agent—perception, planning/reasoning, tools, and memory—and demonstrates how to construct a multimodal agent capable of answering questions about documents containing both text and images, and analyzing charts or diagrams.
- Why Your Agent’s Brain Needs a Playbook: Practical Wins from Using Ontologies - Jesús Barrasa, Neo4j
This talk explores the integration of knowledge graphs and Large Language Models (LLMs) for building robust AI applications, specifically focusing on the Graph RAG architecture. It highlights how ontologies, as formal, implementation-agnostic schemas, can significantly enhance knowledge graph creation and retrieval strategies, leading to more grounded and accurate AI responses. The core thesis is that a model-driven approach using ontologies provides practical wins by improving data quality and enabling dynamic retriever behavior.
- Memory Masterclass: Make Your AI Agents Remember What They Do! — Mark Bain, AIUS
This talk explores the critical role of memory in AI agents, positing that true AI memory encompasses all data, code, algorithms, and hardware, along with any causal changes affecting them. It draws parallels between the principles governing Large Language Models (LLMs), neuroscience, and mathematics, suggesting that asymmetries are necessary for existence and that preserving causal links through relationships is key to solving issues like hallucinations and optimizing hypothesis generation. The presentation highlights the potential of graph databases and agentic systems for building more robust and context-aware AI.
- Graph Intelligence: Enhance Reasoning and Retrieval Using Graph Analytics - Alison & Andreas, Neo4j
This talk explores how graph data science can enhance reasoning and retrieval in AI applications, particularly within the context of Retrieval Augmented Generation (RAG). It emphasizes that graphs provide a way to understand data relationships beyond simple vector similarity, enabling more comprehensive and context-aware responses. The session introduces graph concepts and demonstrates how to leverage graph analytics to manage and improve data quality for AI systems.
- GraphRAG methods to create optimized LLM context windows for Retrieval — Jonathan Larson, Microsoft
This talk introduces GraphRAG, a method for optimizing Large Language Model (LLM) context windows using graph structures for retrieval. The core thesis is that LLM memory with structure is a key enabler for building effective AI applications, and when paired with agents, this combination offers even greater power. The presentation demonstrates GraphRAG's application in code understanding and feature development, alongside the announcement of Benchmark QED, an open-source tool for evaluating LLM systems.
- Agentic GraphRAG: Simplifying Retrieval Across Structured & Unstructured Data — Zach Blumenfeld
This talk introduces Agentic GraphRAG, a method for simplifying data retrieval across both structured and unstructured sources. It proposes using a knowledge graph as an intermediary layer to enhance agentic workflows. By modeling data within a graph, agents can more accurately decompose complex questions, pull relevant information, and perform analytical tasks that go beyond simple semantic search.
- Revenue Engineering: How to Price (and Reprice) Your AI Product — Kshitij Grover, Orb
This talk explores the complexities of pricing AI products, emphasizing that pricing is a form of friction that must be carefully managed to align with product value and audience needs. It moves beyond traditional pricing models to discuss AI-native considerations like predictability, speed of value demonstration, and rapidly changing cost structures. The core thesis is that effective AI product pricing requires a deep understanding of the target audience, value delivery mechanisms, and flexible margin structures, allowing for continuous experimentation and adaptation.
- \"Data readiness\" is a Myth: Reliable AI with an Agentic Semantic Layer — Anushrut Gupta, PromptQL
The talk argues that the concept of "data readiness" for AI is a myth, as data is rarely perfect and constantly changing. Instead of striving for pristine data, the focus should be on building AI systems that can reliably work with messy, evolving data. This is achieved by creating an agentic semantic layer that learns and adapts to the specific business domain and its nuances over time, much like an experienced human analyst.
- Building Agentic Applications w/ Heroku Managed Inference and Agents — Julián Duque & Anush Dsouza
This talk introduces Heroku's Managed Inference and Agents service, designed to simplify the development and deployment of agentic AI applications. The service allows developers to integrate AI models directly within their application's infrastructure, enhancing security and control. It provides primitives for inference, model context protocol (MCP) integration, and a managed PostgreSQL database with PGVector for embeddings, enabling the creation of sophisticated AI-powered features.
- Events are the Wrong Abstraction for Your AI Agents - Mason Egger, Temporal.io
This talk argues that event-driven architecture (EDA) is the wrong abstraction for AI agents, leading to complex, tightly coupled systems. Instead, the speaker proposes durable execution as a superior model. Durable execution, exemplified by Temporal.io, offers crash-proof processing by automatically preserving application state, virtualizing execution across machines, and being time and hardware agnostic. This shift allows developers to focus on core business logic rather than managing the intricacies of event handling and system failures.
- Prompt Engineering is Dead — Nir Gazit, Traceloop
The talk argues that traditional prompt engineering is ineffective and proposes an automated approach to improving AI model responses. The speaker shares a personal experience of enhancing a Retrieval Augmented Generation (RAG) chatbot by building an agent that iteratively refines prompts based on an evaluation system, rather than manual prompt tuning. This method aims to achieve significant improvements without direct, manual prompt manipulation.
- The Eyes Are The (Context) Window to The Soul: How Windsurf Gets to Know You — Sam Fertig, Windsurf
This talk explores how AI coding tools can better understand and assist developers by deeply understanding context. The core thesis is that generating code is no longer the primary challenge; instead, the difficulty lies in creating code that integrates seamlessly into existing, complex codebases, adheres to specific standards, and aligns with individual developer preferences. This is achieved by focusing on the "what" and "how much" of context, rather than simply increasing context window sizes.
- Mastering Engineering Flow with Windsurf - Eashan Sinha, Windsurf
This talk introduces Windsurf's approach to enhancing the developer experience through agentic IDEs, focusing on the concept of "AI flows." The core thesis is that by treating AI coding assistants as collaborative teammates rather than separate tools, developers can achieve a more seamless and productive engineering flow. Windsurf's Cascade agent aims to achieve this by deeply understanding user intent and codebase context, moving beyond simple autocomplete or autonomous agents.
- (possible dupe but better sound) What does Enterprise Ready MCP mean? — Tobin South, WorkOS
This talk explores what it means for Model Communication Protocol (MCP) to be enterprise-ready, moving beyond basic tool-use capabilities to address the complexities of production environments. It highlights the evolution from simple chatbot interactions to sophisticated AI agents that require robust security, management, and scalability, particularly when integrating with internal enterprise systems. The core thesis is that while MCP offers a standardized way for AI to interact with external resources, achieving enterprise readiness involves overcoming significant challenges in authentication, authorization, and operational management.
- CI in the Era of AI: From Unit Tests to Stochastic Evals — Nathan Sobo, Zed
This talk explores the challenges and strategies for testing AI-enabled features, particularly agentic editing in code editors. It emphasizes the shift from deterministic testing to embracing stochastic evaluations when LLMs are involved. The core thesis is that rigorous, empirical software development practices, adapted for the non-deterministic nature of AI, are crucial for building reliable AI-powered products.
- Fun stories from building OpenRouter and where all this is going - Alex Atallah, OpenRouter
This talk explores the founding story of OpenRouter, initially conceived as an experiment to address the burgeoning AI inference market. It delves into the investigation of whether this market would be dominated by a single player or foster a diverse ecosystem. The discussion highlights the evolution from a model exploration platform to a comprehensive marketplace, emphasizing the challenges and innovations in aggregating and standardizing diverse AI models.
- Building AI Agents that actually automate Knowledge Work - Jerry Liu, LlamaIndex
This talk explores how AI agents can automate knowledge work, moving beyond simple chatbots to handle complex tasks involving unstructured data. It outlines a framework for building these agents, emphasizing the need for robust toolkits and well-defined agent architectures to process and act upon diverse data formats like documents and spreadsheets. The core thesis is that by combining advanced document understanding with flexible agent design patterns, AI can significantly enhance efficiency in knowledge-intensive roles.
- RFT, DPO, SFT: Fine-tuning with OpenAI — Ilan Bigio, OpenAI
This talk explores various fine-tuning techniques for OpenAI models, including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Fine-Tuning (RFT). The presenter, Ilan Bigio from OpenAI's developer experience team, emphasizes that fine-tuning is a specialized tool for optimizing models beyond what prompt engineering can achieve, particularly for specific domains or behaviors. The discussion covers the data requirements, use cases, and limitations of each method, offering practical examples and best practices.
- Windsurf everywhere, doing everything, all at once - Kevin Hou, Windsurf
This talk introduces Windsurf, an AI-powered editor designed to revolutionize software creation by extending beyond traditional code completion. The core thesis is that AI should operate on a shared timeline with human developers, ingesting context from all sources a developer uses and taking actions across various platforms, not just within the IDE. This approach aims to enable AI to perform nearly all tasks a human software engineer would, moving towards a future where AI handles 99% of the workflow with human oversight only for final approval.
- Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP - Dan Mason
This talk presents a case study and deep dive into building telemedicine support agents using LangGraph and MCP. The core thesis is that LLM-powered agents can significantly enhance flexibility and capability in healthcare support systems, moving beyond traditional software limitations. The system aims to provide scalable, automated patient support while maintaining human oversight for complex or sensitive situations.
- Veo 3 for Developers — Paige Bailey, Google DeepMind
This talk introduces Google DeepMind's latest generative media models, focusing on V3 for video and audio generation, Imagine 4 for static images, and LIA 2 for music. The presentation highlights how these tools can revolutionize content creation, advertising, and user experiences by enabling the generation of novel content and offering enhanced creative control. The models are designed to improve stylistic and contextual consistency, making them powerful tools for developers and creators.
- Building Agents with Amazon Nova Act and MCP - Du'An Lightfoot, Amazon (Full Workshop)
This workshop introduces building intelligent autonomous AI systems using Amazon Nova Act and the Modern Communication Protocol (MCP). It focuses on a do-it-yourself approach, enabling developers to create agentic systems that can plan, act, and reason to achieve objectives. The session highlights how these agents can leverage tools, knowledge bases, and LLMs to tackle complex tasks, with a particular emphasis on browser automation and integrating various AI components.
- The Web Browser Is All You Need - Paul Klein IV, Browserbase
This talk argues that the web browser is the essential bridge for AI agents to interact with the vast majority of the internet, which lacks modern APIs. It posits that browsers, specifically headless browser MCP servers, are the key to unlocking AI's potential for interacting with legacy websites and services. The presentation explores different approaches to web agents and browser tools, emphasizing the browser's role as a universal integration point.
- Building Protected MCP Servers — Den Delimarsky and Julia Kasper, MCP Steering Committee & Microsoft
This talk addresses the critical need for security in Multi-Call Protocol (MCP) servers, particularly for remote instances. It introduces a new draft specification that simplifies authorization by separating the MCP server's role from that of an authorization server. This approach aims to reduce the burden on developers by allowing them to leverage existing OAuth 2.0 libraries and identity providers, rather than implementing complex authorization logic themselves.
- The State of MCP observability: Observable.tools — Alex Volkov and Benjamin Eckel, W&B and Dylibso
The increasing prevalence of Machine Communication Protocol (MCP) agents is creating an observability blind spot for developers. As agents utilize more tools via MCP, understanding their end-to-end execution becomes challenging. This talk introduces observable.tools as a manifesto to drive a community conversation around standardized, vendor-neutral MCP observability, advocating for the adoption of OpenTelemetry (otel) to address this growing issue.
- The Geopolitics of AI Infrastructure - Dylan Patel, SemiAnalysis
This talk examines the geopolitical landscape of AI infrastructure, focusing on the capabilities of China and the Middle East in contrast to US limitations. It highlights how China, despite sanctions, is advancing its AI chip manufacturing and deployment through innovative engineering and supply chain circumvention. Simultaneously, the Middle East is emerging as a significant hub for AI infrastructure development, attracting substantial investment and GPU resources, partly due to power constraints in the US.
- Remote MCPs: What we learned from shipping — John Welsh, Anthropic
This talk discusses the challenges and solutions for implementing and scaling Message Communication Protocol (MCP) clients within a large organization. It highlights how the rapid evolution of AI models capable of tool calling led to integration chaos, with duplicated functionality and inconsistent interfaces across services. The solution proposed is to standardize on MCP for providing model context, treating it as a plumbing layer for integrations rather than a competitive differentiator.
- MCP: Origins and Requests For Startups — Theodora Chu, Model Context Protocol PM, Anthropic
The Model Context Protocol (MCP) is an open-source, standardized protocol designed to give AI models agency by allowing them to interact with the outside world. Originating from the need to copy context from external sources into LLM context windows, MCP aims to enable models to reach a new level of usefulness and intelligence by facilitating tool calling and broader interaction capabilities. The protocol prioritizes server simplicity and encourages community contributions to evolve its standards and utility.
- How to Build Trustworthy AI — Allie Howe
This talk addresses the critical need for building trustworthy AI systems, emphasizing that responsibility ultimately lies with the user and developer. It defines trustworthy AI as a combination of AI security (protecting AI from external harm) and AI safety (preventing AI from harming the world). The presentation advocates for a shift from traditional DevSecOps to MLSecOps, highlighting the importance of integrating security and safety practices throughout the AI development lifecycle, from build to runtime.
- Exposing Agents as MCP servers with mcp-agent: Sarmad Qadri
This talk introduces the Model Context Protocol (MCP) as a standardized interface for connecting Large Language Models (LLMs) to external tools and resources, aiming to simplify and enhance agent development. It posits that 2025 will be the year agents hit mass production, facilitated by advancements in LLM reasoning capabilities, the widespread adoption of MCP, and simpler agent architectures. The presentation also explores modeling agents as asynchronous workflows and exposing them as MCP servers for greater composability and scalability.
- Supercharging developer workflow with Amazon Q Developer - Vikash Agrawal
This talk demonstrates how Amazon Q Developer, an AI coding assistant, can be integrated throughout the software development lifecycle (SDLC). It highlights Q Developer's capabilities in IDEs, CLIs, and GitHub to assist with planning, coding, testing, documentation, deployment, and even operational debugging, aiming to supercharge developer workflows.
- Just do it. (let your tools think for themselves) - Robert Chandler
This talk addresses the common challenges of unreliability, slowness, and cost associated with current AI agents that interact with external tools. The core thesis is that instead of creating simple, low-level wrappers around APIs, tools should be imbued with more agency, becoming specialized agents themselves. This approach blurs the line between tools and agents, enabling agents to offload complex tasks to more capable, specialized entities for improved performance and reliability.
- Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents - Aparna Dhinakaran
This talk addresses the challenges of building and iterating on AI agents, particularly in evaluating their performance and identifying bottlenecks. It emphasizes the need for systematic evaluation beyond manual inspection of a few examples. The core thesis is that a self-improving stack for AI agents requires not only improving the agent's prompts and models but also continuously refining the evaluation methods themselves.
- MCPs are Boring (or: Why we are losing the Sparkle of LLMs) - Manuel Odendahl
This talk argues that the current focus on Multi-Call Protocol (MCP) and tool calling for LLMs is limiting their potential. The speaker contends that LLMs are powerful code generators and language producers, capable of much more than simply calling predefined functions with fixed schemas. The core thesis is that by treating LLMs as sophisticated code-generation engines and embracing recursive creation, developers can unlock greater flexibility and power, moving beyond the "boring" limitations of current tool-calling paradigms.
- The Many Ends of Programming - Ray Myers
This talk explores the evolving landscape of programming in the age of AI, moving beyond hype to discuss practical scenarios and their implications for software development. It argues that the future of programming isn't a single predetermined outcome but a spectrum of possibilities, emphasizing the need for empathy, careful consideration, and active participation in shaping this future. The core thesis is that while AI tools are rapidly advancing, human engineers play a crucial role in guiding their integration and ensuring positive outcomes.
- Why Bolt.new Won and Most DevTools AI Pivots Failed - Victoria Melnikova
This talk argues that most developer tool AI pivots fail because they incorrectly implement AI features. Instead of simply adding AI to existing workflows or trying to compete with large AI providers, successful products leverage their unique competitive advantages. By identifying what makes a product distinct and then exploring how AI can amplify that specific advantage, companies can create new categories and reinvent user experiences, leading to significant growth.
- Beyond Conversation: Why Documents Transform Natural Language into Code - Filip Kozera
This talk argues that document-based interfaces are superior to chat-based interfaces for complex tasks involving AI. Chat interfaces suffer from context pollution, lack structured iteration, and offer poor version control, leading to degraded model performance. Documents, conversely, provide a mechanism for forced clarity and structured communication, enabling more robust specification of complex systems and paving the way for background agents that can perform tasks autonomously.
- The 4 Patterns of AI Native Development — Patrick Debois
This talk outlines four key patterns emerging in AI-native development, shifting the developer's role beyond traditional coding. As AI tools evolve from simple code generation to complex agent teams, developers are transitioning into roles like managers, specifiers, discoverers, and knowledge curators. This evolution signifies a move towards more strategic and less purely executional work, fundamentally changing the software development lifecycle.
- Breaking the Chain: Agent Continuations for Resumable AI Workflows - Greg Benson
This talk introduces agent continuations, a new mechanism designed to address key challenges in deploying AI agents in production. It enables agents to pause their execution, save their complete state, and resume later, facilitating human approval workflows and robust error handling for long-running processes. This approach aims to make AI agents more reliable and manageable in complex, distributed environments.
- Are MCPs Overhyped? A Rant about MCPs — Henry Mao, Smithery
This talk critically examines the Model Context Protocol (MCP) ecosystem, arguing that despite initial excitement and the promise of standardizing AI agent interactions with services, significant challenges remain. The speaker contends that the current state of MCPs is fragmented and difficult to use, hindering the practical application of autonomous agents. The core thesis is that while the potential for AI agents is vast, the MCP infrastructure is not yet mature enough to fully realize this potential, necessitating further development and standardization.
- Why the Best AI Agents Are Built Without Frameworks (Primitives over Frameworks) — Ahmad Awais, CHAI
This talk argues that building AI agents on primitives rather than frameworks leads to more production-ready and scalable solutions. Frameworks are often bloated, slow, and introduce unnecessary abstractions. Instead, developers should leverage composable, low-level primitives, similar to how cloud services like Amazon S3 function, to create flexible and efficient AI agents.
- GPU-less, Trust-less, Limit-less: Reimagining the Confidential AI Cloud - Mike Bursell
This talk introduces confidential AI as a solution to trust issues in AI development and deployment. It explains how confidential computing, utilizing Trusted Execution Environments (TEs), protects data and models during processing, even from system administrators or hardware providers. This technology enables secure collaboration, monetization, and the use of sensitive data for AI tasks across various industries.
- Grounded Reasoning Systems for Cloud Architecture - Iman Makaremi
This talk explores the development of grounded reasoning systems for cloud architecture, emphasizing the need for AI systems that can understand, debate, and justify architectural decisions rather than simply automate tasks. It highlights the complexity of cloud architecture, the cognitive aspects of architectural decision-making, and the challenges of integrating textual requirements with graph-based architectural data. The core thesis is that multi-agent orchestration is key to building AI copilots capable of complex reasoning in this domain.
- The Agent Native Company — Rick Blalock, Agentuity
This talk introduces the concept of an "agent-native" or "AI-native" company, distinguishing it from merely AI-enhanced businesses. An agent-native company is built from the ground up with AI agents as a core component, fundamentally altering how work is done, teams are structured, and products are developed. This paradigm shift moves beyond AI as an add-on to AI as the central engine driving operations, culture, and product innovation.
- The Voice-First AI Overlay: Designing Conversational Co-Pilots - Gregory Bruss
This talk explores the concept of a voice-first AI overlay designed to enhance human-to-human conversations by providing real-time assistance without directly participating as a third speaker. The core idea is to leverage advancements in AI agent capabilities and voice technology to keep humans informed and on track during interactions, using voice as the most natural interface.
- Arrakis: How To Build An AI Sandbox From Scratch - Abhishek Bhardwaj, OpenAI
Arrakis is an open-source, self-hosted service for creating and managing AI sandboxes, designed for secure code execution and computer use by AI agents. It leverages microVMs to provide isolated environments, enabling AI models to safely utilize tools like code execution and search. The system emphasizes security, speed, and the ability for agents to backtrack and replan using snapshots, facilitating more complex task completion.
- 7 Habits of Highly Effective Generative AI Evaluations - Justin Muller
This talk emphasizes that robust evaluations are the most critical, yet often overlooked, component for successfully scaling generative AI workloads. The speaker argues that evaluations are not merely for measuring quality but are primarily a tool for discovering problems within AI systems. Implementing a strong evaluation framework is presented as the key differentiator between a project that remains a science experiment and one that achieves production-level success and scalability.
- Agents reported thousands of bugs, how many were real? - Ian Butler and Nick Gregory
This talk investigates the effectiveness of AI agents in identifying and fixing bugs within software development, moving beyond traditional feature development benchmarks. The presenters introduce a new benchmark designed to evaluate agents on maintenance tasks, highlighting that while agents can often patch simple bugs, their ability to comprehensively detect them, especially complex ones, is still in its early stages. The research suggests current agents struggle with holistic code evaluation and deep reasoning, leading to missed bugs and high false positive rates.
- The Coherence Trap: Why LLMs Feel Smart (But Aren’t Thinking) - Travis Frisinger
This talk argues that Large Language Models (LLMs) excel due to coherence, not intelligence. While LLMs can produce outputs that feel remarkably insightful and intelligent, they lack genuine understanding, intent, or desire. The speaker proposes that coherence is a system property, not a cognitive one, and explains how LLMs construct meaning on demand through pattern alignment within a high-dimensional latent space, rather than retrieving stored knowledge.
- The Knowledge Graph Mullet: Trimming GraphRAG Complexity - William Lyon
This talk introduces the "Knowledge Graph Mullet," a hybrid approach combining property graph and RDF triple concepts for more versatile knowledge graph management. It advocates for using a property graph model for data interaction and querying, while leveraging the scalability of RDF triples for the underlying storage. The presentation demonstrates this approach using Dgraph, an open-source graph database, and explores its application in building sophisticated Graph RAG (Retrieval Augmented Generation) workflows and AI agents.
- Building Reliable Support Agents Using the Effect Typescript Library - Michael Fester
This talk introduces the Effect TypeScript library as a robust solution for building reliable AI-powered customer support platforms. The core thesis is that while TypeScript provides a good foundation, Effect offers essential tools for managing the complexities of unreliable APIs, non-deterministic LLM outputs, and long-running workflows, thereby enhancing system stability, testability, and maintainability at scale.
- The Robots are coming for your job, and that's okay - Elmer Thomas and Maria Bermudez
This talk explores how AI agents can be used to enhance, rather than replace, human workflows, particularly within a small documentation team. The core thesis is that by automating repetitive, rule-based tasks, AI can free up human engineers to focus on higher-level judgment, clarity, and creativity, ultimately boosting productivity and reducing burnout.
- Blender MCP and The Future Of Creative Tools - Siddharth Ahuja
This talk explores the Blender MCP, an open-source project that enables Large Language Models (LLMs) to interact with and control Blender, a complex 3D creation tool. The core thesis is that by leveraging the MCP protocol and Blender's scripting capabilities, LLMs can significantly lower the barrier to entry for 3D content creation, allowing users to generate intricate scenes and assets through simple text prompts. This approach democratizes creative tools and points towards a future where LLMs act as orchestrators across various creative software.
- Agentic Enterprise - What your CEO must know about AI - Hubert Misztela
This talk explores the concept of agentic enterprise, positing that AI agents could fundamentally reshape organizations within three years. It defines AI agents as autonomous, LLM-based applications capable of planning, tool usage, and dynamic adaptation. The presentation emphasizes that digital assets should evolve to serve as tools or agents themselves, enabling capabilities like multi-step reasoning, adaptability, and computer vision for automation.
- The End of Awkward AI Transcriptions - Travis Bartley and Myungjong Kim
This talk details Nvidia's approach to developing enterprise-level speech AI models, focusing on robustness, coverage, personalization, and deployment efficiency. The core thesis is that a variety of specialized models, rather than a single monolithic solution, best meets diverse customer needs for conversational AI, emphasizing low latency and high efficiency for embedded devices.
- How agents broke app-level infrastructure - Evan Boyle
This talk addresses the challenges of building reliable AI applications, particularly focusing on the compute layer. Traditional web infrastructure is ill-suited for the long-running, non-deterministic nature of AI workflows, which often involve extensive data ingestion and complex agentic processes. The core thesis is that current infrastructure limitations lead to unreliable user experiences and hinder rapid experimentation, necessitating new architectural approaches.
- RAG Evaluation Is Broken! Here's Why (And How to Fix It) - Yuval Belfer and Niv Granot
This talk addresses the shortcomings of current Retrieval Augmented Generation (RAG) evaluation methods, arguing that they are fundamentally broken. The presenters contend that most benchmarks rely on simple, local questions with easily identifiable answers within specific text chunks, which does not reflect real-world data complexity. This leads to a cycle of optimizing for flawed benchmarks, resulting in RAG systems that perform poorly when deployed with actual user data.
- The Demo I Wish I'd Had: OpenAI's Agents SDK... serverless! - Brook Riggio
This talk presents a preferred architecture for building full-stack AI applications in a serverless environment. The core thesis is that by combining specific tools like Next.js, OpenAI's Agents SDK, and Ingest, developers can create resilient, agent-powered applications without managing complex infrastructure. The approach emphasizes integrating AI workflows directly into client applications for a seamless user experience.
- Will Agent evaluation via MCP Stabilize Agent Networks? - Ari Heljakka
This talk explores how the Model Contest Protocol (MCP) can be used to stabilize AI agents and agent networks. The core idea is that by systematically evaluating agent behavior and providing feedback, agents can learn to improve their performance and consistency, especially when tackling complex problems. This approach aims to create more controllable, transparent, and self-correcting agent systems.
- MCP Agent Fine tuning Workshop - Ronan McGovern
This workshop details the process of fine-tuning an AI agent that utilizes the Model Context Protocol (MCP) for tool access. It covers generating high-quality reasoning traces from agent interactions with tools, saving these multi-turn conversations, and using them to fine-tune a language model, specifically demonstrating with a Quen model. The goal is to improve the agent's performance by training it on its own successful interactions.
- From PM at Stripe to Building an AI startup, a recent founder's journey - Mounir Mouawad
This talk chronicles a founder's transition from a product role at a large tech company to establishing an AI startup. It highlights the unique challenges of identifying user problems in the nascent AI space, where problems are emergent rather than clearly defined. The presentation emphasizes the need for rapid iteration, hypothesis-driven development, and the creation of user narratives to navigate this evolving landscape.
- My AI Thinks I'm Eating My Feelings (and Other Nutritional Insights) - Rami Alhamad
This talk introduces Alma, an AI-powered nutrition companion designed to simplify healthy eating. The core thesis is that current nutrition tracking methods are overly complex, providing little value for the effort invested. Alma aims to solve this by making tracking natural and easy, building rich user context, and proactively connecting users with suitable food options.
- Rust is the language of the AGI - Michael Yuan
This talk argues that Rust is the ideal language for Artificial General Intelligence (AGI) development, positioning AI as the primary coder and humans as assistants. The core thesis is that Rust's rigorous compiler and strong type system, while challenging for humans, provide an excellent feedback loop for AI code generation, leading to more correct and efficient code. The Rust Coder project aims to facilitate this by making Rust easier for AI to generate and for humans to work with.
- Real AI Agents Need Planning, Not Just Prompting - Yuval Belfer
This talk argues that current large language models (LLMs), despite advancements, still struggle with complex instruction following. The core thesis is that AI agents require robust planning capabilities beyond simple prompting to effectively tackle intricate tasks, emphasizing the need for a lookahead strategy rather than step-by-step execution.
- The Current State of Browser Agents - Jerry Wu and Wyatt Marshall
This talk explores the current capabilities and limitations of browser agents, which are AI systems designed to control web browsers and perform tasks on behalf of users. The discussion covers the fundamental loop of observation, reasoning, and action that powers these agents, their common use cases like web scraping and form filling, and an evaluation of their performance based on a newly released benchmark dataset. The presentation highlights significant performance gaps between read tasks (information retrieval) and write tasks (state modification), and discusses the underlying reasons for these discrepancies.
- Text-to-Speech Data Preparation and Fine-tuning Workshop - Ronan McGovern
This workshop details the process of preparing data and fine-tuning a text-to-speech (TTS) model, specifically Sesame's CSM 1B model, to produce speech that mimics a target voice. It covers extracting audio from sources like YouTube, transcribing it, and formatting it into a dataset suitable for training. The process leverages the Unsloth library for efficient fine-tuning and demonstrates how to evaluate the model's performance before and after the fine-tuning process.
- Invisible Users, Invisible Interfaces: Accelerating Design Iteration with AI Simulation - Alex Liss
This talk proposes using AI simulation to accelerate design iteration by creating "intelligent twins" that act as invisible users. This approach aims to overcome the current AI trust gap, which stems from poorly implemented AI features, by enabling designers to identify and address user pain points more effectively. The methodology draws parallels to pilot training through simulation, suggesting a shift from traditional data collection to active AI-driven feedback loops within the design process.
- The RAG Stack We Landed On After 37 Fails - Jonathan Fernandes
This talk details the iterative process of building a Retrieval Augmented Generation (RAG) system, highlighting lessons learned from 37 failed attempts. It emphasizes practical choices for different components of a RAG stack, distinguishing between prototyping and production environments. The core thesis is that careful selection and integration of components like orchestration, embedding models, vector databases, and language models are crucial for effective RAG implementation.
- Luminal - Search-Based Deep Learning Compilers - Joe Fioti
Luminal is a deep learning library that aims for radical simplification through search-based compilation. Instead of the complex, multi-million line codebases of traditional libraries like PyTorch, Luminal represents models as simple directed acyclic graphs of a minimal set of core operations. This simplified representation allows for a compiler that uses search to discover highly optimized kernels, including complex ones like FlashAttention, which were previously difficult to implement.
- The Benchmarks Game: Why It's Rigged and How You Can (Really) Win - Darius Emrani
This talk argues that current AI benchmarks are fundamentally flawed and often manipulated, leading to misleading performance claims. The speaker contends that the immense financial and market value tied to benchmark scores incentivizes companies to game the system through various deceptive practices. Ultimately, the presentation advocates for building custom, use-case-specific evaluations rather than relying on public benchmarks.
- Stop Ordering AI Takeout A Cookbook for Winning When You Build In House - Jan Siml
This talk argues against the common practice of adopting complex, cutting-edge AI solutions for internal business needs, likening it to ordering expensive takeout when a simpler, in-house meal would suffice. The core thesis is that building AI solutions internally, when focused on specific, high-value workflows and leveraging existing data, can deliver significant revenue and operational improvements more effectively than off-the-shelf, overly complex systems.
- Buy Now, Maybe Pay Later: Dealing with Prompt-Tax While Staying at the Frontier - Andrew Thomspson
This talk addresses the challenges of building and shipping AI agentic products at the cutting edge of model development. It introduces the concept of "prompt tax," which refers to the unintended consequences and risks associated with integrating rapidly advancing AI models into applications. The core thesis is that to maximize opportunities from new AI capabilities, developers must ship products quickly, embracing the inherent uncertainties and managing them through iterative feedback and progressive rollout strategies.
- Cognitive Shield Real Time Real Smart - Rachna Srivastava
This talk introduces Cognitive Shield, a platform designed to combat the rising tide of AI-driven fraud. It highlights how advanced AI and machine learning can be leveraged not only to detect but also to prevent sophisticated scams, synthetic identities, and deepfake threats that bypass traditional security measures. The core thesis is that the same AI technologies used to perpetrate fraud can be repurposed to build robust defenses, thereby rebuilding and strengthening trust in the digital age.
- Unlocking Africa's Potential with AI — Thabang Ledwaba
This talk explores the transformative potential of Artificial Intelligence for Africa, arguing that the continent possesses immense creativity and a unique perspective that, when combined with AI, can lead to groundbreaking innovations. It challenges the perception of Africa as merely a consumer of technology and advocates for a shift towards becoming a producer and active player on the global stage. The core thesis is that by reimagining how Africa is perceived and by leveraging AI, the continent can overcome its challenges and unlock its full potential.
- Analyzing 10,000 Sales Calls With AI In 2 Weeks — Charlie Guo
This talk details how an AI engineer analyzed 10,000 sales calls in two weeks, a task previously requiring extensive manual effort or teams over months. The core thesis is that modern large language models, when integrated with sound engineering practices, can transform massive amounts of unstructured data into actionable insights, turning potential liabilities into valuable assets. The project highlights the feasibility of single engineers tackling complex data analysis challenges that were once insurmountable.
- Letting AI Interface with your App with MCP — Kent C Dodds
This talk introduces Model Context Protocol (MCP) as a standardized way for AI assistants to interface with various tools and services, aiming to bridge the gap between current AI capabilities and the desire for a universal assistant like Jarvis. It posits that the primary barrier to advanced AI assistants has been the difficulty of building numerous integrations, and MCP offers a solution by establishing a common protocol that any AI assistant can use to interact with any service provider.
- ChatGPT is poorly designed. So I fixed it
This talk argues that ChatGPT's user interface is poorly designed, leading to a confusing experience despite its rapid growth. The presenter demonstrates issues with its voice and text interaction, highlighting how separate functionalities feel disconnected. The core thesis is that by integrating multimodal capabilities and intelligently routing requests to appropriate models, the user experience can be significantly improved using off-the-shelf tools.
- Effective AI Agents Need Data Flywheels, Not The Next Biggest LLM – Sylendran Arunagiri, NVIDIA
This talk argues that building effective AI agents relies on data flywheels rather than simply using the largest available language models. A data flywheel is a continuous cycle of data processing, model customization, evaluation, and safety guardrailing. This process allows agents to refine their performance over time by learning from user feedback and production data, ultimately enabling the use of smaller, more cost-effective models without sacrificing accuracy.
- open-rag-eval: RAG Evaluation without \"golden\" answers — Ofer Mendelevitch, Vectara
This talk introduces open-rag-eval, an open-source project designed for scalable Retrieval-Augmented Generation (RAG) evaluation. It addresses the common challenge of RAG evaluation requiring "golden answers" or "golden chunks," which is often impractical and non-scalable. The project, developed in collaboration with the University of Waterloo, offers a research-backed approach to evaluate RAG pipelines without relying on pre-defined correct answers.
- Designing AI To Scale Human Thought — Jun Yu Tan, Tusk
This talk proposes a paradigm shift in AI interface design, moving from pure automation to augmentation that enhances human capabilities. Instead of AI systems attempting to automate complex tasks suboptimally, the focus should be on using AI to help humans produce higher quality work. This approach emphasizes interaction patterns that help users identify blind spots, foster creativity, and amplify thoughtful decision-making, ultimately aiming for trustworthy human-AI partnerships that grow with the user.
- The Future of Qwen: A Generalist Agent Model — Junyang Lin, Alibaba Qwen
This talk introduces the Qwen series of large language and multimodal models from Alibaba, focusing on their development towards generalist agent models. It highlights recent advancements in Qwen 3, including its hybrid thinking mode, extensive language support, and enhanced capabilities for agents and coding. The presentation also touches upon the future direction of AI development, emphasizing training agents through reinforcement learning and scaling multimodal capabilities.
- Creating Agents that Co-Create — Karina Nguyen, OpenAI
This talk explores the evolution of AI research scaling paradigms, moving from next-token prediction to scaling reinforcement learning on Chain of Thought. These advancements have unlocked new frontiers in product research, enabling rapid iteration cycles and the creation of novel AI capabilities. The future vision is one of AI agents as co-innovators, collaborating with humans to create new knowledge and experiences.
- How to Build Your Own AI Data Center in 2025 — Paul Gilbert, Arista Networks
This talk focuses on the infrastructure required to build and operate AI data centers, specifically addressing the networking challenges and solutions for training and inference. It highlights the significant differences between traditional data center networking and the demands of AI workloads, emphasizing the need for high bandwidth, low latency, and specialized traffic management. The core thesis is that building effective AI data centers requires a fundamental shift in network design and implementation to handle the unique, high-intensity demands of GPUs.
- Function Calling is All You Need — Full Workshop, with Ilan Bigio of OpenAI
This workshop explores the power and versatility of function calling in AI models, arguing that it is a fundamental capability for building advanced AI applications. The talk traces the evolution of language models from simple text completion to instruction following and finally to sophisticated tool use. It demonstrates how function calling enables AI to interact with external systems, fetch data, take actions, and manage complex workflows, forming the backbone of agentic behavior.
- Ensure AI Agents Work: Evaluation Frameworks for Scaling Success — Aparna Dhinkaran, CEO Arize
This talk addresses the critical need for robust evaluation frameworks for AI agents, moving beyond initial development to production readiness. It emphasizes that as AI agents become more sophisticated, particularly with multimodal and voice capabilities, ensuring their reliability and effectiveness in real-world scenarios requires a structured approach to testing and validation. The core thesis is that comprehensive evaluation at every stage of an agent's operation is essential for scaling success and building trust.
- The missing pieces of workflow automation — Shirsha Chaudhuri, Thomson Reuters Labs
This talk addresses the current limitations in achieving comprehensive workflow automation with AI agents. While generative AI, RAG, and agent frameworks have advanced significantly, enterprises are still missing key components to fully reimagine and automate complex business processes. The presentation highlights the gap between existing stable technology stacks, like mainframes, and the potential of agentic workflows, emphasizing the need for better connectors, reliability, standardization, and collaborative user experiences.
- Evaluating Domain Specific LLMs for Real World Finance — Waseem Alshikh, Writer
This talk explores the necessity of domain-specific language models in finance, challenging the notion that general models with high accuracy are sufficient. The evaluation presented reveals that while general models perform well on standard benchmarks, they struggle with real-world complexities like misspellings, incomplete queries, and crucially, grounding answers within provided context. This suggests a significant gap in robustness and reliability, indicating that specialized models and robust system design are still essential for critical applications.
- The Devops Engineer Who Never Sleeps — Diamond Bishop, Datadog
This talk explores Datadog's development of AI agents designed to assist with DevOps tasks, focusing on the "AI on-call engineer" and "AI software engineer." The core thesis is that as AI capabilities advance, platforms like Datadog can evolve from mere observability tools into intelligent agents that proactively manage systems, reduce human toil, and enhance developer productivity. The presentation highlights the challenges and learnings in building, evaluating, and integrating these agents into existing workflows.
- Self Coding Agents — Colin Flaherty, Augment Code
This talk explores the development and capabilities of AI coding agents, focusing on how these agents can contribute to their own creation and improvement. The core thesis is that AI agents are rapidly evolving and will significantly change software engineering, with a notable statistic indicating that over 90% of a 20,000-line codebase for an agent was written by the agent itself under human supervision. The presentation highlights the practical applications, lessons learned, and future implications of this technology.
- Vercel AI SDK Masterclass: From Fundamentals to Deep Research
This talk demonstrates how to build AI agents using Vercel's AI SDK, progressing from fundamental concepts to a practical deep research project. It highlights the SDK's unified interface for model switching, tool-use capabilities for interacting with the external world, and structured output generation for creating type-safe data. The deep research example showcases how to break down complex tasks into agentic workflows, combining web searching, result analysis, and recursive query generation to produce comprehensive reports.
- Frontier Feud: Anthropic, Google DeepMind, Meta FAIR, Thinking Machines — Barr Yaron, Amplify
This talk, Frontier Feud, hosted by Barr Yaron of Amplify, uses a game show format to explore the opinions and priorities of 100 AI engineers. Teams comprised of individuals from Anthropic, Google DeepMind, Meta FAIR, and Thinking Machines competed by guessing survey answers on topics like influential researchers, model selection criteria, and AI buzzwords. The event highlighted differing perspectives on AI development and its future impact.
- AI + Security & Safety — Don Bosco Durai
This talk addresses the critical need for safety and security in AI agents, which are increasingly autonomous and capable of complex actions. It highlights that traditional software security practices are insufficient for agents due to their non-deterministic nature and potential for unknown vulnerabilities. The core thesis is that building secure and reliable AI agents requires a multi-layered approach encompassing rigorous evaluation, robust enforcement mechanisms, and continuous observability.
- Stateful Agents — Full Workshop with Charles Packer of Letta and MemGPT
This workshop introduces the concept of stateful agents in AI, emphasizing that true agentic capabilities require memory and the ability to learn from experience, which current stateless Large Language Models (LLMs) lack. The session explores how to build these stateful agents using the Leta framework, focusing on memory management systems and tool-calling mechanisms to create more human-like and persistent AI interactions.
- Voice Agent Engineering — Nik Caryotakis, SuperDial
This talk explores the engineering challenges and best practices for building reliable voice AI applications, particularly in production environments. It emphasizes the shift from prescriptive to descriptive development in voice user interfaces and highlights the importance of focusing on conversational content and vertical integrations over superficial realism. The speaker advocates for a pragmatic approach, prioritizing reliability and functionality, especially in sensitive domains like healthcare administration.
- Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
This talk addresses the current limitations and challenges in building and evaluating AI agents. While agents are increasingly integrated into products, ambitious visions of their capabilities are far from realized. The core thesis is that AI engineering must prioritize reliability and rigorous evaluation, treating these as first-class concerns to overcome the inherent stochasticity of language models and move beyond misleading benchmarks.
- Building LinkedIn's GenAI Platform — Xiaofeng Wang
This talk details LinkedIn's journey in building a Generative AI platform, emphasizing its critical role in supporting evolving AI product experiences. The platform evolved from simple prompt-based features to sophisticated multi-agent systems, necessitating a unified infrastructure to manage complexity, enforce best practices, and bridge the gap between AI and product engineering.
- Insights on Building AI Teams — Heath Black, SignalFire
This talk by Heath Black of SignalFire explores how AI companies can leverage data to build more effective teams. The core thesis is that traditional recruiting methods are becoming less effective, and a data-driven approach is necessary to identify, attract, and hire top AI talent in a competitive market. The presentation emphasizes moving beyond credentials and focusing on experience, strategic location analysis, understanding timing, and crafting compelling narratives beyond just salary and equity.
- AI Engineers: The Next Generation — Stefania Druga, Google Gemini
This talk explores how open-source multimodal agents can enhance learning and creativity. These agents are proactive, nudging users to explore diverse topics and approaches, while also identifying potential cognitive biases and misconceptions. Examples are provided in math and science learning, demonstrating how these agents can support or challenge existing understanding.
- How to Fail at AI Strategy: Hamel Husain & Greg Ceccarelli
This talk, presented by Hamel Husain and Greg Ceccarelli, offers a satirical guide on how to spectacularly fail at AI strategy. Instead of best practices, the speakers advocate for "worse practices" to ensure projects are torpedoed and teams are alienated. The core thesis is that by inverting conventional wisdom and embracing division, ambiguity, and a focus on tools over processes, organizations can achieve complete AI failure.
- Anthropic in the Enterprise — Alexander Bricken & Joe Bayley
This talk by Alexander Bricken and Joe Bayley from Anthropic focuses on implementing AI effectively within enterprise settings. It emphasizes moving beyond basic chatbot functionalities to solve core business problems with AI. The discussion highlights Anthropic's approach to AI safety, their latest models like Sonnet 3.5, and the importance of interpretability research for understanding and steering AI behavior. The speakers also share practical advice on best practices and common pitfalls encountered when deploying AI solutions, drawing from extensive customer interactions.
- Finetuning: 500m AI agents in production with 2 engineers — Mustafa Ali & Kyle Corbitt
This talk details how Method scaled its AI agent operations to over 500 million agents with a small engineering team. Initially, they faced challenges with manual data aggregation processes and later with the high costs and limitations of using large language models like GPT-4 for parsing unstructured financial data. The solution involved fine-tuning smaller, more efficient models to meet specific production requirements for accuracy, latency, and cost.
- The Agent Development Life Cycle — Zack Reneau-Wedeen, Sierra
This talk introduces the Agent Development Life Cycle (ADLC) as a structured approach to building and improving AI agents, drawing parallels to traditional software development. It emphasizes that agents are products requiring a robust platform for iterative refinement, customer feedback integration, and continuous improvement, much like a mobile app or website. The ADLC aims to leverage the strengths of large language models while incorporating traditional software practices for reliability and efficiency.
- RAG Agents in Prod: 10 Lessons We Learned — Douwe Kiela, creator of RAG
This talk addresses the challenges and lessons learned from deploying Retrieval-Augmented Generation (RAG) agents in production environments, particularly within enterprises. The core thesis is that while large language models are powerful, they are often only a small part of a larger system. True value and business transformation come from effectively managing context, specializing AI for domain expertise, and building robust systems that can handle enterprise-scale data and user needs, rather than focusing solely on model capabilities.
- Trust, but Verify: Knowledge Agents for Finance Workflows - Mike Conover
This talk explores the development and application of knowledge agents designed to process vast amounts of financial data, drawing parallels to the transformative impact of spreadsheets on accounting. The core thesis is that these AI agents can significantly accelerate financial research and due diligence, moving beyond human limitations in handling complex, large-scale information. The presentation emphasizes the need for systems that can reveal their thought processes and allow for human oversight and intervention to ensure accuracy and relevance.
- Building AI Agents with Real ROI in the Enterprise SDLC: Bruno (Booking.com) & Beyang (Sourcegraph)
This talk explores the practical application of AI agents within the enterprise Software Development Life Cycle (SDLC), focusing on achieving tangible Return on Investment (ROI). It highlights the challenges of integrating AI into large organizations, particularly concerning measuring impact and overcoming developer toil caused by bloated codebases. The discussion emphasizes the journey from initial experimentation with AI coding assistants to building sophisticated agents that automate complex tasks and improve developer productivity.
- Anchoring Enterprise GenAI with Knowledge Graphs: Jonathan Lowe (Pfizer), Stephen Chin (Neo4j)
This talk explores the practical application of Generative AI within large enterprises, specifically addressing the challenges of project failure and the strategic integration of AI solutions. It highlights how knowledge graphs can anchor enterprise GenAI initiatives by providing structured, contextual data, thereby improving accuracy and enabling more precise decision-making. The discussion emphasizes the importance of aligning AI projects with clear business use cases and navigating organizational complexities to achieve successful production deployment.
- Personal, Local, Private AI Agents: Soumith Chintala
This talk argues for the necessity of personal, local, and private AI agents. The core thesis is that for AI agents to be truly useful and reliable, especially in handling personal life context, they must operate locally and privately. This approach addresses concerns around control, unexpected behavior, and potential misuse of personal data inherent in cloud-based solutions.
- How We Build Effective Agents: Barry Zhang, Anthropic
This talk focuses on building effective AI agents, emphasizing a practical, non-hyped approach. The core thesis is that agents should be used for complex, valuable tasks where their autonomy is beneficial, rather than as a universal upgrade. Key principles include judiciously selecting use cases, maintaining simplicity in agent design, and adopting the agent's perspective during development to better understand their behavior and limitations.
- Scaling Agents for Gen AI Products - Anju Kambadur, Bloomberg Head of AI Engineering
This talk focuses on the practical challenges and strategies for scaling generative AI agents in product development, particularly within a large financial data organization. It emphasizes the need for robust engineering practices, acknowledging the inherent fragility and evolving nature of LLMs and agentic systems. The core thesis is that building reliable AI products requires moving beyond basic LLM capabilities to implement structured development, resilient architectures, and clear organizational design.
- AI Engineering at Jane Street - John Crepezzi
This talk details Jane Street's approach to integrating large language models (LLMs) into their developer workflow, focusing on overcoming challenges posed by their unique technology stack, particularly their extensive use of OCaml. The core thesis is that building custom LLM solutions, including fine-tuned models and tailored editor integrations, is necessary to maximize value when off-the-shelf tools are insufficient due to specialized languages and internal development environments.
- How Deep Research Works - Mukund Sridhar & Aarush Selvan, Google DeepMind
This talk details the development of Deep Research, a personal research agent integrated into Gemini Advanced. The core thesis is that by removing compute and latency constraints at inference time, AI models can perform extensive web browsing to generate comprehensive answers, addressing the limitations of traditional chatbots that often provide blueprints rather than direct solutions. The presentation highlights the product and technical challenges encountered in building such an asynchronous, long-form research tool within a synchronous chatbot interface.
- Why Agent Engineering — swyx
This talk argues that AI engineering is emerging as a distinct discipline, moving beyond its roots in machine learning and software engineering. The speaker posits that the current wave of advancements, particularly in agent capabilities, is driving this evolution. The core thesis is that the development and widespread adoption of AI agents are reshaping the field and creating new opportunities for builders and engineers.
- Rethinking how we Scaffold AI Agents - Rahul Sengottuvelu, Ramp
This talk proposes a paradigm shift in building AI agents, advocating for systems that scale with compute rather than rigid, fixed architectures. The core idea is that systems designed to leverage increasing computational power will inherently outperform those that do not, drawing parallels to historical advancements in fields like chess and computer vision where brute-force search eventually surpassed human-designed heuristics. This approach suggests that by embracing the exponential improvements in AI models, developers can build more robust and adaptable agent systems with less manual engineering effort.
- Navigating AI’s Frontier in 2025 - Grace Isford, Lux Capital
The AI landscape in 2025 is experiencing exponential growth, with numerous companies releasing increasingly performant and efficient models. While this presents a perfect storm for AI agents, they are not yet fully autonomous or reliable. Cumulative errors in decision-making, implementation, heuristics, and user preferences hinder their effectiveness, leading to unmet expectations despite advancements.
- How Windsurf writes 90% of your code with an Agentic IDE - Kevin Hou, Windsurf
Windsurf is an AI agent-powered IDE designed to significantly increase developer productivity by automating a large portion of code generation. The core thesis is that agents are the future of software development, capable of handling routine tasks and complex problem-solving, allowing developers to focus on higher-level product building and feature creation. Windsurf aims to keep developers in a state of flow by minimizing manual input and maximizing the agent's contribution.
- Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley
This talk explores the potential of reinforcement learning (RL) to advance AI agents beyond current chatbot and reasoner capabilities. It posits that RL offers a path to developing more autonomous systems that can learn and improve through interaction with their environment, moving beyond static prompt engineering and tool-calling approaches. The discussion highlights emerging trends and open-source efforts in this domain, suggesting a future where RL is integral to agent engineering.
- OpenAI for VP's of AI + Advice for Building Agents
This talk discusses strategies for enterprises to build and scale AI use cases, focusing on enabling the workforce, automating operations, and infusing AI into end products. It outlines a phased approach to AI adoption, starting with a clear business strategy, identifying high-impact use cases, and building organizational capability. The discussion also delves into the practicalities of building AI agents, emphasizing a step-by-step methodology for development and deployment.
- Building Agents with Model Context Protocol - Full Workshop with Mahesh Murag of Anthropic
This talk introduces the Model Context Protocol (mCP), an open protocol designed to standardize how AI applications and agents integrate with external systems, tools, and data sources. The core motivation behind mCP is that AI models are only as effective as the context provided to them. By establishing a common interface, mCP aims to reduce fragmentation in AI development, enabling seamless integration between AI clients and servers, ultimately leading to more powerful and personalized AI applications.
- AI Agents, Meet Test Driven Development
This talk introduces a test-driven development (TDD) approach for building reliable AI agentic workflows. It argues that while AI models are improving, success in production hinges on robust engineering practices. The presentation outlines a structured process for experimenting, evaluating, scaling, and continuously monitoring AI solutions, emphasizing the need for a systematic approach to manage complexity and ensure reliability.
- WTF do people use Open Models for??
This talk explores the actual usage patterns of open-source AI models, moving beyond benchmarks to understand how individuals and businesses are employing them. It highlights that while new models emerge rapidly, established, reliable models often persist in production due to the significant effort required to change them. The presentation emphasizes that user preference, or "vibes," plays a crucial role in model selection, especially for creative and companionship use cases.
- This video was edited with AI agent. But how?
This talk introduces an AI agent designed for open-source video editing, addressing limitations of existing tools. The agent leverages a programmatic interface for complex video compositions, enabling LLMs to generate and execute code for editing tasks. This approach is presented as superior to JSON-based tool calling for LLMs.
- Voice Agents: the good, the bad, and the ugly
This talk explores the complexities and challenges of building AI voice agents, using a case study of an automated interview system. It highlights that while LLMs offer powerful capabilities, developing robust voice applications requires overcoming issues like transcription errors, conversational flow, and agent behavior. The presentation emphasizes the need for sophisticated architectural patterns beyond simple prompt engineering to achieve reliable and effective voice AI.
- The Price of Intelligence - AI Agent Pricing in 2025
This talk explores the evolving landscape of AI agent pricing in 2025, emphasizing that effective pricing strategies are not unique to AI but are fundamental business practices. It highlights the importance of understanding target audiences, maintaining simplicity and predictability in pricing, and considering the long-term implications of pricing models on product adoption and cost structures. The discussion also touches upon the necessity of pricing flexibility to adapt to market changes and technological advancements.
- Don't just slap on a chatbot: building AI that works before you ask
This talk challenges the common practice of simply adding chatbot interfaces to products, arguing that this approach is often a superficial solution. Instead, it advocates for building AI that proactively assists users by understanding context and anticipating needs, drawing parallels to the conceptual idea behind Clippy but with modern execution. The core thesis is that AI should integrate seamlessly into natural workflows, offering suggestions and actions without requiring explicit prompts, thereby enhancing productivity and user experience.
- Beyond APIs: How AI Web Agents Are Automating the \"Long Tail\" of Knowledge Work
This talk introduces Retriever, a Chrome extension that functions as an AI web agent to automate knowledge work. It addresses the inefficiencies of manual data copying and pasting, unreliable scraping, and siloed information by allowing users to issue natural language commands for autonomous task execution and structured data extraction across web pages. Retriever aims to be as transformative to the browser as its initial creation.
- Agent Evals: Finally, With The Map
This talk introduces a framework for evaluating AI agents, dividing agent evaluation into semantic and behavioral aspects. Semantic evaluation focuses on how an agent's internal representations align with reality, while behavioral evaluation assesses how an agent's actions and tool usage contribute to achieving its goals. Both aspects are further categorized into single-turn and multi-turn scenarios, providing a comprehensive map for understanding and measuring agent performance.
- Mission-Critical Evals at Scale (Learnings from 100k medical decisions)
This talk outlines the challenges and solutions for building scalable evaluation systems for AI, particularly in mission-critical domains like healthcare where errors have significant consequences. It emphasizes the limitations of traditional human review and offline datasets, advocating for real-time, reference-free evaluation methods to ensure customer trust and enable rapid response to issues.
- Your Evals Are Meaningless (And Here’s How to Fix Them)
This talk argues that traditional evaluation methods for AI systems are often insufficient and can lead to meaningless results. The core thesis is that evaluations must be dynamic and continuously aligned with real-world usage and user expectations, rather than relying on static, generalized criteria.
- OpenLLMetry is all you need
OpenLLMetry extends the open-source OpenTelemetry project to provide observability for Generative AI applications. It standardizes the collection of logs, metrics, and traces for LLM frameworks, foundation models, and vector databases, allowing developers to use their preferred observability platforms without vendor lock-in.
- Privacy First Enterprise AI: Building AI Agents that Never Leave Your Security Boundary
This talk proposes a privacy-first approach to deploying enterprise AI agents by leveraging existing, decades-old enterprise infrastructure rather than building new, parallel systems. The core thesis is that AI agents should operate within established security boundaries, utilize current workflows, and be managed through familiar IT tools, mirroring how human employees are handled. This strategy aims to integrate AI capabilities seamlessly into existing trusted platforms, enhancing them with AI rather than creating new, potentially redundant interfaces.
- Stop Guessing: Build Robust AI with Layered CoT
This talk introduces Layered Chain of Thought (Layered CoT) as a method to build more robust and self-correcting AI systems. It addresses the limitations of traditional Chain of Thought (CoT) prompting, such as sensitivity to prompt phrasing and lack of real-time error correction. Layered CoT integrates a verification step after each reasoning stage, cross-referencing generated thoughts against a knowledge base to ensure accuracy and prevent error propagation.
- How to Improve Your Agents: Academic Lit Review
This talk explores methods for enhancing AI agent capabilities beyond basic chat interactions, focusing on improving reasoning, reflection, and action execution. It delves into academic literature to present techniques that allow agents to learn and improve without direct human supervision, emphasizing the potential for smaller models to achieve greater performance through refined feedback and optimized training processes. The discussion highlights advancements in both self-improvement mechanisms and more complex decision-making strategies like tree search for agents.
- The Model Isn’t Wrong—You’re Just Bad at Prompting
This talk focuses on the nuances of prompt engineering, arguing that effective prompting is crucial for achieving optimal outputs from large language models. It emphasizes that while models are powerful, user-defined prompts significantly influence performance, offering a competitive advantage in product development. The core thesis is that mastering prompt engineering techniques is essential for builders and engineers working with AI.
- Tool Calling Is Not Just Plumbing for AI Agents — Roy Derks
This talk argues that tool calling is a critical, often overlooked component of AI agent development, deserving more attention than the agents themselves. The speaker emphasizes that an agent's effectiveness is directly tied to the quality, reusability, and robustness of its tools. The presentation explores different approaches to tool integration, from traditional, explicit tool calling to more embedded methods, and advocates for architectural patterns that promote separation of concerns.
- Your LLM Ran Out of Knowledge — Now What?
This talk addresses the challenge of applying large language models (LLMs) to domains where structured training data is scarce. The core thesis is that by providing LLMs with explicit rules, heuristics, and guidelines, similar to how junior professionals are mentored, their powerful reasoning capabilities can be effectively leveraged even in low-knowledge areas. This approach aims to bridge the gap between domains with abundant data and those lacking it, enabling LLMs to assist in complex problem-solving across a wider range of professions.
- Lets Build An Agent from Scratch — Kam Lasater
This talk demonstrates how to build an AI agent from scratch, breaking down the core components and their interactions. The presenter aims to demystify agent functionality by starting with a simple LLM call and progressively adding elements like conditional logic, tool usage, and planning mechanisms. The goal is to provide viewers with an intuitive understanding of how agents operate, enabling them to experiment with and adapt the provided code.
- The Hidden Costs of Building Your Own RAG Stack — Ofer Vectara
Building a production-ready Retrieval Augmented Generation (RAG) stack is significantly more complex than initial prototypes suggest. The talk highlights seven hidden costs and pitfalls associated with DIY RAG implementations, including response quality issues, latency, scaling expenses, security and compliance challenges, vendor management complexities, talent acquisition and retention difficulties, and multilingual support limitations.
- Building Multi agent Systems with Finite State Machines
This talk explores how finite state machines (FSMs) can provide a structured and reliable foundation for building complex AI systems, particularly agentic AI. While AI intelligence is advancing, the critical need for predictability, observability, and control in orchestrating autonomous agents is paramount. FSMs, and their extension state charts, offer a robust framework for governing agent decision-making and behavior, complementing the dynamic intelligence of large language models (LLMs).
- Where AI is superhuman: The right jobs to automate with LLMs
This talk explores how Large Language Models (LLMs) can automate tasks in the workplace, focusing on identifying jobs where AI can be superhuman. It posits that LLMs excel at data transformation, synthesis, and reasoning, making them particularly disruptive for high-volume, low-complexity tasks. The presentation suggests that AI automation will shift organizational structures from pyramid shapes to inverted pyramids or diamonds, with fewer entry-level roles and more advanced or managerial positions.
- Your AI Agent Isn't an Engineer: The Art of Thoughtful Anthropomorphism
This talk argues against framing AI agents as direct replacements for human software engineers, advocating instead for thoughtful anthropomorphism in developer marketing. The speaker suggests that marketing AI tools as engineers leads to misleading expectations, alienates developers who are key users, and misses the true value proposition of AI as an augmentation tool. A framework is proposed for authentic AI marketing that builds trust and encourages sustainable adoption.
- Reverse Conway's law and GenAI: How agents will take over the organisation - Patrick Debois
This talk explores how Generative AI and agents will fundamentally reshape organizational structures, moving beyond simple co-pilot roles to become integrated team members and even peers. It examines the potential for AI to automate tasks, influence decision-making, and alter the very nature of work, prompting a re-evaluation of human roles and skills within companies. The discussion touches on the evolution from individual AI assistance to team-level AI integration and the broader implications for organizational design and the future of employment.
- How Coding Agents change Software Development Forever - Hailong Zhang
- The LLM Triangle: Engineering Principles for Robust AI Applications - Almog Baku:
This talk introduces the LLM Triangle, a framework for building robust AI applications by focusing on three core principles: the model itself, engineering techniques, and contextual data. The speaker emphasizes that successful LLM applications are not just about sophisticated models but heavily rely on rigorous experimentation and data-driven engineering, akin to standard operating procedures in manufacturing. The goal is to treat LLMs as capable but inexperienced interns that require clear, step-by-step guidance to ensure consistent quality and performance in production environments.
- Lessons from building GenAI based applications — Juan Peredo
This talk shares practical lessons learned from building Generative AI applications over the past year and a half. It emphasizes that while AI can accelerate development, integrating it introduces significant new complexities. Key challenges include model selection, preventing hallucinations, managing infrastructure costs (especially GPUs), ensuring correct outputs, and understanding the operational overhead of agentic systems. The speaker advocates for careful evaluation, externalizing prompts, and robust observability to navigate these complexities effectively.
- Cohere: Building enterprise LLM agents that work (Shaan Desai)
This talk focuses on the practical challenges and solutions for building enterprise-grade LLM agents. It emphasizes that while agents are a powerful application of generative AI, creating them at scale, safely, and seamlessly is complex due to the vast array of frameworks, tools, and evaluation criteria. The presentation aims to guide builders through critical decision-making processes, sharing key learnings from Cohere's experience in developing these agents.
- Patrick Dougherty: How to Build AI Agents that Actually Work
This talk focuses on practical lessons learned from rebuilding a product around AI agents, emphasizing the importance of enabling agents to reason rather than solely relying on the underlying model's knowledge. Key to this approach is the concept of the Agent Computer Interface (ACI), which involves carefully designing the structure and format of tool calls and their responses. The speaker argues that focusing on ACI iteration and the surrounding ecosystem, including user experience and security, is more valuable than fine-tuning models or over-reliance on abstraction frameworks for production environments.
- Keynote: The AI developer experience doesn't have to suck – why and how we built Modal
This talk addresses the challenges in the AI developer experience, arguing that it doesn't have to be cumbersome. The speaker introduces Modal, an infrastructure platform designed to make writing, deploying, and scaling data, AI, and machine learning applications enjoyable again. Modal focuses on high-code use cases, allowing developers to run arbitrary Python code and containers in the cloud, abstracting away complex infrastructure management.
- Keynote: Why people think \"agent\" is a buzzword but it isn't
This talk challenges the perception of AI agents as mere buzzwords, arguing instead for their significant potential. The speaker defines agents as entities that can perceive and act upon their environment, drawing parallels from historical definitions to modern applications like coding assistants. The core thesis is that while agents are not new, building and deploying them effectively presents substantial challenges that, once overcome, will unlock transformative use cases.
- Personality Driven Development: Exploring the Frontier of Agents with Attitude
This talk explores the concept of personality-driven development in AI agents, where agents are given distinct forms and personalities to enhance user interaction and understanding. The speaker argues that anthropomorphizing AI, while not new, becomes particularly relevant with modern agents, influencing user perception and expectations. This approach aims to simplify complex AI functionalities by mapping them to familiar human roles and traits, thereby improving product development and user experience.
- Optimizing LLMs in Insurance with DSPy: Jeronim Morina
This talk emphasizes a return to first principles for AI engineers, moving beyond superficial prompt engineering and tool tinkering. It argues that AI models are most effective when integrated into larger systems designed to solve real-world problems, particularly in complex domains like insurance. The presentation advocates for a structured approach to development, including clear problem definition, robust evaluation, and modular system design, with DSPy presented as a tool to aid this process.
- Customized, production ready inference with open source models: Dmytro (Dima) Dzhulgakov
This talk focuses on the advantages and practicalities of using open-source models for production AI applications. It highlights how custom-tuned, smaller models can outperform larger, general-purpose proprietary models in terms of speed, cost, and domain-specific accuracy. The presentation also addresses the challenges of deploying open-source models, such as setup complexity and optimization, and introduces Fireworks AI's platform designed to streamline these processes.
- Claude plays Minecraft!
This talk demonstrates an agentic workflow using Claude 3 Haiku to control a Minecraft bot named Rocky. The agent takes chat inputs and uses a set of defined tools to perform actions within the game, such as moving, locating players, digging, and even building structures. The system leverages Amazon Bedrock Agents for a managed agentic workflow, allowing for orchestration of multiple tasks and providing return of control to the user.
- The Adversarial Path to the Personal Assistant: Sumit Agarwal
This talk introduces Ario, a personal assistant AI designed to give users back approximately one hour per day by automating mundane tasks. The core of Ario's system is adversarial ETL, which focuses on extracting user data from various online sources. This data is then used to build a dynamic, evolving profile of the user, enabling more personalized and context-aware assistance than generic AI tools.
- Training Albatross An Expert Finance LLM: Leo Pekelis
This talk details the development of Albatross, an expert finance Large Language Model (LLM). It argues that generalist LLMs are insufficient for complex domains like finance, necessitating custom-trained models and extended context lengths. The approach involves creating a domain-specific finance LLM through automated data curation and a tailored training pipeline, alongside developing models with significantly increased context windows to mitigate hallucinations and enable more sophisticated in-context learning.
- RAG at scale: production ready GenAI apps with Azure AI Search
This talk focuses on scaling Retrieval Augmented Generation (RAG) applications for production use, specifically highlighting the capabilities of Azure AI Search. It addresses the challenges that arise when moving from prototypes to production, such as increased data volume, higher data change rates, and more complex multi-step workflows. The presentation emphasizes how Azure AI Search provides integrated solutions for these scaling dimensions, including advanced vector search, hybrid search, and efficient data ingestion.
- Accelerating Mixture of Experts Training With Rail Optimized InfiniBand Networking in Crusoe Cloud
This talk focuses on optimizing the infrastructure for training large AI models, specifically addressing the bottlenecks in distributed training caused by network communication. Crusoe Cloud, an AI cloud platform powered by renewable energy, highlights its rail-optimized InfiniBand networking solution designed to accelerate Mixture of Experts (MoE) training by reducing the time gpus spend idle waiting for data exchange.
- System Design for Next-Gen Frontier Models — Dylan Patel, SemiAnalysis
This talk discusses the system design challenges and future directions for running and training next-generation large language models. It highlights the significant computational and memory bandwidth requirements for inference, particularly for models with trillions of parameters. The presentation also touches upon the massive scale of infrastructure being built for training these models and the associated engineering hurdles.
- AI Music Generation, From Prompt to Production: Phlo Young
This talk explores the rapidly evolving landscape of AI music generation, demystifying the technology and demonstrating practical applications. It covers various categories of AI music, including voice conversion, text-to-music, and audio-to-music generation, highlighting tools like Suno and Udio. The presentation aims to equip attendees with the knowledge to transform musical ideas into finished songs and discusses the emerging opportunities for artists in this new domain.
- Giving a Voice to AI Agents: Scott Stephenson, CEO, Deepgram
This talk explores the evolution of voice AI, moving beyond the limitations of earlier systems to a new era of conversational agents. The core thesis is that true human-like interaction in AI is not solely about speed or accuracy, but critically depends on the contextual understanding and generation capabilities within the entire voice AI pipeline. This contextual awareness is presented as the key innovation that will make AI agents feel genuinely conversational.
- How to build the world's fastest voice bot: Kwindla Hultman Kramer
This talk explores the engineering challenges and architectural considerations for building extremely fast voice bots. The core thesis is that achieving human-like conversational latency requires a deep understanding of media processing, real-time data pipelines, and strategic model deployment, often necessitating self-hosting and colocation of components to minimize network delays.
- Unveiling the latest Gemma model advancements: Kathleen Kenealy
This talk introduces the latest advancements and additions to the Gemma model family, emphasizing Google DeepMind's commitment to empowering the open-source community. The Gemma models are presented as lightweight, state-of-the-art, open-source models built from the same technology as the Gemini models, designed for responsible AI development, high performance, and extensibility across various platforms and frameworks.
- Fine tune 20 Llama Models in 5 Minutes: Santosh Radha
This talk demonstrates how to fine-tune and deploy multiple AI models directly from Python code without requiring complex infrastructure like Kubernetes or Docker. The core idea is to use a Pythonic interface to abstract away the underlying compute resources, allowing users to specify hardware requirements and execution parameters with simple decorators. This approach simplifies the process of running computationally intensive tasks, such as model training and inference, on various backends, including local machines, on-premise systems, or cloud-based GPU clusters.
- The GenAI Maturity Curve or You Probably Don't Need Fine Tuning: Kyle Corbitt
This talk argues that most AI engineers and builders likely do not need fine-tuning for their current projects. The core thesis is that while fine-tuning can significantly improve model performance and cost-effectiveness, it introduces upfront time investment and reduces flexibility. The presentation aims to help the audience determine if and when fine-tuning becomes a beneficial step in their development process.
- Building an AI assistant that makes phone calls [Convex Workshop]
This talk introduces Floyd, an AI assistant designed to make phone calls, demonstrating how to integrate various technologies to create a more capable personal assistant. The core idea is to leverage real-time speech-to-text, text-to-speech, and AI models to handle conversations with humans, managing tasks like booking appointments or ordering services. The system aims to provide a seamless experience by learning user context and proactively managing interactions.
- LLM Quality Optimization Bootcamp: Thierry Moreau and Pedro Torruella
This talk focuses on optimizing Large Language Model (LLM) quality through fine-tuning, addressing common pain points like high operational costs and the inability to meet production-ready quality standards. It positions fine-tuning within a broader "crawl, walk, run" strategy for LLM quality, emphasizing that it should be considered after prompt engineering and Retrieval Augmented Generation (RAG) have been explored. The presentation outlines a continuous deployment cycle for fine-tuned LLMs, including data collection, model fine-tuning, deployment, and evaluation, aiming to demystify the process for AI engineers.
- Building security around ML: Dr. Andrew Davis
This talk addresses the critical need for robust security measures in machine learning systems. It highlights the inherent fragility of ML models, making them susceptible to various attacks. The discussion covers data poisoning, model theft, adversarial examples, supply chain vulnerabilities, and software exploits, emphasizing proactive strategies and continuous vigilance to protect ML deployments.
- GitHub's AI Powered Security Platform: Sarah Khalife
This talk focuses on how GitHub is integrating AI into its Advanced Security platform to enhance developer productivity and security. The core thesis is that AI can significantly improve the identification and remediation of security vulnerabilities and secrets, making security a more integrated and less burdensome part of the daily development workflow. By embedding AI capabilities directly into the platform, GitHub aims to bridge the gap between security teams and developers, fostering a shared responsibility for application security.
- Agentic Workflows on Vertex AI: Rukma Sen
This talk explores the concept of AI agents as the primary interface for human interaction with generative AI. It defines an agent as a system designed to achieve specific goals through environmental interaction, comprising a reasoning model, tools for action, and orchestration for memory and state management. The discussion emphasizes the responsibility that comes with building these agents, focusing on ethical considerations, safety, cybersecurity, and data privacy.
- Insights from Snorkel AI running Azure AI Infrastructure: Humza Iqbal and Lachlan Ainley
This talk focuses on Snorkel AI's approach to data development for enterprise AI models, emphasizing the challenges of achieving enterprise-grade quality, latency, and cost requirements with off-the-shelf large language models. Snorkel AI specializes in developing data to fine-tune these models, enabling them to meet specific business needs. The discussion also touches upon the infrastructure and best practices for running large-scale AI training and inference workloads on Azure.
- RAG and the MongoDB Document Model: Ben Flast
This talk explores the integration of Retrieval Augmented Generation (RAG) with MongoDB's document model and Atlas Vector Search. It highlights how combining a flexible document database with vector search capabilities enables more sophisticated and context-aware AI applications. The presentation emphasizes that modern AI applications require more than just generic LLMs, necessitating the augmentation of prompts with relevant, up-to-date data.
- GitHub Next Explorations: Rahul Pandita
GitHub Next explores the future of software engineering, particularly how generative AI can transform developer workflows. The team operates independently to experiment with new ideas, similar to the early exploration phase of electricity's adoption in manufacturing. Their process involves rapid prototyping, internal testing, and releasing tech previews to gather user feedback before potentially integrating successful concepts into mainstream products.
- [Full Workshop] Llama 3 at 1,000 tok/s on the SambaNova AI Platform
This workshop demonstrates SambaNova's AI platform, highlighting its capability to achieve over 1,000 tokens per second inference speed with Llama 3. The platform offers a full-stack solution from chip design to software, aiming to simplify the process of fine-tuning, pre-training, and deploying AI models. It addresses enterprise needs for scalability, security, and cost-efficiency by integrating various expert models behind a single endpoint and providing orchestration and access control.
- [Full Workshop from Microsoft] Github Copilot - The World's Most Widely Adopted AI Developer Tool
This workshop provides a comprehensive guide to GitHub Copilot, positioning it as the world's most adopted AI developer tool. It covers Copilot's functionality as an AI pair programmer, its integration within Integrated Development Environments (IDEs), and best practices for maximizing its utility. The session emphasizes how Copilot synthesizes code based on context from open files, comments, and direct chat queries, aiming to streamline the software development lifecycle.
- [Full Workshop] How to add secure code interpreting in your AI app: Vasek Mlejnsky
This workshop demonstrates how to integrate secure code interpretation capabilities into AI applications, similar to Anthropic's AI Artifacts feature. The process involves building a system that allows an AI model, specifically Claude's Sonnet, to generate and execute Python code within a secure sandbox environment. The workshop guides participants through setting up the necessary tools, defining code execution logic, and displaying the results, including visualizations, directly within the application's user interface.
- Substrate Launch: the API for modular AI
Substrate is presented as an API for building modular AI applications, moving away from monolithic models. The core thesis is that complex AI tasks are best accomplished by systems of multiple inference runs orchestrated logically, rather than relying on a single foundation model. This modular approach offers benefits in legibility, debuggability, extensibility, and easier evaluation due to explicit decision trees.
- GitHub Copilot: The World's Most Widely Adopted AI Developer Tool
GitHub Copilot has evolved from an AI pair programmer generating code snippets to a comprehensive tool integrated across development workflows. It now offers chat functionalities within IDEs and on GitHub.com, enabling code explanation, debugging, test generation, and interaction with enterprise knowledge bases.
- Build, Evaluate and Deploy a RAG-Based Retail Copilot with Azure AI: Cedric Vidal and David Smith
This talk demonstrates how to build, evaluate, and deploy a Retrieval-Augmented Generation (RAG) based copilot for a retail environment using Azure AI services. The core concept involves creating a chatbot that can answer customer questions by accessing product information from a vector database and customer data from a relational database, all orchestrated through Azure AI Studio's Prompt Flow.
- Lessons from the Trenches: Building LLM Evals That Work IRL: Aparna Dhinkaran
This talk focuses on the practical challenges and solutions for building effective LLM evaluation systems in real-world applications. It distinguishes between model evals, which rank models against benchmarks, and task evals, which assess whether an LLM application is functioning correctly for its intended purpose. The core thesis is that robust task evals, especially those providing explanations for failures, are crucial for iterating and improving deployed LLM applications.
- Accelerate your AI journey with Azure AI model catalog: Sharmila Chokalingam
This talk introduces the Azure AI model catalog as a comprehensive platform for accelerating AI development. It highlights how the catalog provides access to a wide range of foundation models, including flagship LLMs and smaller models, alongside tools for prototyping, optimizing, and operationalizing generative AI applications. The platform emphasizes ease of model switching, enterprise-grade security, and data privacy to support production workloads.
- AI Templates: Gabriela and Aishwarya
This talk introduces AI templates as a rapid method for deploying AI applications, focusing on how Microsoft for Startups and its Founders Hub platform can support builders. The core thesis is that these templates, particularly for complex tasks like Retrieval Augmented Generation (RAG), significantly reduce the time and complexity of getting started with AI development, enabling faster prototyping and deployment.
- Multi model multimodal and multi agent innovations in Azure AI: Cedric Vidal
This talk showcases advancements in Azure AI, focusing on multimodal and multi-agent capabilities. It highlights how new models and tools, such as GPT-4o and Phi-3 Vision, can process and reason across text, vision, and speech. The presentation emphasizes practical applications and the integration of these technologies within Azure AI Studio for building sophisticated AI solutions.
- Which Jobs Can Be Replaced Today: Fryderyk Wiatrowski and Peter Albert
This talk explores the potential for AI agents to replace human tasks, moving beyond simple prompting to autonomous operation. The core thesis is that agents can handle low-leverage, reactive tasks, freeing humans to focus on high-leverage, proactive activities. The discussion outlines a path from basic prompting to more complex agentic workflows, emphasizing the eventual goal of agents operating independently within broader workflows.
- Creating and scaling your own custom copilots with Azure AI Studio: Hanchi Wang
This talk introduces Azure AI Studio and Prompt Flow, a suite of tools designed to streamline the development, evaluation, and deployment of AI applications. The core thesis is that while large language models are powerful, they require integration with domain-specific knowledge and tools, along with careful filtering, evaluation, and continuous monitoring to be effective in enterprise settings. These tools aim to make AI application building efficient and trustworthy.
- Ionic Launch: Opening the economy to AI agents
This talk introduces Ionic's mission to enable AI agents to interact with the economy, starting with e-commerce. The core thesis is that the current digital world is built for advertisements, not for actions agents can take. Ionic aims to bridge this gap by providing agents with enriched, dynamic product data and facilitating direct transactions, thereby creating a new economic model where merchants pay for direct access to consumers matched with suitable products.
- EyeLevel Launch: Your RAG is Tripping, Here's the Real Reason Why
This talk addresses common accuracy issues in Retrieval Augmented Generation (RAG) applications, particularly with complex enterprise documents. The core thesis is that RAG failures are typically due to content ingestion problems rather than LLM or prompt issues. The presented solution focuses on a novel ingestion pipeline that preserves crucial context lost during traditional chunking and vectorization, leading to significantly improved accuracy.
- The Rise of the AI Software Engineer: Jesse Han
This talk introduces the concept of a personal AI software engineer designed to understand and augment developers throughout the entire software development lifecycle. The core thesis is that AI will evolve to handle complex engineering tasks, making sophisticated coding assistance accessible to everyone.
- How to evaluate a model for your use case: Emmanuel Turlay
This talk addresses the challenge of evaluating language models for specific use cases, highlighting that generic metrics and benchmarks often fall short. It proposes a method where another language model is used to grade the output of the model being evaluated, allowing for the creation of specialized metrics tailored to an application's needs. This approach enables data-driven decisions when selecting LLMs.
- Building efficient hybrid context query for LLM grounding: Simrat Hanspal
This talk introduces a method for building efficient hybrid context queries for Retrieval Augmented Generation (RAG) applications, specifically demonstrated with an e-commerce product search use case. It highlights the limitations of traditional keyword search and the need for natural language understanding. The core idea is to leverage Large Language Models (LLMs) by providing them with relevant data context, enabling more accurate and grounded responses. The presentation showcases how a unified data API can handle semantic, structured, and hybrid queries, enhancing RAG pipelines.
- Best Practices for Evaluating Large Language Model Applications with llmeval: Niklas Nielsen
This talk introduces llmeval, a command-line tool designed to help teams ship reliable large language model (LLM) applications. It addresses the challenge of defining and measuring "good" in generative AI, which is crucial for evolving applications, changing prompts, or considering different models. llmeval provides a structured approach to testing and evaluating LLM outputs, enabling more confident development and deployment.
- BotDojo Launch: Enhancing AI Assistants with Evaluations and Synthetic Data
This talk introduces BotDojo, an AI enablement company focused on helping businesses deploy AI applications to production. The demonstration highlights how to use synthetic data generation combined with evaluations to improve the performance of chatbots. It showcases a low-code editor for building AI flows, emphasizing the importance of detailed tracing for debugging and the integration of evaluations to identify issues like hallucinations and insufficient information retrieval.
- No-code fine-tuning: Mark Hennings
This talk introduces a no-code approach to fine-tuning large language models, enabling specialized task training without traditional programming. Fine-tuning offers advantages over prompt engineering, including faster and cheaper execution, reduced prompt length, better handling of edge cases, and inherent resistance to prompt injection. The presented method aims to lower the barrier to entry for fine-tuning, making it accessible beyond just developers.
- Prompt Engineering Tactics: Dan Cleary
This talk introduces three practical prompt engineering tactics designed to improve the accuracy and reliability of responses from large language models (LLMs). The presenter emphasizes that while basic prompts often yield good results, advanced techniques are crucial for integrating AI into products, ensuring consistent user experiences, and meeting high user expectations for speed, accuracy, and freedom from hallucinations. These methods are applicable for both casual use of LLMs and for developers building AI-powered applications.
- Scaling AI in Education: A Khanmigo case study: Shawn Jansepar
This talk details Khan Academy's journey in scaling AI for educational purposes, focusing on their AI-powered tutor and teacher assistant, Khanmigo. The core thesis is that AI can democratize one-on-one tutoring and teaching assistance at scale, addressing learning gaps and transforming educational experiences. The presentation highlights how Khan Academy shifted to an AI-first organization by rapidly prototyping, iterating, and integrating AI deeply into their platform and content, while also addressing key technical challenges and future directions.
- Cohere for VPs of AI: Vivek Muppalla
Cohere is an enterprise AI company focused on building trustworthy AI models for real-world business use cases. The company emphasizes practical application, efficiency, and scalability over simply having the largest models. Key offerings include generative models like Command R and R Plus, and advanced retrieval models such as embeddings and rerankers. Cohere prioritizes enterprise-specific performance, customization, data privacy, and deployment flexibility across various cloud and on-premise environments.
- Mastering LLM Inference Optimization From Theory to Cost Effective Deployment: Mark Moyou
This talk focuses on optimizing Large Language Model (LLM) inference, emphasizing the practical challenges and cost-effectiveness of deploying these models at scale. It delves into the technical intricacies of the LLM inference workload, from tokenization and attention mechanisms to the crucial role of the KV cache in managing GPU memory and performance. The presentation aims to provide attendees with a deeper intuition about these processes and actionable strategies for measuring and improving deployment efficiency.
- Navigating Challenges and Technical Debt in LLMs Deployment: Ahmed Menshawy
This talk addresses the practical challenges and technical debt encountered when deploying Large Language Models (LLMs) in enterprise environments. It highlights the shift from structured to unstructured data in AI applications and emphasizes that current LLMs augment human productivity rather than replace jobs. The discussion also touches on the limitations of LLM foundations and the importance of focusing on present AI risks over speculative future ones.
- LLM Safeguards: Security Privacy Compliance Anti Hallucination: Daniel Whitenack
This talk addresses the practical challenges of deploying secure, accurate, and compliant AI systems, particularly focusing on open-access large language models (LLMs) in enterprise environments. It outlines common risks such as hallucinations, supply chain vulnerabilities, data breaches, and prompt injection, and proposes mitigation strategies. The core thesis is that by understanding and proactively addressing these risks with existing and emerging technologies, organizations can build more trustworthy AI applications.
- Cooking with fire without burning down the kitchen: Dominik Kundel
This talk discusses how a team at Twilio navigated the rapidly evolving landscape of AI, focusing on disruptive innovation without compromising customer trust. The core thesis is that companies must embrace emerging AI technologies, particularly agents, by adopting principles of customer obsession, shipping early and often, fostering a curious and problem-solving team culture, and sharing learnings both internally and externally. This approach allows for rapid prototyping and iteration while managing customer expectations regarding quality and reliability.
- E-Values Evaluating the Values of AI: Sheila Gulati and Nischal Nadhamuni
This talk addresses the critical and evolving landscape of AI evaluations, particularly as systems become more agentic and automated. It argues that current evaluation methods are often simplistic and prone to "benchmark hacking," failing to capture the true performance and values embedded in AI systems. The discussion emphasizes the need for more robust, multifaceted evaluation strategies that consider real-world scenarios, user experience, and the underlying values driving AI development to ensure these systems align with human goals.
- Hiring & Building an AI Engineering Team: Dr. Bryan Bischof
This talk focuses on the practicalities of building and hiring for AI engineering teams, emphasizing a shift from pure ML research to production-ready AI product development. It argues that AI engineering requires a blend of software engineering, product thinking, and data intuition, and that the hiring process should reflect these needs. The core thesis is that successful AI teams are built by understanding the evolving stages of AI product development and hiring individuals with specific, complementary skill sets.
- RAG for VPs of AI: Jerry Liu
This talk focuses on Retrieval Augmented Generation (RAG) for enterprise applications, emphasizing the challenges of moving from prototype to production. It highlights the critical role of data processing and quality in achieving accurate, production-ready AI applications. The speaker, Jerry Liu, co-founder and CEO of LlamaIndex, discusses how their platform aims to empower developers to build robust LLM applications over their data, addressing issues like data silos, accuracy, and scalability.
- AI Platform Engineering: Patrick Debois
This talk discusses the engineering principles and platform considerations necessary for building and scaling Generative AI applications within an organization. It draws parallels with the evolution of DevOps, emphasizing the need for structured platforms, enablement, and governance to manage the complexities of AI development and deployment effectively. The core thesis is that successful AI adoption requires a dedicated platform engineering approach, similar to how DevOps matured, to bridge the gap between AI capabilities and traditional software development.
- Real ROI: Lessons from Enterprises that have already succeeded with LLMs at Scale: Raza Habib
This talk focuses on how enterprises can achieve real return on investment (ROI) from large language models (LLMs) and generative AI products. It emphasizes that the initial hype phase is over, and companies are now generating tangible revenue and cost savings. The core message is that success hinges on a strategic approach to team composition, robust evaluation methods, and appropriate tooling, rather than solely on complex model architectures.
- Understanding AI Stakes to Break Production Code: Philip Rathle
This talk explores the relationship between the stakes of an AI application and the obstacles encountered during production. Higher stakes, defined by potential impacts on finances, reputation, health, or safety, correlate with increased challenges. The discussion contrasts vector-based retrieval-augmented generation (RAG) with knowledge graphs, suggesting that structured data and reasoning from knowledge graphs can enhance reliability for high-stakes use cases.
- The ROI of AI: Why you need Eval Framework - Beyang Liu
This talk addresses the challenge of measuring the return on investment (ROI) of AI tools, particularly for developers. The core thesis is that while AI can significantly enhance developer productivity, demonstrating its business impact requires moving beyond anecdotal evidence or simple engagement metrics to more structured evaluation frameworks. The speaker emphasizes that the difficulty in measuring AI ROI is akin to the NP-hard problem of measuring developer productivity itself, suggesting that a precise, universally applicable solution is unlikely, but practical frameworks can be developed.
- AI Frontiers in Trust and Safety Combatting Multifaceted Harm on Tinder at Scale: Vibhor Kumar
This talk explores the application of AI, particularly Large Language Models (LLMs), to enhance trust and safety measures on platforms like Tinder. It addresses the challenges posed by generative AI, such as content pollution and sophisticated scams, while highlighting the opportunities LLMs present for detecting and mitigating multifaceted harm at scale. The presentation details the end-to-end process of using LLMs for violation detection, from data preparation and fine-tuning to production deployment.
- Enhancing Quality and Security in CI: Gunjan Patel
This talk explores a "ghost pilot" system designed to enhance code quality and security within the CI/CD pipeline. Unlike quick, just-in-time coding assistants, this system performs deliberate, iterative analysis. It aims to automate tedious tasks like improving code comments, adding unit tests, identifying security vulnerabilities, and suggesting fixes, thereby freeing up developers to focus on core coding tasks.
- Iterating on LLM apps at scale Learnings from Discord: Ian Webster
This talk discusses learnings from building and scaling LLM applications at Discord, focusing on the challenges of safety, security, and legal considerations. The core thesis is that robust evaluation (evals) and a strong eval culture are crucial for mitigating risks and enabling the successful deployment of LLM products at scale. The speaker emphasizes a pragmatic, developer-first approach to evals, treating them like unit tests to ensure fast, reliable feedback loops.
- Decoding Mistral AI's Large Language Models: Devendra Chaplot
This talk by Devendra Chaplot from Mistral AI details the company's approach to developing and releasing open-source large language models. Mistral AI's mission is to democratize frontier AI by making powerful models accessible to developers. They emphasize principles of openness, portability, performance-to-speed optimization, and customizability, releasing models like Mistral 7B, 8x7B, and Code 22B. The presentation also covers the three core stages of LLM training: pre-training, instruction tuning, and learning from human feedback, highlighting the trade-offs in data and compute requirements for each.
- The AI emperor has no DAUs why most devs still don't use code AI: Quinn Slack
This talk argues that despite the hype, the vast majority of developers do not actively use code AI tools beyond basic autocomplete. The speaker, Quinn Slack, emphasizes that the current low adoption rate poses a significant risk to the AI industry's financial sustainability, which relies on widespread enterprise adoption. The core thesis is that the industry needs to shift focus from hype to building genuinely useful, deeply integrated tools that developers will use daily.
- A Practical Guide to Efficient AI: Shelby Heinecke
This talk focuses on practical techniques for making AI models more efficient, bridging the gap between demos and production-ready AI. Efficiency is presented as crucial for deploying AI at scale, especially given resource constraints in cloud, on-premise, and edge device environments. The speaker introduces five dimensions for building and deploying efficient AI: efficient architectures, pre-training, fine-tuning, inference, and prompting.
- Moondream: how does a tiny vision model slap so hard? — Vikhyat Korrapati
This talk introduces Moondream, a small, open-source vision language model with under two billion parameters, designed to be efficient and runnable on various devices. Despite its size, Moondream achieves performance comparable to much larger models on vision benchmarks. The development focused on creating a tool for developers, prioritizing accuracy and avoiding hallucinations, rather than general world knowledge.
- Navigating RAG Optimization with an Evaluation Driven Compass: Atita Arora and Deanna Emery
This talk focuses on optimizing Retrieval Augmented Generation (RAG) systems by employing an evaluation-driven approach. It highlights the common challenges encountered in RAG implementations, from data processing and retrieval to response generation, and proposes systematic methods for improvement. The core thesis is that continuous, data-driven evaluation is essential for effectively refining RAG performance and achieving desired outcomes.
- How Zapier Builds AI Products and Features with the Help of Braintrust: Ankur Goyal & Olmo Maldonado
This talk details Zapier's approach to building AI-powered products, emphasizing an iterative development process heavily reliant on robust evaluation and observability. The speakers share their journey of integrating AI features, highlighting the challenges and successes encountered when developing and refining tools like the AI Zap Builder and Zapier Copilot. Their strategy involves close collaboration between product and engineering teams, continuous testing, and leveraging platforms like Braintrust to ensure product quality and performance.
- What It Actually Takes to Deploy GenAI Applications to Enterprises: Arjun Bansal and Trey Doig
This talk addresses the challenges of deploying Generative AI applications to enterprises, focusing on the critical need for accuracy and trust. It highlights how traditional methods of analyzing customer interactions through manual review or scripted analysis are insufficient at scale. Generative AI offers a solution by enabling 100% coverage of conversations, surfacing unknown insights, and transforming the process of understanding customer needs and business operations.
- Knowledge Graphs & GraphRAG: Techniques for Building Effective GenAI Applications: Zach Blumenthal
This talk explores building effective Generative AI applications using Knowledge Graphs and Graph Retrieval Augmented Generation (GraphRAG). It demonstrates how to combine vector search with graph traversal and graph embeddings to create more personalized and contextually relevant AI-driven responses. The session focuses on practical implementation using Neo4j, LangChain, and OpenAI, showcasing a fashion recommendation email generation use case.
- AI Engineering Without Borders — swyx
This talk challenges the conventional, often arbitrary, boundaries within AI engineering. It posits that AI, by its nature, disregards human-made borders like language, copyright, and even ground truth. The core thesis is that AI engineers should embrace this border-disrupting nature, moving beyond rigid definitions and disciplines to focus on the fundamental laws and utility of AI for humanity's benefit.
- State Space Models for Realtime Multimodal Intelligence: Karan Goel
This talk introduces State Space Models (SSMs) as a promising architecture for real-time multimodal intelligence, contrasting them with traditional batch-oriented AI systems. The core thesis is that SSMs, by efficiently modeling long contexts and compressing information, can enable faster, cheaper, and more ubiquitous AI applications, particularly in areas requiring instant responses like conversational interfaces and robotics.
- Second Order Effects of AI: Cheng Lou
This talk explores the unpredictable second-order effects of AI, moving beyond immediate consequences to consider broader societal and individual impacts. It suggests that as AI automates tasks, the focus may shift from efficiency to personal learning and skill development, leading to new forms of interaction and creative expression. The discussion also touches on how AI can widen information bandwidth and personalize user experiences, potentially transforming communication and interface design.
- Build an AI Research Agent: Apoorva Joshi
This talk introduces the fundamental concepts of AI agents, detailing their components and use cases. It emphasizes that agents are best suited for complex, multi-step tasks requiring the integration of various capabilities, such as data aggregation, visualization, and reasoning, or when personalization and adaptive responses are necessary. The session also provides a hands-on guide to building a research agent from scratch, incorporating tools, memory, and reasoning patterns.
- The Multimodal Future of Education: Stefania Druga
This talk explores the transformative potential of multimodal AI in education, addressing critical needs for improved literacy, bridging learning gaps, and upskilling. It highlights how AI can serve as a powerful tool for students and educators alike, fostering deeper understanding and engagement. The presentation emphasizes the importance of cultivating AI literacy and critical thinking from a young age.
- Productionizing GenAI Models – Lessons from the world's best AI teams: Lukas Biewald
The core thesis is that while Generative AI models are easy to demo, productionizing them presents significant challenges. The talk emphasizes that the AI development process is fundamentally experimental and non-deterministic, unlike traditional software development. This necessitates a robust approach to tracking learnings, ensuring reproducibility, and building comprehensive evaluation frameworks to move AI applications from demo to production successfully.
- Code Generation and Maintenance at Scale: Morgante Pell
This talk focuses on the practical application of AI agents for large-scale code generation and maintenance, moving beyond simple autocompletion or generating entirely new applications. The core thesis is that AI agents are most effective when used to augment the capabilities of experienced engineers, allowing them to tackle complex modifications across vast codebases more efficiently. The presentation emphasizes the need for specialized tools that operate at a higher level of abstraction than traditional IDEs to manage the increasing volume of AI-generated code.
- The Hierarchy of Needs for Training Dataset Development: Chang She and Noah Shpak
This talk addresses the critical importance of training data development for large language models, emphasizing that the quality and format of data directly impact model performance. It proposes a hierarchy of needs for dataset development, starting with clean data and progressing through evaluations, dataset management, and advanced techniques like synthetic data generation and quality scoring. The discussion highlights the challenges posed by the increasing scale and multimodality of AI data and introduces the Lance format as a solution for efficient data handling.
- No more bad outputs with structured generation: Remi Louf
This talk addresses the inherent unreliability of large language model (LLM) outputs, which often fail to adhere to desired formats, leading to errors like JSON decoding failures. The core thesis is that structured generation, a technique for guiding LLMs to produce outputs in specific formats, can resolve these issues. The open-source library Outlines is presented as a tool to implement structured generation, enabling more robust and predictable LLM applications.
- Realtime Data Connectivity for AI: Tanmai Gopal
This talk addresses the challenge of connecting Large Language Models (LLMs) to real-time data and business logic. The core thesis is that LLMs, while capable of complex tasks like coding, often struggle to interact intelligently with structured and unstructured data sources. The proposed solution involves making live data and business logic available to LLMs as tools, enabling them to perform data-driven actions.
- Architecting and Testing Controllable Agents: Lance Martin
This talk focuses on architecting and testing reliable AI agents, moving beyond traditional chains and open-ended agents that suffer from poor reliability. The presenter introduces Lang Graph as a solution for building controllable agents that balance flexibility with robustness by expressing control flows as graphs with nodes and edges, incorporating state for memory.
- Breaking AI's 1-GHz Barrier: Sunny Madra (Groq)
This talk draws a parallel between the historical achievement of microprocessors breaking the 1 GHz barrier and the current rapid advancements in Large Language Models (LLMs). It posits that LLMs are innovating at a pace exceeding Moore's Law, leading to transformative capabilities. The core thesis is that the increasing speed and efficiency of LLMs will fundamentally change how we build, run, and scale software, potentially making LLMs the core of future computing paradigms, akin to the industrial revolution's impact on manufacturing.
- Making Open Models 10x faster and better for Modern Application Innovation: Dmytro (Dima) Dzhulgakov
This talk focuses on the advantages of using open-source AI models for modern applications, emphasizing how they can achieve significantly better performance and speed compared to proprietary models. The presentation highlights the challenges in productionizing open models, such as complex setup, optimization, and achieving enterprise-scale reliability, and introduces Fireworks AI's solutions for these issues. The core argument is that by optimizing serving stacks and enabling fine-tuning, open models can be made 10x faster and more cost-effective for a wide range of applications.
- We accidentally made an AI platform: Jamie Turner
This talk introduces Convex as a platform that accidentally became an AI development tool. Initially designed to replace traditional backend engineering with a high-level, functional interface similar to Firebase, Convex's pervasive data flow tracking and reactive paradigm proved exceptionally well-suited for generative AI applications. The platform's ability to seamlessly sync state between backend processes and the application frontend has led to over 90% of its projects being AI-related since the post-ChatGPT boom.
- Building and Scaling an AI Agent Swarm of low latency real time voice bots: Damien Murphy
This talk introduces a new voice agent API that consolidates speech-to-text, large language model (LLM) processing, and text-to-speech into a single audio-in, audio-out interface. This approach simplifies development by abstracting away the complexities of integrating these components, enabling developers to build real-time voice applications more efficiently. The presentation demonstrates how to create a voice AI agent capable of function calling and discusses strategies for scaling these agents.
- The era of unbounded products: Designing for Multimodal IO: Ben Hylak
This talk explores the evolution of product design, moving from traditional screen-based interfaces to "unbounded products" that transcend typical input methods. It emphasizes the challenges of unpredictability and user confusion in these new paradigms and proposes design structures to create clarity and familiarity. The discussion covers historical precedents, current AI product patterns, and future interface directions.
- Everything you need to know about Fine-tuning and Merging LLMs: Maxime Labonne
This talk explores the practical aspects of fine-tuning and merging Large Language Models (LLMs). It outlines the LLM training lifecycle, distinguishing between pre-training, supervised fine-tuning (SFT), and preference alignment. The discussion emphasizes when fine-tuning is necessary, often driven by the need for customization and control beyond what prompt engineering can achieve, and introduces various libraries and techniques for these processes.
- LLM Scientific Reasoning: How to Make AI Capable of Nobel Prize Discoveries: Hubert Misztela
This talk explores how Large Language Models (LLMs) can be leveraged to accelerate scientific discovery, moving beyond simple question-answering to tackle complex reasoning tasks. It highlights a historical paradox in biology that took years to resolve, suggesting that LLMs could potentially speed up such processes by analyzing scientific literature. The core idea is to enhance Retrieval Augmented Generation (RAG) systems with reasoning capabilities, both before and after retrieval, to handle semantically complex questions and uncover novel insights.
- How to Construct Domain Specific LLM Evaluation Systems: Hamel Husain and Emil Sedgh
This talk outlines a systematic approach to constructing domain-specific LLM evaluation systems. It emphasizes the importance of moving beyond initial "vibe checks" and prompt engineering to establish a robust framework for consistent AI improvement. The core thesis is that a well-defined evaluation system, built on foundational principles, is crucial for developing production-ready AI applications and unlocking advanced capabilities like fine-tuning.
- From model weights to API endpoint with TensorRT LLM: Philip Kiely and Pankaj Gupta
This talk introduces TensorRT-LLM, an NVIDIA SDK designed for high-performance deep learning inference on NVIDIA GPUs. It focuses on optimizing large language models (LLMs) to achieve higher throughput and lower latency, crucial for production environments. The presentation covers building, configuring, benchmarking, and deploying TensorRT-LLM engines, emphasizing practical application through live coding and detailed explanations.
- Build enterprise generative AI apps using Llama 3 at 1,000 tokens/s on the SambaNova AI platform
This talk introduces SambaNova's full-stack AI platform, highlighting its capability to achieve over 1,000 tokens per second inference speed for Llama 3. The platform integrates hardware, system software, and AI models to simplify the development and deployment of enterprise-grade generative AI applications. It aims to combine the broad capabilities of large monolithic models with the adaptability and control of smaller, open-source models.
- Going beyond RAG: Extended Mind Transformers - Phoebe Klett
This talk introduces Extended Mind Transformers (EMTs), a novel approach to enhance language model performance by integrating a retrieval mechanism directly into the Transformer's attention mechanism. Unlike traditional methods like long context windows or Retrieval-Augmented Generation (RAG), EMTs allow the model to dynamically retrieve and attend to relevant information from a memory store during generation, without requiring fine-tuning. This method aims to improve accuracy, enable more granular citations, and reduce hallucinations.
- Judging LLMs: Alex Volkov
This talk uses a courtroom drama format to highlight common pitfalls in AI engineering, particularly concerning Large Language Model (LLM) development and deployment. The core thesis is that rigorous evaluation, logging, and prompt iteration are crucial for successful LLM projects, and neglecting these can lead to significant problems. The presentation emphasizes the importance of a human-in-the-loop approach for effective LLM judging and evaluation.
- Pydantic is STILL all you need: Jason Liu
This talk argues that Pydantic remains the essential tool for building reliable applications with large language models (LLMs). The core thesis is that by leveraging Pydantic's schema definition and validation capabilities, developers can move beyond the unreliability of unstructured text outputs and program with data structures, similar to classical software development. This approach enhances composability, reliability, and developer experience when interacting with LLMs.
- Hyperspace More Nodes Is All You Need: Nicolas Schlaepfer
Hyperspace is developing a decentralized AI network that leverages community-contributed computing resources without relying on centralized GPUs. Their product, Hyperspace, aims to provide a superior AI experience by integrating diverse, fine-tuned models rather than relying on a single large model. This platform combines prompt engineering, visual react flow, Python execution, and retrieval-augmented generation (RAG) for advanced AI workflows.
- GraphRAG: The Marriage of Knowledge Graphs and RAG: Emil Eifrem
This talk introduces Graph RAG, a method that combines knowledge graphs with Retrieval Augmented Generation (RAG) to enhance the accuracy and capabilities of LLM applications. It posits that by integrating structured knowledge graph data with unstructured text retrieval, developers can achieve more precise and comprehensive answers, moving beyond the limitations of traditional RAG.
- Building State of the Art Open Weights Tool Use: The Command R Family: Sandra Kublik
This talk introduces the Command R family of open-weight models, highlighting their capabilities in retrieval augmented generation (RAG), tool use, and sequential reasoning. The models are designed to be competitive with leading proprietary models like GPT-4 Turbo and Claude Opus, while being significantly smaller and more cost-effective. The presentation emphasizes the engineering decisions behind these models, focusing on addressing challenges in prompt sensitivity, model bias, and steering knowledge to improve performance in RAG and tool-use scenarios.
- Git push get an AI API: Ryan Fox-Tyler
This talk demonstrates how to build and iterate on AI features within applications by leveraging a platform that integrates models with traditional programming paradigms. It showcases practical examples, starting with a game called Hyper Categories to illustrate AI-powered scoring and uniqueness validation, and then progresses to applying these concepts to real-world problems like triaging GitHub issues and enabling natural language search for similar issues. The core thesis is that combining AI models with functions and data sources, managed through a cohesive development environment, simplifies the creation of powerful AI-driven applications.
- Hypermode Launch: Kevin Van Gundy
This talk introduces Hypermode, a runtime and toolset designed to simplify AI integration into applications. The core thesis is that rapid iteration is key to success in software development, a principle that also applies to AI. Hypermode aims to reduce the friction and fear associated with getting AI wrong, enabling developers to experiment, integrate, and observe AI functions in production with ease.
- Disrupting the $15 Trillion Construction Industry with Autonomous Agents: Dr. Sarah Buchner
This talk introduces Trun Tools, a generative AI provider focused on the construction industry. It highlights the immense complexity and data volume in construction projects, where millions of pages of documentation are common for a single skyscraper. The core problem addressed is data discrepancies and rework, which cost the industry an estimated $1.5 trillion annually. Trun Tools aims to solve this by creating a centralized knowledge base, referred to as the "brain behind construction," and deploying AI agents on top of it to assist human decision-making and resolve issues.
- 10x Development: LLMs For the working Programmer - Manuel Odendahl
This talk explores how programmers can leverage Large Language Models (LLMs) to significantly boost their productivity, aiming for a "10x" improvement. The core thesis is that LLMs should be treated not as autonomous reasoning agents, but as powerful translation engines and world simulators. By understanding how to decompose problems into language translation steps and by creatively framing prompts, developers can unlock new levels of efficiency and innovation.
- Building Reliable Agentic Systems: Eno Reyes
This talk explores practical lessons learned from building autonomous software engineering systems, referred to as droids. It defines agentic systems by three core characteristics: planning, decision-making, and environmental grounding. The discussion emphasizes strategies for enhancing reliability and effectiveness in these systems, drawing inspiration from control systems, robotics, and symbolic AI.
- Building with Anthropic Claude: Prompt Workshop with Zack Witten
This talk, Building with Anthropic Claude: Prompt Workshop with Zack Witten, focuses on practical prompt engineering techniques for using Anthropic's Claude models. The session involved live testing and iteration of user-submitted prompts, demonstrating how to refine prompts for better accuracy, conciseness, and desired output formats. Key themes include structuring prompts, utilizing specific formatting like XML, managing response length, and employing advanced techniques like pre-fills and stop sequences.
- Running AI Application in Minutes w/ AI Templates: Gabriela de Queiroz, Pamela Fox, Harald Kirschner
This talk demonstrates how to quickly deploy AI applications using AI templates and Azure services. It highlights the benefits of Microsoft's startup programs, including Azure credits and access to developer tools, and introduces AI templates as a way to accelerate development with pre-built skeletons. The session focuses on practical, hands-on deployment of various AI applications, including chat and Retrieval Augmented Generation (RAG) models.
- Decoding the Decoder LLM without de code: Ishan Anand
This talk provides a deep dive into the inner workings of Large Language Models (LLMs) by dissecting GPT-2 small and reconstructing its functionality within a Microsoft Excel spreadsheet. The presentation aims to demystify LLMs for individuals without formal machine learning degrees, illustrating how text generation is fundamentally a complex mathematical problem. It covers the model's anatomy, its thought process through a virtual MRI, and concludes with a demonstration of AI "brain surgery" to alter its behavior.
- Using agents to build an agent company: Joao Moura
This talk explores the practical application of AI agents in building a company, emphasizing their ability to handle complex automations and adapt in real-time. The speaker argues that the software development paradigm is shifting from strictly defined inputs and outputs to a more fluid, AI-driven approach. The core thesis is that AI agents are not just a future concept but a present reality, rapidly transforming how businesses operate and build products.
- What's new from Anthropic and what's next: Alex Albert
This talk draws a parallel between the early adoption of electricity in factories and the current integration of AI and LLMs. It argues that true innovation comes not from simply replacing old technologies with new ones, but from redesigning systems from the ground up with the new technology at their core. The presentation highlights Anthropic's latest advancements, particularly Claude 3.5 Sonnet and the new Artifacts feature, as steps towards enabling this paradigm shift in AI product development.
- How Codeium Breaks Through the Ceiling for Retrieval: Kevin Hou
This talk explores how Codeium addresses limitations in traditional retrieval methods for AI agents, particularly within code generation. The core thesis is that current embedding-based retrieval hits a ceiling due to its inability to effectively reason over multiple contextual elements simultaneously. Codeium proposes a vertically integrated approach, leveraging custom infrastructure and models to enable a more powerful and efficient retrieval system called M query.
- Emergence Launch: AI Agents and the future enterprise: Dr. Satya Nitta
Emergence is developing AI agents and infrastructural platforms to enable complex enterprise workflows. The core thesis is that AI agents, capable of acting, planning, and verifying, are poised to drive significant productivity benefits, particularly within enterprise environments. The company is focused on advancing the science of AI agents through self-improvement, planning, and reasoning, with an emphasis on agent-oriented programming for stitching agents together.
- Low Level Technicals of LLMs: Daniel Han
This talk delves into the low-level technical aspects of Large Language Models (LLMs), focusing on practical implementation details and common pitfalls. The speaker, Daniel Han, shares insights gained from analyzing and debugging open-source models like Gemma and NVIDIA's NeMo, highlighting issues in tokenization, model architecture, and training methodologies. The presentation aims to equip AI engineers with the knowledge to identify and fix bugs, optimize fine-tuning processes, and understand the underlying mathematical principles of LLMs.
- Fixing bugs in Gemma, Llama, & Phi 3: Daniel Han
This talk addresses common bugs and issues encountered when fine-tuning and deploying open-source large language models, specifically focusing on Gemma, Llama 3, and Phi 3. It provides practical solutions and best practices to ensure successful model training and inference, highlighting the importance of careful attention to tokenization, model templates, and export formats.
- Copilots Everywhere: Thomas Dohmke and Eugene Yan
This talk explores the evolution and integration of AI-powered coding assistants, focusing on GitHub Copilot. It posits that AI tools are shifting from simple autocompletion to comprehensive development partners, aiming to enhance developer productivity and democratize access to coding. The core thesis is that AI should augment human capabilities, not replace them, by handling tedious tasks and facilitating exploration within the codebase.
- Unlocking Developer Productivity across CPU and GPU with MAX: Chris Lattner
This talk introduces MAX, an AI framework designed to enhance developer productivity and performance across both CPU and GPU hardware. It addresses the fragmentation and complexity in the current AI development landscape, aiming to provide a unified, high-performance solution that allows developers to own and control their AI models and data. MAX focuses on inference and aims to simplify the deployment of PyTorch models and generative AI applications.
- From Software Developer to AI Engineer: Antje Barth
This talk outlines five practical steps for software developers transitioning to AI engineering. It emphasizes that while deep ML research is no longer a prerequisite, understanding AI fundamentals, leveraging AI developer tools for productivity, and actively prototyping with AI models are crucial. The presentation highlights the evolving role of AI engineers and the tools available to facilitate this transition, including AI-powered assistants and managed services for model experimentation.
- Lessons From A Year Building With LLMs
This talk, delivered by a collective of six AI engineers, distills a year of practical experience building with Large Language Models (LLMs). The core thesis is that successful LLM application development hinges not on proprietary models, but on strategic product building, robust operational processes, and meticulous tactical execution. The speakers emphasize a continuous improvement loop, drawing parallels to established software engineering and machine learning practices, to navigate the complexities and uncertainties inherent in LLM development.
- Open Challenges for AI Engineering: Simon Willison
The talk addresses the evolving landscape of AI models, highlighting the democratization of advanced capabilities previously exclusive to GPT-4. It emphasizes that the barrier to high-quality AI models has been significantly lowered, with multiple organizations now offering comparable performance. However, the core challenges shift from model access to effective and responsible utilization, including navigating complex tool interfaces, building trust, and mitigating risks like prompt injection and the proliferation of unreviewed AI-generated content.
- Llamafile: bringing AI to the masses with fast CPU inference: Stephen Hood and Justine Tunney
Llamafile is an open-source project from Mozilla aiming to democratize AI access by packaging AI models into single, executable files. These files run on various operating systems and hardware, including CPUs and GPUs, without requiring installation. The project emphasizes improving CPU inference speed and enabling local, private AI applications.
- The Future of Knowledge Assistants: Jerry Liu
This talk explores the evolution of knowledge assistants, moving beyond basic retrieval-augmented generation (RAG) to sophisticated single-agent and multi-agent systems. The core thesis is that building production-grade knowledge assistants requires advancements in data processing, agentic query flows, and multi-agent orchestration to handle complex tasks and interactions effectively.
- The Making of Devin by Cognition AI: Scott Wu
This talk introduces Devin, an autonomous AI software engineer developed by Cognition AI. It highlights Devin's capabilities through demonstrations, including building a mobile-friendly website and contributing to Cognition AI's own codebase by adding a search bar. The presentation emphasizes the shift from text completion AI to autonomous agents capable of decision-making and problem-solving within a software engineering context.
- From Text to Vision to Voice Exploring Multimodality with Open AI: Romain Huet
This talk explores the evolution and future of AI, focusing on OpenAI's advancements in multimodality, enabling more natural human-computer interactions. It highlights the journey from text-based models to integrating vision, audio, and video, culminating in the GPT-4o model. The core thesis is that by embracing these multimodal capabilities and focusing on developer experience, builders can create the next generation of AI-native products.
- The Code AI Maturity Model and What It Means For You: Ado Kukic
This talk introduces a Code AI Maturity Model, conceptualized as six distinct levels across three categories: human-initiated, AI-initiated, and AI-led code development. This model aims to categorize the evolving capabilities of AI in software development, drawing parallels to the SAE levels of vehicle autonomy to illustrate increasing AI involvement and independence. The core thesis is that AI's role in coding is progressing from simple assistance to full autonomy in the software development lifecycle.
- How to Become an AI Engineer from a Fullstack Background - Reid Mayo
This talk presents a syllabus designed to guide full-stack engineers into AI engineering roles, assuming no prior AI/ML background. It emphasizes leveraging foundational models and new techniques to deploy AI solutions, moving beyond traditional ML expertise and extensive data collection. The approach focuses on understanding fundamentals, efficient learning strategies, and practical application through a structured curriculum.
- Using AI to Build an Infinite Game: Jeff Schomay
This talk details the process of building a game with entirely AI-generated content, focusing on creating unique, procedurally generated scenes for an infinite exploration experience. The core idea is to leverage AI for both text descriptions of game scenes and the accompanying visual assets, ensuring each playthrough is distinct.
- GPT Web App Generator - 10,000 apps created in a month: Matija Sosic
This talk introduces Mage, a GPT-powered full-stack web application generator that enabled the creation of over 10,000 applications in a single month. It details how Mage works, its underlying technology, and the reasons behind its performance and affordability, positioning it as a novel approach to SaaS starters.
- Storyteller: Building Multi-modal Apps with TS & ModelFusion - Lars Grammel, PhD
This talk details the creation of Storyteller, an application designed to generate short audio stories for preschool children. The system utilizes TypeScript and the ModelFusion AI orchestration library, accepting a voice input to produce approximately two-minute audio narratives. Key technical challenges addressed include achieving responsiveness, ensuring quality, and maintaining consistency in the generated content.
- Open Questions for AI Engineering: Simon Willison
This talk reflects on the past year of AI engineering, highlighting the rapid evolution of the field and posing open questions for its future. The speaker emphasizes how large language models (LLMs) uniquely enable building previously impossible things and doing them faster. The discussion covers the impact of user interfaces beyond chat, the challenges of AI safety and security, the democratization of model development, and the potential for LLMs to empower individuals to automate tedious tasks.
- Trust, but Verify: Shreya Rajpal
This talk introduces a new programming paradigm for generative AI applications called Trust, but Verify. It addresses the inherent non-deterministic nature of machine learning models, which leads to issues like hallucinations and unreliability, especially when moving from prototyping to production. The proposed solution involves implementing a verification suite around LLM outputs to ensure correctness and build trust in AI applications.
- Harnessing the Power of LLMs Locally: Mithun Hunsur
This talk introduces lm.rs, a Rust library designed for local inference of large language models (LLMs). It highlights the advantages of running LLMs locally, such as greater control, privacy, and potential cost savings compared to cloud-based solutions. The library aims to provide a flexible, Rust-native experience for developers to integrate various LLM architectures into their applications.
- The Weekend AI Engineer: Hassan El Mghari
This talk focuses on building successful AI applications rapidly, often within a weekend, and scaling them to millions of users. The speaker shares personal projects and lessons learned, emphasizing simplicity, leveraging off-the-shelf APIs, and the importance of user experience. The core thesis is that even basic AI tools can achieve significant traction if built with a focus on user needs and iterative development.
- 120k players in a week: Lessons from the first viral CLIP app: Joseph Nelson
This talk details the creation and viral success of paint.wtf, an AI Pictionary game where users draw prompts generated by GPT-3, and CLIP judges the submissions based on image-text similarity. The game attracted 120,000 players in its first week, handling seven requests per second at its peak. The presentation covers the technical implementation, lessons learned from user interactions, and the potential of multimodal AI in building new types of applications.
- Building Production-Ready RAG Applications: Jerry Liu
This talk focuses on building production-ready Retrieval Augmented Generation (RAG) applications, addressing the limitations of naive RAG implementations. It emphasizes that while RAG is powerful for querying data, common issues like poor retrieval accuracy and low response quality hinder production use. The presentation outlines strategies for improving RAG performance across the entire pipeline, from data ingestion to synthesis, with a strong emphasis on evaluation and iterative optimization.
- Retrieval Augmented Generation in the Wild: Anton Troynikov
This talk explores the limitations of basic retrieval-augmented generation (RAG) loops and argues for more sophisticated memory systems to power advanced AI applications. It highlights the need for RAG systems that can incorporate human feedback, self-update based on agent interactions, and dynamically adapt to evolving data. The core thesis is that future powerful AI applications will require retrieval systems far beyond simple search indexes.
- Domain adaptation and fine-tuning for domain-specific LLMs: Abi Aryan
This talk provides a comprehensive overview of domain adaptation and fine-tuning techniques for large language models (LLMs). It aims to consolidate existing literature to help users make informed decisions for their specific enterprise or hobby use cases. The core thesis is that while general-purpose LLMs are powerful, adapting them to specific domains is crucial for optimal performance, and various methods exist to achieve this efficiently.
- Pragmatic AI with TypeChat: Daniel Rosenwasser
This talk introduces TypeChat, a library designed to bridge the gap between powerful language models and traditional applications. It addresses the challenge of reliably extracting structured data from natural language outputs, proposing a type-driven approach for both guidance and validation. The core thesis is that by leveraging type definitions, developers can effectively control and verify the data received from AI models, simplifying complex integrations.
- Building Reactive AI Apps: Matt Welsh
This talk introduces AI.JSX, an open-source framework designed to simplify the development of reactive AI applications, particularly for TypeScript and full-stack developers. It aims to abstract away the complexities of LLM app development, such as vector databases and context window limits, allowing developers to focus on building AI-powered user experiences. The framework enables the creation of sophisticated AI applications through a component-based, declarative approach similar to React.
- AI Engineering 201: The Rest of the Owl
This talk explores the engineering challenges and emerging patterns in building AI systems beyond just the inference engine. It posits that the current wave of AI development is focused on creating language user interfaces, analogous to the shift from text-based terminals to graphical user interfaces. The presentation delves into various architectural patterns, including retrieval-augmented generation (RAG) and structured outputs, and discusses the critical need for robust monitoring, observability, and evaluation in the AI engineering lifecycle.
- Move Fast Break Nothing: Dedy Kredo
This talk proposes a GAN-like architecture for code generation, emphasizing a two-component system: one for code generation and another for code integrity analysis. This critic component analyzes generated code, identifies edge cases, and aims to produce high-quality code that aligns with developer intent. The approach focuses on behavior coverage as a more valuable metric than traditional code coverage.
- The AI Evolution: Mario Rodriguez, GitHub
This talk by Mario Rodriguez, VP of Product at GitHub, explores the evolution of AI in software development, focusing on GitHub Copilot's journey and future potential. It highlights the importance of user experience, rapid iteration, and responsible AI practices in building impactful developer tools. The presentation also looks ahead to a future where AI assists developers in achieving goals and constraints rather than just executing procedures, transforming the entire software development lifecycle.
- The AI Pivot: With Chris White of Prefect & Bryan Bischof of Hex
This talk features Chris White of Prefect and Bryan Bischof of Hex discussing their companies' approaches to integrating AI. Both emphasize that AI is becoming a foundational element for data platforms, crucial for staying competitive and enhancing user experience. They share insights on strategic decision-making, resource allocation, and the practical challenges of building and deploying AI features within existing products.
- [Workshop] AI Engineering 201: Inference
This workshop focuses on the engineering challenges and considerations for deploying AI inference workloads, moving beyond the initial development of AI-powered applications. It breaks down the process into two main parts: understanding and optimizing inference, and then exploring broader architectural patterns, monitoring, and evaluation for AI applications. The core thesis is that while the capabilities of AI models are rapidly advancing, the engineering behind robust, efficient, and scalable inference is crucial for successful product deployment.
- The Hidden Life of Embeddings: Linus Lee
This talk explores the concept of latent spaces within AI models, particularly focusing on embeddings, as a means to gain more direct control and understanding of model behavior. Instead of relying solely on indirect prompting, the speaker advocates for interacting with the internal representations of models to achieve more precise manipulation and insight into their decision-making processes.
- [Workshop] AI Engineering 101
This workshop, AI Engineering 101, aims to provide a foundational understanding of key concepts and tools for aspiring AI engineers. It covers programmatic interaction with Large Language Models (LLMs), the underlying principles of embeddings and tokens, and practical applications like code and image generation. The session emphasizes hands-on experience, guiding participants through building a Telegram bot that integrates these AI capabilities. The core thesis is that by understanding these fundamentals, engineers can effectively leverage AI to build innovative products and augment their development workflows.
- Supabase Vector: The Postgres Vector database: Paul Copplestone
This talk makes a case for using PostgreSQL with the PG Vector extension as a robust solution for storing and querying vector embeddings. It highlights how this approach integrates seamlessly with existing PostgreSQL infrastructure, offering a powerful and extensible option for AI applications, particularly for developers building with open-source tools and services.
- Climbing the Ladder of Abstraction: Amelia Wattenberger
This talk explores how AI can enhance user interfaces by applying the concept of the "ladder of abstraction." Instead of viewing AI as purely automation or augmentation, the speaker proposes that augmentation is built upon smaller, automated tasks. The core idea is to use AI to generate and navigate different levels of detail for information, making complex tasks more manageable and intuitive, much like how digital maps adjust detail based on zoom level.
- The Intelligent Interface: Sam Whitmore & Jason Yuan of New Computer
This talk explores the future of human-computer interaction, moving beyond traditional text and voice inputs to embrace a more intuitive, context-aware, and multimodal interface. The core thesis is that computers should adapt to human circumstances and context, rather than forcing users to adapt to rigid, deterministic interfaces. This involves leveraging a wider range of sensory inputs and expressive outputs to create a more natural and intelligent interaction loop.
- Building Blocks for LLM Systems & Products: Eugene Yan
This talk outlines essential building blocks for developing effective Large Language Model (LLM) systems and products. It emphasizes the critical role of evaluations, retrieval augmented generation, guardrails, and feedback collection in the LLM development lifecycle. The core thesis is that a structured, iterative approach focusing on these components is key to building robust and reliable LLM applications.
- Pydantic is all you need: Jason Liu
This talk introduces Pydantic as a powerful tool for building applications with language models, focusing on structured prompting. The core thesis is that by defining data structures with Pydantic, developers can move beyond unreliable string-based interactions with LLMs and achieve more robust, maintainable, and type-safe code. This approach enhances data validation, simplifies prompt engineering, and enables LLMs to integrate more seamlessly with existing software systems.
- Building Context-Aware Reasoning Applications with LangChain and LangSmith: Harrison Chase
This talk explores building context-aware reasoning applications by integrating language models into larger systems. It emphasizes that while language models are powerful, they require external context and capabilities to perform complex tasks. The presentation outlines various approaches for providing context and enabling reasoning, highlighting the engineering challenges and tooling needed to develop these sophisticated AI applications.
- See, Hear, Speak, Draw: Logan Kilpatrick & Simón Fishman
This talk explores the burgeoning field of multimodal AI, moving beyond text-based interactions to incorporate vision and audio. While current applications often treat different modalities as separate "islands" connected by text, the future points towards unified models capable of processing and generating across various inputs and outputs simultaneously. The presentation highlights practical patterns and demos for building with existing multimodal capabilities, anticipating future advancements.
- The Age of the Agent: Flo Crivello
This talk explores a future where AI agents significantly reduce the friction of starting and running businesses, enabling individuals to achieve impact comparable to large corporations. It draws parallels to the internet's democratization of media creation, suggesting that AI will similarly empower individuals with unprecedented capabilities, leading to more creative and diverse business ventures.
- The Future of Work: Toran Bruce Richards, Silen Naihin et al
This talk focuses on the evolution and future potential of AI agents, particularly emphasizing the impact and growth of the open-source Auto-GPT project. It highlights how AI agents can significantly enhance productivity by automating menial tasks, freeing up human potential for creative work. The presentation underscores the importance of community-driven development and the ongoing efforts to improve agent capabilities, safety, and reliability.
- Building AI For All: Amjad Masad & Michele Catasta
This talk focuses on the evolution of programming tools and how AI is fundamentally changing software development. It highlights the historical progression from punch cards to text-based programming and the significant productivity gains achieved through advancements like IDEs and language servers. The core thesis is that AI is not just an add-on but needs to be deeply integrated into the entire programming workflow, aiming to empower a billion new developers.
- The 1,000x AI Engineer: Swyx
This talk argues that current AI engineers are positioned at a pivotal moment in technological history, comparable to mathematicians with the invention of zero or physicists in the early 20th century. The speaker emphasizes that this is the right time to be involved in AI, suggesting that the field is entering a significant growth phase driven by increasing compute power and advancements in language models. The core message is to embrace this opportunity and aim for higher orders of magnitude in ambition and impact.
- Announcing the AI Engineer Network: Benjamin Dunphy
This talk announces the launch of the AI Engineer Network, a new initiative aimed at fostering connections and knowledge sharing within the AI engineering community. It highlights the growth of AI engineering as an industry and introduces a new mobile event app, also named Network, designed to enhance attendee interactions at conferences. The app features AI-driven matching to connect individuals with shared problems and solutions.
- Principles for Prompt Engineering - Karina Nguyen (Claude Instant @ Anthropic)
This talk by Karina Nguyen from Anthropic focuses on the principles of effective prompt engineering for large language models, particularly Claude. The core thesis is that prompting is a form of creative writing that requires clarity, conciseness, coherence, consistency, and direction to guide models toward desired outputs. By understanding why prompting is challenging and applying structured writing principles, users can significantly improve model performance without retraining.