Infrastructure
24 sessions
- Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon)
This talk explores privacy-preserving intelligence, focusing on the infrastructure and methods required to build AI systems that respect user privacy. It delves into the technical challenges and potential solutions for developing intelligent agents and applications without compromising sensitive data. The core thesis revolves around enabling powerful AI capabilities while maintaining robust privacy guarantees.
- Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs
This talk addresses the economic challenges of relying on rented AI inference infrastructure, particularly for agents. The core thesis is that while renting models is useful for initial learning and finding product-market fit, long-term operation requires owning the inference infrastructure to manage costs effectively. The speaker shares personal experience moving agents off a paid API to owned infrastructure after encountering unsustainable expenses.
- WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar
This talk argues that while AI model intelligence has rapidly advanced, its real-world usefulness is hampered by a lack of business-specific context. The core thesis is that a "context layer" is the missing infrastructure for production-ready AI agents, analogous to how humans learn and operate within a company's knowledge, skills, and norms. This layer aims to transform implicit knowledge into machine-usable context, enabling AI to perform effectively in business environments.
- From fork() to Fleet: Designing an Agent Sandbox Cloud — Abhishek Bhardwaj, OpenAI
This talk explores the design principles behind creating secure and scalable agent sandbox clouds. It begins by explaining why AI models need the ability to execute code or use tools to handle tasks with verifiable rewards, such as math and coding problems. The presentation then delves into the evolution of sandbox technologies, from basic process isolation to advanced virtualization, emphasizing the critical need for robust security to protect against malicious or overzealous code execution in both research and product environments.
- Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
This talk addresses the critical challenge of building reliable infrastructure for non-deterministic AI agents, which are increasingly moving beyond simple chatbots to perform complex tasks like planning, tool coordination, and decision-making affecting production systems. The core thesis is that while AI models are inherently probabilistic, the infrastructure supporting them must be deterministic to ensure reliability, safety, and cost-effectiveness at scale.
- GPU Cloud Deployment Without Leaving Your IDE — Audry Hsu, RunPod
This talk introduces RunPod's Flash, a Python SDK designed to streamline the development and deployment of AI models on GPU infrastructure. The core thesis is that developers can significantly reduce iteration time by deploying functions directly to the cloud from their local IDE, eliminating the need for manual commits, Docker builds, and server configuration. This allows for rapid testing and iteration on GPU-accelerated workloads.
- Building safe Payment Infrastructure for the autonomous economy — Steve Kaliski, Stripe
This talk addresses the critical need for secure payment infrastructure in an increasingly autonomous economy, where AI agents will act as economic actors. It highlights the inherent risks of agents transacting online, such as purchasing from incorrect vendors, selecting the wrong items, or spending unintended amounts. The core thesis is that while discovery and exploration benefit from non-determinism, financial transactions require strict determinism to ensure safety and prevent fraud.
- What if the network was the sandbox? — Remy Guercio, Tailscale
This talk proposes a shift in how we conceptualize sandboxes for AI agents, suggesting the network itself can serve as the sandbox. Instead of relying on traditional methods like API keys or OAuth within a VM or container, the approach leverages network-level identity and permissions. This allows for granular control over agent access and behavior directly at the network layer, enabling more secure and observable AI agent deployments.
- Scaling Agents on Kubernetes with acpx and ACP — Onur Solmaz, OpenClaw
This talk explores the challenges and solutions for scaling open-source AI agents, particularly within the Kubernetes ecosystem. It introduces ACP (Agent Client Protocol) as a standardized way for human users to interact with agents, aiming to reduce duplicated effort in building client integrations. The discussion also covers developing automated workflows for tasks like code review and error reporting, emphasizing the potential for agents to handle complex, multi-step processes.
- The Small Model Infrastructure Nobody Built (So We Did) — Filip Makraduli, Superlinked
This talk addresses the gap in infrastructure for small model inference, particularly for AI search and document processing. The speaker, Filip Makraduli, shares his journey from overlooking inference to co-founding the open-source Superlinked Inference Engine (SIE). The core thesis is that effective inference for agentic workflows requires a holistic approach combining robust model support with scalable infrastructure, enabling developers to efficiently deploy and manage diverse small models.
- Why, and how you need to sandbox AI-Generated Code? — Harshil Agrawal, Cloudflare
This talk addresses the critical need to sandbox AI-generated code, emphasizing that such code should be treated as untrusted. The core thesis is that while AI can accelerate development, running its output without proper security measures is akin to executing code from an anonymous internet source, posing significant risks. The presentation advocates for applying established sandboxing principles, particularly capability-based security, to mitigate these threats.
- Infra that fixes itself, thanks to coding agents — Mahmoud Abdelwahab, Railway
This talk introduces a system where an AI coding agent monitors application infrastructure for issues and automatically generates pull requests to fix them. Instead of just alerting developers to problems like memory leaks or high error rates, the system aims to detect, diagnose, and propose solutions, significantly reducing manual debugging and speeding up the resolution process. The core idea is to shift from reactive alerting to proactive, automated self-healing infrastructure.
- Infrastructure for the Singularity — Jesse Han, Morph
This talk introduces Morph's Infinibranch technology, a virtualization, storage, and networking solution designed for advanced AI agents. It posits that as AI develops intelligence and personhood, it requires infrastructure that can match its speed and complexity. Infinibranch aims to provide a substrate for a cloud environment where AI agents can operate with zero latency, enabling reversible actions, parallel exploration of possibilities, and a more dynamic interaction with the digital world.
- Serving Voice AI at $1/hr: Open-source, LoRAs, Latency, Load Balancing - Neil Dwyer, Gabber
This talk details the experience of hosting an open-source voice AI model, Orpheus, for real-time consumer applications. The core challenge addressed is achieving low latency and high fidelity voice generation at a cost viable for widespread consumer use, which often requires near-free operation. The presentation highlights the technical hurdles in serving these models efficiently and presents solutions involving model fine-tuning, optimized infrastructure, and load balancing.
- Introduction to LLM serving with SGLang - Philip Kiely and Yineng Zhang, Baseten
This talk introduces SGLang, an open-source framework designed for high-performance serving of large language models (LLMs) and large vision models (LVMs). It emphasizes SGLang's production readiness, speed, and strong community support, highlighting its ability to provide day-zero support for new model releases and allow users to contribute to its development. The framework is presented as a valuable tool for optimizing LLM inference.
- Continuous Profiling for GPUs — Matthias Loibl, Polar Signals
This talk introduces continuous profiling for GPUs, a method to monitor and analyze GPU performance. It highlights the importance of profiling for improving application performance and reducing operational costs by enabling more efficient resource utilization. The approach leverages Linux eBPF for low-overhead, always-on profiling in production environments without requiring application instrumentation.
- What every AI engineer needs to know about GPUs — Charles Frye, Modal
This talk explains why AI engineers need to understand GPUs, shifting focus from API-based development to leveraging hardware capabilities. It draws an analogy to database usage, where developers don't build databases but must understand how to query them effectively. Similarly, AI engineers will increasingly need to understand GPU architecture, particularly tensor cores, to optimize performance for tasks like language model inference.
- Serving Voice AI at Scale — Arjun Desai (Cartesia) & Rohit Talluri (AWS)
This talk addresses the challenges and advancements in serving voice AI at scale, focusing on the critical need for low latency and high quality in real-time interactive applications. It introduces state space models (SSMs) as a more efficient alternative to transformers for handling long sequences, enabling faster and more natural voice interactions across various devices. The discussion highlights how these advancements are crucial for enterprise voice AI, customer support, gaming, and content creation.
- The Geopolitics of AI Infrastructure - Dylan Patel, SemiAnalysis
This talk examines the geopolitical landscape of AI infrastructure, focusing on the capabilities of China and the Middle East in contrast to US limitations. It highlights how China, despite sanctions, is advancing its AI chip manufacturing and deployment through innovative engineering and supply chain circumvention. Simultaneously, the Middle East is emerging as a significant hub for AI infrastructure development, attracting substantial investment and GPU resources, partly due to power constraints in the US.
- GPU-less, Trust-less, Limit-less: Reimagining the Confidential AI Cloud - Mike Bursell
This talk introduces confidential AI as a solution to trust issues in AI development and deployment. It explains how confidential computing, utilizing Trusted Execution Environments (TEs), protects data and models during processing, even from system administrators or hardware providers. This technology enables secure collaboration, monetization, and the use of sensitive data for AI tasks across various industries.
- Arrakis: How To Build An AI Sandbox From Scratch - Abhishek Bhardwaj, OpenAI
Arrakis is an open-source, self-hosted service for creating and managing AI sandboxes, designed for secure code execution and computer use by AI agents. It leverages microVMs to provide isolated environments, enabling AI models to safely utilize tools like code execution and search. The system emphasizes security, speed, and the ability for agents to backtrack and replan using snapshots, facilitating more complex task completion.
- How agents broke app-level infrastructure - Evan Boyle
This talk addresses the challenges of building reliable AI applications, particularly focusing on the compute layer. Traditional web infrastructure is ill-suited for the long-running, non-deterministic nature of AI workflows, which often involve extensive data ingestion and complex agentic processes. The core thesis is that current infrastructure limitations lead to unreliable user experiences and hinder rapid experimentation, necessitating new architectural approaches.
- Insights from Snorkel AI running Azure AI Infrastructure: Humza Iqbal and Lachlan Ainley
This talk focuses on Snorkel AI's approach to data development for enterprise AI models, emphasizing the challenges of achieving enterprise-grade quality, latency, and cost requirements with off-the-shelf large language models. Snorkel AI specializes in developing data to fine-tune these models, enabling them to meet specific business needs. The discussion also touches upon the infrastructure and best practices for running large-scale AI training and inference workloads on Azure.
- Unlocking Developer Productivity across CPU and GPU with MAX: Chris Lattner
This talk introduces MAX, an AI framework designed to enhance developer productivity and performance across both CPU and GPU hardware. It addresses the fragmentation and complexity in the current AI development landscape, aiming to provide a unified, high-performance solution that allows developers to own and control their AI models and data. MAX focuses on inference and aims to simplify the deployment of PyTorch models and generative AI applications.