World's Fair 2025
Foundry Local: Cutting-Edge AI experiences on device with ONNX Runtime/Olive — Emma Ning, Microsoft
Overview
Foundry Local is a Microsoft solution designed to enable developers to build cross-platform AI applications that run directly on user devices. It addresses the need for local AI by providing reasons such as offline access, enhanced privacy and security for sensitive data, cost efficiency for high-volume inference, and real-time latency requirements. The platform leverages ONNX Runtime for accelerated on-device inference across various hardware, integrating with Azure AI Foundry for model management and on-demand downloads.
Who should watch
- AI Engineers
- Product Managers
- Developers building AI-powered applications
- Those concerned with data privacy and security
- Developers needing real-time AI inference
- Users experiencing unreliable network connectivity
Key takeaways
- Local AI is necessary for applications requiring offline access, data privacy, cost efficiency at scale, and low latency.
- Advances in hardware and optimized AI models make on-device inference a practical reality.
- Foundry Local integrates Microsoft assets like Azure AI Foundry and ONNX Runtime for a seamless on-device AI experience.
- The platform offers a CLI and SDKs for easy model exploration and integration into custom applications across Windows and macOS.
- Foundry Local optimizes performance across various hardware accelerators through collaboration with vendors like Nvidia, Intel, AMD, and Qualcomm.
- Early customer feedback highlights ease of use and strong performance for on-device model deployment.
- The platform supports building local agents by combining models with MCP servers, with features like an OCR agent demonstrated.
- While local models may not match the full capability of cloud models, they unlock significant potential for new applications.
Notable quotes
*Many companies work with very sensitive data such as legal documents and patient information. They need to process that data entirely locally without anything ever leaving the device.*
*Foundry local is a perfect solution for these scenarios as it allows us to easily run Genai models locally.*
*If you're looking to deploy ondevice models, you can't go wrong with Foundry Local.*
Unofficial community note. Prefer the recording for nuance.