World's Fair 2024
Customized, production ready inference with open source models: Dmytro (Dima) Dzhulgakov
Overview
This talk focuses on the advantages and practicalities of using open-source models for production AI applications. It highlights how custom-tuned, smaller models can outperform larger, general-purpose proprietary models in terms of speed, cost, and domain-specific accuracy. The presentation also addresses the challenges of deploying open-source models, such as setup complexity and optimization, and introduces Fireworks AI's platform designed to streamline these processes.
Who should watch
- AI Engineers looking to deploy custom models efficiently.
- Product Managers evaluating cost-effective AI solutions.
- Builders seeking to leverage open-source models for specific use cases.
- Developers facing latency or cost issues with proprietary models.
- Teams interested in agentic architectures and function calling.
Key takeaways
- Open-source models offer significant advantages in domain adaptability, allowing for fine-tuning to achieve superior performance in narrow tasks compared to large, general-purpose proprietary models.
- Customizing open-source models can lead to substantial improvements in latency and cost, with some fine-tuned models performing up to 10 times faster for specific use cases like function calling.
- Deploying open-source models involves challenges like complex setup, optimization, and achieving enterprise-scale reliability, which platforms like Fireworks AI aim to solve.
- Fireworks AI provides a custom serving stack optimized for speed and efficiency, supporting various modalities including LLMs and image generation models like SDXL and SD3.
- The platform enables fine-tuning of models and offers serverless inference, allowing thousands of model variants to be served on a single GPU with pay-per-token pricing.
- Modern AI applications are increasingly becoming compound systems, integrating models with tools, RAG, and function calling, with agentic architectures at their core.
- Fireworks AI offers specialized models for function calling, such as FireFunction V, designed to translate user queries into complex reasoning and external tool interactions.
- The Fireworks AI platform provides a tiered offering from serverless prototyping to on-demand and enterprise-level solutions, supporting API compatibility with popular frameworks like LangChain and LlamaIndex.
Notable quotes
*You don't really need a customer support chatbot to know about 150 Pokemons or be able to write your poetry.*
*The challenges really come from three areas: complicated setup and maintenance, optimization, and getting it production ready.*
*We really focus on customizing this service tech to your needs, which basically means for your custom workload and for your custom cost and latency requirements we can tune it for those settings.*
Unofficial community note. Prefer the recording for nuance.