← Browse

World's Fair 2024

Build enterprise generative AI apps using Llama 3 at 1,000 tokens/s on the SambaNova AI platform

Overview

This talk introduces SambaNova's full-stack AI platform, highlighting its capability to achieve over 1,000 tokens per second inference speed for Llama 3. The platform integrates hardware, system software, and AI models to simplify the development and deployment of enterprise-grade generative AI applications. It aims to combine the broad capabilities of large monolithic models with the adaptability and control of smaller, open-source models.

Who should watch

Key takeaways

Notable quotes

*We are building the full stack from the ground up so that means we build our own chip.*
*We are actually integrating things into a very seamless experience from deciding on what Chip is going to work with what compute.*
*Our underlying platform is delivering the means to actually achieve the scale of a trillion parameters plus.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.