Europe 2026
FLUX, Open Research, and the Future of Visual AI — Stephen Batifol, Black Forest Labs
Overview
This talk introduces FLUX, an open-source visual AI model developed by Black Forest Labs, and discusses its evolution and the future of visual AI. The presentation highlights FLUX's capabilities in text-to-image generation, image editing, and its progression towards visual intelligence. It also delves into a novel training methodology called Self Flow, designed to improve multimodal generative models by integrating representation learning directly into the generation process, eliminating the need for external encoders.
Who should watch
- AI Engineers interested in the latest advancements in visual AI and generative models.
- Product Managers and Builders looking for cutting-edge AI tools for product development.
- Researchers exploring new approaches to training multimodal AI systems.
- Developers working with image generation, editing, or video synthesis.
- Anyone interested in the future of AI in robotics and physical automation.
Key takeaways
- Black Forest Labs has developed a series of FLUX models, starting with FLUX 1 for text-to-image, FLUX Context for image editing, and FLUX 2 for advanced visual intelligence.
- FLUX 2 demonstrates state-of-the-art performance in both text-to-image generation and image editing, producing highly realistic outputs.
- FLUX 2 Klein offers near real-time image generation and editing, with latencies significantly lower than comparable models.
- Traditional generative model training often relies on external encoders for representation alignment, which can lead to scaling limitations and modality-specific issues.
- The Self Flow research introduces a new training approach that combines representation learning and generation within a single flow, using a student-teacher model architecture without external encoders.
- Self Flow has shown improved performance across multiple modalities including audio, images, and video, and converges faster than baseline methods.
- This new methodology also addresses common artifacts in generated content, such as incorrect text rendering and anatomical inaccuracies.
- The research points towards a future of visual AI that includes real-time interactive engines, world models for simulating physical interactions, and applications in robotics and automation.
Notable quotes
*Our first operating principle is to release state-of-the-art models. This is what we want to focus on.*
*We want to raise the bar on quality with every release we do.*
*This is where we believe this is the future.*
Unofficial community note. Prefer the recording for nuance.