← Browse

Europe 2026

From MCP to Scale: Pipelines That Build Themselves — Rafael Levi, Bright Data

Rafael Levi , Bright Data

Overview

This talk explores building self-maintaining data collection pipelines using Large Language Models (LLMs) and specialized tools. It addresses the challenge of collecting large-scale data from websites, particularly those with anti-bot measures, by shifting from direct LLM parsing of every page to an agent-driven approach that builds and maintains scrapers. The core idea is to empower LLMs with the ability to create and manage the tools needed for data extraction, thereby saving significant computational resources and development time.

Who should watch

Key takeaways

Notable quotes

*Instead of telling, Hey, LLM, can you go and parse this for me? build a scraper that's going to parse it for me.*
*The whole session is how do you build pipelines, right? With LLM.*
*Public data is public data. It doesn't matter how you collect it.*

Watch on YouTube →

Unofficial community note. Prefer the recording for nuance.