Trails

Start here

Four short paths through 946 talks. Pick the problem you have — not the hype cycle.

Path 1

Evals first

If agents are non-deterministic, measurement is the product. Start here before you scale prompts or skills.

  1. 1

    The stake in the ground: don’t ship agent skills without a way to measure them.

    Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
  2. 2

    Hands-on: what evals look like when the unit of work is an agent, not a completion.

    Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize
  3. 3

    The category error — why unit-test instincts fail for stochastic systems.

    Evals Are Not Unit Tests — Ido Pesok, Vercel v0
  4. 4

    Field notes from running evals in production: what actually breaks.

    Five hard earned lessons about Evals — Ankur Goyal, Braintrust
  5. 5

    Where the practice is headed once measurement is table stakes.

    The Future of Evals - Ankur Goyal, Braintrust

Path 2

Coding agents in production

Harnesses, skills, and taste — how teams actually ship with Cursor, Claude Code, and friends.

  1. 1

    Compression as craft: when a skill beats a codebase of ceremony.

    Replacing 12K LoC with a 200 LoC Skill — David Gomes, Cursor
  2. 2
  3. 3

    How Composer was built: taste under the hood of a coding agent.

    Building Cursor Composer – Lee Robinson, Cursor
  4. 4

    What breaks when you run coding agents against real repos.

    Hard Won Lessons from Building Effective AI Coding Agents – Nik Pash, Cline
  5. 5

    The human skill that remains: wielding agents, not typing faster.

    The emerging skillset of wielding coding agents — Beyang Liu, Sourcegraph / Amp

Path 3

MCP & tools

From protocol origins to production MCP — what survives contact with real users.

  1. 1

    Protocol intent from the people shaping MCP’s next chapter.

    The Future of MCP — David Soria Parra, Anthropic
  2. 2

    Remote MCP in the wild: what shipping taught Anthropic.

    Remote MCPs: What we learned from shipping — John Welsh, Anthropic
  3. 3

    The skeptic’s cut — what’s still broken before you bet the architecture.

    MCP Is Not Good Yet — David Cramer, Sentry
  4. 4

    Skills + MCP together: closing the context gap in practice.

    Combine Skills and MCP to Close the Context Gap — Pedro Rodrigues, Supabase
  5. 5

    Security is not optional: insecure MCP dies in production.

    Your Insecure MCP Server Won't Survive Production — Tun Shwe, Lenses

Path 4

Ship products, not demos

For builders who care about judgment, UX, and getting something into someone else's hands.

  1. 1

    The north star: shipping to a real someone beats polishing the demo.

    Shipping something to someone always wins — Kenneth Auchenberg (ex. Stripe, VSCode)
  2. 2

    Product judgment when capabilities keep moving under you.

    Shipping Products When You Don't Know What they Can Do — Ben Stein, Teammates
  3. 3

    Leaving the prototype — quality code after the vibe spike.

    Beyond the Prototype: Using AI to Write High-Quality Code - Josh Albrecht, Imbue
  4. 4

    Broken AI UX is usually interaction design, not the model.

    Why Your AI UX Is Broken (and It's Not the Model's Fault) — Mike Christensen, Ably
  5. 5