Lenny’s Podcast · Podcast episode · Building AI products · Metrics, data & experimentation

The 100-person AI lab that became Anthropic and Google's secret weapon

Lenny Rachitskywith Edwin ChenDec 7, 202597 min▶ 61K
SourceLenny’s Podcast
KindPodcast episode
PublishedDec 7, 2025
Readers▶ 61K
Originalyoutube.com ↗
N:

Edwin Chen, founder and CEO of Surge AI, a bootstrapped data company that supplies human-generated training data and evaluations to frontier AI labs, discusses how it grew to over $1B in revenue with under 100 people. The conversation covers what high-quality AI training data means, why benchmarks and leaderboards may be steering models in the wrong direction, the rise of reinforcement learning environments, and his contrarian views on startups and building AI responsibly.

01Key learnings

  • Define quality in rich, specific terms for each domain (e.g., what makes a poem Nobel-worthy, not just whether it has eight lines) before trying to source data for it.
  • Gather thousands of signals on each worker and task, from background and expertise to performance, and use them like an ML problem to route the best people to the right projects.
  • Don't trust benchmarks at face value: many contain wrong answers, and optimizing for them can make models worse at messy real-world tasks.
  • Evaluate models with deep human expert testing rather than quick preference votes, since casual raters reward flashy formatting and hallucinations alike.
  • Reinforcement learning environments simulate real workplaces with tools and multi-step tasks; reward trajectories, not just final answers, since inefficient or reward-hacked paths can still hit the right answer.
  • Be deliberate about the objective function you optimize for; engagement-maximizing targets can produce sycophantic, time-sucking models instead of genuinely useful ones.
  • Build a company around one idea only you could build, and resist pivoting, hype, and hiring for resume-padding; bootstrapping and focus can work without VC money.
  • Expect AI models to differentiate as their creators' values shape their behavior, so decide what behavior you actually want, such as efficiency over endless iteration.
“I'm worried that instead of building AI that will actually advance us as a species, curing cancer, solving poverty, understand the universe, we are optimizing…”Edwin Chen · Lenny’s Podcast · 00:23:14
“It's literally optimizing your models for the types of people who buy tabloids at the grocery store.”Edwin Chen · Lenny’s Podcast · 00:24:05
“The values that the companies have will shape the model.”Edwin Chen · Lenny’s Podcast · 00:48:59

02Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.