Source: raw/Why_the_people_building_AI_can_t_tell_you_what_s_next_Dianne_Penn_Anthropic.md

Creator: Lenny’s Podcast (Lenny Rachitsky) with Dianne Penn | Platform: YouTube | Feed: Lenny’s Podcast | URL: https://www.youtube.com/watch?v=tivaWTTVRhY

Dianne Penn joined Anthropic in 2023 as its first technical product manager (five product engineers, one engineer on the entire API business) and now heads product for the research and labs teams — she has helped ship every model from Claude 2 through Fable and helped incubate Claude Code, MCP, skills, and Claude Design. The durable payload of this interview is a reusable PM methodology: her team’s saying that “evals are the new PRDs” — user pain becomes a measurable eval set that researchers can act on and re-run against every model version. Around it she describes why frontier products and frontier models need each other, and why nobody inside the labs can tell you exactly what’s next.

Key Takeaways

  • “Evals are the new PRDs.” For research PMs, the way to drive user value is figuring out the right user feedback and encoding it as evals — “a personification of that user need” — rather than defaulting to the document formats of the last two decades.
  • The methodology is repeatable: read failed-trajectory transcripts deeply → classify the failure theme (hallucination? overconfidence? wrong tool call? right document, wrong facts?) → build a ~30–40-example eval set with golden answers → hand to research → re-run on every model version to measure whether the pain point is actually improving. Lenny’s label, which Penn accepts: “test-driven development for PMs… you write the test first.”
  • The worked example: Claude-2-era users said “Claude is not good at following instructions.” Digging into the actual trajectories showed ~80% of what people meant was Claude not writing correct JSON. That became a 30–40-example schema-following eval; it now scores “always 100% or like 99.9” and is no longer a pain point — and schema-following became foundational to Claude working as an agent at all.
  • “You need frontier products in order to have frontier models.” Opus 4.5 landed the way it did because Claude Code existed as the vehicle: “Opus 4.5 wouldn’t have had that moment without a product like Claude Code, and Claude Code wouldn’t have had that type of adoption accelerated without Opus 4.5.”
  • Product overhang and user overhang are real: there is a lot still unexplored on current models — capability jumps are discontinuous and emergent, so you need evals even to notice a new capability exists.
  • PRDs are not dead. They survive for two jobs: aligning large groups on a source of truth (every model still gets a PRD for product surfaces, engineering, legal, safety) and exploring ambiguous zero-to-one opportunities where user pain points don’t exist yet (e.g., pre-launch computer use).
  • “Sweat the tokens as much as you sweat the pixels.” Transcript-reading is the new user-flow walkthrough; managers and tenured PM hires get the same hands-on onboarding as juniors, because you can’t judge good AI product work you haven’t built.

The Evals-as-PRDs Loop, Step by Step

  1. Access the pain point differently. Vague feedback (“Claude hallucinated”) is not actionable for researchers. Look at the consented trajectory behind the feedback: should Claude have called a tool at that moment? Did it read the right document but pull the wrong facts? Each answer routes to a different owner — tool use vs search/synthesis vs alignment.
  2. Build a sustained description of the pain point — the failure theme with its nuance, reproducibility, and size (“is it a big enough problem?”).
  3. Generate the eval set: start with 30–40 examples of the failure, each a prompt plus a golden answer. Check the eval is on-distribution — it should capture where the model fails and where it should not fail.
  4. Hand to research in their day-to-day language, then re-run the eval on each new model version to verify the area is improving. Penn: “you can’t improve what you can’t measure,” and the work stays tactile and judgment-based.
  5. Penn says she pioneered this concept within Anthropic, and that it is increasingly the PM skill set elsewhere too, because products now live at the intersection of models × harnesses × context × users.

On Living Inside the Exponential

  • In 2024 Anthropic shipped four model series in the whole year; it shipped more than that volume in Q2 of this year alone.
  • Nobody can predict the exact moment or model where a capability lands — so the premium skills are adaptability (re-decide when new information arrives rather than keeping the plan) and first-principles reasoning about “what’s the so-what.”
  • Her forward-compatibility prompt for the team: “Let’s say Claude 8 comes around. What changes in what users do? What does that mean for how you’re building today?” Be stubborn about the area, loose about the exact approach.
  • Model-release reality shifted with Fable/Mythos: more scrutiny and gating pre-release, so Anthropic built fallback UX so flagged users still get a great response from Opus 4. She frames the evolving “model safeguards package” as an active product surface, and says the goal is to keep general-purpose technology as inclusive as possible rather than restricted-access.
  • Labs’ thesis: identify discontinuous large bets outside the core roadmap (Claude Code, skills, Claude Design, MCP came from there), keep pods small — some bets start with one engineer — and treat prototypes that fail as learning to revisit “in one to two model generations.”

The Garry Tan Token-Maxing Exchange

  • Lenny relays Garry Tan’s claim: “If you are willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live” — an alpha opportunity to live in the future.
  • Penn’s reframe: token spend is the input; the output you should orient goals around is experimentation. Internally, the most creative thinkers and best prototypers spend heavy time with every new research model — “there’s no substitute” — but experimentation is not an individual sport: Anthropic’s early all-company Slack channel of people testing Claude in public produced communal use-case discovery within ~10 requests of someone posting an idea.

Working With Claude Without Losing Your Brain

  • Anti-brain-rot rule: form your own point of view first, then use Claude as a sparring partner; keep your sense and tone. For standardized artifacts (monthly business reviews) she wants the opposite — delegate writing fully to Claude and act as reviewer/verifier, because there the writing is “asymmetrically less valuable than the thinking.”
  • The lens that matters is shifting from who’s writing to who’s verifying and signing off.
  • Claude’s pushback (rooted in alignment/character work) is what makes it a useful thinking partner: “a thinking partner doesn’t just agree with you… the hero goal” is coming away with better ideas, not 10%-better ideas.
  • She built a personal skill from the book Crucial Conversations that coaches her through difficult manager conversations — an example of using Claude to raise EQ, not just IQ.
  • Where humans stay valuable: judgment (accumulated nuance the systems haven’t experienced), persistence, proactivity, and picking which of the many buildable things to build. Writing remains a jagged edge Claude teams are actively training on; tone and character are a stated priority.

Try It

  • Steal the loop for any AI product surface: collect 30–40 real failure examples with golden answers before writing any strategy doc; re-run the set on every model/prompt revision. Treat that eval file as the PRD for that pain point.
  • Triage vague feedback by trajectory, not by quote: decide whether each failure was tool-choice, retrieval, synthesis, or alignment before assigning it.
  • If you manage PMs: carve out one to two hands-on work streams for yourself per model cycle, as Penn does, to keep a live theory of mind about model behavior.
  • Run her Claude 8 question against your current roadmap as a forward-compatibility audit.

Open Questions

  • The interview references “$50 billion in ARR” (Lenny’s recalled figure, hedged as “the latest number I saw”) — not confirmed by Penn in the episode.
  • How eval ownership divides between research PMs and researchers at scale — the episode covers the loop, not the org design.
  • Which parts of the “model safeguards package” evolution she teases actually shipped “in the coming weeks and months.”