Source: raw/The_AI_factory_playbook_for_engineering_teams.md — How I AI (host Claire Vo), guest Zach Lloyd, CEO of Warp, youtube.com/watch?v=4_SHhSMHzNo (fetched 2026-09-29). raw/What_product_looks_like_when_coding_is_solved_Geoff_Charles_Ramp_CPO.md — Lenny’s Summit talk by Geoff Charles, CPO of Ramp, published on Lenny’s Podcast, youtube.com/watch?v=ZG8Mf3P9xzI (fetched 2026-09-29).
Warp’s CEO and Ramp’s CPO describe the same shift, from code being the bottleneck to everything around it, from opposite ends. Lloyd shows the engineering side: a “factory” defined in code that takes requests from Slack, runs the ticket-to-QA sequence, and records every run so it can be scored, replayed and improved. Charles shows the product side: each time one bottleneck is automated the next one appears, and he walks through the internal agent Ramp built for each step, from customer insight to coordination. Every number below is self-reported by the company that owns it, and Warp sells factories as a product.
Key Takeaways
- A factory is configuration, not a chat window. Lloyd defines it as “an actual noun … a bunch of repos, a bunch of like MCP servers, like a bunch of configuration and then a bunch of agents … and it’s all defined in code.” Because it is code, you can freeze its state, re-run tasks after a change, and let coding agents edit the factory itself.
- Work starts in public. At Warp you tag the factory (named “Wilson”) in a public Slack channel. It triages the request, opens a Linear issue, implements, opens the GitHub PR, and runs computer-use QA that records a video of the finished feature with keystrokes. Sentry crash reports start the same flow automatically.
- Measure how often a human has to step in. Warp’s dashboard tracks human interactions per PR: re-prompts in Slack, comments in Linear, corrections in code review. Lloyd compares it to a car plant: “how many times do you have to stop the line?” Vo says she had not seen the metric before.
- The human is still the slowest step. Vo, reading Warp’s dashboard: kickoff to PR is 35 minutes, but “PR to first human review, it’s 3 and 1/2 hours”, with over 2,000 PRs in the last month. Every PR still gets human review; Warp dropped the rule that a second person must review, so the person who prompted the agent can review its code.
- Model choice is the biggest cost lever. Warp’s per-PR agent cost rose during adoption, then fell after model-configuration changes. Lloyd: model is “the biggest”, and “context management matters as a secondary thing.”
- Score every run, then let an observer propose fixes. LLM-as-judge scorers run across all recorded runs; a second loop has an observer agent read failed runs and propose edits to the factory’s agent definitions, but only after “maybe 20 25 failed runs” so it does not overcorrect.
- Ramp’s rule: remove a bottleneck, then find the next one. Charles walks Identify → Define → Build (code, review, test) → Coordinate → Improve, with an internal agent at each step (table below). His summary: “the bottleneck has now shifted to us”, meaning product managers.
- The PM job splits three ways (Charles): a technical PM who builds the factory, a “tastemaker” who holds the quality bar, and a PM who grows into a general manager owning the business outcome.
Warp: a closed-loop factory (Zach Lloyd)
- Two personas. For the builder, working with the factory feels much like using a local coding agent, except the work is in a shared Slack thread where others can watch or join. For the manager, everything is centralised in the cloud, so automation rate, velocity, cost and quality can be measured across the team.
- Inputs from anywhere. Requests can start from a person in Slack, from an external system such as Sentry, or from Linear or GitHub.
- Code review as risk management. Vo describes her “Merge Mommy” agent (built on Vercel Eve): risk-score every PR, give low and extra-low risk PRs an agent approval so a human can merge them, and send medium and high risk to full human review. Lloyd: “code review becomes an exercise in risk management.” Full build: PR Review Risk-Scoring Agents.
- Scoring in practice. Lloyd’s example scorer checks for redundant tests, a common agent failure. He picks “a kind of medium smart judge model” so scoring does not get expensive, and sets a sampling rate rather than scoring every run.
- Replay on your own tasks. Because every task is recorded, you can curate a set (say, front-end tasks), replay it under other model configurations, score the results with the same judges, and plot a cost/quality Pareto chart. The result feeds a model-routing strategy. Lloyd’s illustration: GLM 5.3 instead of Opus on front-end tasks, “Yes, definitely”; Gemini 3.7 Flash, “You’re going to take a quality hit.” No numbers were shown.
- Vo’s advice for teams on local agents: capture session-level telemetry (sessions, tool calls, MCP calls, test failures, computer use) for every coding session. She says some teams copy every local session to S3 and run their own evals on it.
- Process that did not change. Warp still dogfoods heavily, runs user interviews, watches people use the product, and holds design jams before building “to make sure that we’re tackling the right user stories.” Lloyd says tool usage follows a power law, and working in public threads lets the heaviest users teach everyone else.
- Naming. Lloyd dislikes the word “factory” (“a little bit sort of like dehumanizing”) but uses it because the industry has standardised on it. For the argument against running a company as a factory, see Karri Saarinen in Lenny’s Summit 2026.
Ramp: remove the bottleneck, find the next one (Geoff Charles)
| Step | Bottleneck | What Ramp built | Reported result |
|---|---|---|---|
| Identify | Customer pain spread across Gong, Zendesk, LogRocket, surveys and angry emails; “a 1 million token window, that’s less than 0.5% of Gong transcripts at Ramp” | A daily Slack “hate channel” of customer quotes, then a customer-insight agent (ETL pipelines, vector search, clustering) served through a Slack agent, an HTML dashboard and a daily “hate podcast” | Traceable data tells PMs which customers to call |
| Define | PMs asking engineers “Is this possible? Would this break something?” | Glass, an agent connected to Snowflake, user research, product strategy, the spec format and the codebase: “now AI is your tech lead”. It builds a prototype inside the real design system | The handoff to engineering is evidence, agent-ready requirements and a working prototype |
| Build: code | Coding | Inspect, Ramp’s own coding agent in Slack; starts “under 5 seconds” and returns a deployed preview | ”a million sessions”; “75% of our PRs”; “a thousand of those PRs over the last month was submitted by a non-engineer” |
| Build: review | Engineers combing through agent code | Review Buddy: knows Ramp’s quality and security checks, picks the human reviewer, and sees the prompts that produced the code | ”93% of our PRs are now automatically handled by Review Buddy” |
| Build: test | Manual QA environments and fixtures | Testo, a browser QA agent that runs the product “in 100 different combinations based on the actual production data” and reports bugs plus design feedback | 425 bugs caught in the last 30 days |
| Coordinate | Human attention | ”Every question is an API”: an agent (captions: “gadget”) reads the Notion roadmap, specs and call notes, Slack and Linear; answers status and sales questions, updates the roadmap, pings late owners, drafts help-centre articles, blog posts and customer emails | ”85% of questions that are being asked to PMs now are fully answered with AI” |
| Improve | PMs drifting to small, easy fixes | An automated loop for small issues: route, dedupe against the Linear backlog, count, rank, plan, check with a human in Slack, code, run CI, update the knowledge base | ”60% of UX issues … are fixed within 24 hours” |
- Architecture first. Charles: coding agents “are really good when you have a strong architecture and a strong codebase”, which widens how much of the company can use them.
- On budgets. “Embrace your constraints, but not your bottlenecks”: pick the one bottleneck you think could 10x the company and start there.
- On measuring speed. He has no metric; his answer is to hire someone “that knows what speed looks like” and let them challenge you.
Beyond code: Lloyd’s go-to-market uses
Lloyd runs the same coding agent (Warp; he says Claude Code or Codex would work) for CEO work. These are the most directly reusable parts for a marketing team.
- Slide edits through Figma MCP. He dictates a change to a Figma Slides deck and tells the agent to “duplicate the existing slide rather than making changes directly to it.”
- Sales-call FAQ from meeting notes. Through the Granola MCP: “look back over my last four weeks of sales meetings and try to build up a list of the top 10 frequently most asked questions”, anonymised. The top theme was buy-versus-build, then security, workflow, measurement and cost. He uses the list to check the deck, website and FAQ against what prospects actually ask.
- Cold-lead rediscovery. Through a Google Workspace command-line tool (captions: “GOG CLI”): find emails and calendar events from the past six months with possible enterprise leads, put them in a Google Sheet, and share the link without printing names in the thread.
- Model choice for this work. He has been using Grok, which he calls “pretty good from like a cost and speed and quality trade-off.”
Try It
- Add human interactions per PR and PR → first human review to your engineering dashboard before adding more agents.
- Risk-score PRs and let an agent approve only the lowest tier; route everything else to a human (how to build it).
- Write one LLM-judge scorer for your most common agent failure (redundant tests is Warp’s example), use a mid-tier judge model, and sample runs rather than scoring all of them.
- Hold self-improvement edits until you have about 20 failed runs of the same kind.
- Before switching models, replay 10–20 past tasks of one type on the cheaper model and score them with the same judges.
- Marketing: point an agent with a meeting-notes connector at the last month of sales calls, ask for the top 10 questions (anonymised), and compare them with your deck and FAQ.
- Write your pipeline as Ramp’s seven steps, name the current bottleneck, and automate only that one.
Open Questions
- Every figure is self-reported. Ramp’s “75% of PRs”, “93% automatically handled” and “85% of questions” have no stated denominator or quality check; “automatically handled” is not defined (approved, or only routed?).
- Warp builds and sells factories, and the dashboard shown is its product.
- Lloyd says human interactions per PR “could definitely be refined”; no target values were given.
- The GLM 5.3 and Gemini 3.7 Flash comparison was an illustration with no scores shown.
- Agent names (Wilson, Glass, Inspect, Review Buddy, Testo, “gadget”) come from auto-captions and may be misspelled.
Related
- Vercel Eve) — the Merge Mommy pattern in full
- Replit’s “Self-Driving Company” — another six-month, self-reported org-wide agent rollout
- Anthropic’s AI-Native SDLC Playbook — the first-party version of stage-by-stage bottleneck removal
- Running an AI-native engineering org (Fiona Fung) — Anthropic’s own “the bottleneck has shifted” account
- The Loop Is the Unit of Work — self-improving loops across frameworks
- Picking the Right Model — Evals for Model Selection — replaying tasks to choose models
- Orchestrator + Cheap-Worker Routing — the routing decision the replay data feeds
- Lenny’s Summit 2026 — Product Work After Code Got Cheap — the rest of the summit Charles spoke at