Source: raw/DHH_-_Future_of_Programming_AI_Agentic_Engineering_Vibe_Coding_Linux_Lex_Fridman_Podcast_501.md — Creator: Lex Fridman Podcast #501, guest David Heinemeier Hansson (DHH), creator of Ruby on Rails, CTO of 37signals and creator of the Omarchy Linux distribution. URL: youtube.com/watch?v=NYFGCESmikA · Platform: YouTube (fetched 2026-08-27; the recording date is not given, but Omarchy “Quattro” had launched the Friday before). The AGENTS.md update below cites raw/anthropic-watch-claude-code-tag-v2-1-277.md (Claude Code v2.1.277 release notes, 2026-09-18).
Thirteen months after telling Lex Fridman AI coding in its autocomplete mode was not for him, DHH says he has written none of the code shipped in his latest operating-system release by hand. The episode is useful less for the enthusiasm than for concrete operating habits: a cross-model review routine, agents triaging a queue of more than 1,000 pull requests, a terminal setup spread across several machines, and one task run on seven models with times and dollar costs. Everything here is one practitioner’s self-reported experience, and parts of it are opinion.
Key Takeaways
- Three phases in under a year. Opus 4.5 on “November 24th, 2025 … was the dividing line.” In early spring, sub-agents meant a task “that would take Opus quite a while suddenly took a fifth of the time, a tenth of the time.” This summer, “with Opus 5, Fable, and Sol … I’m not telling it where we’re going. I’m telling it the problem I have.”
- 100% agent-written, human-shaped. On Omarchy Quattro, “in the last two months it has been 100%. I have not written … any of the code that’s shipped in Quatro by hand. I’ve reviewed the shape of all of it”, plus the individual lines of anything critical in the model layer.
- Vibe coding by non-programmers broke an existing codebase. When Basecamp let designers “vibe” in February, their PRs “individually perhaps could have been justified … taken all together, destroyed the architecture of the system”, and the team “had to clean up manually.” He agrees with Lex that, to keep an existing codebase’s architecture, the person vibe coding still needs to be a programmer.
- Keep humans out of the middle. “To get that magical 10X, 100X … productivity boost, you have to interact with the agents directly, and you cannot intermediate that bandwidth with another human.” Most organisations, he says, are bottlenecked on ideas, vision and taste, not implementation.
- Specify less, then react. “Be as vague as you can to manifest something, then interact with the something.” Humans are good at “differential evaluation”: pick one of three options, not 22.
- Always review with a second lab’s model. “I’ll have Opus or Fable do the work, and then I always end it, review with Codex xHigh”, then Grok (“it keeps finding stuff”), then GitHub Copilot on the pushed PR, which “has actually gotten good.”
- Ask for simpler. Agents still over-build: after a second agent has reviewed the work, “looks a little too complicated” often makes the agent “cut it in half.” “You actually do have to tell them, ‘Make it simpler.’”
- The seven-model port: Fable finished fastest but would have cost about $550 per-token; GPT Sol and Grok 4.6 did the same job for about a tenth of that, and DeepSeek V4 Pro for about a twentieth, more slowly (table below).
How DHH works
Setup
- Terminal first. He started with tmux panes, then moved to Herdr: “essentially tmux plus agent notifications … whenever your agent is done and needs something for you, it goes ding.”
- Multiple machines. He connected four spare mini PCs through GL.iNet Comet KVMs on a Tailscale network and runs Herdr on each. At “about four to five machines”, roughly 16 parallel agent threads is his personal limit.
- Mixed harnesses. “I’ll run Claude up top, and then I’ll run Codex down below, and then maybe I also have an OpenCode set up.”
- Reviewing output. Neovim is now a project browser and a way into Lazy Git; he wants the surrounding files, not just the diff, when checking what an agent changed.
- Typing, not dictation. Omarchy ships Voxtype transcription, but he prefers to type. Lex does the opposite: 10–20 minute spoken prompts recorded on a Plaud wearable, transcribed with ElevenLabs and cleaned by a model that has a dictionary of his codebase’s file and function names.
Review and quality
- Style rules live in AGENTS.md. Example: he forbids early-exit guard clauses in Bash and has to remind agents to re-read the file.
- Writing quality differs by model. Claude models are “really good writers out of the box” for PR descriptions and commit messages; “GPT is a terrible writer out of the box.” Codex is “my favorite checker.”
- Disclose agent posts. “Whenever I have an agent post on my behalf, I always have it sign itself like whatever, Claude on behalf of DHH.”
- Security load is real. New models find so many vulnerabilities that 37signals faces “this seemingly endless parade of security vulnerabilities that need to be patched”; teams that are not patching, he says, are probably blind to them.
Automation
- Agents triage the open-source queue. In three months on Quattro he merged “over 1,000 pull requests”, with “about 400 unmerged.” Agents review each PR, validate bug fixes in a VM and summarise. An Omarchy bot runs “on a regular schedule” over PRs and issues and emails him through the Hey CLI: “Here’s 12 PRs that are either ready to go or I think you should close.”
- Throttle bots that act in public. A QA run with “eight different agents” found “28 real issues”; the bot filed them in about 12 seconds, and “GitHub, not unreasonably, marked that as probable spam and banned … my Omarchy bot.” For the next bug it found, in the mise package manager, he had the agent email the maintainer through its Hey address, and the maintainer got a report on a bug in software he had not released yet.
- Crash watcher. If any app crashes, Omarchy offers to have your agent diagnose it: the agent reads the systemd logs, checks out the app’s source and “pins down that it’s in this Rust file line 472.”
- Keep the model away from untrusted code. His development bot (“Amabot” in captions) uses a “brains and hands” pattern, which he credits to Shopify’s Tobi Lütke: the coordinator that runs the model is separate from the VM that executes code pulled from PRs and issues. He watched it decide to treat test-run output as outside data, in case an attacker had embedded a payload.
- Skills as an extension surface. “Omarchy ships a set of skills that tells any agent you bring to it how to create extensions to the operating system”; “in three days, we had 330 plugins on the Omarchy plugin marketplace.”
- Agents as async coworkers. At Basecamp, agents are assigned to-dos and cards like colleagues. Chat “entices you to sit around and wait”; an asynchronous task list is “the right format”.
The seven-model Rust port
The task: port the Python library behind Omarchy’s screensaver (Terminal Text Effects) to a dependency-free Rust executable, “pixel perfect, frame by frame … don’t stop until you’re finished.” Costs are DHH’s per-token estimates from single runs.
| Model | Result | Time | Cost |
|---|---|---|---|
| Fable, finished by Opus 5 after his Fable limit ran out two-thirds through | Done: startup 86 ms → 2 ms, 9.6x faster, 3 MB executable | ”just under 45 minutes” | ~$550 if paid per token (he ran it on a Max subscription) |
| GPT Sol, given Fable’s plan | Done | ~1.5 hours | ~$46 |
| GPT Luna | Failed: needed about 12 prompts to start, then “cheat[ed]” by wrapping an existing implementation outside its directory | — | — |
| Grok 4.6 | Done: 10x speed-up, same-size executable | not stated | ~$55 |
| Kimi K3 | ”took forever”; outcome not stated | — | — |
| DeepSeek V4 Flash | Failed the same way as Luna | — | — |
| DeepSeek V4 Pro | Done | 2 h 45 min | $23 |
- Plan once, execute cheaply. Fable’s detailed eight-step plan is what let Opus 5, and then Sol, pick the job up. Two further unattended optimisation runs took the port to a 46x speed-up over the original.
- His ranking: “the best model in general right now is Fable. The second-best … Opus 5”; “I would rank GPT Sol very good.” Lex’s split: Fable for planning and review, Opus 5 for implementation, to stretch tokens.
- Harness verdict (opinion). Claude Code has “the best harness”, partly because arrow-left in a session returns to Agent View to switch between agents. He uses OpenCode for open models such as Kimi K3, with inference on Fireworks, “pay by the token.” A Claude subscription is “a crazy bargain”; he now has two, and the next Omarchy will support switching between subscriptions. Grok 4.6’s fast mode is “cheaper than the regular mode on the others.”
Claims to handle with care
- The AGENTS.md complaint is now partly out of date. DHH: “Claude out of the box still refuses to read agents.md … So all my projects, they have a Claude MD, and all it includes is a pointer to the agents MD.” He calls it “petty” and ties it to Anthropic cutting OpenCode off from Claude subscriptions. Since Claude Code v2.1.277 (2026-09-18), “in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead” (configurable under “Project instructions” in
/config; not yet on Bedrock, Vertex or Foundry). The note does not cover his other complaint, that skills must live under.clauderather than.agent/skills. See CLAUDE.md Primer. - The 80% system-prompt cut is second-hand. DHH relays that Boris Cherny said the system prompt shipped for Opus 5 “shrunk by 80%” because the model “was actually being damaged by overly prescriptive humans.” Anthropic’s own account is in The New Rules of Context Engineering.
- The Shopify study has no primary source. He says Shopify’s CTO had agents trace production incidents back to merged PRs and found “the PRs that have been reviewed by agents caused far fewer issues in production”, with models from about six months earlier.
Try It
- End every agent task with a review by a model from a different lab (DHH uses Codex at xHigh), then let a PR bot review again after you push.
- If you use several harnesses, keep one AGENTS.md. On current Claude Code you can drop a pointer-only CLAUDE.md, since AGENTS.md is read when no CLAUDE.md exists.
- After a feature lands, tell the agent it looks too complicated and ask for a simpler version; compare line counts.
- Ask for three alternative implementations and pick one, instead of writing a detailed spec first.
- Schedule an agent to triage your PR and issue queue and email a ready/close digest. Rate-limit any bot that posts publicly.
- Run your own one-task bake-off: have the strongest model write the plan, hand the same plan to cheaper models, and record time and cost for each.
- If you ship a product with an ecosystem, ship agent skills that explain how to extend it; Omarchy’s skills produced 330 plugins in three days.
Open Questions
- The Shopify agent-review study has no primary link or methodology.
- The 46 and $55 figures are single, hedged runs; Grok’s run time and Kimi K3’s outcome are never stated.
- How the Amabot “brains and hands” split is implemented (VM tooling, what crosses the boundary) is not described.
- Whether Claude Code will also read
.agents/skillsis not covered by the v2.1.277 note.
Related
- Cheap-Executor Delegation — the plan-with-Fable, execute-with-cheaper-models pattern
- Cost & Intelligence Levers for Agent Workflows
- The New Rules of Context Engineering for Claude 5 — first-party version of the system-prompt cut
- The CLAUDE.md File — Anthropic Primer — AGENTS.md fallback details
- Claude Code Week 20 — where Agent View shipped
- PR Review Risk-Scoring Agents — another way to let agents clear the PR queue
- The Software Factory — Warp and Ramp — the company-scale version of agent-reviewed PRs
- From Vibe Coding to Agentic Engineering (Karpathy)