Source: ai-research/claude-blog-ai-native-sdlc-playbook-2026-08-21.md — Anthropic first-party blog post “The AI-Native SDLC playbook”, Author: Louis Claxton (Applied AI), URL: https://claude.com/blog/the-ai-native-sdlc-playbook, Published: 2026-08-21. Discovery pointer and community framing: raw/reddit-1vzl6kk.md (r/ClaudeAI, u/Forward_Mind6886, 75 score, 46 comments, 2026-08-27).

Anthropic’s Applied AI team argues that once agents write most of the code, the bottleneck moves to the human-speed steps on either side of the build: planning, review and test, and deployment. The playbook rebuilds all six stages of the software development lifecycle (Plan, Design, Build, Test, Deploy, Maintain) so that each stage ends by committing an artifact the next stage reads, and the chain of commits doubles as the audit trail. Line-by-line human review gives way to policy encoded as skills, agentic review passes and deterministic hooks, with humans approving at named gates. Each “play” comes with prerequisites, steps, governance notes and a leading and a lagging indicator, so a team can adopt it one stage at a time.

Key Takeaways

  • Build is no longer the constraint. “Human-speed stages keep their length while build collapses to hours.” Security review is the worked example: a team sized for human output either builds a queue or ships under-reviewed code once agents multiply output.
  • Every stage commits an artifact, and each commit triggers the next stage:
    • Plan → intent.md (problem, proposed outcome, affected users and systems, constraints, open questions). An accepted intent.md triggers the design pass.
    • Design → spec.md, produced by Claude from intent.md under the organisation’s brand, security, compliance and UX skills. An approved spec.md triggers plan mode.
    • Build → plan.md from plan mode, plus a version-controlled CLAUDE.md and skills as institutional knowledge.
    • Test → a feedback loop in every session, plus continuous evals in CI.
    • Deploy → the PR with its review findings, gated by hooks. A merged PR triggers the pipeline.
    • Maintain → a breached control band in production writes the next intent.md, and the loop restarts.
  • What replaces line-by-line review: “Reviewing each line by hand made sense when a person had written it, but it can’t keep up once agents write most of the diff.” The replacement is identical agentic review passes on every PR, findings ranked by severity, and human attention moved up to “whether the change does what the plan intended and whether the risk is acceptable.” Human review is reserved for regulated and critical code.
  • Skills are advisory; hooks are deterministic. “The skill makes violations rare and the hook makes them close to impossible.” Any policy that must always hold needs a hook or a PR-time re-check behind the skill.
  • Hooks can allow, ask or block. “A hook can also ask, pausing the action until a specific person approves, which is what release gating needs.” Ask-hooks belong at Deploy, because an approval prompt during Build “puts a person back on the critical path of all the sessions running in parallel.”
  • The agent may act up to the production gate and not past it. Branch protection turns anything the agent writes into a PR; a production-deploy hook blocks release until a named release manager authorises it.
  • Separation of duties survives: “the agent that wrote the code has no way to approve it.”
  • Legacy systems stay. For each artifact, name one source of truth: the repo, the legacy tool (Jira, ServiceNow, a requirements tool) with markdown as working copies, or at minimum two-way linkage (record ID in the file, commit SHA in the record).

The plays, stage by stage

Plan: capture as intent.md.

  • The originator brainstorms with Claude until the idea is concrete; Claude asks the analyst’s questions (scope, users, constraints, success).
  • Claude writes intent.md from the organisation’s template, which can be encoded as a skill. The originator corrects it; the product owner accepts or closes it.
  • Non-engineers work from claude.ai or Cowork; a version-control connector lets Claude commit the markdown for them. The simplest home is an intent/ folder in the product repo.
  • Indicators: time from first conversation to committed intent.md (expected to fall from weeks to hours); share of intents accepted.

Design: requirements and design in one session.

  • The product owner attaches intent.md and asks for a spec that applies the organisation’s skills and flags “any areas of concern, especially where you cannot satisfy contradicting policies.” Front-end work can go through Claude Design (beta) first.
  • Run it by hand, then as an org-level slash command, then as a non-interactive job that fires on the intent.md merge and opens spec.md as a PR.
  • Indicator to watch: spec.md commits dated after the first plan.md commit (requirements rework).

Build: plan mode, CLAUDE.md, skills, hooks, parallel sessions.

  • Plan mode by default. Claude interviews the engineer against spec.md; iterate “until an engineer who has never seen the conversation could implement the change from the plan alone.” Commit plan.md; update it in the same commit when implementation departs, optionally enforced by a hook.
  • Auto mode becomes the default for routine work once the guardrails mature: a tuned CLAUDE.md, policy skills, blocking hooks and a runnable test suite.
  • CLAUDE.md: run /init, cut it to what a new joiner needs on day one, check it in, keep it under a page. Working rule: “When Claude makes a mistake twice, the correction goes into CLAUDE.md.” See the CLAUDE.md primer.
  • Skills: pick one policy enforced inconsistently today, write it as a skill from the policy owner’s source of truth, ship it in .claude/skills/ or an org plugin, and test that it triggers.
  • Build-time hooks: block edits to protected paths, run formatter and linter after edits, keep credentials out of the diff. Keep them fast and scoped to the changed file; heavier checks belong at commit or PR.
  • Parallel sessions: one task per worktree; start with two or three sessions and add more “only while review is keeping up.” Turn repeated jobs into subagents in .claude/agents/ (the post’s example is a report-only verifier that runs the app and checks behaviour against plan.md).

Test: a feedback loop, then continuous evals.

  • Give every session a way to verify its own work: one-command tests and build, a quantified target, a browser or screenshot tool for UI.
  • For bug fixes, write the failing test first and commit it, then have Claude make it pass “without editing the test.” A hook that blocks edits to test files during a fix protects the loop.
  • Continuous evals: collect 20–50 real tasks with accepted outcomes, write each as prompt plus checks, and run the suite non-interactively on a schedule and on any change to CLAUDE.md, skills or hooks. Gate configuration changes on the pass rate. Every production incident becomes a permanent eval.

Deploy: review, gates, managed settings, CI/CD.

  • PR review: the managed Code Review service (research preview) or claude-code-action in your own CI. The tech lead writes REVIEW.md with passes (bugs, security, compliance against spec.md and plan.md), defines Important vs Nit, caps nits (the example: at most five per review) and excludes generated paths. Findings do not approve or block on their own; branch protection still requires a code owner.
  • Tagging @claude on a review comment gets the fix pushed. A mistake flagged twice in review goes into CLAUDE.md.
  • Approval-gate hooks: leadership lists the human gates that must survive; each becomes a PreToolUse script that allows, asks or blocks (exit code 2 blocks and sends the reason to Claude). Team hooks go in .claude/settings.json; non-negotiable ones in managed settings that engineers cannot switch off. “A block should explain itself.”
  • Managed settings for a regulated enterprise: the post gives a full example (deny reads of .env* and secrets, deny WebFetch/curl/wget, sandbox with a domain allowlist and credential-file denies, managed-only hooks, permissions and MCP servers, an approved plugin marketplace, a minimum Claude Code version). It calls this “a starting point to tailor, rather than a recommendation to copy.”
  • CI/CD: start with read-only claude -p steps (triage a failed build, draft a changelog), then write steps behind existing gates, sandboxed with short-lived scoped tokens and no production credentials. Expose deploy, status and rollback through MCP, scoped per environment. Tier autonomy: free in dev, gated in production. Make rollback “the most rehearsed path in the pipeline.”

Maintain: close the loop.

  • A deterministic detection script watches one metric with a stable baseline (CI failure rate, post-deploy 5xx rate, PR cycle time). Response tiers live in version-controlled config: 1σ log, 2σ invoke Claude read-only to diagnose, 3σ Claude may act, only by opening a PR or triggering a pre-approved runbook. The diagnosis is written as a new intent.md.
  • Recurring codebase scans: Claude Security runs scheduled scans on Claude Mythos 5 in Anthropic’s infrastructure, validates each finding and attaches a confidence rating. It is in public beta for Claude Enterprise and needs the Anthropic GitHub App, Claude Code on the Web, Extra Usage with a spend limit, premium seats and an admin toggle; scans bill at Mythos 5 rates. Weekly is the suggested default for active services. Bounded findings go through the PR gate; wider ones become intent.md.
  • On call with Claude Tag (public beta, Slack): Claude joins incident channels under its own identity, diagnoses in-thread, verifies the metric is back at baseline over MCP and writes the post-mortem to a version-controlled lessons file.

Community framing (not in the Anthropic post)

  • u/Forward_Mind6886 (r/ClaudeAI): “the bottleneck moved from writing code to deciding what to write and checking what came back.” The post compares this to 1957 arguments that compilers could never match hand-written assembly.
  • Statistics the Redditor added, not from Anthropic and not verified here: Faros AI telemetry (10,000 developers, 1,255 teams) reportedly shows high-AI-adoption teams merging 98% more PRs, with review time up 91% and average PR size up 154%; DORA 2025 reportedly finds throughput up and stability down. The same post cites GitHub’s Spec Kit (MIT), AGENTS.md in “60k+ repos”, and an unnamed repo that packages the playbook as a Claude Code and Codex skill. None of these figures appears in the Anthropic post.
  • The thread’s open question is the useful one: “what actually replaced line-by-line review for you? Evals in CI, a verifier subagent with a fresh context, hooks on protected paths, something else? And what still slips through?”

Try It

  1. Start with a play that has no prerequisites. intent.md capture, CLAUDE.md, the feedback loop and hooks-as-approval-gates all list “None.”
  2. Add an intent/ folder and a template skill to one product repo. Measure the git timestamps from first conversation to committed intent.
  3. Write REVIEW.md with three tagged passes, a definition of Important, a nit cap and an exclusion list, then run it through claude-code-action or the managed Code Review service.
  4. Convert one “never” rule into a hook. Pick a policy that currently lives only in a skill or CLAUDE.md (a frozen package, a protected path) and back it with a PreToolUse hook that explains its block. See Steering Claude Code for which method fits which rule.
  5. Stand up a 20-task eval suite that runs on changes to CLAUDE.md, .claude/** and nightly, and add one eval per incident.
  6. Pick one metric and write bands.yaml with log/diagnose/propose tiers before letting any agent act on production signals.

Open Questions

  • The dependency graph is not in the saved text. The post refers to a diagram of play dependencies (“Start with any clay play — nothing points into it”); only the per-play prerequisites survived extraction.
  • No outcome data. The playbook defines leading and lagging indicators for every play but reports no measured results from Anthropic or its customers.
  • Linked companion posts were not fetched: “securing an AI-native SDLC at Anthropic” and “how Claude Tag runs on-call for CI/CD at Anthropic.”
  • The Faros AI and DORA 2025 numbers are community-cited only; verify against the primary reports before reuse.
  • Which repo packages the playbook as a skill? Unnamed in the Reddit post.
  • The managed-settings example pins requiredMinimumVersion: "2.1.193"; whether every key shown is available at that version is not stated.
  • Claude Code Hooks — the allow/ask/block mechanism the playbook’s gates are built on
  • Steering Claude Code — CLAUDE.md vs rules vs skills vs hooks, the same advisory-vs-deterministic split
  • Verifier-First Loops — the feedback-loop and verifier-subagent discipline behind the Test stage
  • OpenSpec — a community spec-driven workflow with the same intent → spec → plan shape
  • GSD — another plan-first, artifact-per-phase framework for Claude Code
  • Agent Guardrails — hooks, permissions and sandboxing in more depth
  • Agent Workflow Patterns — the general orchestration patterns the Maintain loop instantiates