Source: raw/Build_Sell_with_Codex_5+_Hour_Course.md — Nate Herk, “Build & Sell with Codex 5+ Hour Course”, YouTube (youtube.com/watch?v=X-pbJWKmwi0), transcript fetched 2026-09-29. Dates Nate reads out on screen (September 7 and September 15) place the recording in September 2026. One claim is checked against the first-party Claude Code release note raw/anthropic-watch-claude-code-tag-v2-1-277.md.
Creator: Nate Herk (AI Automation Society). Duration: 5+ hours. Platform: YouTube. Skill level: operator, non-technical.
Nate Herk’s five-hour follow-up to his one-hour Codex course, recorded with GPT-6 Astra as his main model. It covers the same Codex desktop app, but the new material is his six-step skill method with a live model walk-down, branded deliverables, browser and computer use, video editing with HyperFrames, hosting automations outside the Codex subscription, voice mode, a 15-task Astra vs Fable 5.1 cost comparison, and a pricing module. His Four Cs AIOS framework and Bike Method already appear in his Claude Code course; this article covers only what is new or Codex-specific. The pricing material has its own article: Pricing and Selling AI Automations on Value.
Model names in this course. Astra = GPT-6. Sol, Terra and Luna = the GPT-5.6 family (“5.6 Sol”, “5.6 Terra”). The course predates GPT-6 Sol and Luna. Auto-captions render Sol as “Soul” throughout.
Key Takeaways
- Walk every skill down the model list, then down the effort list. Build and prove a skill on the strongest model, then rerun it on each cheaper one: “There’s no reason to be running a skill with Astra if you could run it with Luna and get the exact same output.” Once a model holds, start effort at high and step down. His simple YouTube-description skill runs on Luna; his X-article skill stays on Astra.
- The live walk-down did not come out in price order. On his X-article skill (all runs on high), Luna took 28m33s and impressed him; Terra took 26m and was worse than Luna (“it just gave up or it stopped and thought it was done”); Sol took 38m and was “not too bad”. Across repeated runs on other videos “Astra just seems to be the best”, so he runs it on Astra low with fast mode. One test is not enough to pick a model.
- A QA report is the part to keep. Every model’s run produced a QA report; the one he reads out checks the title, body, 76 non-empty text blocks, 11 inline screenshots, seven H2 headings and a visual and privacy review. Nate: “that’s the most important part”. He estimates verification lifts first-run quality to “more like 95% rather than… 75 or 80”.
- Don’t host recurring automations on the agent. Codex scheduled tasks inject a prompt into a chat, so they “eat” the weekly limit. He builds automations in Codex, pushes them to a private GitHub repo and hosts them on trigger.dev. A simple agent routine of his “after about a month… started to just go rogue”, posting to the wrong DM and to team channels “because it was an agent on the back end… interpreting the message different every time”. The fix was a script with the channel IDs hard-coded.
- Automation order of preference: API, then a deterministic macro, then browser use. Browser use is for the cases with “an element of vision or reasoning”. He estimates APIs cover “95% of your automation things”.
- Astra vs Fable 5.1 over 15 of his own tasks: Astra won 10, Fable 5. Totals: Fable 9h35m45s and 326.98 — “Astra was $186 cheaper but took an hour and 43 minutes longer.” Fable won the consulting deck, sales letter, 3D-brain visual, HTML explainer and site clone. Details below; these are one run each, graded by his taste.
- He does not default to Astra for everything. “GPT 5.6 Sol is so so good still… for the majority of my knowledge work”; for most knowledge work he would “chuck it on Sol” at medium, and Astra for an email is “overkill”.
- Plan economics (Nate’s figures). He says the 14,000 worth of inference every month” if he uses his full weekly limit. Astra’s API price is 50 output per million tokens. Over the weekly limit you buy Codex credits. Fast mode is “1.5 speed” but uses more of the plan.
- Claude Code now reads AGENTS.md. Nate reports “Thariq from Claude Code… tweeted… we’re adding support for agents.md”. The first-party release note confirms it: v2.1.277 (2026-09-18) — “in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead” (
raw/anthropic-watch-claude-code-tag-v2-1-277.md). See the CLAUDE.md primer.
The 18 concepts (what’s new since the one-hour course)
Nate groups 18 concepts into four parts. Items already in the one-hour course are listed by name only.
- Foundations: projects;
agents.md; the agent loop;/goal— “the more objective a goal, the easier it is to prove”, e.g. “pull in 257 sources and then write the report”. His motion-graphics goal defined one intro scene and 18 cards and ended with “I finished and I visually checked”. For low-risk work: “give them a goal and… get out of the way”. - Environments: local vs cloud (cloud chats run in a sandbox where you can no longer pick the model or permissions; he rarely uses them, and instead drives desktop Codex from the ChatGPT phone app’s Remote view); worktrees (rarely needed for knowledge work); the
.codexfolder (project scope vs user/global scope, holding config, memories, sessions and automations). - Control: models; effort — he names low, medium, high, extra high and ultra, and “ultra I have to allow full access”; permissions — Full Access (“Codex’s yolo mode”), Ask for approval (asks before editing external files or using the internet), and “approve for me” (asks only when an action is detected as unsafe, such as deletes); skills and the
.agentsfolder (project or global); plugins. - Tools and scale:
- Browser: logins saved in the Codex browser persist, so Codex can act in sites that have no API.
- Sites:
/siteor “turn that into a site” hosts a build on OpenAI’s domain, public or login-only, with analytics, a custom domain and a database. He says it can replace Vercel or Squarespace for him (opinion). - Subagents with pinned models: “delegate a bunch of sub agents out… I want all of these sub agents to be using 5.6 Terra”, so the expensive main thread is not spent on research.
- Scheduled tasks: a scheduled task injects a prompt into a Codex session, local or cloud, in a chosen project and thread. He runs seven or eight trading tasks with “$10,000 of my real money”.
- Voice mode: one voice chat spins up and coordinates several threads; threads can send each other messages and files. It also works from the phone.
AIOS setup in Codex (the Codex-specific parts)
agents.md= core rules + a routing map. Most of his file is routing (“if you need business stuff, you go to the wiki… if you need his voice, you go here”) so the agent can find things in a project of thousands of files. Codex also keeps amemory.mdand reads past chats.- Migrating from Claude Code took one request. He copied
CLAUDE.mdtoagents.md, skills into the.agentsfolder and settings over — by asking Codex to “make this Codex ready”. His onboarding skill also writes a.claudefolder andCLAUDE.md, so the same folder works in both tools. With v2.1.277, a project that only hasAGENTS.mdwould no longer need the duplicate. - Resource pack skills: onboard, audit (scores the Four Cs; a fresh demo scored 30/100 and saves each report to an
audits/folder), link (routing), level up (turns an audit into next steps), and 3D brain (renders the knowledge base as a navigable graph). - Grill Me — inspired by Matt Pocock’s skill of the same name (see Matt Pocock’s skills) — interviews you relentlessly and saves each session as a file, to move knowledge out of your head.
- Karpathy’s LLM wiki builds the relationships between notes; he keeps separate vaults for YouTube videos, business knowledge and meeting transcripts. See Karpathy’s LLM-wiki techniques.
- His three tests of a working AIOS: it answers a teammate’s question better and faster than you, with sources; you stop switching to other tabs; you stop trying to remember things because you trust retrieval.
Six-step skill method
- Reverse-engineer from the output. Start from a finished deliverable you like and walk back through how it was made; asking for “a skill for X” from scratch gets you “a chicken sandwich” when you wanted chicken parm.
- One job, one trigger, set in the YAML frontmatter. Break a role into task-level “leaves” rather than a skill that runs a whole team.
- Freedom level. Deterministic processes get strict step-by-step instructions; judgment-heavy ones (his X-article skill) get room to decide.
- Verification. Objective checks (counts, rules) and subjective checks via LLM-as-judge; have the agent or subagents verify before you see V1. “If you assigned that task to a human, what would you do to basically give it the stamp of approval? … explain that to the agent.”
- Walk it down the models, then the effort levels (see Key Takeaways).
- Bike method: the skill is never finished; give good and bad feedback after every run and tell it to update the skill.
Other build segments
- Branded deliverables. A
brand-assetsfolder with logos and a README that tells agents where things live; Codex generated a brand-guidelines image and a markdown version with hex codes and font names. Skills then bake the brand in (a student resource guide, internal memos, a YouTube report sheet in his brand font, a slide deck with a backing chart sheet, an HTML explainer). He prefers Google Drive outputs because the team can edit and share them directly. His team shares an “AI phrase kill list” (em dashes, “honestly”, “it’s not just X, it’s Y”). - Images and video. Codex makes images natively but not video. He connects the new pay-per-use Higgsfield API through a key in
.env, never pasted into chat (sponsored segment; details and break-even figures in Higgsfield Overview). - Web design. He rates Astra above Fable 5.1 for one-shot site design. His ingredients: his “scrollcraft” skill (layering, scroll, typography, spacing), inspiration from godly.design, 21st.dev and awwwards.com sorted by category, and the three Ps (pain, person, promise). A 3D scrolling site is wrong for “a 70-year-old man… looking for medication”.
- Browser and computer use.
- Annotate: click an element in the in-app browser and describe the fix.
- “Try your best to break it” QA: found invalid contact data getting through and a UK country code resetting to US, then switched to mobile view and found responsiveness failures.
- Bank-statement download: a one-shot skill; when the bank logged him out, he imported a password-manager CSV (name, url, username, password, notes) instead of pasting credentials into chat. For a bank, watch the skill run “at least 10 times” and keep the script strict.
- X articles through the browser: the API “charges you” and mishandles formatting and media, so the X-article skill posts through the browser instead.
- Computer use is a separate plugin, which he installed from Plugins in a couple of seconds. It tints the screen edges blue while active and stops at admin-password or security prompts.
- Video editing with HyperFrames. Install the HyperFrames repo into a project, then run a loop: transcribe → cut → plan beats → generate HTML → verify, repeating the last two. He transcribes with an ElevenLabs key restricted to speech-to-text (“Whisper… it’s just slower”). One detailed
/goalprompt cut a 1m05s intro to 28s in 18 minutes; a second/goalrevision took 10 minutes. The local HyperFrames studio handles small manual tweaks. Once an output is good, “turn that into a skill” so the next prompt is short. See HyperFrames. - Sizzle reel. From 152 GB of event footage, Codex produced a 60-second recap in about 35 minutes “in only two prompts”. He says his attempts with Sol and Fable 5.1 “did not go nearly as well”.
- Hosting off-subscription (trigger.dev), three shapes:
- Scheduled: a 6am weekday calendar brief to a ClickUp DM, about 1.33 cents per run under a 25-cent cap. A “morning brief enabled” environment variable stayed false until the hosted test passed, and duplicate-send prevention blocked a second run the same day. Codex flagged that trigger.dev’s free plan can delay a 6am start by up to an hour.
- Webhook: a form submission produces a ClickUp DM plus an AI-written outreach draft. He notes the notification alone needs no AI.
- Codex SDK: a full agent loop, for when you need Codex programmatically at scale. His example is a trading research loop every 30 minutes that emails a Grok-based bot to place trades, because “OpenAI models and Claude models don’t actually let you trade anymore”. API billing costs more than the subscription, so keep agent loops on the subscription unless you need scale.
- He prefers API keys in
.envover plugins, because plugin connections “will not transfer over to trigger.dev”. “90% of business use cases” are deterministic flows with one or two AI steps (his estimate).
- Voice-mode gotcha. Voice mode first started “regular chats” outside his project, which lacked his context; telling it to work inside his project fixed it.
Astra vs Fable 5.1 — 15 tasks
One run per task, judged by Nate. ”—” means the figure was not read out or the caption is garbled.
| # | Task | Winner | Time (Fable / Astra) | Cost (Fable / Astra) | Note |
|---|---|---|---|---|---|
| 1 | Consulting-style deck | Fable | 37 / 23 min | 12 | Nate preferred Fable’s deck; Astra asked clarifying questions first |
| 2 | Sales letter | Fable | ~4 / ~3 min | ~$4 / — | ~2,800 vs ~1,300 words; Fable answered buyer objections |
| 3 | Taxes (Q1–Q2) | Astra | 22 / 40 min | 22 | Astra asked 7 questions before starting |
| 4 | Subscription audit from email | Astra | 17 / 30 min | ~12.50 | Fable left one sheet empty |
| 5 | Meeting analysis | Astra | — | 5.50 | Astra read 79 meetings, Fable 58 |
| 6 | Event recap reel (150+ GB) | Astra (slight) | 50 / 30 min | 16 | Both “really well” |
| 7 | Sizzle reel | Astra | similar | — | Astra pulled real screenshots |
| 8 | Browser game | Astra | similar | Astra ~$8–9 less | Subjective |
| 9 | Eval SaaS app | Astra | 1h20m / 2h25m | Astra about half | |
| 10 | 3D knowledge-brain visual | Fable | 24 / 40 min | about equal | |
| 11 | HTML explainer from a video | Fable | — | Astra “much cheaper” | Fable took more, better-annotated screenshots |
| 12 | Paint a portrait in Canva (browser) | Astra | — | Astra cheaper | Fable’s output “laughable” |
| 13 | Upload course drafts to Skool (browser) | Astra | similar | Astra slightly less | Fable could not upload the videos |
| 14 | Clone a website’s feel | Fable | — | — | Fable closer to the original |
| 15 | 12-month YouTube review | Astra | 17 / 42 min | 27 |
- His team reports “they are liking the GPT models lately better for writing”, although in this test he preferred Fable’s sales letter.
- He says Codex and GPT models are “significantly better with browser use”, based on this test and his own Skool work.
Try It
- Walk one working skill down. Rerun it on the next-cheaper model with a QA report required; keep stepping down while the report and your review hold, then step effort down from high. Run it on more than one input before switching.
- Write
/goalwith an objective finish line (“18 cards finished and visually checked”, “257 sources, then the report”). - Before moving a recurring automation off the agent, list every value the agent currently interprets (channel IDs, recipients, thresholds) and hard-code them in the script.
- Pin subagent models in the delegation prompt (“use 5.6 Terra for the subagents”) so research does not run on your most expensive model.
- Keep integrations in
.env, not only in plugins, so an automation can move to another host. - If you run both Claude Code and Codex on one folder, check whether an
AGENTS.md-only setup now covers both (Claude Code v2.1.277+).
Open Questions
- The $14,000-per-month inference figure is Nate’s estimate; how he computed it is not shown.
- Astra vs Fable totals: Nate says both models produced the same summary numbers, but the per-task costs he reads out do not all appear, and some are caption-garbled (task 2’s Astra cost, task 7, task 8).
- Effort names: a “light” setting is mentioned once in passing; whether it is a separate level is unclear.
- Codex SDK automation: the build ran 44 tests and created four trigger.dev tasks (check, install plan, pre-flight, proof); only the check task runs on schedule. How the SDK run is billed per check is not stated.
- Refusals to place trades: his claim that OpenAI and Claude models refuse to place trades is anecdotal and undated.
Related
- Master 97% of Codex in One Hour (Nate Herk) — the earlier course this one builds on
- Build & Sell Claude Code Operating Systems (Nate Herk) — Four Cs, Three Ms and Bike Method in full
- Pricing and Selling AI Automations on Value — the course’s pricing module
- Model vs. Effort (Claude Code) — the Claude-side view of choosing a model and effort level
- Orchestrator + Cheap-Worker Routing — pinning cheaper models to workers
- Cost & Intelligence Levers for Agent Workflows
- Wire It or Loop It — scripts vs agent loops, which the “rogue routine” illustrates
- HeyGen HyperFrames — the video-editing framework in the course