Source: raw/Agentic_Loops_for_Knowledge_Workers.md — Creator: Nufar Gaspar (presenter; name captioned “Newfar”), webinar hosted by Nathaniel Whittemore (NLW) on The AI Daily Brief · URL: https://www.youtube.com/watch?v=dLiXiD8hOAI · Platform: YouTube (yt-podcast) · Published: around Labor Day 2026 (NLW’s intro); Gaspar dates her advice “as of August 2026”. The webinar closes with a plug for the host’s paid training programs.
Most loop material in this topic assumes code, where tests and compilers verify the work for free. Gaspar’s webinar is the first in the wiki aimed at non-coding knowledge work: her central move is that knowledge workers have to design the referee themselves, as a “goal card” with a boring, machine-checkable finish line. It then covers when one looping agent stops being enough and how to build a multi-agent work graph without writing code.
Key Takeaways
- Framing (Gaspar quoting NLW): “a loop is basically a job and a graph as an organization.” She places both at the end of a lineage: prompt engineering (what you say to the model) → context engineering (what it knows) → harness engineering (where it runs and what tools it can touch) → loop engineering (how long it runs alone) → graph engineering (how many agents work together). “Each evolution is about giving the AI more independence at a bigger scale.”
- Every agentic tool already runs a loop. Cowork, ChatGPT Work, Codex and Cursor all plan, act through tools, check and adjust under the hood. Those built-in loops are generic, which is why you often have to nudge them. The “advanced” loop is the one where you set the end goal, via
/goalin Claude Code or the loop command in Cursor. - A loop is not a schedule. A schedule answers when (a clock or an event); a loop answers until (“it stops when the work meets the bar”).
- Coders get verification free; knowledge workers must “manufacture the referee.” “Is this report good enough to present to management? … Nobody’s compiler answer these questions for you.” If you cannot design a clear, verifiable finish line, “the answer is don’t loop it.”
- “Knowledge work only succeeds on loop exactly when you invent a very boring very checkable finish line… boring here by the way is a compliment.” “200 verified data points” or “every competitor covered, every claim cited, summary under 150 words” can converge; “make it insightful” cannot.
- Loops are among the most token-hungry things you can run, so reserve them for work where the value is there. For most tasks, one shot with a strong model is the right call.
- Fan out into a graph only on specific signals (see below). “If none of this is correct, stay in the loop or stay in the agent and don’t over complicate things.”
- Knowledge workers may design graphs better than engineers, because a graph is “a confession of how your work really flows, who really owns what and where quality really gets decided”, and that is knowledge they already hold.
Is This Task Loop-Worthy?
Gaspar’s criteria. The first two “have to come together”; she calls them “the most important pair”.
| Loop it when… | Don’t loop it when… |
|---|---|
| It is long-running: one prompt does not get a good-enough result | It is short and one pass does it |
| You can check whether results are good enough or heading the right way | Your judgment is the work and cannot be handed to a referee |
| You would send it off overnight or over lunch and want a finished result, not a draft | It is executive communication, hiring or strategy (“autonomy… does not have a taste on its own”) |
| It should run until a bar is met, or watch indefinitely for something | |
| You tried it one-shot with the smartest model available and it fell short | |
| You need the model to work harder than “one polite pass” (research is the classic case) | |
| It has a natural retry-and-improve shape |
Choosing a normal conversation with your agent for the right-hand column “is the smart move, not the copout.”
Example domains people loop: deep research; ad and campaign optimization (click-through rate is verifiable; run indefinitely or cap the attempts); competitive scans; content audits; result verification; compliance.
Three build requirements: (1) a checkable finish line; (2) a bounded sandbox where loop mistakes are cheap, so not live campaigns at the highest stakes; (3) a task that can converge. Research converges (“more sources, fewer gaps”); “make it better” does not.
The Goal Card
A goal card is what you append after the loop command. Its parts:
- Objective, clear and machine-readable. Her example: “the definitive token efficiency playbook”, as of August 2026.
- Output: here a file with an executive summary and the data.
- Judging and stopping criteria (“what makes or breaks everything”):
- more than 200 unique data points on token efficiency and agentic AI work
- no two points stating the same fact
- each point with a URL, date and type
- a source mix of at least 40 vendor docs, 40 practitioner sources and 20 benchmarks
- Stages or gates (optional). Sometimes you want to leave the agent room to decide how to hit the criteria.
- Fail-safes, in case the target is unreachable (there may not be 200 unique points on the internet): “try up to 30 turns and sandbox only”. You can also cap by time, or with whatever limits the tool offers.
- A cycle log. She asks the model to be verbose about what it is doing and which cycle it is on, and recommends this while you are new to loops.
Her run: Claude Code /goal, Opus rather than Fable, auto mode, “20 30 minutes or more”. Cycle 2 reached 56/200 data points, cycle 3 reached 90/200, and the run overshot the target. Other runs went as high as 300, so “it’s not always as disciplined”, which is why the 30-turn cap matters.
Failure Modes and Fixes
| Failure | What it looks like | Fix |
|---|---|---|
| Runaway spend | It just keeps going | Hard cap |
| Stuck cycles | No progress; the conditions cannot converge | Stop it manually or with the tool’s stop command, or tell it up front to stop and report |
| ”Done but mediocre” | It met the letter of the finish line and the result is still bland | Not a loop failure: your goal definition is wrong, usually on quality and taste |
| Wrong task looped | It should never have been a loop | Turn cap |
When to Fan Out Into a Graph
A graph is “dots and arrows”: nodes are agents or tasks, edges are the work or information flowing between them. A node pointing back to itself is the textbook loop, so “a loop is probably the simplest form of a graph.” What changed this year is the node. It used to be “one fragile LLM call”; now it is a whole agent with two gears, a one-pass task or a full loop. She separates three things people blur: knowledge graphs (stored facts and relations), LangGraph (a developer framework), and the work graph she is teaching (who does what, in what order, and what flows between them).
The progression runs from one agent, once → an agent on a loop → a work graph (several agents on one job) → an “org graph” (a standing team). “You don’t graduate there.” Move on only when you are unhappy with what the single agent produces. The signals:
- Rubber stamp. The loop says done and you keep finding issues it should have caught. Models tend to agree with themselves: GPT verifying GPT is likelier to pass the work than Claude verifying GPT.
- Context overflow / too many hats. One agent is both objective researcher and creative designer, so the research turns creative too early.
- Parallelizable work that you are waiting on serially.
- The finish line keeps changing mid-run: you keep rewriting the goal card because “it’s really two jobs wearing one card.”
- Quality flatlined despite your best effort.
Her research example as a graph: parallel agents for vendor docs, practitioner sources and benchmarks → a synthesizer → a citation verifier with fresh context → a human or second agent check → a designer agent for the output.
Five Ways to Build a Graph
Ordered by increasing complexity. She is explicit that this is “not a competition”, and most knowledge work lives in the first four.
- Draw it on paper or a whiteboard. Drawing forces you to confront what is not well defined, and “until you can agree upon a work graph for something you cannot automate that.” A photo of the drawing can go straight to an agent to build.
- Let the harness improvise. Tools already spawn subagents and “divide and conquer” on complex prompts; that is a work graph you do not control.
- Prompt the graph in words. “Research these five competitors in parallel with separate sub-agents. Then have a fresh context reviewer check the merged results against this rubric.” That one sentence is a seven-node graph.
- Persistent workers plus a skill. Create subagent files (or “folder agents”) you can summon, then a skill that describes the graph: phases, which worker to summon in each, and which model each agent uses. Her research set is a benchmark collector, a citation verifier and a report visualizer. In the demo she handed the finished playbook to the citation verifier with fresh context (“it knows nothing about how the document was produced”) and told it to sample 10 random data points and check each against its real URL, then count totals and the source mix. The verdict was “ship with one required fix”. The report visualizer then turned it into a self-contained HTML page following rules kept in its own file.
- Canvas or code. Her n8n version runs on a schedule: collector agents → a synthesizer (ideally a strong model; collectors can be cheaper) → a citation verifier (“as much as possible needs to be a different model”) → a loop until quality is met → a human verifier → the report visualizer → email. LangGraph is the code option and can visualize its own graph.
Six Habits of Graphs That Aren’t Expensive
- Match the model to the node. Cheap and fast for mechanical steps and yes/no verdicts; stronger models where judgment is needed.
- Pass the contract, not the conversation. Send a draft and a rubric, or findings in a set format, between nodes, never the whole transcript.
- Spend where verification pays. Fan-out costs tokens. Summarize between nodes and cap turns per node. Built-in deep-research features “can easily take millions of tokens.”
- Verify early. In bigger graphs, mistakes compound.
- Put humans at the right gate, at the end or early on to approve the plan, and bring them in on purpose.
- Don’t copy human process limits. Human workflows are shaped by attention span, time and bandwidth that agents do not share; design around the job to be done, not the current process.
Surface caveat. Commands get renamed or deprecated, so check the current command for your tool before relying on it. Gaspar’s example: in Claude Code, subagent tasks were “only accessible in the CLI and not in the desktop” (her claim as of August 2026).
Try It
- Pick one research deliverable you would normally get as a single “polite pass” from your agent.
- Write a goal card with:
- a count target and a source mix
- per-item fields (URL, date, type)
- zero duplicates
- a 30-turn cap and sandbox-only scope
- a cycle log
- Run it with
/goalin the Claude Code CLI. Watch the cycle log the first few times. - If the result is “done but mediocre”, rewrite the card, not the loop.
- Add one verifier node. A fresh-context subagent (ideally a different model) spot-checks 10 random items against their sources and reports totals against the card.
Open Questions
- The webinar shows cycle counts but no cost or token figures for the 30-turn research run. How much did a run like this cost?
- The claim that subagent tasks work only in the CLI and not the desktop app is dated August 2026 and unverified here; check before relying on it.
- The webinar references OpenAI usage data showing agentic tokens overtaking chat around April–May 2026; see OpenAI Enterprise Signals for the first-party figures.
- Where the four-condition test requires automated verification before you loop, Gaspar’s answer for knowledge work is to design a checkable finish line yourself. The webinar gives no evidence on how often designed finish lines hold up compared with test-based ones.
Related
- Graph Engineering (Greg Isenberg) — the diamond-shaped agent graph and the knowledge-graph vs agent-graph distinction Gaspar also draws.
- Should You Build a Loop? The Four-Condition Test — the coding-side decision test and cost math.
- Verifier-First Loops — why the checker must sit outside the maker, which is Gaspar’s “rubber stamp” signal.
- goal` Walkthrough — the command her demo runs.
- Looper — review-gated loops that default to a different model family as judge.
- Claude Code Subagents — the persistent-worker files behind her build option 4.
- The Loop Is the Unit of Work — how loops map across Claude Code and other frameworks.