Loop Engineering: Gate the Sub-Agents, or They Burn Your Usage Cap
Loop engineering replaces you as the prompter with a system that runs agents on a schedule. Here is the six-part anatomy, the /goal stop condition, and how to gate sub-agent spend.

For the first couple of years the way you got work out of an AI coding agent was to hold it: write a prompt, read what came back, write the next one. You were the loop. Loop engineering is deciding to stop being the loop — you design a small system that finds the recurring work, hands it to agents, checks the result, records what is done, and decides what happens next.
This guide covers that outer system: its six parts, how /goal lets it stop itself, and how to keep sub-agents from quietly burning your usage cap.
The practice is only a few months old and not settled. Peter Steinberger, the creator of OpenClaw, posted the line that started it on June 7, 2026:
“You shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.”
Two days later Boris Cherny, who heads Claude Code at Anthropic, described the same shift:
“I don’t prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops.”
Addy Osmani’s June 9 essay gave the practice its name and anatomy. Everything below comes from those sources and the enterprise critique that followed, with caveats attached.
What a loop actually is
Osmani’s definition is the cleanest available:
“Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead. A loop here can be thought of a recursive goal where you define a purpose and the AI iterates until complete.”
Strip it down and a loop is three things: a goal written as a stopping condition, success criteria the agent can measure its own output against, and recurrence. The chain you are replacing, and the one you are building:
You → Agent → You → Agent → … (you drive, every turn)
Schedule → Triage → Execute → Verify → Deliver → Schedule → … (the system drives)
Osmani places loop engineering “one floor above” harness engineering. The harness is the environment one agent runs inside; the loop decides which tasks run, when, and what counts as done. The single-agent cycle this sits above — observe, plan, act, check — is covered in the discipline guide.
The six parts
The anatomy is five primitives plus one place to remember things. Both major coding tools now ship all six, which is a strong signal that the shape is general rather than one person’s hack:
- Automations — discovery and triage on a schedule. Codex app: an Automations tab whose findings land in a Triage inbox. Claude Code: scheduled tasks, cron,
/loop. - Worktrees — isolate parallel runs. Codex app: a worktree per thread. Claude Code:
git worktreeor--worktree. - Skills — codify project knowledge once. Both use a
SKILL.mdfolder, invoked with$nameor implicitly. - Connectors — reach your real tools. Both speak MCP; Codex also bundles them as plugins.
- Sub-agents — one has the idea, a different one checks it. Codex app: TOML files in
.codex/agents/. Claude Code: subagents in.claude/agents/. - State — track what is done. Markdown files or a Linear board, held outside the conversation.
State is the part people skip and then regret: the model forgets between runs, so memory has to live on disk, not in the context window. A loop with no state file re-derives the same conclusions every cycle. Sub-agents are the part that decides your bill.
The stop condition is the product
/goal is the primitive that lets a loop terminate on its own. Codex CLI added it in v0.128.0; early write-ups describe it as “run until condition met.” Claude Code’s version layers the maker/checker split onto the stop condition — Osmani describes it as “a separate small model checks whether you are done” after every turn, so the agent that wrote the code is not the one grading it.
Verifiability decides whether any of this works. ZeroFutureTech’s editorial spectrum puts code tests, math proofs, and physics simulations at the machine-checkable end, and writing and strategy at the end where a human has to judge. If a machine can decide done-or-not-done, the loop runs while you sleep. If it cannot, you have built a scheduler that wakes you up.
One early ancestor of the pattern is the Ralph Loop, which Geoffrey Huntley described in summer 2025:
while true; do
cat prompt.md | claude
done
Each pass reads the same prompt file, changes code, and starts fresh, so the next iteration gets a clean context window and nothing but filesystem state. Cruder than /goal, but it sets the same two rules: state lives on disk, and the loop needs a reason to stop.
How to gate sub-agent spend
Sub-agents split the writer from the checker, which Osmani calls “the most useful structural thing in a loop, by far” — the model that wrote the code grades its own homework too kindly. They are also the fastest way to burn budget, because every sub-agent does its own model and tool work and usually inherits the coordinator’s model. Four controls, in order:
- Pin every sub-agent’s model explicitly. Never let one inherit the coordinator’s default. The expensive model is a deliberate choice for a verifier, not the starting point.
- Set ceilings before the first unattended run. A step limit and a token cap per run, with the loop stopping and logging when it hits either.
- Default cheap, opt in to expensive. Research and exploration sub-agents rarely need the top model.
- Write cost into the state file. Tokens by sub-agent, per run. Without that line, a leak stays invisible until the invoice arrives.
Build your first loop
This assumes a task that already repeats by hand. Do not automate a task you have not run manually a few times.
- Write the stop condition first. “All tests in test/auth pass and lint is clean” is a stop condition; “make it better” is not. Checkpoint: can you express done as one command whose exit code decides the loop? For example, done when
npm test auth && npm run lintexits 0. - Pick a task that repeats. A loop amortizes its design cost across runs. For a one-off, a good prompt is faster.
- Put the project knowledge in a skill. Without one, the loop re-derives your conventions from zero every cycle.
- Isolate the workspace. A git worktree gives each run its own checkout, so parallel work cannot collide.
- Split the maker from the checker. A different agent, ideally cheaper, checking against the spec and tests, not approving.
- Set the budget. Step ceiling, token cap, explicit model per sub-agent. TrueFoundry’s summary of why: “an unattended loop also makes mistakes and spends money unattended.”
- Put state on disk. A progress file recording what was tried, what passed, and what is open. Checkpoint: run the loop twice and confirm the second run does not redo the first run’s work.
Then test it like a production job. Around five steps each reliable 95% of the time complete cleanly only about three-quarters of the time, and unattended mistakes compound into the state the next run reads.
Where loops leak
The sub-agent burn trap. On a September 2026 daily-brief episode, the host handed an agent a routine research task and found “a major percentage of our current usage cap” gone, because the five or 10 sub-agents it spun up were all on the most advanced model — “completely not required for the task at hand.” That is a default, not a one-off: the coordinator picks a top model and every sub-agent inherits it. Labs may route this automatically eventually; the host’s own view is that model choice “will still be something that people need personal agency around.”
Failure stacking. Every extra step adds another chance to fail, and the mistakes that survive get written into state. This is why the maker/checker split is essential, not optional.
Comprehension debt — the growing gap between what is deployed and what you actually understand. Osmani: “The faster the loop ships code you did not write, the bigger the gap between what exists and what you actually get.”
Cognitive surrender — letting the loop decide so you do not have to think about the work. Osmani’s mechanism is exact: designing the loop is “the cure when you do it with judgement and the accelerant when you do it to avoid thinking — same action, opposite result.”
Verification stays yours. “A loop running unattended is also a loop making mistakes unattended.” The verifier split is what makes the loop’s done mean something, and even then done is a claim rather than a proof.
Anti-patterns
- A loop whose stop condition only a human can judge.
- The same agent writing and grading the work.
- Spawning a sub-agent without setting its model.
- State that lives only in the conversation.
- Adding a schedule before the manual version has worked more than once.
- Automating a one-off because automation feels like progress.
Done means
Before you schedule an unattended run, these should all be true:
- One repeating task has one loop, with a stop condition a machine checks.
- The maker and the checker are different agents, and the checker is cheaper.
- Every sub-agent has an explicit model, and the loop has a step and token ceiling.
- Project knowledge sits in a skill, not in a prompt pasted into a schedule.
- A state file records what was tried, what passed, and what is next.
- You have read one run end to end.
What this article does NOT cover
This guide does not walk through the Codex app or Claude Code configuration screens, and does not compare the two products. It does not cover the agent’s inner loop, which is a separate guide, or fleet runtimes that stack many loops under one orchestrator. It takes no position on enterprise governance beyond noting that the runtime, not the loop, decides credentials, agent identity, approval gates, and spend attribution. Check each command and flag name in the current docs before you copy anything; the tools are still shipping.
Related guides
- Loop Engineering — the single-agent cycle this guide sits above
- Harness Engineering — the environment a single agent runs inside
- Context Engineering
- Monothread AI Workflows
Research basis: shared/abs-research-briefs/aidb/loop-engineering/loop-engineering-research-2026-09-20.md (verified 2026-09-20; URL gate PASS 23/23).
Sources
- Steinberger, P. — X post coining loop engineering (June 7, 2026)
- Osmani, A. — Loop Engineering (June 9, 2026)
- Osmani, A. — Comprehension Debt
- Osmani, A. — Cognitive Surrender
- ZeroFutureTech — Stop Prompting, Start Building Loops (June 19, 2026)
- TrueFoundry — Loop Engineering at Enterprise Grade (June 16, 2026)
- OpenAI — Scheduled tasks, ChatGPT docs
- Anthropic — Claude Code overview
- Huntley, G. — Ralph Wiggum as a Software Engineer
- The AI Daily Brief — Episode 2026-09-20 (transcript)
Sources
- Steinberger, P. — X post coining loop engineering (June 7, 2026)
- Osmani, A. — Loop Engineering (June 9, 2026)
- Osmani, A. — Comprehension Debt
- Osmani, A. — Cognitive Surrender
- ZeroFutureTech — Stop Prompting, Start Building Loops (June 19, 2026)
- TrueFoundry — Loop Engineering at Enterprise Grade (June 16, 2026)
- OpenAI — Scheduled tasks, ChatGPT docs
- Anthropic — Claude Code overview
- Huntley, G. — Ralph Wiggum as a Software Engineer
- The AI Daily Brief — Episode 2026-09-20 (transcript)



Submit a take
Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.