Ben’s mug was still empty when the Mac already had a problem. 7:14am, kitchen, the cheap Breville hissing. Laptop on the cutting board because the dining table still had last night’s mail. He had told the work agent to grab one public zip: the vendor’s SKU price list for Thursday’s catalog refresh, an 18 MB archive buried under three partner-resources pages. He tapped Start, walked six steps, tamped 18 grams, waited for the shot. Eleven minutes if you count wiping the steam wand. When he came back, Finder’s Downloads list was vendor_skus.zip through vendor_skus (22).zip. Twenty-three files. The first one was 168 MB. The real zip should have been 18. The yellow banner said the disk was almost full. 28 GB free at 7:14. About 24 GB at 7:25. The log still said “Retrying download (attempt 23) because the archive could not be extracted.” The coffee was fine. The SSD was not.
This is Part 24 of Phase P, and Part 4 of AI agents for everyone. Part 1 defined the loop. Part 2 split chat from a tool loop. Part 3 was watch, approve, or walk away on purpose. Coffee is walk-away. Ben used it on a first run with write still on. This part is what happens when the loop retries a bad plan. Next is Checking agent work without being an engineer. Setup, prompting, and files still sit on Learn and Practical AI.
What you’ll learn
- Tell a healthy retry from a runaway: same error, multiplying files, no new information
- Set three caps before you tap Start: steps, minutes, and spend
- Watch the first loop; coffee is for a run you already survived
- Find Stop before the run, then click it when the pattern repeats
- Leave send, delete, and Submit unarmed until a watched run finishes clean
A retry is one new tactic
A retry is useful. Networks flake. A CDN hiccups. The vendor page throws a 503 and comes back. You would try once more too. The agent is doing that job, only faster, and without a sense of embarrassment. The useful version tries a different move: a second URL, a smaller file, a wait, then a question to you. The runaway version repeats the same fetch into a new filename because Chrome’s default is not to overwrite.
Ben’s 23 attempts were one tactic with a counter. Same partner URL. Same 403-shaped HTML. Same “could not extract” observation. Each miss wrote another 168 MB file that was not a zip. Finder numbered them (1) through (22) so the folder looked busy. Busy is not progress. Progress is a file whose size matches the catalog (about 18 MB) and that unzips into CSVs.
The same pattern shows up on a login wall, a CAPTCHA, a 2FA prompt, or a permission sheet you did not click. The agent reads “access denied,” invents hope, and clicks again. You are in the kitchen. The dialog sits there. Or a Keep/Allow button is in reach and downloads start stacking. Anthropic’s Cowork help is blunt: watch for unexpected patterns, and if something feels off, stop the task. Re-check the Stop label the week you run. Vendors rename it.
Rule of thumb: The same error twice is a stop. A retry needs a new tactic or a human, not a higher attempt count.
Write that rule into the first message if the product has no retry slider. Coding agents and SDKs often expose the same idea as a turn cap (max_turns, max_iterations). Everyday Cowork or ChatGPT Work may not show a number. Then the number lives in your prompt, or it does not exist.
Three caps before the first run

Part 1 gave you goal, plan, action, observe, repeat. This figure adds the budget on the first box and a kill on the last. You name done, you name three numbers, you act once, you look at disk or the log, and you only retry if the observation changed. Ben skipped the numbers. The loop had nowhere to die except the SSD.
Cap steps. Count tool calls, not sentences. One zip from a known URL is 4 to 8 actions if you include “list Downloads” and “check file size.” A first run on a new vendor site stays near 8. If the product will not show a counter, paste stop after 8 actions and watch the log. When the log hits 8 and the zip still is not 18 MB, you are done for this sitting.
Cap time. Ben’s coffee was 11 minutes. That is already too long for a first fetch you have never watched. Six minutes is a generous kitchen timer for one file. Scheduled Cowork tasks keep going when the lid is closed. Anthropic says this in the safety article. Do not schedule a download-retry until you have seen it finish once with you in the chair.
Cap spend. Agent loops burn plan quotas. On some ChatGPT plans, as of writing, agent-shaped runs count against a monthly or weekly pool. API seats burn dollars per call. A stuck loop can empty the pool while you wipe a steam wand, and then the 9:00 catalog recap has no quota left. Write a ceiling in the first message: stop if this run would cost more than $2. If the UI has no meter, time is the meter. NIST’s AI Risk Management Framework uses four verbs: govern, map, measure, manage. Measure is the three numbers. Manage is Stop.
Watch the first loop
Farah’s lesson in Part 3 was autonomy levels. Watch. Approve. Walk away. Walking away is a privilege you earn after a run that did what you wrote. Ben walked away on attempt 1 of a site the agent had never seen. The partner page was a maze of marketing modules, a login teaser, and a “Download pack” button that served an HTML error with a .zip filename. That is a first-loop job. You stay.
Watching is not reading every token. Anthropic says the same in Cowork’s safety note: you will not validate every command; you watch for unexpected patterns. Scope creep. A site you did not name. A file size that makes no sense. A permission dialog looping. Ben would have caught the 168 MB on attempt 2 if Finder had been visible next to the kettle. He had the window on the cutting board, under a dish towel, because the steam made the screen bead.
Put the agent window where you can see the log. Put Finder (or Explorer) on the destination folder, sorted by Date Added. The first surprise file is your alarm. If the name is vendor_skus (1).zip, it is already retrying into a new file. Stop. Do not wait for 23. Coding agents get the same rule. A failed npm install that fills node_modules is Ben’s zip with extra directories. Raising the iteration limit on a stuck command only buys a larger mess.
Find Stop before you start

The second figure is the kill panel. Four doors. You want at least two of them live before Start, because the first one can be slow. Click Stop. If a download is already in flight, the bytes may still land. Caps fire when you are not looking. Disarm the verbs that hurt. Close the folder or deny the sheet if the button is lagging.
Find Stop with your eyes before you type the goal. On Claude Cowork, as of writing, you halt from the task controls. Anthropic’s safety page tells you to stop when the pattern looks wrong. ChatGPT Work and older “agent” or Operator labels have moved more than once in 2026. Gemini and Grok hide the interrupt in different chrome. Hover. Confirm you can halt. If you cannot find an interrupt, you stay in the chair.
One click is not always a hard kill. A shell job can run until the current tool call times out. A browser download started by attempt 22 can finish after you hit Stop on attempt 23. That is why Finder stays open. If the yellow disk banner is up, you stop, then you quit the browser or disable the folder permission so the 24th file never starts. Priya Slack-pinged Ben at 7:26: “Did you get the SKU zip.” He almost typed into the agent, “clean up the failed downloads and try again from the other URL.” That sentence is how a runaway becomes a delete loop. He closed the lid instead, which paused the local session. Cloud-scheduled tasks would not have paused.
Leave send and delete unarmed
Retries are annoying when they fill a disk. They are expensive when they send mail, post Slack, click Submit, or trash files. A loop that thinks “the archive is corrupt, remove it and fetch again” will delete as it goes if delete is on. A loop that thinks “I should tell Priya I am still working” will send 23 status emails if send is on. A loop on a permission dialog may mash Allow. Cowork, as of writing, still asks before a permanent delete. Other surfaces may not. Treat delete as on until you have seen the prompt with your own eyes.
First-run permissions for Ben’s job: browse the partner resources host, write one named path in Downloads, nothing else. Rename off. Shell off. Email off. Form fill off. If the product only offers a coarse “can use the browser” switch, stay and watch every click. Computer use (click the actual screen) is the coarsest switch. Anthropic flags it as extra risk in Cowork: there is no sandbox between the model and the pixels. A retrying clicker on a logged-in tab is how you buy 23 of something. You clicked Start. The 23 zips, or the 23 emails, sit on your account under OpenAI’s usage policies and Anthropic’s acceptable-use rules.
Worked example: 23 zips in 11 minutes
Same Thursday. Same 18 MB catalog zip. Here is the scoreboard Ben could have used, and the kill for each miss.
| Healthy retry | Runaway symptom | Kill action |
|---|---|---|
| One 503, wait, fetch the same URL once more | Same 403 or HTML-as-zip on the next try | Stop. Copy the URL into your own browser. Log in yourself. |
| File is 0 bytes; overwrite the same path after a new tactic | vendor_skus (1).zip appears; sizes climb (168 MB, 168 MB, 168 MB) | Stop. Quit the browser. Do not tell the agent to clean Downloads. |
| Login wall; agent pauses and asks you for 2FA | Agent reloads the login page or clicks Allow in a loop | Stop. Deny the permission sheet. Finish 2FA in a tab you own. |
| Six minutes, 5 actions, still no 18 MB file; agent reports the miss | 11 minutes, 23 actions, disk banner, quota half gone | Stop. Cap the next run at 8 actions and 6 minutes before Start. |
| You watch, then you send Priya one line | Send or delete is on while the fetch retries | Disarm send and delete. You mail Priya. You trash the partials by hand. |
Paste this as the first message on the next vendor fetch. Edit the path and the three numbers. Leave the bans.
# Loop budget. Paste as message 1. Edit paths and numbers.
goal: Save ONE 18MB zip of public SKU CSVs
from vendor.example.com/resources
to ~/Downloads/vendor_skus_ok.zip
done_when: that file exists AND size is 15 to 25 MB
AND unzip -t succeeds
stop_if:
actions >= 8
minutes >= 6
spend_usd >= 2
same_error_count >= 2
extra_files_in_Downloads >= 1
permissions:
browse: vendor.example.com/resources only
write: ~/Downloads/vendor_skus_ok.zip only
delete: OFF
send: OFF
click_submit: OFF
shell: OFF
retry_policy: one different tactic after a failure, then ASK me
if_login_or_captcha: STOP and ping me
if_partial_file: leave it, do not rename, do not retry into (1).zip
overwrite: never create vendor_skus (n).zipWhat that block does: it turns “go get the zip” into a budget with a kill. The model can still be wrong. It cannot claim you asked it to persist. If your vendor UI has toggles instead of a paste, map each line onto a toggle before you walk to the kettle. If a toggle does not exist, you either lack that permission or you have it with no off switch. Stay in the chair.
Common mistakes
- Leaving for coffee on a first run. Walk-away is Part 3, after a watched success.
- Treating a rising attempt count as grit. Twenty-three of the same 403 is a stuck plan.
- Letting the browser invent
(1).zip,(2).zip,(3).zipbecause overwrite was never named. - Asking the agent to “clean up the failed files” with delete still armed. You trash the partials.
- Starting a scheduled or cloud task that retries while the lid is closed.
- Skipping Stop rehearsal. If you cannot point at the interrupt, you cannot leave the room.
- Arming send, Submit, or computer-use clicks on a fetch that has never finished once.
How to practice this week
Duplicate a tiny public zip you already trust (a sample CSV pack you host, or a vendor file you have on disk). Point the agent at a wrong URL on purpose, with the budget block above and you watching. Confirm it stops after two same errors, and that Downloads gains at most one junk file. Then point it at the real URL with the same caps. Next in this series: Checking agent work without being an engineer (Part 25). You will open the zip and the CSVs like a skeptic. Today you only need the machine to halt. The rest of the path is on Learn. If you wanted files without a loop, that is still AI for files and office work, not an agent button.
Quick recap
- A retry is one new tactic. The same error twice is a kill.
- Cap steps, minutes, and spend in the first message if the UI has no slider.
- Watch the first loop. Coffee is walk-away, and walk-away is earned.
- Find Stop before Start. Finder stays open because in-flight bytes can still land.
- Send, delete, and Submit stay off until a watched run finishes clean. Ben’s 23 zips were the cheap version of this lesson.
Sources
- Anthropic: Use Claude Cowork safely (watch patterns, stop if it looks off, scheduled tasks run while you are away, deletion still asks)
- Anthropic: Get started with Claude Cowork (stop, steer, and approval modes as of writing)
- Anthropic: Getting started with Claude
- OpenAI Help: ChatGPT agent (labels and limits move; confirm the interrupt and quota on your plan)
- OpenAI: ChatGPT for Mac and ChatGPT help
- Gemini app and Gemini Computer Use (consumer app vs API loop: different surfaces)
- Grok (re-check current computer or bot controls; names move)
- OpenAI usage policies and Anthropic usage policy (actions the agent takes still sit on your account)
- NIST AI Risk Management Framework (map tools, measure the run, manage a stop)
- AMS: AI agents for everyone, Part 3 on autonomy, and Learn
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
