Skip to content
,

Autonomy levels: watch, approve, walk away

11 min read
Featured image: Watch, approve, walk away. Editorial illustration for Analytics Made Simple.

Farah put the mug down, started the agent, and walked to the kitchen. Eleven minutes. When she came back, Gmail’s Sent folder already showed a recap to northline-ops@, the client’s shared alias with three names on it. She had typed save. The agent treated send as the next step, because Tuesday’s thread in the same chat still said send when ready. The dollar figure in the recap was $48,200, the old quote. The revised number, $41,900, sat in a yellow comment on slide 4, which the agent never opened.

This is Phase P23, and Part 3 of AI agents for everyone. Part 1 defined what an agent is. Part 2 sorted chat vs agent vs work mode vs coding agent. Here the question is how much rope you give it, on purpose. Next is loops, retries, and runaway tasks. If you still need a first account, start with AI setup from zero (including accounts and Free vs paid) and the chooser.

What you’ll learn

  • The three levels you pick before a run: watch, approve, walk away
  • Why approve is the default for anything that leaves the machine
  • When you stay in the chair, even if the product will keep going
  • How to match a job type to a level (and why send is never walk away)
  • A paste-ready checklist so the level is written before Sent pings

Three levels, picked on purpose

An agent is a chat that can take steps. Autonomy is the dial on those steps: how far they may go before a human looks. You set that dial. If you do not, the product’s keep-going habit sets it for you, which is how Farah’s save became a send.

Watch: you sit there while it acts. Approve: it works, then pauses before a step you cannot cheaply undo. Walk away: you leave, and only for work you can reverse, with no money, no delete, and no mail. Vendors invent extra names. Your note can stay at three words.

Four cards: watch while it acts, approve before irreversible steps, walk away only for undoable local work, and the default to approve when the job leaves the machine
Four cards: watch while it acts, approve before irreversible steps, walk away only for undoable local work, and the default to approve wh…

Rule of thumb: If it can send, pay, delete, or post, you do not walk away. Approve at minimum. Watch if you have not done this job before.

NIST’s AI Risk Management Framework is a voluntary map for organizations, not a consumer settings screen. The useful borrow is still simple: match oversight to what the system can break. A scratch recap.md and a live Gmail send are not the same risk.

Watch: stay in the chair

Watch means your body is there. Screen on. You can see the next click, the next file, the next draft line. You know where Stop is. If you are making tea, you are not watching, even if the laptop is open in the other room. Farah’s mug was the whole failure mode.

Use watch when the agent can speak as you: mail, Slack, a live ticket comment, a form on a site where you are signed in. Use it the first time you connect Gmail, the first time Chrome is in the loop, the first time a folder holds anything you would miss. Use it for money. OpenAI’s 2025 write-up of ChatGPT agent (as of writing the longer jobs live under ChatGPT Work; names move) described two gates that still teach: confirmation before purchases, and a watch-style lock on sending email. That is “this action is you.”

Read the plan before the first tool call. If the plan says send and you asked for save, stop. If it opens a PDF you did not name, ask why. If it heads for Sent or a checkout, take the keyboard back. A 14-minute sit for a first client mail is cheaper than an 11-minute kitchen trip plus a 5:02 apology. After two clean runs of the same job, you can drop to approve. You do not skip the sit because the model sounded confident.

Approve: pause before it leaves the machine

Approve is the everyday default. The agent may list files, draft, rename copies, fetch a public page. Then it stops before a step that leaves the machine or is hard to reverse. You read that step. Allow or Deny. You are not staring at every keystroke. You are the lock on the door out.

“Leaves the machine” is the test. Sending mail. Posting. Paying. Deleting. Sharing a link with a client. Pushing a git branch. Filling a live form. Writing into a shared Drive you do not own. If someone else can see it, or it is expensive to undo, it should pause. A local recap.md in a folder you made for the agent does not leave. A Gmail send does.

As of writing, Claude Cowork names this Manually approve (it used to say Ask before acting). Claude pauses on connector actions; you Allow or Deny. Automatically approve is not walk away: Auto keeps moving, screens actions for safety (prompt injection, data leaving), and still bounces some sensitive moves to you. Extra folders, deletes, and scheduled tasks are the usual bounce list; confirm on the Help Center the week you click. Skip all approvals is walk away with no extra check. Anthropic’s safety article is blunt: for money, messages sent as you, and important files, stay close or switch back to Manual. Deletion still needs an explicit Allow in every mode. Labels move.

ChatGPT Work (as of writing) keeps plan, interrupt, and approvals; look for those, because Help articles have already renamed the older agent surface once. Coding agents (Claude Code, Codex, Cursor-style apply) often pause on shell and diffs: treat accept as approve, and treat push to main as watch until you trust the branch. Gemini, Grok, and browser clickers mix “ask before this site” with “let it click.” If you cannot find a pause, you supply it by staying in the chair. You remain responsible for outbound actions. Anthropic says that in so many words for Cowork. OpenAI’s usage policies still bind the account holder. A progress bar is not a legal shield.

Approve fails in two boring ways. The pause is off, so send never asks. Or the pause is on, and you click Allow because the summary said “save draft” while the tool call was messages.send. Read the tool, not the vibe. If you cannot tell what the step is, Deny and ask it to show the recipient, the dollar figure, and the file path in one sentence.

Walk away: only work you can undo

Walk away is a privilege you earn. You leave the room. The agent keeps going. You come back to files. That only fits jobs with a cheap undo and no outbound action: public research into a scratch folder, a first-pass outline, sorting copies you duplicated first, a rename on a zip you still have, a table you will open with your own eyes.

Two gates, both required. One: it cannot spend money, delete for keeps, or send (mail, Slack, calendar, a tweet). Two: you can undo the mess in about ten minutes, or you do not care if the scratch folder is trash. If either gate fails, this is not walk away. It is neglect with a progress spinner.

Scheduled tasks are walk away by the calendar. Anthropic’s Cowork safety note is clear: do not schedule send, purchase, or sensitive-file work. Start with summaries. Review each run. Pause jobs you are not using. ChatGPT-style work that keeps going after you close the laptop is the same rule. Closed laptop means connectors off first.

Graduate slowly. Watch run one. Approve run two. Walk away on run three only if the folder is still scratch and mail and pay are still disconnected. The day the job grows a “and then send the recap,” you drop back to approve, and you send. Farah’s run was run one with Gmail on. That is a watch job wearing a walk-away costume.

Match the job to the level

Put the job on one line. Then pick the level with the four questions in the figure: does it leave the machine, can you undo it fast, is this the first time with this connector, and did you write the word down. If you cannot write the word, the answer is watch.

Four numbered questions that pick an autonomy level: leave the machine, undo in ten minutes, first time with this connector, then write watch approve or walk away in the brief
Four numbered questions that pick an autonomy level: leave the machine, undo in ten minutes, first time with this connector, then write w…

Toy jobs, labeled as examples, shaped like Farah’s Northline week:

Job typeRecommended levelWhy
Public research into a scratch folderWalk away after one watched runOutput is a local file you can delete; no send
Rename copies in a backup folderApprove (watch the first time)Cheap undo if the zip is still there
Draft a client recap, do not sendApprove, Gmail connector offDraft stays in the folder; mail write is the leak
Send the recap or Slack itWatch, then you hit sendIrreversible, your name, Farah’s $48,200 problem
Fill a form or buy while signed inWatch, and you click submit or buyThe browser is you; money is you
Delete files or empty trashApprove always (often forced)Hard to undo; Cowork still prompts Allow
Git push or open a PRApprove on apply; watch the first pushLocal commits revert; remotes are public enough
Scheduled weekly summaryNo send, review each runUnattended by design; start with files only

If two rows apply, take the stricter one. “Draft and send” is send. “Rename and delete originals” is delete. The agent will gladly chain them. Your level is for the scariest step in the chain, not the cute one at the start.

Worked example: Farah’s 11-minute send

Friday 4:47pm, 14-person agency, Northline Bio recap. Farah used a work-style agent in ChatGPT (as of writing, ChatGPT Work; confirm the label) with Gmail connected from a Wednesday demo. Goal she typed: save the draft and tidy recap.md. Goal still sitting in Tuesday’s messages: send when ready. She put the kettle on.

ClockWhat she thoughtWhat ran
4:47Save a draft, tidy recap.mdGmail on; old “send when ready” still in thread
4:51KettlePulled $48,200 from the March PDF; skipped slide 4’s yellow $41,900
4:58Honey in the teaSent to northline-ops@ (three people on the alias)
5:02Call the clientApology, correct number, please ignore the recap

Fix, in order. Disconnect Gmail before any job that is not mail. Put watch on the first line of the brief. Tell it the recipient is nobody until you say send. Tell it the dollar figure is $41,900 and to refuse any other amount. Sit. When the draft looks right, you send from your own window. The agent does not get the Send button on run one, or run two.

The leftover Tuesday line is the ugly part. Agents read the whole thread as instructions. A human “send when ready” from three days ago is still a goal. Start a new chat for a new level, or paste a one-line override at the top: no outbound mail, no payments, write only to /scratch/northline/. Then watch anyway, because connectors do not read your hopes.

A permission checklist you can paste

Copy this into the same note as the goal. Fill it before you press run. If a line is blank, the level is watch.

Autonomy check (fill before run)
Job (one sentence):
Leaves the machine?  send / post / pay / delete / share / push / none
Undo path (how, in minutes):
Money, email, or delete?  yes -> not walk-away
Level:  watch / approve / walk away
If watch: I stay until stop. Stop control is:
If approve: irreversible steps must pause. Setting confirmed:
If walk away: scratch folder only, no mail/pay connectors, timer:
Connectors ON:
Connectors OFF (mail, calendar, payments, production folders):
Max minutes / max files / review-before-send:
Recipient is nobody until I type SEND.
Dollar figures to use (or "none"):
Owner in the room:

Farah’s card would have failed at “Leaves the machine” (Gmail on) and at “Recipient is nobody.” Either line would have kept her in the chair. Put your name on Owner. If you walk to the kitchen, you are not the owner in the room. Narrow folders still matter: see files and office work and privacy. Autonomy is the extra lock after the folder is already scratch.

Common mistakes

  • Walking away on run one because the plan looked tidy.
  • Leaving Gmail, Slack, or calendar connected “for later” on a file-only job.
  • Trusting a summary that says save while the tool call sends.
  • Reusing a thread that still contains send when ready, buy it, or delete the old ones.
  • Treating Auto or Skip as a smarter model. They are unattended modes with different nets.
  • Scheduling a send. Recurrence is walk away by the calendar.

How to practice this week

Pick one real job under 30 minutes. Write the checklist. Run it at watch even if you think it is small. Then pick a second job that is allowed to be walk away: public pages into a scratch folder, mail connector off. Compare the two notes. Next is loops, retries, and runaway tasks (Part 4), then checking the work without being an engineer. Related: writing, coding, Prompting for everyone, Learn, and Practical AI.

Quick recap

  • Write watch, approve, or walk away before you run.
  • Watch: sit there. Mail, money, signed-in forms, first runs.
  • Approve: default when it can leave the machine. Read the tool call, not the vibe.
  • Walk away: undoable local work, no send, no pay, no delete. Review when you return.
  • New thread for a new level. Old “send when ready” is still a goal.

Sources