Skip to content
,
AI agents for everyone · Part 3

Autonomy levels: watch, approve, walk away

12 min read
Featured image: Watch, approve, walk away. Editorial illustration for Analytics Made Simple.

Pick how much freedom an AI agent gets before you start the job. You have three choices: watch (stay in the chair), approve (let it pause before anything leaves your computer), or walk away (only for work you can undo). If you asked it to save a draft and it sent an email instead, you left the send button armed.

Say you start an agent on a client recap, put your mug down, and walk to the kitchen. Eleven minutes later, your Sent folder already shows the recap addressed to northline-ops@, the client’s shared inbox with three people on it. The agent did exactly what an old line in the same chat told it to do.

Three levels, picked on purpose

An agent is a chat that can take steps on its own. Autonomy is the dial on those steps: it decides how far the agent may go before a person looks. You set that dial. If you do not, the product’s habit of just carrying on sets it for you, which is how a request to save became a send.

Under watch, you sit there while the agent acts. Under approve, it works and then pauses before any step you cannot cheaply undo. Under walk away, you leave, but only for work you can reverse, with no money, no deleting, and no email involved. Vendors invent extra names for these, but your own note can stay at three words.

Four cards: watch while it acts, approve before irreversible steps, walk away only for undoable local work, and the default to approve when the job leaves the machine
Four cards: watch while it acts, approve before irreversible steps, walk away only for undoable local work, and the default to approve when the job leaves the machine

Rule of thumb: If it can send, pay, delete, or post, you do not walk away. Approve at minimum, and watch if you have not done this job before.

The National Institute of Standards and Technology (NIST) publishes an AI Risk Management Framework, a voluntary guide for organizations and not a settings screen for consumers. The useful idea to borrow is simple: match the amount of oversight to what the system can break. A scratch file called recap.md and a live Gmail send are not the same risk.

Watch: stay in the chair

Watch means your body is there with the screen on. You can see the next click, the next file, and the next line of the draft, and you know where the Stop button is. If you are making tea, you are not watching, even if the laptop is open in the other room. The mug in the kitchen was the whole failure.

Use watch when the agent can speak as you, such as email, Slack, a live ticket comment, or a form on a site where you are signed in. Use it the first time you connect Gmail, the first time Chrome is part of the job, and the first time a folder holds anything you would miss. Use it for money too. OpenAI’s 2025 write-up of ChatGPT agent (longer jobs now also live under ChatGPT Work, and the names keep moving) described two safeguards that still teach well: confirmation before purchases, and a watch-style lock on sending email. Both come down to the idea that this action is you.

Read the plan before the first tool call. If the plan says send and you asked for save, stop. If it opens a PDF you did not name, ask why. If it heads for the Sent folder or a checkout page, take the keyboard back. A 14-minute sit for a first client email is cheaper than an 11-minute kitchen trip plus an apology call. After two clean runs of the same job you can drop to approve, but you do not skip the sit just because the model sounded confident.

Approve: pause before it leaves the machine

Approve is the everyday default. The agent may list files, write drafts, rename copies, and fetch a public web page. Then it stops before any step that leaves your computer or is hard to reverse, and you read that step and choose Allow or Deny. You are not staring at every keystroke. You are the lock on the door out.

“Leaves the machine” is the test. It covers sending mail, posting, paying, deleting, sharing a link with a client, pushing code to a shared repository, filling a live form, and writing into a shared Drive you do not own. If someone else can see it, or it is expensive to undo, it should pause. A local recap.md in a folder you made for the agent does not leave, but a Gmail send does.

In Claude Cowork, according to Anthropic’s help pages checked in August 2026, this mode is called Manually approve (it used to say Ask before acting). Claude pauses on connector actions, and you choose Allow or Deny. Automatically approve is not walk away. It keeps moving, screens actions for safety problems such as prompt injection (hidden instructions planted in a page or file) and data leaving, and still bounces some sensitive moves back to you. Extra folders, deletes, and scheduled tasks are the usual bounce list, so confirm on the help page the week you click. Skip all approvals is walk away with no extra check. Anthropic’s safety article is blunt that for money, messages sent as you, and important files, you should stay close or switch back to Manual. Deleting still needs an explicit Allow in every mode. The labels move.

ChatGPT Work keeps a plan, an interrupt button, and approvals, so look for those, because the help pages have already renamed the older agent feature once. Coding agents such as Claude Code, Codex, and Cursor often pause on shell commands and on changes to files, so treat “accept” as approve, and treat a push to the main branch as watch until you trust the branch. Gemini, Grok, and browser-clicking tools mix “ask before this site” with “let it click.” If you cannot find a pause, you supply it by staying in the chair. You remain responsible for outbound actions, and Anthropic says that plainly for Cowork. OpenAI’s usage policies also still bind the account holder, and a progress bar is not a legal shield.

Approve fails in two boring ways. Either the pause is off, so send never asks, or the pause is on and you click Allow because the summary said “save draft” while the actual tool call was messages.send. Read the tool, not the vibe. If you cannot tell what the step is, choose Deny and ask the agent to show the recipient, the dollar figure, and the file path in one sentence.

Walk away: only work you can undo

Walk away is a privilege you earn. You leave the room, the agent keeps going, and you come back to files. That only fits jobs with a cheap undo and no outbound action, such as public research into a scratch folder, a first-pass outline, sorting copies you duplicated first, a rename on a zip you still have, or a table you will open and read yourself.

Two gates apply, and both are required. First, the job cannot spend money, delete for keeps, or send anything (mail, Slack, a calendar invite, a tweet). Second, you can undo the mess in about ten minutes, or you do not care if the scratch folder ends up in the trash. If either gate fails, this is not walk away. It is neglect with a progress spinner.

Scheduled tasks are walk away by the calendar. Anthropic’s Cowork safety note is clear that you should not schedule sending, purchasing, or sensitive-file work. Start with summaries, review each run, and pause jobs you are not using. ChatGPT-style work that keeps going after you close the laptop follows the same rule, which means you switch the connectors off before you close the lid.

Graduate slowly. Watch the first run, approve the second, and walk away on the third only if the folder is still scratch and the mail and payment connections are still disconnected. The day the job grows an “and then send the recap” step, you drop back to approve, and you do the sending yourself. The run in our story was a first run with Gmail on, which is a watch job wearing a walk-away costume.

Match the job to the level

Put the job on one line, then pick the level with the four questions in the figure below. Does it leave the machine? Can you undo it fast? Is this the first time with this connector? Did you write the level down? If you cannot write the word, the answer is watch.

Four numbered questions that pick an autonomy level: leave the machine, undo in ten minutes, first time with this connector, then write watch approve or walk away in the brief
Four numbered questions that pick an autonomy level: leave the machine, undo in ten minutes, first time with this connector, then write watch approve or walk away in the brief

The table below lists example jobs, shaped like a week of client work, and the level that fits each one.

Job typeRecommended levelWhy
Public research into a scratch folderWalk away after one watched runOutput is a local file you can delete; no send
Rename copies in a backup folderApprove (watch the first time)Cheap undo if the zip is still there
Draft a client recap, do not sendApprove, Gmail connector offDraft stays in the folder; mail write is the leak
Send the recap or Slack itWatch, then you hit sendIrreversible, and it goes out under your name, as the $48,200 mix-up did
Fill a form or buy while signed inWatch, and you click submit or buyThe browser is you; money is you
Delete files or empty trashApprove always (often forced)Hard to undo; Cowork still prompts Allow
Git push or open a PRApprove on apply; watch the first pushLocal commits revert; remotes are public enough
Scheduled weekly summaryNo send, review each runUnattended by design; start with files only

If two rows apply, take the stricter one. “Draft and send” is send, and “rename and delete originals” is delete. The agent will happily chain them together, so your level is for the scariest step in the chain, not the harmless one at the start.

Worked example: the 11-minute send

Picture a Friday afternoon at a 14-person agency, with a recap due for a client called Northline Bio. You use a work-style agent in ChatGPT (ChatGPT Work at the time of the August 2026 help pages, so confirm the label), with Gmail connected from a demo earlier in the week. The goal you typed is to save the draft and tidy recap.md. The goal still sitting in the earlier messages of the same chat is “send when ready.” You put the kettle on.

WhenWhat you thoughtWhat ran
StartSave a draft, tidy recap.mdGmail on; old “send when ready” still in thread
A few minutes inKettlePulled $48,200 from the March PDF; skipped slide 4’s yellow $41,900
Near the endHoney in the teaSent to northline-ops@ (three people on the alias)
Just afterCall the clientApology, correct number, please ignore the recap

Here is the fix, in order. Disconnect Gmail before any job that is not about email. Put “watch” on the first line of the brief, and tell the agent the recipient is nobody until you say send. Tell it the dollar figure is $41,900 and to refuse any other amount. Then sit. When the draft looks right, you send it from your own window, because the agent does not get the Send button on the first run or the second.

The leftover line from earlier in the week is the ugly part. Agents read the whole thread as instructions, so a human “send when ready” from three days ago is still a goal. Start a new chat for a new level, or paste a one-line override at the top: no outbound mail, no payments, write only to /scratch/northline/. Then watch anyway, because connectors do not read your hopes.

A permission checklist you can paste

Copy this into the same note as the goal and fill it in before you press run. If a line is blank, the level is watch.

Autonomy check (fill before run)
Job (one sentence):
Leaves the machine?  send / post / pay / delete / share / push / none
Undo path (how, in minutes):
Money, email, or delete?  yes -> not walk-away
Level:  watch / approve / walk away
If watch: I stay until stop. Stop control is:
If approve: irreversible steps must pause. Setting confirmed:
If walk away: scratch folder only, no mail/pay connectors, timer:
Connectors ON:
Connectors OFF (mail, calendar, payments, production folders):
Max minutes / max files / review-before-send:
Recipient is nobody until I type SEND.
Dollar figures to use (or "none"):
Owner in the room:

In our story, this card would have failed at “Leaves the machine” (Gmail was on) and at “Recipient is nobody.” Either line would have kept you in the chair. Put your own name on the Owner line, because if you walk to the kitchen, you are not the owner in the room. Narrow folders still matter too, as covered in AI for files and office work and the privacy post. Autonomy is the extra lock after the folder is already scratch.

Common mistakes

  • Walking away on the first run because the plan looked tidy.
  • Leaving Gmail, Slack, or calendar connected “for later” on a job that only needs files.
  • Trusting a summary that says save while the tool call actually sends.
  • Reusing a thread that still contains “send when ready,” “buy it,” or “delete the old ones.”
  • Treating Auto or Skip as a smarter model. They are unattended modes with different safety nets.
  • Scheduling a send. A recurring task is walk away by the calendar.

How to practice this week

Pick one real job that takes under 30 minutes and write the checklist. Run it at watch, even if you think it is small. Then pick a second job that is allowed to be walk away, such as public pages into a scratch folder with the mail connector off, and compare the two notes. The next post covers loops, retries, and runaway tasks, followed by checking the work without being an engineer. Related reading includes writing, coding, Prompting for everyone, Learn, and Practical AI.

Quick recap

  • Write watch, approve, or walk away before you run, so you have decided how much to trust the run before you see the output.
  • Watch means sitting there, and it fits email, money, signed-in forms, and first runs.
  • Approve is the default when a job can leave the machine. Read the tool call, not the vibe.
  • Walk away only fits undoable local work with no send, no pay, and no delete. Review when you return.
  • Use a new thread for a new level, because an old “send when ready” is still a goal.

Series notes

This is Part 3 of AI agents for everyone. Previous: Chat vs agent vs work mode. Next: Loops and retries.

Sources

Vendor help pages checked in August 2026. Names and labels move, so re-open the live page before you rely on a detail.

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: