,

First agentic task with approvals

16 min read
Featured image: First agentic task with approvals

You finished Part 1 with a sticky note that says “Work, not Chat” and a vague hope that the product will just know what you mean. It will not. The first good Work run is boring on purpose: a low-risk deliverable from notes you own, only the connections the job needs, approvals you actually use, and a review before anything leaves your desk. That is this part.

This is Part 2 of the ChatGPT Work tutorial. Part 1 sorted what Work is for (multi-step office work with apps, files, decks, sheets, docs) and what it is not (one-shot rewrites, shipping code). Here we complete one real agentic task under approvals. If Chat vs Work vs Codex still blur, re-read Part 1 or the map-level overview in the ChatGPT product map. Orientation and paste rules still live in Learn ChatGPT from scratch. Everyday Chat writing without agent blast radius lives in the ChatGPT everyday tutorial.

Menus move. Treat steps below as a July 2026 field guide. Re-check OpenAI Help: ChatGPT Work and Codex, chatgpt.com/work, and openai.com/chatgpt-work the week you enable anything for a team.

What you’ll learn

  • Why the first task must be low risk and built from materials you own
  • A preflight checklist: plan, surface, capacity, policy, inputs
  • How to state an outcome (not a 40-step micro-script)
  • How approvals fit the loop: plan, act, checkpoint, review
  • A copyable first-task recipe: owned notes → short status deck + risk table
  • What to watch while Work runs, and how to stop scope creep
  • A review checklist before anyone else sees the pack
  • First jobs that teach the product without career-ending risk

The first-task rule

Your first Work task should produce a deliverable you can throw away without emailing a client, charging a card, or renaming a shared drive. The classic mistake is “prove the agent” with the noisiest, most connected job you have. That is how people learn approvals from an incident report.

First-task rule: Low risk deliverable from notes (or files) you own. Not auto-send. Not money. Not the only copy of anything important. Review is part of the job, not a ceremony after the fact.

Why notes you own? Because you can verify them. You know which number is a guess. You know which person left the company. You know which date is still a draft. Agents invent plausible details under pressure. Your own notes give you a ground truth without asking legal for a sandbox dataset on day one.

Why not auto-send? Because outbound messages and calendar invites are high blast radius. Product demos love “find emails and draft replies.” Fine later. First run: draft and leave unsent, or better, do not touch mail at all. Build a pack. Review the pack. Then decide if mail is ever in scope.

First agentic task with approvals: pick low risk goal, connect only what is needed, set approvals, run and watch, review output
First agentic task with approvals: pick low risk goal, connect only what is needed, set approvals, run and watch, rev…

Before you click anything

Work is agentic multi-step office work. Same family of ideas as “hand off a job,” different surface from Chat. You still need a few preconditions or the first session turns into support archaeology.

Plan and capacity

As of writing, desktop Work is available more broadly, including limited Free access for orientation. Web and mobile Work are more complete on paid plans such as Plus, Pro, Business, Enterprise, and Edu. Multi-step Work burns more usage than Chat. Free can be enough to learn the shape of one careful task. Weekly agent habits usually want a paid budget. Confirm what your account actually has this week, not what a screenshot from last quarter promised.

Where you will work

  • Desktop (macOS or Windows): best first deep run when local files, apps, or side-by-side browser workflows matter. Chat, Work, and Codex live together in the desktop app.
  • Web: fine when files already live in connected cloud tools and install is blocked.
  • Mobile: useful for check-in and light steering once a task is running, not for inventing your whole permission model.

For this first task recipe, prefer desktop if you have it. If you only have web on a paid plan with Work, that can still work. Just stay honest about what the surface can and cannot touch.

Policy and data

If this is work data, know your company rules before you grant folders, plugins, or connectors. Personal tax PDFs and client contracts do not belong in a learning sandbox on a home Free account. When policy is silent, ask before the first run that touches real customer information. Paste rules from Learn ChatGPT still apply: less private data is better until you know the trust box you are in.

Inputs for this first run

Prepare a tiny pack you own. Example:

  • A short meeting-notes file (your words, no secrets)
  • A simple open-items list (owners you can verify)
  • Optional: a one-page partner or project list with non-sensitive fields only

Keep originals elsewhere. Work from copies in a dedicated folder if local files are involved. Name the folder something obvious, like Work-Sandbox-2026-status-drill, so you never confuse it with the shared drive finance treats as source of truth.

Approvals in plain English

Approvals are how you stay in control when an agent can act across tools. Exact UI labels move (plan approval, step approval, app action prompts, “allow once,” workspace agent settings on Business/Enterprise). The idea does not: you decide how much the agent can do without asking, and you stay close when stakes rise.

Approval levels idea: manual approve steps for high stakes, auto with review for faster safer work, skip only for trusted low risk
Approval levels idea: manual approve steps for high stakes, auto with review for faster safer work, skip only for tru…
PostureWhen it fitsWhat you do
Manual / high touchMoney, outbound messages, sensitive files, first ever run in a new toolApprove plans and important actions; watch the run; refuse scope creep
Faster + reviewRepeated low-risk pack building after you trust the patternLet more steps proceed; still review the final artifact before share
Loose / skipOnly trusted, low-risk, reversible jobs you have done many timesStill sample outputs; never confuse “I trust this pattern” with “I skip review forever”

For task one, live in the left column. Manual is not puritanical. It is how you learn what the agent tries to do before you give it a wider leash. Stay closest on money, messages, and sensitive files. If the agent asks for a connection that is not required for this status deck, deny it and restate the scope.

Plan before long runs

When the product offers a plan-style step (gather context, propose steps, wait for your OK), use it. Approving a plan is cheaper than undoing a half-built mess after the agent already spent usage. Ask for:

  • What inputs it will read
  • What it will produce (file types and rough structure)
  • What it will not do (no send, no delete, no new connectors)
  • Where intermediate checkpoints sit if the job is long

If the plan includes “email the stakeholders” and you did not ask for that, cut it. Scope creep is how first tasks become second jobs.

Start a Work session (any surface)

Chat and Work share a home. That is a feature and a trap. The trap is typing a multi-step job while still in Chat and wondering why nothing built a deck.

  1. Open ChatGPT on desktop (preferred), web, or mobile.
  2. Select Work (not Chat, not Codex).
  3. Describe the task as an outcome, not a 40-step micro-script.
  4. Attach or point only at the inputs this job needs.
  5. Review the approach or plan. Grant only the access the job needs. Refuse extras.
  6. Let it run. Watch progress. Answer questions. Redirect if it wanders.
  7. Review the deliverable before anyone else sees it.

UI chrome may say Work next to Chat, or use a mode control that is clearly task-oriented. If you cannot find it, confirm plan eligibility, update the desktop app, and re-check Help for your surface. Web and mobile rollouts have been plan-phased; desktop is often the most complete place to learn.

Write the outcome like a manager, not a script

Agents handle goals better than micromanaged click lists, as long as the goal is concrete. Bad first prompts sound like “make this better” or “do something smart with these files.” Good first prompts name the artifact, the audience, the constraints, and the stop lines.

Outcome template

Goal: [one sentence deliverable]
Audience: [who will open this]
Inputs you may use: [only these notes/files]
Must produce:
  - [artifact 1]
  - [artifact 2]
Must not:
  - send email or messages
  - delete or rename outside the sandbox
  - invent metrics not in the inputs
  - add new connectors I did not approve
Format rules: [dates, currency, slide count, tone]
When unsure: leave a blank and a question, do not guess.
Stop when the pack is ready for my review.

Filled example for the first task

Goal: Build a short internal status pack from my own notes for a 10-minute team sync.
Audience: My teammates this week. Not clients. Not leadership board.
Inputs you may use: only the three files I attached (meeting notes, open items, project list).
Must produce:
  - A 5-7 slide status outline (title, wins, risks, decisions needed, next steps, open questions)
  - A one-page risk table with columns: risk, owner, due date if known, confidence (high/med/low)
Must not:
  - send any email, Slack, or calendar invite
  - open tools I did not attach or approve
  - invent revenue, headcount, or customer names not in the inputs
  - delete files
Format rules: dates as YYYY-MM-DD; plain language; no jargon I did not use in the notes.
When a field is missing: write "unknown" and add it to open questions.
Stop when both artifacts are ready for my review only.

That prompt is longer than a Chat subject-line request on purpose. Work is not free. Clear stop lines save usage and rework.

First-task recipe: notes you own → status pack

This is the drill. Copy it, swap the filenames, keep the risk profile.

1. Build a sandbox (if local files)

mkdir -p ~/ChatGPT-Work-Sandbox/2026-status-drill
cd ~/ChatGPT-Work-Sandbox/2026-status-drill
mkdir inputs outputs notes
# copy only non-sensitive files you own into inputs/
# keep a short brief in notes/brief.txt if you like standing rules

Example notes/brief.txt you can paste as standing rules for this folder only:

Status drill rules for this folder only.
Audience: internal team sync.
No external send.
No inventing metrics not present in inputs.
Write finished artifacts to outputs/.
If a date or owner is missing, mark unknown.
Do not delete files without asking.

2. Stage tiny inputs

You do not need a beautiful dataset. You need three short files you can verify by eye. Example content you might write yourself in five minutes:

# inputs/meeting-notes.txt
Week of 2026-07-14 team sync.
Wins: finished partner FAQ draft; fixed login copy on pricing page.
Risks: design review slipped to Friday; still waiting on legal one-pager.
Decisions needed: ship FAQ without legal one-pager, or wait?
Owners: FAQ = Sam; pricing copy = Jordan; legal chase = Alex.
Open: confirm webinar date with marketing (unknown as of notes).
# inputs/open-items.txt
1. Legal one-pager - Alex - due unknown - blocked on counsel
2. Design review for pricing - Jordan - due 2026-07-18
3. Webinar date - Sam - due unknown - needs marketing
4. FAQ final pass - Sam - due 2026-07-17
# inputs/project-list.txt
Project: pricing page refresh - status: in progress - owner: Jordan
Project: partner FAQ - status: draft ready - owner: Sam
Project: Q3 webinar - status: date not set - owner: Sam
Project: legal one-pager - status: waiting - owner: Alex

These files are boring. Boring is the point. You can check every claim in sixty seconds.

3. Open Work and paste the filled outcome

Select Work. Attach the three inputs (or point at the sandbox folder if your surface supports local folder access). Paste the filled outcome prompt. When the agent proposes a plan, read it. Cut anything that looks like send, connect, or expand beyond the three files.

4. Approve only what this job needs

  • Allow reading the attached notes
  • Allow creating the deck/doc/sheet artifacts you asked for
  • Deny mail, chat, calendar, finance systems, and random plugins
  • If it asks for “all files” or a broad connector “to be thorough,” refuse and restate the inputs list

5. Watch and steer

Do not leave for a two-hour meeting on run one. Watch the first few steps. If it invents a revenue number, stop and redirect: “Remove any metric not present in the inputs.” If it starts a twelve-slide epic, stop and redirect: “Stay at 5-7 slides.” If usage indicators look hot and the plan is wandering, pause and narrow.

6. Review before share

Open both artifacts. Run the checklist in the next section. Only then decide if this pack is good enough for an internal sync, or if it stays a practice run in your sandbox.

What to watch while Work runs

Agent runs feel productive because something is always moving. Your job is to watch for the wrong kind of motion.

SignalWhat it might meanWhat you do
Asks for a new connector you did not planScope expandingDeny; restate allowed inputs
Invented metric or personHallucinated fill-inStop; require unknown over guess
Slide count ballooningPretty over usefulCap structure; demand cuts
Wants to send or scheduleBlast radius risingHard no on first tasks
Usage meter climbing with little outputVague explorationPause; tighten the one-sentence goal
Stuck re-reading the same fileAmbiguous notesAnswer the clarifying question briefly; do not dump your whole week

Steering mid-task is normal. You are not “interrupting the AI.” You are managing work. Say what to drop, what to keep, and what “done” looks like now that you have seen the first draft of the plan.

Review loop before anyone else sees it

Work can produce polished packs. Polish is not truth. Use a fixed loop so review does not depend on mood.

  1. Scope: Did it build what you asked for, and only that?
  2. Numbers: Does every figure appear in your inputs or stay marked unknown?
  3. Names and dates: Owners and due dates match your notes?
  4. Commitments: Did it promise a decision or ship date nobody made?
  5. Tone: Internal sync voice, not press-release fluff?
  6. Actions: Confirm nothing was sent, scheduled, or deleted outside the sandbox.
  7. Human bar: Would you put your name on slide three in a room with your manager?

If any answer is no, fix it in Work with a short redirect, or export the useful parts and finish in Chat or your normal office app. Do not “send and hope.” Hope is not a review method.

Worked walkthrough: Sam’s first status pack

Sam is an ops lead who already uses Chat for email rewrites (everyday tutorial habits). On Wednesday, Sam wants a 10-minute status for the team sync. Old path: open notes, rebuild slides by hand, forget one open item, get called out in the meeting. New path: Work under approvals.

Sam copies three non-sensitive files into ~/ChatGPT-Work-Sandbox/2026-status-drill/inputs/, opens desktop ChatGPT, selects Work, pastes the filled outcome prompt, and attaches only those files. The plan says: read three files, draft a six-slide outline, build a risk table, leave everything for review. It also suggests “email the team a summary.” Sam cuts that line before approving the plan.

Mid-run, Work invents “ARR impact: medium” on a risk row. Sam stops it: “No ARR language. That is not in the inputs. Use confidence only.” The final pack has a missing webinar date marked unknown, four open items that match the list, and a deck that is slightly long. Sam asks for a cut to six slides, reviews again, and saves the outputs. Nothing was sent. Total human time was still less than the old copy-paste grind, and Sam now knows how the approval prompts look on this account.

That is a successful first task. Not because the agent was perfect. Because Sam stayed close, refused scope creep, and owned the review.

First jobs that teach without career risk

After the status pack, pick from this list. Keep the same rules: owned or scrubbed inputs, no auto-send, review always.

First jobWhy it teaches WorkKeep out of scope
Status pack from your notesMulti-artifact, easy verificationClient send, finance systems
Meeting notes → agenda + action listStructure and ownersCalendar invites auto-created
Survey themes from a sheet you ownTable + brief combinationPublishing raw quotes externally
Launch brief from product doc + timeline you wroteMulti-source synthesisPartner emails
Project tracker draft from a planning docSheet structure habitsWriting into the live team tracker
FAQ outline from your own draft answersDoc polish without tool sprawlLegal claims you cannot stand behind

Notice what is missing: “clean the shared drive,” “reconcile production CRM,” “send all late training reminders,” “open the repo and ship.” Those are later problems, different modes, or higher approval bars. First jobs teach the loop. They are not a resume of every tool the product can touch.

Common mistakes on the first run

Starting with auto-send

Drafting mail can be useful later. First run should end in a file you review, not a message that already left. Prefer “leave unsent” forever for learning tasks. Prefer “do not touch mail” for task one.

Granting every plugin “just in case”

Each new connection is a new blast radius and often more usage. Connect only what this job needs. You can add tools on run five after you trust the loop.

Using the only copy of important files

Work from copies in a sandbox. If renames go wrong, you want a boring recovery story, not a restore ticket.

Skipping the plan approval

If the product offers a plan step, read it. Thirty seconds of reading beats thirty minutes of undoing.

Treating pretty output as finished work

Run the review checklist. Numbers, names, commitments, and “would I put my name on this?” still sit with you.

Opening Codex or Chat by accident

Wrong door, wrong tools, wrong expectations. Confirm Work before you attach the folder. Confirm Chat if you only need a paragraph. Confirm Codex only for software.

Burning the weekly limit on a vague goal

Work burns more than Chat. One-sentence outcome first. If the agent explores for ten minutes without an artifact shape, stop and rewrite the goal.

When to stay in Chat instead

Not every office task deserves Work. Stay in Chat when:

  • You need a one-shot rewrite or a short plan you will execute yourself
  • The blast radius should stay at “wrong paragraph,” not “wrong multi-file pack”
  • You are still learning audience and structure (everyday tutorial skills)
  • Usage is tight and the job is language-only

Part 4 of this series goes deeper on when Work beats Chat. For now, if you catch yourself opening Work for a subject line, close it. That is Chat’s job.

How this connects to the rest of AMS

Practice for this week

  1. Create a sandbox folder and three tiny input files you own (or scrub).
  2. Write the filled outcome prompt with hard must-not lines (no send, no invent metrics).
  3. Open Work on desktop if you can. Attach only those inputs.
  4. Approve the plan only after you cut extras.
  5. Watch the run. Steer once on purpose even if it is going well.
  6. Run the seven-point review. Save outputs. Do not send anything.
  7. Write one note to yourself: which approval prompt appeared, and what you denied.

Quick recap

  • First task: low risk deliverable from owned notes, not auto-send
  • Preflight plan capacity, surface, policy, and sandbox inputs
  • State an outcome with must-produce and must-not lines
  • Use manual approvals on run one; stay close on money, messages, sensitive files
  • Connect only what the job needs; refuse scope creep
  • Watch the run; review with a fixed checklist before share
  • Work burns more usage than Chat; vague goals waste limits
  • Not for shipping code (Codex); not for one-shot rewrites (Chat)

What’s next

Part 3: Multi-step jobs and checking results expands the loop for longer Work runs: intermediate checkpoints, result checks that catch invented numbers, and habits when usage spikes. After that, Part 4 settles when Work beats Chat and when you should walk back to conversation on purpose.

Sources

Official product and help materials used for this article (re-check the week you set policy; UI and plan packaging move):