Tuesday, 11:15 a.m. You start a “quick” Work job: turn three PDFs, a CSV export, and last week’s notes into a six-slide status deck for the 2 p.m. leadership sync. At 11:22 the agent is still reading files. At 11:41 it has a full outline, a draft sheet of “key metrics,” and two slides already laid out. At 11:58 it asks whether to pull last quarter’s revenue from a connected sheet you barely remember approving. You are mid-slack thread. You tap Yes. At 1:47 you open the deck. It looks finished. The chart is wrong. The customer name on slide 3 is a near-miss from a different account. The “risks” slide invents a vendor SLA that was never in your PDFs. Nobody lied. The run was multi-step, fluent, and under-checked. That is the failure mode this part is built to prevent.
This is Part 3 of the ChatGPT Work tutorial. Part 1 covered what Work is for (and what it is not). Part 2 walked a first agentic task with approvals. Here we stay in Work and go longer: multi-step jobs, intermediate checkpoints, mid-run steering, usage awareness, and a result checklist that catches workslop (polished, fluent, wrong) before the file leaves your desk. If product doors still blur, keep the ChatGPT product map open. Everyday Chat habits (structure, verification, no auto-send) live in the ChatGPT everyday tutorial. Account and safety foundations stay in Learn ChatGPT from scratch.
What you’ll learn
- The multi-step loop: goal, plan, act (tools), checkpoint, done or steer
- How to write a goal that survives a long run without scope creep
- When to demand a plan before the agent burns usage and files
- How to interrupt and redirect mid-run without starting from zero every time
- A result checklist for numbers, names, dates, scope, citations, and “would a human send this”
- How workslop shows up in decks, sheets, and docs, and how to detect it fast
- Usage and limit habits so one enthusiastic afternoon does not empty the tank
UI labels move. Plan menus move. The loop does not. Treat official Help Center pages as the source of truth for buttons and quotas the week you train a team.
What a multi-step job actually is
In plain Chat, a “job” is often one request and one reply: rewrite this email, outline this doc, explain this error. You stay in the thread. You copy the useful bits. The blast radius is mostly words on a screen.
In Work, a job is a goal that needs a sequence: gather context, open files, use connected apps or tools, draft intermediate artifacts, revise, and leave finished deliverables (docs, sheets, decks, reports, sometimes Sites). OpenAI’s product framing for Work is longer, more involved work across files and apps, with progress you can follow, questions you can answer, direction you can change, and important actions you approve. That is agentic office work, not a longer chat bubble.
Multi-step does not mean “dump everything in one paragraph and walk away.” It means you accept that the agent will plan, act with tools, and produce intermediate state. Your job shifts from writing every sentence to specifying the goal, approving the path, and verifying the result.
The loop: goal, plan, act, checkpoint, done or steer
Run every non-trivial Work job through the same loop. Make it boring. Boring is how you avoid the 1:47 p.m. disaster.

1. Goal
Write the outcome, audience, constraints, sources, and non-goals in one block before tools fire. If you cannot finish the sentence “done looks like…,” you are not ready for a long run.
WORK_GOAL (paste as the top of the job)
Outcome: [what file or artifact must exist when we stop]
Audience: [who will read or use it; what they already know]
Sources allowed: [files you attach / apps you enable / URLs you name]
Sources forbidden: [guessing, web unless I say so, other clients]
Constraints: [length, format, tone, deadline, brand rules]
Must include: [exact sections, metrics, dates]
Must not: [send email, edit live systems, invent numbers]
Success test: [how I will check this before anyone else sees it]
Open questions for me: [list anything missing before you act]Example for the Tuesday deck job:
Outcome: 6-slide PPTX status deck + 1 tab sheet of metrics
Audience: leadership sync; 10 minutes; no deep methodology
Sources allowed: three PDFs attached, CSV export attached, notes.txt
Sources forbidden: inventing SLA language; pulling other accounts
Constraints: North America only; as-of date on every chart; plain English
Must include: progress vs plan, top 3 risks with owners, next 2 weeks
Must not: send calendar invites; email anyone; change shared drive files
Success test: I can recompute the 3 headline numbers from the CSV
Open questions: confirm whether "pipeline" means open ops or weighted $2. Plan
For anything longer than a tiny tidy-up, ask for a plan first. Product surfaces often support plan-style review before work begins: context, questions, step list, then your approval. Even when the UI does not force it, you can still say: “Propose steps only. Do not create files yet.”
A usable plan names tools, intermediate files, and decision points. A vague plan (“research, draft, polish”) is not a plan. Reject it.
| Plan quality | What you see | What you do |
|---|---|---|
| Weak | “I’ll analyze the docs and make a great deck.” | Stop. Demand steps, sources, and stop rules. |
| OK | Numbered steps, named files, one approval gate before write. | Edit scope; approve if sources match. |
| Strong | Steps + intermediate checkpoints + explicit “I will ask you before X.” | Approve, then watch the first real write. |
3. Act (tools)
This is where Work earns its cost: reading attachments, using enabled apps, browsing when allowed, drafting sheets and slides, iterating. Keep the tool surface small. Every extra connector is more data the agent can touch and more ways a confused step can wander.
While it acts, you are not a spectator of theater. You are a supervisor of a junior analyst who works very fast. Watch which files it opens. Watch which apps it calls. If it starts “helpfully” expanding to a second client folder you did not name, stop it.
4. Checkpoint
Checkpoints are deliberate pauses. Insert them yourself if the agent does not. Good places:
- After source inventory (“list every file you will use and what each contributes”)
- After metric definitions (“write the formula for each headline number in plain English”)
- After first draft structure (slide titles or doc outline only)
- Before any high-impact action (share, send, overwrite, publish)
- After final package (files closed, checklist ready)
A checkpoint is not vibes. It is a short artifact you can inspect: a bullet outline, a metric table, a file list. If the agent only says “looks good so far,” you do not have a checkpoint.
5. Done or steer
At each checkpoint you either continue, stop, or steer. Steering is normal. Multi-step work without mid-course correction is how wrong definitions get polished into leadership art.
Steering mid-run without chaos
Official agent guidance for ChatGPT has long stressed that you can interrupt, clarify, take over sensitive browser steps, or stop a task. Work inherits the same human-in-control idea: follow progress, answer questions, change direction, approve important actions. Use that power early, not after the wrong deck is done.
What good steer messages look like
Bad steer: “Make it better.” The agent will polish. Polish is not truth.
Good steer is specific, scoped, and negative when needed:
STEER (mid-run)
Stop creating new slides.
Keep slides 1-3 as drafted.
On metrics: pipeline = open opportunities in USD, weighted, as of 2026-07-01.
Remove any SLA claims not present in the three PDFs.
Do not open additional apps.
Show me the revised metric table only, then wait.When the agent is mid-tool-call and something feels off, stop the run rather than stacking five corrective paragraphs. A clean stop plus a corrected goal beats a long argument with a confused partial state.
When to restart instead of steer
Restart when the wrong source set was used, the wrong client was mixed in, or the definition of the primary metric was wrong for more than one step. Steering a contaminated intermediate sheet often leaves residue: old columns, old chart series, old footnotes. Restart costs usage. Shipping contaminated work costs trust. Trust is more expensive.
| Signal | Steer | Restart |
|---|---|---|
| Tone or slide order off | Yes | No |
| One number formula wrong, rest solid | Yes (recompute + replace) | Maybe if many dependents |
| Wrong client files mixed in | No | Yes |
| Agent invented sources | No for narrative polish | Yes after scrubbing claims |
| Scope doubled without ask | Hard scope cut | If files already sprawled |
Usage: multi-step jobs are not free chat
Agentic runs burn more than a quick Chat rewrite. Older agent-mode help documented monthly message-style limits that counted each user-initiated agent request (with intermediate clarifications often not counting the same way). Work is a newer product surface and quotas can differ by plan and rollout. The durable habit is the same: treat long jobs as a budget item, not as infinite coffee.
- Batch small rewrites in Chat; save Work for multi-step deliverables.
- Ask for a plan first so you do not pay for three wrong full drafts.
- Disable apps you do not need for this run (less tool thrash, less risk).
- Prefer one well-scoped job over five overlapping experiments.
- On team plans, assume someone owns spend; do not be the surprise invoice.
If you are training a team, put the usage rule in the same one-pager as the approval rule. People who think Work is “Chat with more polish” will empty limits on subject-line rewrites and then have nothing left for the real deck at 1 p.m.
Result checklist (run before anyone else sees the file)
When the agent says it is done, you are not done. Run this checklist in order. Skip steps only for truly disposable drafts, and even then run a shortened version so your standards do not rot.

1. Numbers match sources you control
Pick the numbers that would change a decision. Recompute at least three from the CSV, export, or system of record you trust. Do not only re-read the agent’s summary of the numbers. Open the cells. Open the PDF page. Open the sheet tab.
NUMBER_CHECK
1. Row count of source filter vs rows implied in the summary
2. Sum of amount (or count) vs warehouse / finance total for same filter
3. One rate = numerator / denominator with the definition written out
4. As-of date and timezone if the metric moves daily
5. Units and currency (thousands? millions? local currency?)
If any fail: quarantine the deliverable. Do not "fix the narrative" first.2. Names and dates correct
People, companies, products, ticket IDs, file names, “as of” stamps, meeting dates. Near-miss names are a classic multi-step error: the agent saw a similar string two files ago and carried it forward with confidence. Search the draft for every proper noun you care about. Compare to the source spelling character by character when stakes are high.
3. Scope matches the ask
If you asked for six slides and got fourteen, that is not a free gift. Extra scope often means diluted attention and invented sections. If you asked for North America and the chart includes EMEA, stop. Scope failures are cheaper to catch on titles and filters than after the meeting.
4. No invented citations
Require sources for claims that sound factual. “According to industry benchmarks…” without a real source is decoration. “Per PDF page 4…” you can verify. If the agent cannot point to an attached file, a named URL you allowed, or a sheet cell, treat the claim as untrusted.
5. A careful human would send this
This is the social test. Would you put your name on the email that attaches this file? Would you defend every number if Finance joins the call? If the answer is “mostly, if nobody digs,” you still have workslop risk. Fix it or label it draft with explicit open questions.
| Check | Pass looks like | Fail looks like |
|---|---|---|
| Numbers | Three recomputes match source | Pretty chart, untraceable cells |
| Names / dates | Exact match to source spelling | Near-miss client, wrong quarter label |
| Scope | Deliverable matches goal block | Extra regions, extra claims, extra apps |
| Citations | Claims map to files you own | Vague “research shows” |
| Human send | You would defend it live | You hope no one asks |
Workslop detection for multi-step outputs
Workslop is fluent, structured, wrong. It is not “AI writing with a personality.” Tone is a style choice. Workslop is a truth and fitness problem. Agents make it easier to ship because each step can polish the previous step’s error. A misread PDF becomes a cell. The cell becomes a chart. The chart becomes a bold claim in an email subject.
Fast detectors that do not require redoing the whole job:
- Empty sections that look full: a Risks slide with three vague bullets and no owners.
- Methodology with no method: “we analyzed the data” without filter, grain, or as-of date.
- Perfect round numbers where the source is messy (unless you know the source is rounded).
- Citations that point nowhere: footnotes without pages, “internal research” with no file.
- Scope inflation: you asked for status; you got strategy, market, and a product roadmap.
- Consistency breaks: slide 2 says 12%, appendix table says 9% for the same filter.
Plain test: If the only review you did was “does this sound like something we would send,” you reviewed the outfit, not the body.
Analysts already know a cousin of this risk from AI-written SQL: confident clauses, wrong grain. Same muscle. If you want that parallel on AMS, keep How to check AI-written SQL nearby.
Worked example: multi-step status package
Walk a compact version of the Tuesday job so the loop is concrete.
Goal block (you write this)
Six-slide deck + one metrics tab. Leadership, 10 minutes. Only the three PDFs, the CSV, and notes. No email. No extra apps. Success = three headline numbers recompute from CSV.
Plan you demand first
- Inventory sources; list metrics available in CSV columns.
- Propose metric definitions; wait for your yes.
- Build metrics tab only; checkpoint.
- Draft slide titles only; checkpoint.
- Fill slides from approved metrics; no new metrics.
- Package files; run result checklist with you.
Where you steer
At step 2 the agent proposes “pipeline = total open opportunities unweighted.” You steer: weighted USD, as of the CSV export date printed in the filename. At step 4 it proposes a competitive landscape slide. You cut it: not in scope. Those two steers prevent hours of pretty wrong work.
Result check (you do this)
You recompute weighted pipeline from the CSV filter. You fix one client spelling. You delete a “benchmark” bullet that had no source. You attach the deck to the meeting agenda yourself. The agent never “sent” anything. That is success, not a limitation.
Common mistakes on long Work runs
| Mistake | What happens | Fix |
|---|---|---|
| Goal is a vibe | Agent invents success criteria | Write outcome + must/must-not |
| No plan gate | Usage burns on wrong path | Plan only, then approve |
| All apps on | Scope and privacy wander | Enable only this job’s tools |
| Approve while distracted | High-impact Yes you did not mean | Pause notifications for approvals |
| Review only the final PDF | Intermediate wrongness is hidden | Checkpoint intermediate files |
| Polish before numbers | Workslop gets prettier | Numbers first, design second |
| Silent restart loops | Same bad goal, more spend | Rewrite goal, then re-run once |
Practice this week
- Pick one real multi-step job you own (status deck, folder tidy + summary, sheet from messy notes). Use only files you are allowed to use.
- Write a full WORK_GOAL block. Force a plan. Approve only after you edit it.
- Insert at least two checkpoints. Steer once on purpose, even if the draft is fine, so you practice the muscle.
- Run the five-point result checklist. Write down what failed. Fix before anyone else sees the file.
- Note rough time and whether you would have been faster in Chat. Save that note for Part 4’s decision table.
Do not practice high-impact sends, production CRM writes, or customer secrets on an unapproved personal setup. Skill building does not require a compliance incident.
How this fits the AMS ChatGPT path
You are in the Work tutorial. Part 1 was chooser energy: when Work is the right door. Part 2 was a first supervised task. This part is the long run: loop, steer, check. Part 4 closes with an explicit decision table for when Work beats Chat and when it does not, plus the path into Codex and Custom GPTs.
- Learn ChatGPT from scratch for accounts, plans, memory/Projects basics, privacy, and judgment
- ChatGPT product map for Chat vs Work vs Codex and related splits
- ChatGPT everyday tutorial for human-led Chat skills without agentic blast radius
- ChatGPT Work tutorial (this series)
- Learn for the wider curriculum hub
FAQ
Should every Work job use the full loop?
Use a shortened loop for tiny tasks, but keep goal + final checklist. The full plan-and-checkpoint version is for multi-file deliverables, anything with numbers leadership will quote, or anything that can touch shared systems.
Is interrupting rude to the agent?
No. Interrupting is the job. The product expects you to clarify, redirect, and stop. Waiting politely while it digs a deeper hole is how workslop scales.
Can I trust intermediate files if the final looks good?
Not automatically. Final polish can hide intermediate errors. If a chart depends on a metrics tab, check the tab. If a summary depends on extracted quotes, check the quotes against the PDF.
What if usage runs out mid-project?
Save your goal block, plan, and any verified intermediate files outside the chat if you can. Finish verification in Chat or by hand for the last mile if needed. Do not lower the checklist because the meter is empty. Re-check current plan limits in OpenAI’s Help Center for your account type.
Quick recap
- Multi-step Work = goal → plan → act → checkpoint → done or steer.
- Write success tests before tools fire.
- Demand a real plan; reject vibes.
- Steer with specific constraints; restart when sources or core definitions are wrong.
- Budget usage; Work is not free chat.
- Result checklist: numbers, names/dates, scope, citations, human-send test.
- Workslop is polished wrong. Catch it before the meeting.
- Next: when Work beats Chat (and when it does not), then Codex and Custom GPTs.
Sources
Research and further reading used for this article:
- OpenAI Help: ChatGPT Work and Codex (Chat vs Work vs Codex product split; Work as multi-step work and finished deliverables)
- OpenAI Help: Creating and editing documents, spreadsheets, and presentations with ChatGPT Work (deliverable creation and editing from instructions and sources)
- OpenAI Help: ChatGPT release notes (ChatGPT Work introduction: longer tasks, apps and files, follow progress, change direction, approve actions, scheduled tasks)
- ChatGPT Work product page (goal to plan, side-by-side iteration, review approach before work begins)
- OpenAI: Introducing ChatGPT agent (agentic control themes: permission for consequential actions, interrupt, steer mid-task; historical product context as surfaces evolved toward Work)
- OpenAI Help: ChatGPT agent (agent mode retired in favor of Work for multi-step tasks; usage and safety notes still useful for agentic habits)
- Analytics Made Simple: ChatGPT Work tutorial (this series)
- Analytics Made Simple: Learn ChatGPT from scratch (orientation and safety foundations)
- Analytics Made Simple: ChatGPT product map (modes and product orientation)
- Analytics Made Simple: ChatGPT everyday tutorial (Chat skills without agentic blast radius)
- Analytics Made Simple: Learn (curriculum hub)
