When an AI coding agent says a change is done and shows green checkmarks, you still have to read what it actually changed. Green checkmarks show that some automated checks passed. They do not show that the change is what you asked for, or that it is safe to ship.
Picture a Friday afternoon. Codex, OpenAI’s coding assistant, has just finished a “small” fix: the filter on your sales dashboard (the screen of live charts your team checks) was ignoring old, closed accounts. Its message in the chat sounds cheerful, and a row of green checkmarks sits in the terminal (the text window where you type commands) like a participation trophy. Your finger hovers over the button that sends your changes to the shared copy of the code, and you are already thinking about the weekend. Then you open the diff, which is the list of every line that changed. It shows forty-three changed files and a whole set of automated tests deleted. It also shows a new line holding what looks like a real password for an outside service. The one-line fix you asked for is buried among changes to dozens of files you never asked it to touch. The model is not evil, but shipping blind is how small fixes become Monday emergencies.
This post is part of the ChatGPT Codex / coding tutorial. The previous post covered skills, tasks, and scheduled help, which you should adopt carefully. This post is the brake that makes those extras safe. You will learn to read the diff, run the automated checks yourself, look for passwords that slipped in, confirm it only changed what you asked for, and keep a person in charge of saving the change. If you still need to sort out when to use Chat, Work, or Codex, use the ChatGPT product map. Everyday habits for checking AI output sit in the ChatGPT everyday tutorial, and the multi-step result checks for office agents live in the ChatGPT Work tutorial. The basics stay in the beginner series on learning ChatGPT.
Permission screens and default modes keep changing, so treat the git commands below as patterns you can carry anywhere. Re-check your team policy and the current Codex docs before you write company-wide rules.
Why green checkmarks are not a review
Codex is good at producing plausible sets of changes, and that is exactly what it is built to do. Plausible is not the same as correct, complete, or limited to what you asked for. A change can:
- Pass a unit test (a small automatic check of one piece of code) that never covered the real edge case.
- Fix the bug and also “clean up” three unrelated modules you did not ask for.
- Delete a flaky test instead of fixing the flake.
- Hardcode a secret “just for local use” that later lands on a shared branch.
- Rename something globally in a way that breaks a plugin you forgot existed.
If you only look at the chat summary, you review the story the agent told about the work, but if you look at the diff, you review the work itself. Those are different jobs. Analysts already know this pattern from AI-written SQL (the language for asking a database questions), where the explanation is fluent and the answer is wrong. Code carries the same kind of risk. For the SQL version of this habit, see How to check AI-written SQL before you ship it.
Green checkmarks are especially seductive because they look objective. A passing suite is evidence about the cases the suite covers. It is not evidence that the change matches what the product should do, that no secret slipped in, or that the agent stayed inside the ticket. Treat continuous integration (CI), the automatic checks that run on every change, as a second opinion and not as an official stamp.
The review loop (do not skip steps)
After Codex finishes a change that touches many files, run this loop in order. Make it boring, because a boring routine is one you will actually repeat.

1. Read the diff
Start with the file list, not the first hunk. Ask: does this list match the ticket? Open each file that sits outside the area you expected to change, and ask why it is there. Then read the pieces of the diff that touch business logic, sign-in, money, personal data, database migrations, and settings. Skim pure formatting only after you know nothing dangerous hides inside a commit labeled as style.
These git commands show you the changes. Run them yourself, and do not rely only on the agent’s paraphrase:
# What changed?
git status
# File list + short stats
git diff --stat
# Full unstaged patch
git diff
# Staged only
git diff --cached
# Against main (adjust branch name)
git diff main...HEADIn a code editor, the same idea appears as a side-by-side view of each changed file. A terminal or a visual tool are both fine. What matters is that a human looks at the actual patch before the commit leaves your machine, or at least before it reaches the shared main branch.
2. Run the tests that matter
“The agent said tests passed” is only a rumor until you run them in your own terminal or see the automatic checks on the pull request (PR, a proposed change waiting for review). Use the project’s documented commands from its readme, AGENTS.md, package scripts, or your team wiki. If the project has a fast suite and a slow suite, always run the fast one, and run the slow one when the change touches risky areas such as sign-in, payments, data migrations, or shared libraries.
# Examples only. Use your repo’s real scripts.
npm test
npm run typecheck
npm run lint
# Or
pytest -q
go test ./...
cargo testIf tests fail, fix the code or roll the change back, and do not plan to make it green later. If tests pass but you never understood them, you still have not reviewed the logic. Tests reduce risk, but they do not replace reading the sign-in change.
Also watch for a sneaky pattern where the suite got greener only because coverage shrank. If the stats show deleted test files, assume something is wrong until proven otherwise.
3. Check for secrets
Search the diff for keys, tokens, connection strings, private URLs, and customer data samples. Agents sometimes invent placeholders that look real, or paste values from local settings files into the source code. Both are bad, because a placeholder that looks real confuses the next person, and a real secret in git is a security incident.
# Crude but useful smell search on the patch
git diff | rg -i 'api[_-]?key|secret|password|token|BEGIN (RSA |OPENSSH )?PRIVATE|AKIA[0-9A-Z]{16}'
# Also check newly added files
git status --shortIf something secret landed in your working files, remove it before you commit. If it was already committed to a shared branch, follow your security process, which means replacing the secret with a new one and cleaning history only together with people who know what they are doing. Do not casually rewrite shared history by yourself late on a Friday.
4. Check scope
Scope is the question: “Did we only do the job we asked for?” Agents love helpful extras: rename for consistency, extract a helper, update docs, reformat imports, upgrade a dependency “while we’re here.” Helpful extras can be good on a dedicated cleanup PR. On a bug-fix PR they hide the real change and make the review much bigger.
A practical rule is that if the ticket was “fix filter on archived accounts,” the PR should mostly be the filter, tests for the filter, and maybe a one-line comment. If the PR also upgrades a programming language version, that is two tickets wearing one jacket.
5. Human commit
You, or a teammate who owns the code, write the commit message and decide what lands. Let Codex draft a message if you want, then edit it until it is true. “Assorted improvements” is not a message. “Fix archived-account filter in dashboard query; add regression test” is a message. (A regression test checks that an old bug stays fixed.)
git add path/to/relevant/files
git commit -m "Fix archived-account filter in dashboard query"
# Push a feature branch, not necessarily main
git push -u origin fix/archived-filterThe default habit is a feature branch plus a pull request. Pushing directly to main is a decision for your team’s policy, not something a model’s permission prompt should settle. The next post in this series covers git habits for beginners in more depth, and for now you should assume you still own the commit button.
Rule of thumb: The chat summary is a sales pitch, the diff is the evidence, and the tests are a second opinion. Your commit is your signature on all of it.
Diff warning signs: stop and reassess
Use this checklist when something feels off, or as a routine gate for larger agent runs. Any one item is enough to pause, and two items mean you almost certainly need to split the change, revert it, or ask again with a tighter goal.

| Smell | What it often means | What to do |
|---|---|---|
| Unrelated files changed | The agent wandered beyond the task, or misread the cause | Restore those files and ask again for a narrower edit |
| Deleted tests without a clear reason | The agent turned a red result green by removing the proof | Restore the tests, then fix the code or mark the test as skipped with a ticket |
| Hardcoded secrets or keys | Env confusion or bad example code | Remove it, replace the key if needed, and use a settings file or secret store |
| Huge rewrite for a tiny bug | The agent guessed too much from a vague prompt | Reset, and restate the bug with the file and line it lives in |
| “Just trust me” commit message | Nobody can review intent later | Rewrite the message, and if you cannot, you do not understand the patch |
| New dependencies you did not ask for | Convenience over control | Demand a justification, and check the license and where the package comes from |
| Permission or sign-in code touched “by accident” | A mistake here can spread widely | Slow down, get a second reviewer, and add tests |
| Force-push suggested | Rewriting history as a way to clean up | Refuse on shared branches, and prefer a revert |
When you find a smell, say it out loud to the agent in the next turn: “Revert changes outside src/filters/ and the related test. Do not reformat. Do not touch package.json.” (package.json is the file that lists which outside code the project depends on.) Tight constraints after a bad run work better than scolding the model for being creative.
Undo without making it worse
Undo is a skill of its own, and the wrong undo turns a local mess into a team mess. Learn the safe defaults first.
Discard uncommitted edits (usually safest)
If Codex edited files and you have not committed yet, you can throw away the changes one file at a time or for the whole project.
# Throw away unstaged edits to one file (modern git)
git restore path/to/file.ts
# Throw away all unstaged edits in the working tree
git restore .
# Unstage without discarding contents
git restore --staged path/to/file.ts
# Older equivalent people still use
git checkout -- path/to/file.tsThe git restore command is the clearer modern way to make your working files match the last commit, or another source you name. Prefer it when your git version supports it. Official reference: git-restore.
Undo a local commit carefully
If you committed only on your machine and have not pushed (or only pushed to a personal branch you fully control), you can move the branch back to an earlier point.
# Soft: undo commit, keep all changes staged
git reset --soft HEAD~1
# Mixed (default): undo commit, keep changes unstaged
git reset HEAD~1
# Hard: undo commit AND discard those changes (destructive)
git reset --hard HEAD~1The --hard option is a shredder, so use it only when you are sure you want the agent’s work gone. Prefer soft or mixed when you still want to salvage pieces. Official reference: git-reset.
Force-push is dangerous
Force-push rewrites history on the remote. On a shared branch it can erase a teammate’s commits, break open PRs, and confuse the automatic checks. Even on a personal feature branch it can surprise anyone who already pulled. Treat git push --force and git push --force-with-lease as tools for known, deliberate recovery, not as the default cleanup after an agent mistake.
If Codex suggests force-pushing to “clean up” a bad commit on main, that is a red flag, so stop. Prefer a reverse commit or a new fix commit, and use a controlled reset only if your team’s process explicitly allows it and you understand who else has the branch.
# Safer pattern for a bad commit already on a shared branch:
# add a reverse commit instead of rewriting history
git revert HEAD
git pushWorked example: the “small filter fix”
Imagine you asked:
In src/dashboard/query.ts, include archived accounts only when
includeArchived is true. Add a unit test. Do not reformat other files.
Do not change dependencies.Codex returns a cheerful “done,” and you run the review loop.
$ git diff --stat
src/dashboard/query.ts | 12 ++++--
src/dashboard/query.test.ts | 28 +++++++++++++
src/utils/format.ts | 140 ++++++++++++++++++----------------
src/utils/dates.ts | 88 +++++++++++---------
package.json | 2 +-
tests/legacy/dashboard.spec.js | 210 --------------------------------
.env.example | 1 +
7 files changed, 200 insertions(+), 281 deletions(-)The stats tell you a lot before you open any of the changes:
query.tsandquery.test.tslook on-mission.format.tsanddates.tslook like unrelated rewrites, so they are probably out of scope.package.jsonwas not requested, so it is out of scope.- A large test deletion is a classic smell.
.env.examplemight be fine or might hide a bad pattern; open it.
The recovery path looks like this:
# Keep the good files, drop the rest of the agent’s surprise
git restore src/utils/format.ts src/utils/dates.ts package.json tests/legacy/dashboard.spec.js
# Inspect remaining diff carefully
git diff
# Run tests
npm test -- src/dashboard/query.test.ts
# Commit only the intentional paths
git add src/dashboard/query.ts src/dashboard/query.test.ts
git commit -m "Respect includeArchived in dashboard query"Then tell Codex what you did and why, so the next turn does not re-apply the junk:
I restored format.ts, dates.ts, package.json, and the legacy dashboard spec.
Keep only query.ts + query.test.ts for this task. Do not reformat utils.
Do not delete tests. Confirm with git diff --stat before more edits.That conversation pattern matters as much as the git commands. Agents will happily widen the scope again if you delete files quietly and keep chatting as if the bigger plan still stands.
Permissions help, but review still wins
Codex and the ChatGPT desktop coding tools have approval steps, sandboxes (walled-off areas where an agent can work safely), and features tied to your plan, and these change over time. Those protections are real, but they are not a substitute for reading the patch. Approving a write to three files still requires you to understand what was written. Modes that skip the prompts raise speed and risk together.
If the earlier post on scheduled help tempted you to run multi-file jobs while you sleep, here is the answer. Schedule a job only when its worst failure is a draft, and still run the review loop before anything merges. Autonomy without a merge gate is only pretend safety.
A 15-minute practice drill
Do this on a throwaway branch this week, even if you already know git.
- Ask Codex for a deliberately vague change in a sample project, such as “improve error messages a bit.”.
- Stop when it claims to be done, run
git diff --statonly, and write down three warning signs you can see without opening the changes. - Open the full diff, and mark one change that is on task and one that is off task.
- Restore the off-task files, then ask again using the limits from the earlier task template.
- Run the real tests, and commit only the files you meant to change, with a message that is true.
The point is to build muscle memory when the stakes are low, so a busy Friday afternoon feels automatic instead of heroic.
Where this sits in the ChatGPT path
- Learn ChatGPT from scratch
- ChatGPT product map
- ChatGPT everyday tutorial
- ChatGPT Work tutorial
- ChatGPT Codex / coding tutorial (this series moves from opening Codex, to exploring, to skills, tasks, and scheduling, to review, and then to git habits).
- A matching brake pedal on Anthropic’s side: Claude Code tutorial (review and undo patterns transfer).
Next in this series come git-friendly habits for beginners, covering branches, small commits, pull request culture, and never force-pushing blind. After Codex, the ChatGPT track moves on to Custom GPTs (your own saved versions of ChatGPT, set up with standing instructions), which handle repeating playbooks that have nothing to do with code.
Quick recap
- The loop is to read the diff, run the tests, check for secrets, check scope, and make a human commit.
- Green checkmarks are a second opinion and not a ship button.
- Warning signs include unrelated files, deleted tests, secrets, huge rewrites, vague messages, surprise new packages, and talk of force-pushing.
- On shared branches, prefer
git restoreand honest reverts over history rewrites. - Tell the agent what you restored, so the next turn does not bloat the PR again.
Sources
Research and further reading used for this article:
- OpenAI: Codex product page (coding agent product context)
- OpenAI: Introducing the Codex app (agent workflows and review-oriented surfaces evolve; re-check current UI)
- Git: git-diff (inspecting patches and stats)
- Git: git-restore (discarding or unstaging safely)
- Git: git-reset (moving branch pointers carefully)
- Git: git-revert (undoing published commits without rewriting history)
- AMS: How to check AI-written SQL before you ship it (same “fluent but wrong” habit in SQL)
- AMS: ChatGPT Codex / coding tutorial (this series)
- AMS: ChatGPT Work tutorial (result checks for office agents)
- AMS: Learn (full series index)
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
