Skip to content
,
ChatGPT · Part 11

Which ChatGPT model should you use? A plain-English guide

11 min read
Which ChatGPT model should you use? A plain-English guide

Pick a ChatGPT model by testing it on your own work, not by its name. Start with the model ChatGPT picks for you. When an answer is weak, first check whether the cause is missing information, a vague request, or the wrong part of ChatGPT, because a stronger model fixes none of those. For a task you repeat every week, a ten-minute side-by-side test tells you which model to use far better than any ranking online.

Say you write a short note to your manager every month explaining why spending was over or under budget. Last month ChatGPT’s draft was vague, so this month you open the model picker, the menu at the top of the chat, and see several names you do not recognize. A friend says the slowest one is the smartest, and a video says a different one is best for writing. You could pick at random. Or you could spend ten minutes finding out which one writes your note well, and then stop thinking about it for the rest of the year.

This post continues the ChatGPT product map. The earlier posts covered where to work (Chat, Work, or Codex) and when a Custom GPT, a saved version of ChatGPT with its own instructions, beats plain chat. This one is about the model menu inside those places. If you are new to ChatGPT, start with the beginner series on learning ChatGPT and come back when the menu starts to bother you.

Why can’t the model name tell you which one to use?

A model’s name tells you where it sits in OpenAI’s lineup, not how well it does your task. OpenAI renames and replaces models often, and what you see in the picker depends on your plan, your country, and whether you are in Chat, Work, or Codex. A list of names copied from an article or a screenshot is out of date within months, and it was written for someone else’s work anyway.

What stays useful is the short description ChatGPT shows next to each name in the picker, which usually says whether a model is built for quick answers or for thinking longer on hard problems. Read those descriptions, then let your own tests decide. OpenAI’s Help Center keeps the current list of models for each plan, so check it when the picker changes.

Before you switch: four reasons an answer is weak

Most weak answers have a cause that switching models will not fix. Before you open the picker, go through these four in order, because each one is cheaper to fix than the one after it.

CauseHow it shows upFix in ChatGPT
Missing informationThe answer is generic, guesses numbers, or says “typically”Attach the file or paste the facts, or keep them in a Project so every chat can use them
A vague requestThe answer is the wrong length, tone, or shapeSay who it is for, what it is for, and the format you want
The wrong part of ChatGPTIt cannot open your files, run steps, or touch your codeMove the job to Work for multi-step office tasks or Codex for code
A real limit of the modelA clear, complete request still gets shallow reasoning or drops a requirementNow try a model built for harder thinking, with the same request

Here is how that plays out with the budget note. Last month’s request was “explain why we were over budget,” with nothing attached. The draft was vague because ChatGPT had no numbers, so the first fix is to attach the budget sheet. The second draft has the numbers but runs to five paragraphs, so the second fix is to say “four sentences, for my manager, biggest cause first.” The third draft is close. Only if it still misses something you can point to, such as mixing up a one-time cost with a monthly one, is it worth trying a stronger model.

That order matters for a second reason. If you switch models at the first weak answer, you never learn whether the model or the request was the problem. You end up using the slowest model for everything, and the requests stay just as vague.

The ten-minute side-by-side test

For a task you do often, test two models once and write down the result. The test works best on a real example from your own work, because that is the only kind of task you care about. Here are the steps:

  1. Pick one real example of the task, such as last month’s budget sheet, and write the best request you can, using the four checks above.
  2. Open two new chats, choose a different model in each, and paste the same request and file into both. New chats matter, because an old chat carries earlier messages that change the answer.
  3. Copy both answers into a document as “A” and “B,” without noting which model wrote which. If you know which one is the expensive model, you will tend to like it more.
  4. Score each answer on three things you can check: are the facts right, did it follow the request, and how much would you still need to edit?
  5. Only now look at which model wrote which, and note how long each one took to answer.

A simple score sheet keeps the test honest. Copy this into a note and fill it in:

Task: monthly budget note for my manager
Request: same text and same file in both chats

             Facts right?  Followed request?  Edits needed   Wait
Answer A     yes           yes                one sentence   fast
Answer B     yes           yes                one sentence   slow

A was: the default model     B was: the thinking model
Decision: use the default for this task. Re-test if the picker changes.

If both answers score the same, pick the faster one, because you gain nothing by waiting. If the slower model wins on facts or on following the request, use it for this task and note what it got right. A tie is a useful result too. It tells you the request was doing the work, not the model.

Keep a model log for your repeat tasks

After a few tests, write the results in one place, so you stop deciding again every time. Each line needs the task, the model, the reason, and the date you checked. A log for an ordinary month might look like this:

TaskWhereModelWhy
Monthly budget noteChat, in a Project with the budget sheetsDefaultTied with the thinking model; faster
Tidy meeting notes into actionsChatDefaultSimple format, many per week
Plan a three-month project with dependenciesChatThinking modelDefault dropped one dependency twice with a clear request
Sort a folder of receipts into a summaryWorkWork’s own choiceThe mode mattered more than the model
Fix a bug in a small scriptCodexCodex’s defaultRe-test only if it gets stuck

Re-run a test when the picker changes or when a task starts going wrong, not every week. The log turns a confusing menu into a short list of decisions you already made, and it is easy to share with a teammate who asks which model to use.

A test the stronger model wins

Ties are common for writing tasks, so it helps to see what a real win looks like. Say you are planning an office move over three months, and you ask ChatGPT for a week-by-week plan from a list of fourteen tasks. Some tasks depend on others: the movers cannot be booked until the lease is signed, and the internet cannot be installed until the movers have a date. You paste the list, say which tasks depend on which, and ask for a plan that never schedules a task before the one it depends on.

Run the side-by-side test and check each plan against your list, one dependency at a time. In a test like this, the quicker model’s plan can look tidy and still book the movers a week before the lease is signed. The model built for harder thinking is more likely to keep all fourteen tasks in order, because the job is to hold many rules in mind at once. That is a difference you can point to on the score sheet, under “followed the request,” and it is the kind of result that justifies the slower model for that task.

Notice what made the test fair. The request was complete, the dependencies were written down, and you checked the answers against your own list instead of trusting either plan. If the quicker model had failed only because you never mentioned the lease, the fix would have been the request, and the test would have told you that too.

Your own results may differ from this example, and that is the point of testing. Models change, and so do your tasks. A plan that one model got wrong in the spring may come out right after an update, so a test result is worth a date in the log, not a permanent rule.

Sharing model choices with a team

On a team plan, the model log is worth sharing, because everyone faces the same menu and each person tends to repeat the same handful of tasks every week. Keep one short shared note, either in a team Project or in your usual documents, with the same four columns: task, where, model, and why. Ask the person who does a task most often to own its line and to re-test it when the picker changes.

Two rules keep a shared log useful. First, every line needs a reason that came from a test, not from a preference, so “tied, faster” or “default dropped a dependency” rather than “feels smarter.” Second, the log covers the model only, never what data is allowed in. Rules about customer data, contracts, and personal details belong in your company’s AI policy, because they apply to every model in the menu.

What a stronger model costs you in ChatGPT

In ChatGPT, a stronger model does not change your monthly price, but it still costs you in two ways. The first is time, since models that think longer can take much longer to answer, which adds up across a day of small tasks. The second is your plan’s limits, because each plan allows only so much use of the heavier models in a given period, and Work and Codex can draw on the same allowance. Spend it on the tasks your log says need it.

A higher plan, such as Plus or Pro, mainly buys more of that allowance and access to more models. It does not make answers true. Check the numbers, names, and sources in any answer you share, whatever model wrote it. The current limits for each plan are on ChatGPT’s pricing page, and they change often enough that it is worth checking before you upgrade.

When the model does not matter at all

Some questions should not go to any model. If you need last quarter’s exact revenue, open the finance system or the official dashboard, because ChatGPT does not have your company’s numbers unless you give them to it. If you need a legal, medical, or tax decision, ask the person who is licensed to make it. A stronger model writes a more convincing answer in these cases, which makes it more risky, not less.

Privacy works the same way. What happens to your data depends on your plan, your workspace settings, and what you paste in, not on which model you picked. A Business or Enterprise workspace has its own data rules, and those apply to every model in it.

Practice: 25 minutes

  1. Open the model picker and write down the names and the short descriptions you see today.
  2. Choose one task you do at least twice a month, and take a real example of it.
  3. Write the request using the four checks: attach the facts, say who it is for and in what format, and confirm Chat is the right place.
  4. Run the side-by-side test with the default and one model built for harder thinking, and fill in the score sheet.
  5. Start your model log with that first line.

Quick recap

  • Model names change and depend on your plan, so test on your own work instead of trusting a ranking.
  • When an answer is weak, check missing facts, a vague request, and the wrong mode before you switch models.
  • For repeat tasks, run a blind side-by-side test once and score facts, fit, and edits.
  • Keep a short model log so you decide once, and re-test only when something changes.
  • A stronger model costs time and plan limits, and it never makes an answer true by itself.

Your next step

Pick the task you did most often last week and run the side-by-side test on it today. Whatever the result, you will know whether that task needs a better model or a better request, and the answer goes straight into your log.

Series notes

This post is part of the ChatGPT product map. The next post closes the map with the API, the separate service your own software uses to call OpenAI’s models, and explains why a ChatGPT plan does not cover it. After the map, the deeper tracks are the everyday tutorial, the Work tutorial, the Codex tutorial, and the Custom GPTs tutorial, all listed on the Learn page.

Sources

Research and further reading used for this article. Model names and plan details change, so check the live pages before you buy or set a team rule.

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: