Choosing an AI model for work is a business decision with limits on speed, cost, privacy, and how much a wrong answer hurts. It is not a beauty contest. Public rankings measure something real, but they rarely measure your own workflow.
Say your team opens a spreadsheet of model names, pastes last week’s “arena” rankings into a chat thread, and asks which one should power customer support summaries, SQL drafts (SQL is the language analysts use to pull numbers out of a database), and a new internal bot. Someone has a favorite model, and someone else has a budget. Nobody has written down what “good” means when the answer turns out to be wrong late on a month-end close day.
This guide is a one-shot playbook for analysts, analytics engineers, and team leads who need a default set of models without chasing every release note. If the vocabulary still feels fuzzy, start with What are LLMs, ChatGPT, generative AI, and more. If the model will write SQL, pair this post with How to check AI-written SQL. For broader hands-on AI habits, see the Practical AI series on the learning paths at Analytics Made Simple Learn.
Leaderboards measure a different game
Public comparisons are useful the way a restaurant rating is useful. They tell you something about average taste on someone else’s menu. They do not tell you whether the kitchen can handle your rush hour, your allergies, or your budget for 10,000 meals a day.
Leaderboards have five blind spots for analytics work, and each one can flip which model is really best for you.
- The task is different. A high Elo score (a chess-style rating built from people picking favorite answers) on open-ended chat says little about a job like “rewrite this dbt test without changing what one row means.”
- Your setup is different. Your instructions to the model, the documents you feed it, and the tools you connect change results more than a half-point swap in rank.
- Speed and cost are ignored. A “smarter” model that takes 12 seconds and costs 8 times as much can lose to a mid-tier model with a tighter prompt.
- Data rules are ignored. The winner may not be allowed to see customer rows under your company’s policy.
- Models change over time. Yesterday’s champion can shift its behavior quietly tomorrow, and nothing tells you.
Use public scores to build a shortlist, and do not treat them as a purchase order. When the leaderboard and your own test set disagree, your own test set should win.
The work axes that actually matter
Write these down before anyone says a model name. If you cannot fill in the table, you are not ready to choose a model, though you are ready to choose the scope of a small pilot.
1. Task shape
Separate the jobs, even when they share the same chat window. There are five common kinds.
- Drafting SQL, Python, email, or documents, where a human is expected to edit the result.
- Classification and extraction, such as labeling support tickets, pulling fields out of text, or matching free text to a fixed list of categories.
- Summarizing meeting notes, long tickets, or incident timelines.
- Answering questions over private knowledge, for example “what does metric X mean?”, by first looking up your own documents.
- Agents with tools, which means a model that can call your warehouse, your ticket system, or a code runner on its own.
A model that is great at long prose may be mediocre at strict JSON, which is the neat labeled format software uses to pass data around. A model that is great at code may over-explain and ignore your rules about when to refuse. Because each job fails in its own way, the kind of task decides how you should test.
2. Latency and interactivity
Latency is the wait between asking and getting an answer. Assistance inside a code editor, known as an integrated development environment (IDE), needs answers in a second or a few seconds. A batch job that enriches records overnight can wait minutes. Products fail when one slow model sits behind a promise of a “snappy” screen.
Measure p50 and p95 on your real prompt sizes, and not on demo prompts of three sentences. The p50 is the typical wait, and the p95 is the wait that only the slowest one in twenty requests exceeds.
3. Cost at your volume
Price lists are quoted per million tokens, and a token is a small chunk of text, roughly three quarters of a word. Your bill is tokens times volume times retries times tool calls. A cheap model that needs three repair turns can cost more than a stronger model that gets the structure right the first time. Include human time as well, because minutes of analyst cleanup are real money even when the invoice is small.
4. Context, tools, and multimodality
Ask what you really need. Do you need a large context window (how much text the model can hold at once) to fit long database layouts? Do you need tool calling that stays reliable under your JSON format, image input for chart screenshots, streaming answers, or structured outputs? Pick the minimum capabilities first. Anything below that minimum is not a candidate, no matter its rank.
5. Policy: privacy, retention, region
Ask legal and security early, and get the answers in writing. Does the vendor train on your prompts? Is there an option for zero data retention? Where is your data processed? Which kinds of personal information (PII) are banned from prompts? Which data must stay inside your own virtual private cloud (VPC), a walled-off network, or go only to approved vendors? A model that cannot legally see the data is not “almost fine.” It is disqualified.
6. Failure cost
A wrong product description draft costs little. A wrong patient summary, or a wrong finance total in a board pack, costs a lot. When failure is expensive, you need stronger tests and a human check before anything ships. Very often the safer choice is a conservative model plus a verification step, and creativity should not be the goal.

The diagram is the artifact to bring to the meeting. Fill the cells with constraints and leave opinions out. If two models both clear the minimum, you then run a head-to-head test on your own examples.
Build a tiny test before you fall in love
You do not need a research lab for this. You need 20 to 50 frozen examples that look like the pain you actually have in production, meaning a fixed set you never edit between runs, so that every model gets the identical exam. Analysts call this an eval set. Each example needs four things.
- A real prompt, or a realistic one with sensitive details removed.
- The expected shape of the answer, such as SQL that runs, the right JSON keys, a summary of the right length, or a refusal.
- Hard cases, such as unclear definitions of what one row means, missing columns, conflicting instructions, and queries that return nothing.
- Pass criteria that a second person could grade without heroics.
Score each model in exactly the same way. Track automatic checks, such as whether the JSON parses, the query runs, and the latency stays in range. Add human checks too: is each row counted at the right level, are there no invented columns, and is the tone useful? Re-run the set whenever a model or a prompt changes, and keep it under version control like any other test file.
For SQL-heavy work, your test should ask “did it invent a column?” and “does the join match how the business counts things?”, not only “does it look like SQL?”. That is where checking AI-written SQL turns into team policy instead of personal preference.
Worked example: three jobs, one scorecard
Imagine a mid-size analytics team with three near-term uses for AI.
- Job A is code-editor help for analysts who draft SQL and small Python transforms.
- Job B is a nightly batch that sorts support tickets into a fixed set of categories and pulls out product names.
- Job C is an internal question-and-answer bot over metric definitions and runbooks. It may look up documents, but raw customer tables stay out of its context.
The team has already agreed on four constraints.
- No customer personal information goes into prompts unless it has been redacted through an approved process.
- Interactive help should feel faster than about 5 seconds at p95 for typical prompts.
- The batch job can spend more per item if accuracy goes up and human review goes down.
- The bot must cite the snippets it retrieved, and inventing policy text counts as a fail.
They shortlist three model tiers. The names are left out on purpose so the example stays about process and does not turn into a sales pitch.
- Fast-cheap is quick and low cost, but weaker at hard reasoning.
- Mid is solid at code and structure, at moderate cost.
- Heavy has the strongest reasoning, along with the highest cost and the longest waits.
After a one-week head-to-head on 40 frozen cases per job, the scorecard looks like this. The numbers are toy values for teaching, so replace them with yours.
| Job | Primary pick | Why | Backup |
|---|---|---|---|
| A: code-editor help for SQL and Python | Mid | Best pass rate on join and row-meaning cases, and the p95 wait is acceptable | Fast-cheap for tiny renames |
| B: ticket sorting | Fast-cheap | F1 score (a measure of labeling accuracy) within 2 points of Heavy at one sixth the cost, and the JSON output stays stable | Mid for the low-confidence group |
| C: metric Q&A | Mid plus document lookup | Heavy was not worth the wait, and refusing when there is no citation improved once the documents were split into better chunks | Heavy only for hard disputes across several documents |

Notice what the scorecard did not do. It did not crown a single model for the whole company. Many mature setups use a cheap model first and pass only the hard cases to a stronger one, which is good operations and not indecision.
Sample scorecard fields you can copy
job: ticket_classify_v1
candidates: [fast-cheap, mid, heavy]
n_eval: 40
metrics:
json_valid: 0.98 / 0.99 / 0.99
label_f1: 0.81 / 0.83 / 0.84
p95_latency_ms: 900 / 2200 / 6100
cost_per_1k_items_usd: 0.40 / 1.10 / 3.80
policy_ok: yes for all (redacted text only)
decision: fast-cheap default; mid if conf < 0.6
owner: analytics-platform
review_date: +90 daysStore that card next to the prompt and the document-lookup settings. When someone later asks why you are not switching to the new shiny model, you can point at the card and the test set. A screenshot of a public chart does not answer the question the same way.
Routing, not monogamy
Routing means sending each request to the model that suits it. Four patterns work well for data teams.
- Default and escalate: a cheap model drafts, and a mid or heavy model steps in only when a check fails or confidence is low.
- A task map, with different defaults for SQL, prose, and classification.
- Offline versus online: batch work can afford heavier models, while interactive screens stay lean.
- A human check on high stakes: the model proposes, and a person ships the finance or customer-facing numbers.
Routing needs a record of what happened: which model ran, which prompt version, how long it took, and whether a person overrode the answer. Without those logs you cannot improve the scorecard, because you have nothing to compare against.
Policy is part of the architecture
Treat model access the way you treat warehouse access. Ask who can send production table layouts to a model, who can turn on tools that run queries, where prompts end up in logs, how long they are kept, and which vendors are approved.
A practical minimum has four parts.
- A written list of allowed use cases and the kinds of data each may touch.
- Redaction, or made-up substitute data, for demos and training material. The site’s writing on AI and data quality also covers synthetic data.
- No secret keys pasted into notebooks. Use a managed secrets tool instead.
- A kill switch, so you can turn off one model path without redeploying the whole company.
Risk frameworks such as the Risk Management Framework (RMF) for AI from the US National Institute of Standards and Technology (NIST) do not pick a model for you. They do remind you that “it worked in a demo” is not the same as accepting the risk. Write down the leftover risk the same way you would for a flaky data pipeline that still runs in production.
Common mistakes
- Picking one model to rule them all, even though different jobs have different minimums and different costs of failure.
- Shopping from leaderboards without running your own tests, which means you are tuning for someone else’s exam.
- Ignoring p95 latency, since averages hide the one slow turn that ruins a meeting demo.
- Counting only API dollars, when human cleanup and rework often dominate the true cost.
- Skipping policy until after launch, because adding retention and PII rules later is expensive and awkward.
- Never re-testing after model upgrades. Silent regressions are real, so pin versions when you can and re-run the frozen set.
- Confusing “feels smart” with “passes the structure checks.” Pretty prose that invents columns is a fail for data work.
- Having no owner, when someone has to own the scorecard, the test set, and the review date.
Practice: run a one-week model decision
Pick one real job this week, and only one. Write the constraint table: the task, the latency target, a monthly volume estimate, the kinds of data involved, and the cost of failure. Build 20 frozen examples with pass criteria, then shortlist two or three candidates that clear your policy. Run the same set on each, fill in the scorecard, and write a three-line decision naming the primary model, the backup, and the rule for escalating. Put a 90-day review on the calendar so the decision does not go stale.
If you cannot find 20 examples, you do not understand the job yet. Shadow the people who do the work and collect their messy prompts. Those are the real test material, and textbook exercises are not.
When SQL is in scope, add at least five cases where the right behavior is to refuse or ask for clarification because columns are missing. Pair the habit with the SQL checking tutorial linked earlier, so a person still owns the verification.
Quick recap
- Leaderboards help you shortlist, while your own tests and constraints decide.
- The axes that matter are task shape, latency, cost at your volume, capabilities, policy, and cost of failure.
- Build a small frozen test before you fall in love with a model name.
- Route by job with a default, a backup, and an escalation rule. A single company-wide “best model” is usually a myth.
- Write the scorecard, give it an owner, re-run it on upgrades, and treat policy as part of the design.
- For data work, structure and honesty beat vibes every time.
Sources
Vendor pages checked in September 2026. Model capabilities and prices change, so verify the current pages before you decide.
- OpenAI, models and pricing documentation (capabilities and token pricing patterns change; verify current pages): https://platform.openai.com/docs/models
- Anthropic, Claude models overview: https://platform.claude.com/docs/en/about-claude/models/overview
- Google, Gemini models documentation: https://ai.google.dev/gemini-api/docs/models
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0): https://www.nist.gov/itl/ai-risk-management-framework
- LMSYS Chatbot Arena (public preference leaderboards; use as shortlist signal only): https://chat.lmsys.org/
- Analytics Made Simple, Practical AI series: https://analyticsmadesimple.com/series/practical-ai/
- Analytics Made Simple, How to check AI-written SQL: https://analyticsmadesimple.com/tutorials/how-to-check-ai-written-sql/
Keep going
Same lessons in your feed
Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.
