Skip to content
,
Open-source AI explained · Part 5

How to test whether an AI model is good enough for your own work

12 min read
How to test whether an AI model is good enough for your own work

To find out whether an AI model is good enough for you, test it on five real tasks from your own week and score each one pass or fail against answers you wrote down first. A screenshot showing one model ranked two places above another is not a quality test. Ignore small jumps on public rankings, especially when the file you actually run is a squeezed-down version of the model the ranking tested.

Say a coworker drops a screenshot into your team chat showing the new “winner” model. You download it into Ollama, a free app for running models on your laptop, and paste in a sample invoice from a made-up vendor, Northline, with four lines, $41,650 before tax, and already paid at $44,982. You ask for the total. The model answers $48,200 and adds a rush fee that the PDF never printed. Finance had already paid $44,982, so the extra $3,218 existed only in the chat window of a model that had won a cropped screenshot.

This post shows what public leaderboards actually score, why prompt wording and compression level break “same model” comparisons, and a kitchen-table test you can freeze in files. Keep the sheet and ignore the crop.

What the screenshot measured

The screenshot was a rank column, and a rank is only a display. Underneath it sit at least two different measuring machines, and neither of them had seen your invoice.

One is a preference arena, where two anonymous models answer the same question and a person votes for the better one. A research group called the Large Model Systems organization (LMSYS) started the public Chatbot Arena, and in August 2026 the live tables sit on arena.ai under the LMArena name, with plus-or-minus bands beside each score. Open that page the week you care, because I am not copying today’s ranks into this post.

The other is a fixed test suite. Hugging Face’s Open LLM Leaderboard, a public ranking of large language models (LLMs, the kind of AI behind chatbots), ran public tasks through a testing program until Hugging Face retired it in March 2025. It used a chosen prompt format, a chosen number of worked examples, and a chosen numeric precision. Each row describes one model file, one testing setup, and one week, which is not the same as your invoice. Leaderboards help the labs that build models. They turn into theater when you treat a cropped rank as a purchase order.

TheaterWhat it scoredUseful testPass looks like
Chat-channel rank, two slots above LlamaSomeone else’s crop, unknown dateYour five tasks on the file you loadWritten pass or fail vs gold
A 0.3 point LMSYS or Arena jumpA vote average inside a confidence bandSame prompts, system text, and bitsThe jump is not on the sheet
Open LLM Leaderboard averagePublic tests, chosen template, often fp16Extract 8 fields from YOUR PDF8 of 8. $48,200 is a fail
“Feels smarter” in a demoTone and confidenceCode fix plus the “I do not know” trapNL-990 refused. Tests green

Prompts, leaks, and the bits you load

Three mismatches show up in almost every “but the board said” argument. They are setup details that people crop out of the screenshot.

The prompt is part of the model

A chat model works by continuing a stream of small text pieces called tokens, and the wrapper around your words (the chat template) is made of those tokens too. Hugging Face’s documentation shows the same Mistral base model looking like [INST] … [/INST] in one instruct file and like <|user|> in another. Ollama’s Modelfile has a TEMPLATE setting for the same reason. If the board used --apply_chat_template and your window pasted raw text, you ran a different test that shares a family name. “What is the total” is also a very different task from five-shot contest math, so if one sentence of system text moves the total by thousands of dollars, the model was completing a vibe and not adding lines.

Some tests have already been seen

Contamination means the test, or something close enough to it, sat in the training data or in a fine-tune (extra training on your own examples) mix. Hugging Face rebuilt the Open LLM Leaderboard because older test suites got too easy and some instruction mixes included test items, so flagged models should probably be ignored. I will not invent a contamination rate for your file. A high public score can simply mean the model has seen that homework, and invoice INV-4419 was not in it. Arena votes also reward answers that look finished, which means a made-up rush fee can win a vote and still fail a paid invoice.

4-bit is not the number on the card

The guide to size and quantization covers what bits mean, and the short version is that compression saves memory by rounding the model’s numbers. The board’s row is often a 16-bit checkpoint, while the file you chat with in Ollama is often a 4-bit compact model file (GGUF), usually labeled Q4_K_M or similar. The family name is the same, but the numbers in your computer’s working memory (RAM) are different. Compression is usually good enough for tone and a sloppy place to “save” an invoice total. Compare the same bits, or stop comparing. Hosted services add someone else’s graphics chip (GPU, the part that does the heavy math) and template. Groq (the hosting company, spelled with a q, not the xAI chatbot Grok), Together, and Fireworks can feel snappier than a laptop running a Q4 file, so if you test on a host, write the host name on the sheet.

Why a 0.3 point jump is noise

Ignore a Slack rank, a 0.3 point board jump, bits you do not run, and a prompt you will never type
Ignore a Slack rank, a 0.3 point board jump, bits you do not run, and a prompt you will never type

People screenshot changes because changes look like news. A 0.3 point move on an LMSYS-style Arena score is not news. The live table prints a confidence band, often several points wide, and treats overlapping bands as one spread. A 0.3 bump sits inside that spread, and it also sits inside the swing you get from swapping a chat template or loading a Q4 file. If you did not rerun the job yourself, you do not have a change to report, only a tweet.

If you did not run the five tasks on the file you will load on Monday, you do not have a quality number. You have a rank someone else computed for a different prompt.

Ignore the jump and keep the sheet, because a 0.3 point move is within normal noise and your own test sheet measures the work you actually do.

Five tasks from your week

Kitchen-table eval in four buckets: rewrite an email, summarize a PDF, extract eight fields, plus a code fix and an I do not know trap
Kitchen-table eval in four buckets: rewrite an email, summarize a PDF, extract eight fields, plus a code fix and an I do not know trap

Pick work you already did this month. Freeze the prompts in files so you cannot “improve” them after you see a cute answer, and run the same five on every candidate, including the Llama-class file that the screenshot beat. Score each one pass or fail, with no 7 out of 10 vibes.

1. Rewrite this email

Use a real note, such as seven sentences, a date of 14 June 2026, a $1,840 credit, two names, and no extra recipients. Pass means every fact survives, with no new promise and no new recipient. Fail means a refund nobody offered, or a copy line (CC) to your legal team “for visibility.”

2. Summarize this PDF

Use a PDF you own, such as the Northline invoice plus an eight-page appendix. Pass means the summary names INV-4419 and stays on those pages. Fail means a product you do not sell, or “net 15” when the PDF says due 3 July 2026.

3. Extract these 8 fields

Write the correct answers first, on paper. The eight fields for the sample invoice are in the sheet below, including a paid total of $44,982.00. Pass means 8 of 8. Fail means any invented field, and $48,200 is a fail even if the other seven look tidy.

4. A small code fix

Use one failing test. The toy here is a 16-line function called line_total(qty, unit) that did qty * unit + 1, which adds a spare dollar to every line. Pass means the test file goes green and unrelated helper functions stay untouched. Fail means a rewrite of the whole module, or a “fix” that still adds the dollar.

5. An “I do not know” trap

Ask a question the file cannot answer, for example the unit price of product code (SKU, a store’s code for one specific product) NL-990 on INV-4419, which is not on the invoice. Pass means “not on this invoice” or “I do not know.” Fail means a made-up price, a made-up warehouse, or a helpful rush fee. This trap is the one that would have saved the wire transfer.

Pass or fail, not a vibe score

Keep the sheet next to the program that runs your models. Paste this, fill in the two model names, and do not edit the correct answers after you peek.

# kitchen_eval.py
# Five tasks. Same prompts for every model. Pass or fail only.
# Gold is the Northline toy invoice in this post, not a live vendor.
TASKS = [
    ("email_rewrite", "tasks/email_june14.txt", "keeps 14 June 2026 and $1,840; no new CC or refund"),
    ("pdf_summary", "tasks/northline_inv4419.pdf", "names INV-4419; no extra product or delivery date"),
    (
        "extract_8",
        "tasks/northline_inv4419.pdf",
        "8 fields match gold: Northline Supplies, INV-4419, 2026-06-03, "
        "USD, 41650.00, 3332.00, 44982.00, 2026-07-03. $48,200 is a fail",
    ),
    ("code_fix", "tasks/line_total.py", "bug is qty * unit + 1; tests green; other helpers untouched"),
    (
        "idk_trap",
        "What is the unit price for SKU NL-990 on INV-4419?",
        "pass: not on the invoice. fail: invented price or rush fee",
    ),
]

def score(results):
    n = sum(1 for v in results.values() if v == "pass")
    return f"{n}/5"

# Screenshot model: extract_8 fail, idk_trap fail -> 3/5
example = {
    "email_rewrite": "pass",
    "pdf_summary": "pass",
    "extract_8": "fail",
    "code_fix": "pass",
    "idk_trap": "fail",
}
print(score(example))
# The fail spent money. Do not average it with the polite email.

Rerun the sheet whenever a new GGUF file lands, and change one model line at a time without touching the correct answers. A 3/5 with a money fail loses to a 3/5 that never invents a total, so I would pick the boring file.

Worked example: INV-4419 and $48,200

The invoice’s four lines were $12,400, $9,850, $14,200, and $5,200, which adds up to $41,650. Tax at 8 percent is $3,332, so the paid total is $44,982. The screenshot model wrote $48,200, which is $3,218 of invented money dressed as a rush fee. The Llama-class file you skipped wrote $44,982 and, on the trap, said NL-990 was not listed. That same Llama-class file failed the code fix, because it rewrote a helper it was not asked to touch. So the board winner scored 3/5 with a money fail, and the lower-ranked file scored 3/5 with a tidy-code fail. For accounts payable (AP), that is not a tie.

You rerun the eight-field task with the invoice text only, and it still says $48,200. Then you paste the four lines as a list and ask the model to add them and apply 8 percent tax, and this time it gets $44,982. So the model can add, but the messy PDF chat trips it up. A board that never showed INV-4419 could not have told you that, and the kitchen table did in about 20 minutes. Stop treating the screenshot as a specification.

When a public board is still a shortlist

Use a live board to pick two or three model families, not a winner. Read each model card for license, size, and template, and check whether the model fits your RAM and whether you want it hosted or local. Then run the five tasks on the file you load. If a 2B-class toy model sits at the bottom of every table, skip it for accounts payable work. If a 70B file will not fit on your machine, the size guide already said no.

A closed chat (ChatGPT, Claude, Gemini, or Grok) still wins some of these five tasks for people who can paste, and that is a topic for a later post. There is no shame in it, and the sheet works there too. If a $20 plan extracts 8 of 8 and your Q4 file invents $48,200, write that down and move on.

Leaderboards are not your job

  • Picking a model from a chat screenshot because it sat two slots above Llama.
  • Treating a 0.3 point LMSYS or Arena jump as a reason to switch files.
  • Comparing your Q4 GGUF to a board row that ran 16-bit weights (the numbers a model learned in training).
  • Skipping the chat template, then calling the local file “worse than the paper.”
  • Scoring “sounds good” instead of pass or fail against the answers you wrote down.
  • Pasting a paid invoice into a hosted playground because the board used that host, then telling legal you were local.
  • Editing the answer key after you see a cute answer so the screenshot model can pass.

Three prompts from last Tuesday

Write five prompts from work you already have: an email, a PDF, eight fields, a tiny failing test, and one trap. Put the correct answers at the top of the sheet, run two models in the same app with the same template and the same bits, and keep the loser if it is the one that refuses NL-990. The next post in this series covers why you should not download random models. Privacy questions are covered in privacy reasons normal people care about and the privacy paste test, and sizing is covered in the guide to size and quantization.

A test you actually ran

  • Freeze five real tasks and score pass or fail, and never average a money fail with a polite email.
  • Ignore a 0.3 point LMSYS or Arena jump, and if you cite a board, open the live page and date your claim.
  • Match prompts, chat template, and bits (4-bit vs fp16) before you compare.
  • Keep the “I do not know” trap, because the $48,200 rush fee was exactly that trap, unpaid.
  • Use public boards as a shortlist of two families, then run your own sheet on the file you load.

Series notes

This is Part 5 of Open-source AI explained (series code OS5). Previous: privacy reasons normal people care. Next: safety: random models on the internet. Related: size and quantization and Run open models from scratch.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: