BenchmarksLiveAInews
The briefing
Practical guides · ORIGINAL GUIDE

Choose an AI model by testing your actual work

Build a small, repeatable comparison that measures useful output, review effort, and the errors you cannot accept.

By 4 min read

Define one decision and one finished result

Begin with a specific choice: which candidate should help with this recurring task? Write the deliverable in ordinary language. For invoice extraction, it might be a table containing supplier, invoice date, currency, and total, with missing information explicitly marked. For meeting notes, it might be decisions and assigned actions supported by the transcript. Avoid testing several unrelated jobs under one vague quality score.

List requirements before seeing the outputs. Separate preferences, such as concise wording, from conditions that every acceptable result must satisfy. A fluent response that changes a total should not recover its score through good formatting. Also define what a human reviewer must still check.

Make a small test set you can understand

Choose permitted, suitably redacted examples from the work you expect to do. Include typical cases, awkward formatting, missing information, and cases where the correct answer is to flag uncertainty. Record the expected answer or an explicit grading rubric for every item. Keep a few examples for improving the prompt and a separate set that you do not inspect while tuning.

Anthropic’s engineering guidance distinguishes a task, a trial, and a grader, and emphasizes checking outcomes. NIST’s ARIA program separates model testing, red-teaming, and field testing. Together, these sources explain why a small offline comparison is one layer of evidence rather than a complete deployment assessment. The following worksheet is a practical starting method, not a validated benchmark.

Run the comparison without moving the target

Give each candidate the same task instructions and source material. Record the model version, settings, available tools, date, and any retries. If one product requires different instructions, keep both the shared-prompt result and the adapted result; otherwise an improvement in your instructions may look like an improvement in the model.

Use a simple row for each case: case ID, pass or fail, error type, review time, and observed response time. Where outputs vary, repeat selected difficult cases and preserve every attempt. Blind the model names during subjective review when practical. If you change a grading rule after seeing a surprising answer, regrade all candidates under the revised rule.

Choose the smallest justified rollout

Compare useful completed work, including correction effort. A fast draft that takes longer to repair may save less time than a slower accurate answer. Keep observed usage costs separate from hypothetical future volume; measure both before making a cost claim. Check account, data-handling, and integration requirements as separate suitability constraints.

Treat a narrow win as permission for a narrow next experiment. Try a supervised pilot, collect failures, and add genuinely new cases before expanding. Save the test set and rerun it when a model, prompt, or workflow changes. A small sample cannot establish rare-failure rates, and repeated reuse can make the test less representative of unseen work.

WORKED EXAMPLE

Hypothetical example: the higher pass rate loses

Suppose you have 30 redacted invoices. Use 10 to clarify the instructions and reserve 20 for comparison. Candidate A gets 18 completely correct; its two failures have correct values but invalid column formatting. Candidate B gets 19 correct; its remaining answer changes the invoice total. These are invented results, not tests of real products.

Your prewritten rule says an incorrect monetary value blocks automatic use. B’s 95% overall result therefore cannot justify automation. A’s 90% result also needs repair. You could evaluate A in a supervised extraction pilot, with every value checked, while investigating both failure types. Neither result supports unattended processing or a claim of superior general intelligence.

TAKE THIS WITH YOU

Your decision checklist

  • Is the task narrow enough to define a correct finished result?
  • Did I write critical failure rules before reviewing outputs?
  • Are prompt-development cases separate from comparison cases?
  • Have I counted retries, correction time, and all failed attempts?
  • Does the next rollout stay within what the evidence supports?

Primary sources & further reading

These sources support the research context and methods discussed in this guide. Worked examples are hypothetical illustrations by LiveAInews.

  1. Anthropic — Demystifying evals for AI agents
  2. NIST — ARIA: Assessing Risks and Impacts of AI

An original guide by LiveAInews. Found an error or a source that has changed? Report a correction.

Turn the reading into a test.

Build your own evaluation worksheet in the Model Decision Lab, or explore published benchmark results with their original sources.

Open Model Decision Lab Explore benchmarks