Start with the task

A leaderboard can be a useful map, but it is not the destination. Before comparing two scores, identify the task: answering factual questions, editing a repository, extracting a table, or operating a computer. These are different abilities. A strong result on one can justify an experiment without establishing that a model will handle every neighboring task.

Read the conditions

Look for the model version, the tools it could use, the reasoning budget, and how many attempts were allowed. An agent with a browser and several retries is not being tested under the same conditions as a model answering once from memory. If these details are missing, the score is harder to interpret. Save them with the result so a later update can be compared on the same basis.

Use more than accuracy

The HELM research program is a useful background example. Its authors evaluate models across scenarios and several dimensions, including accuracy, robustness, and efficiency, explicitly documenting gaps in coverage. Our practical interpretation is to keep a small scorecard for your own work: correctness, time, expense, and the human effort required to repair the result. A system can improve on one dimension while becoming less useful on another.

Build a small reality check

Collect representative tasks before trying the new model, including ordinary examples and a few known failure cases. Define what counts as success before looking at its answer. Repeat uncertain cases and inspect the failures rather than hiding them in an average. A modest, stable set of real tasks will not rank the whole industry. It can answer a narrower and more valuable question: whether this change improves the work you actually do.

Sources & authors

  1. Holistic Evaluation of Language Models
    Stanford CRFM / arXiv · November 16, 2022