Reading benchmarks safely
Model A scores 92% on the leaderboard, Model B 89%. So A is the better choice for you, right?
The idea inside
A single benchmark score hides traps like contamination: the leaderboard is not your task.
After this lesson
You can interpret public benchmark scores, contamination, saturation, jagged capability, domain transfer, and why one number can't answer “is this good for my task?”.
Where it leads
Benchmarks probe what models can do, but what can they never do?
Inside this lesson
That's the real lesson stage, paused. Claim your pass to operate it.
See how AI actually works, end to end.
This lesson is one stop on the full arc. Unlock all of it, and keep it for life.
What you get
- The 34-lesson main path, a finishable route from a word to agents
- Goal tracks for using AI at work and building AI features
- Boss labs that make you apply a whole act, not just recognize it
- Spaced recall that brings each idea back before you forget
- Course memory: every term defined, with links to where it first appears
- A shareable capability card when you finish the main path
- Lifetime access on every device, every future lesson included
Not videos to watch. You predict, operate the machine, then prove it. That is why it stays.
99 interactive lessons and challenges. No videos, no code.
Free launch pass: lifetime access, no card needed
New here? The first lessons are free to try. Start with lesson 0.1
What this lesson shows
A single benchmark score hides traps like contamination: the leaderboard is not your task.
The question it opens with
Model A scores 92% on the leaderboard, Model B 89%. So A is the better choice for you, right?
The walkthrough, in the lesson's own words
- Drag the overlap slider up: leak more of the test into Model A's training, and watch A's score climb past B.
- One number, four hidden traps. Click each chip to inspect the score.
- So before you pick a model from a leaderboard, ask these.
- Same model, same skill, only the leak changed. Slide overlap up and A's number climbs past B on memorized questions; slide it down toward fresh, unseen ones and the lead collapses, B is ahead. A single number can be inflated, noisy, or measuring a skill that isn't yours. Next, take the score apart and watch “A wins” wobble four more ways.
- Pick a chip above. Each one is a real, documented way a single score misleads.
- All percentages shown here are illustrative, not real model scores, but the four traps are real, documented phenomena.
- A leaderboard ranks models on its questions, not yours. Treat a public score as a shortlist filter, never the final answer.
- The only benchmark that measures YOUR task is the eval set you build for it (lesson 7.3). Shortlist with public scores; decide with your own evals on your own data.
- A vendor pitches their model because it tops a popular coding leaderboard by a few points. Before you switch, what should you check?
- Treat that number as a shortlist filter, not a verdict. A few points near the top can be noise (saturation), the questions may have leaked into training (contamination), and overall rank is jagged: the leader can still trail on the specific skill and data you need. Decide it with your own eval set on your own inputs.
- Model A leads, but only on memorized, leaked questions.
- On fresh questions the lead flips: Model B is ahead.
- All within a few points of the ceiling, re-run it and the order can flip.
- A is higher on average, yet B is the one you'd actually want.
Key takeaway
You can read a benchmark number skeptically and ask the follow-ups that actually predict fit.
What you can do after this lesson
You can interpret public benchmark scores, contamination, saturation, jagged capability, domain transfer, and why one number can't answer “is this good for my task?”.
Where it leads: Benchmarks probe what models can do, but what can they never do?
This is the written summary. The lesson itself is interactive: you predict, drag and operate the mechanism above, and the reveal answers you.