Evals ML

Evals ML: LLM evaluation methodology and benchmarks

Model answers, measured.
Evals ML
Independent · 2026
Isometric dark scene of a robotic arm picking sample blocks from a tray of candidates, lined up and fed into a glowing cyan test fixture, evoking pass@k sampling Featured
Metrics

Pass@k Metrics Explained: Formula, Estimator, Pitfalls

Pass@k is the chance at least one of k samples is correct. How the unbiased estimator works, why pass@1 is the honest number, and when pass^k wins.

Read the article →

Latest guides

Start here

Evals ML is a reference for measuring language models: what the standard benchmark suites actually contain, how much confidence a published score carries, which harness fits which kind of task, and how to keep a model grader honest. Every figure quoted here is attributed to the paper or repository it came from.

Pass@k and evaluation cost calculator

Generation count, token spend and worst-case margin of error for a given dataset size and sampling budget. Runs in the browser, no signup.

Open the calculator

All guides · Browse by topic · How these articles are researched