About Evals ML
LLM Evaluation Frameworks, Pass@k Metrics, Hallucination Scoring & Benchmark Datasets
What this site covers
Evals ML is a reference for evaluating language models. It covers what the standard benchmarks measure, how much confidence a reported score actually carries, and how to build a grading harness that answers a question you care about.
- What the standard benchmark suites contain: MMLU, HumanEval, GSM8K, and the newer sets built to replace them
- Pass@k, sampling variance and confidence intervals on a reported score
- Evaluation harnesses and frameworks, and which task shapes each one fits
- Model graders: rubric design, judge bias, and calibration against human labels
- Hallucination and groundedness scoring for retrieval-augmented systems
- Token cost and sampling budget for a planned evaluation run
6 articles are published so far. New ones are announced on the RSS feed.
Where to start
- LLM benchmarks explained — what MMLU, HumanEval and GSM8K actually measure
- LLM evaluation design — building an eval set that answers your question
- Eval frameworks compared — choosing a harness for the task you have
- LLM-as-a-judge bias — detecting and correcting grader error
The site also publishes a Pass@k and evaluation cost calculator, which works out the generation count, token spend and worst-case margin of error for a planned evaluation run. It is free, needs no signup, and runs entirely in the browser.
How these articles are produced
Articles here are researched from primary sources: vendor and project documentation, published standards and specifications, research papers and preprints, and measurements published by whoever took them. Drafts are produced with AI assistance and then edited against those cited sources before anything is published.
No article on this site is based on first-hand testing in a private lab, and nothing here should be read as a measurement report of its own. Where a number appears, it comes from a source that is named, so you can check the original instead of taking this site's word for it.
Everything is published under a single editorial byline. That byline is a publishing identity for the site, not a claim about a named individual, and it does not carry professional credentials.
Corrections
Corrections are welcome. If something here is wrong, out of date, or attributed to the wrong source, email editor@evalsml.com with the page address and what it should say. Substantive corrections are made in the article itself rather than quietly dropped.
How this site is funded
This site currently runs no affiliate links, sponsored posts, display advertising or paid placements. If that changes, the disclosure page will say so.
Contact
Email: editor@evalsml.com
Site: evalsml.com
Published by: Evals ML Editorial
See also the privacy policy, the terms of use, and the editorial disclosure.