pass@k
-
Pass@k Metrics Explained: Formula, Estimator, Pitfalls
Pass@k is the chance at least one of k samples is correct. How the unbiased estimator works, why pass@1 is the honest number, and when pass^k wins.
-
LLM Benchmarks Explained: MMLU, HumanEval, GSM8K
What MMLU, HumanEval and GSM8K actually contain, how each score is computed, and the point where a headline benchmark number stops being useful.
-
LLM Evaluation Design: From Task to Regression Suite
How to design an LLM evaluation: task-grounded eval sets, grader choice, sampling variance, and the contamination traps that mislead benchmark scores.