Benchmarks
-
Pass@k Metrics Explained: Formula, Estimator, Pitfalls
Pass@k is the chance at least one of k samples is correct. How the unbiased estimator works, why pass@1 is the honest number, and when pass^k wins.
-
GUI-Primitives: Spatial Reasoning Failures in Vision Models
GUI-Primitives tests spatial reasoning in vision-language models for GUI grounding, finding weak localization and 32% peak strict point-in-box accuracy.
-
LLM-as-a-Judge Bias: How to Detect and Correct It
Position, verbosity and self-preference bias in model graders, how to measure judge agreement against humans, and the calibration steps that reduce each.
-
LLM Benchmarks Explained: MMLU, HumanEval, GSM8K
What MMLU, HumanEval and GSM8K actually contain, how each score is computed, and the point where a headline benchmark number stops being useful.
-
LLM Evaluation Design: From Task to Regression Suite
How to design an LLM evaluation: task-grounded eval sets, grader choice, sampling variance, and the contamination traps that mislead benchmark scores.