Evals ML
LLM Evaluation & Benchmark Engineering

LLM Benchmark Pass@k & Token Evaluator

Calculate Pass@k accuracy metrics for code generation and reasoning benchmarks (HumanEval / GSM8K), total API evaluation cost, and sample size confidence limits.

Total Generated Samples
820 Generations
Total Evaluation Cost
$4.20 USD
95% Margin of Error (worst case)
± 7.7%

Cost assumes 250 prompt tokens are re-sent with every generated sample, billed at the input rate, and charges the completion tokens at the output rate. Prompt caching, where a provider offers it, cuts the input half of that. Published rates change and the ones listed here are list prices for the named models; check the provider's current pricing page before budgeting. The margin of error is the 95% normal-approximation interval at the worst-case pass rate of 0.5, so it is an upper bound on the width of the interval around a score from a set of this size. It describes the sampling error of the eval set only, not grader error.

How to read these numbers

  • Generations scale with k, not with the dataset. HumanEval at pass@10 is 1,640 model calls, not 164.
  • Cost is dominated by output tokens on every mainstream pricing schedule, so the completion-length field moves the total far more than the dataset choice does.
  • Margin of error is what decides whether a comparison is real. On a 164-item set the 95% interval is roughly ±8 points, which is wider than most of the gaps quoted between models.