Evals ML
Isometric dark scene of a robotic arm picking sample blocks from a tray of candidates, lined up and fed into a glowing cyan test fixture, evoking pass@k sampling
Metrics

Pass@k Metrics Explained: Formula, Estimator, Pitfalls

Pass@k is the chance at least one of k samples is correct. How the unbiased estimator works, why pass@1 is the honest number, and when pass^k wins.

By Evals ML Editorial · · 4 min read

Pass@k is the probability that at least one of k sampled completions for a problem passes its unit tests, averaged over problems. That is pass@k metrics explained in one line. Estimate it with the unbiased 1 - C(n-c, k) / C(n, k) from n ≥ k samples, and report pass@1 unless something checks candidates before users see them.

The failure this prevents is a quiet one. A candidate model wins on pass@10 in the eval dashboard, gets promoted, and the product serves one completion at low temperature. Nothing regressed; the team measured a sampling budget production never spends.

What does pass@k actually measure?

Pass@k measures the fraction of problems where at least one of k independent samples is functionally correct, judged by executing it against unit tests rather than comparing its text to a reference. The Codex paper credits the metric to Kulal et al. (2019) and made it the standard for HumanEval.

Execution is the point: the Codex authors show that BLEU scores for correct and incorrect solutions are not cleanly separable.

k is a sampling budget, not a capability tier. OpenAI’s own table for Codex-12B on HumanEval, a vendor-reported figure, reads 28.81% at pass@1, 46.81% at pass@10 and 72.31% at pass@100: one model, three budgets. The AlphaCode paper calls pass@k “an upper bound metric for using k samples”, because it assumes every sample can be checked against the hidden tests for free. Its n@k variant counts a problem as solved only if one of n submissions picked from k samples passes, which is closer to a contest with a submission limit. What HumanEval itself contains is covered in LLM benchmarks explained: MMLU, HumanEval, GSM8K.

The metric that matters: the unbiased estimator

Generate n ≥ k samples per problem, count the c that pass, and average this over problems:

pass@k = mean over problems of [ 1 - C(n - c, k) / C(n, k) ]

C(n−c, k)/C(n, k) is the probability that a random k-subset of your n samples contains only failures. The Codex paper used n = 200 and reported k up to 100.

It beats both obvious alternatives. Drawing exactly k samples and checking whether any passed is unbiased but noisy, because each problem contributes a single 0 or 1. Plugging the per-sample pass rate into 1 − (1 − p̂)^k is biased: the paper says it “results in a consistent underestimate”, which follows from that function being concave in p (Jensen’s inequality).

A worked example, one problem with n = 20 and c = 3:

  • pass@1 = 3/20 = 0.150, since pass@1 always reduces to c/n
  • pass@5 = 1 − C(17, 5)/C(20, 5) = 1 − 6188/15504 ≈ 0.601
  • naive 1 − 0.85^5 ≈ 0.556, about 4.5 points low

Binomials at n = 200 overflow floats, so use the paper’s product form, shown below.

Wiring it up

The script samples with vLLM, scores with the unbiased estimator, and logs to MLflow. SamplingParams(n=...) returns n outputs per prompt request, so n = 20 costs one request per problem, not twenty.

import numpy as np
import mlflow
from vllm import LLM, SamplingParams

N_SAMPLES = 20        # n: must be >= the largest k reported
KS = (1, 5, 10)
TEMPERATURE = 0.8     # pin it; pass@k across temperatures is not comparable
assert N_SAMPLES >= max(KS)


def pass_at_k(n: int, c: int, k: int) -> float:
    """Unbiased pass@k, numerically stable form (Chen et al., 2021)."""
    if n - c < k:
        return 1.0
    return float(1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1)))


def pass_hat_k(n: int, c: int, k: int) -> float:
    """pass^k: probability that all k trials succeed (Yao et al., 2024)."""
    if c < k:
        return 0.0
    return float(np.prod(np.arange(c - k + 1, c + 1) / np.arange(n - k + 1, n + 1)))


def passes_tests(problem: dict, completion: str) -> bool:
    """Run problem["test"] in an isolated sandbox: no network, no secrets."""
    ...


def evaluate(problems: list[dict], model: str) -> tuple[dict[str, float], list[int]]:
    llm = LLM(model=model)
    params = SamplingParams(n=N_SAMPLES, temperature=TEMPERATURE,
                            top_p=0.95, max_tokens=512, seed=1234)
    outputs = llm.generate([p["prompt"] for p in problems], params)
    counts = [sum(passes_tests(p, o.text) for o in out.outputs)
              for p, out in zip(problems, outputs)]
    metrics = {}
    for k in KS:
        metrics[f"pass_at_{k}"] = float(np.mean([pass_at_k(N_SAMPLES, c, k) for c in counts]))
        metrics[f"pass_hat_{k}"] = float(np.mean([pass_hat_k(N_SAMPLES, c, k) for c in counts]))
    return metrics, counts


problems = [...]  # dicts with "prompt" and "test" keys
with mlflow.start_run(run_name="codegen-candidate"):
    mlflow.log_params({"n": N_SAMPLES, "temperature": TEMPERATURE, "top_p": 0.95})
    metrics, counts = evaluate(problems, "your-org/your-model")
    mlflow.log_metrics(metrics)
    mlflow.log_dict({"per_problem_c": counts}, "per_problem_counts.json")

The assert is not decoration. The reference estimator returns 1.0 whenever n − c < k, so asking for a k larger than n silently reports a perfect score on every problem, including ones where nothing passed.

Per-problem counts go to an MLflow artifact, not a metric. You need them later for paired comparisons, and emitting them as Prometheus series labeled by task ID is a cardinality problem with no payoff.

If you would rather not own the loop, OpenAI’s human-eval harness runs evaluate_functional_correctness samples.jsonl --k=1,10 (default k is 1, 10 and 100) and records a passed, failed or timed-out result per sample. Its execution call ships disabled, because it runs untrusted model-generated code. Hugging Face’s code_eval metric requires HF_ALLOW_CODE_EVAL="1" for the same reason and defaults to a 3-second timeout per candidate. Gating a deploy on either is covered in LLM eval CI: reproducible gates before deployment.

Which k should you report?

Report the k your system actually spends. If users see one completion, pass@1 is the honest number. If an agent executes candidates against tests before surfacing one, pass@k at that candidate count describes the system. If you choose among candidates without tests, measure your selector instead.

That last gap is large. Codex-S solves 37.7% of HumanEval with one sample and produces at least one correct function within 100 samples on 77.5% of problems, but picking the sample with the highest mean log-probability passes on 44.5%. The oracle number is 77.5%; the ranker number is what ships.

Temperature has to be pinned per k. On a 679M-parameter model the Codex paper found the best temperature was 0.2 for pass@1 and 0.8 for pass@100. Model A’s pass@10 at 0.8 against model B’s at 0.2 compares decoding configs, not models. Log temperature, top_p, n and max_tokens on every run and refuse comparisons where they differ; the guide to LLM evaluation design covers building that into a regression suite.

When is pass^k the better metric?

Use pass^k when every interaction has to succeed, which describes most customer-facing agents. τ-bench defines it as the mean over tasks of C(c, k)/C(n, k), the probability that all k i.i.d. trials succeed. pass@k rewards a model that sometimes gets it right; pass^k punishes one that sometimes gets it wrong.

τ-bench comes from Sierra’s research team, so treat it as vendor research rather than an independent benchmark. It reports that GPT-4o succeeds on fewer than 50% of tasks and that pass^8 drops below 25% in the retail domain. Same trials, opposite question. The script above logs both from one set of counts.

What you’ll see

Plot pass@k against k on a log x-axis, one line per candidate model.

Good: a curve that climbs fast and flattens, with pass@1, the high-k tail and pass^k all moving the same direction between releases.

Crossing curves: Yue et al. (2025), an independent academic study, found RLVR-trained models beat their base models at small k, while “the base models achieve a higher pass@k score when k is large.” A fine-tune can lift pass@1 by narrowing the output distribution. That is a win for single-shot serving and a loss for sample-and-verify pipelines, and a dashboard that tracks one k hides it.

A wide pass@k to pass^k gap: high pass@10 with low pass^3 means the model can solve the task and does so inconsistently. Behind a test filter, that is fine. For an agent issuing refunds, that is the incident.

Jitter bigger than the win: on HumanEval’s 164 problems, each problem is worth about 0.61 points of pass@1. A three-point gain is five problems.

Caveats

Cost scales with n, not problem count. HumanEval at the Codex paper’s n = 200 is 32,800 completions per model per decoding config. The pass@k and eval cost calculator sizes generations, tokens and margin of error before you book GPU hours.

Weak tests inflate every k. An under-specified test suite passes wrong code, and more samples mean more chances to hit the loophole. Model graders have the same problem: pass@k needs only one false pass among k, so any judge false-positive rate compounds as k grows. Calibrate the grader first; see LLM-as-a-judge bias: how to detect and correct it.

Error bars. Miller’s Adding Error Bars to Evals recommends CLT standard errors, clustered errors when questions come in related groups, resampling answers to cut within-question variance, and question-level paired differences when comparing two models. Sampling n > 1 is the resampling; the per-problem counts artifact gives you the pairing.

Contamination. Public problems end up in pretraining corpora, and a private held-out set is the only durable fix.

Offline is not online. pass@k gates a release; it does not monitor one. After deploy, production success rates and input drift take over, which is the territory SentryML’s model monitoring coverage picks up.

FAQ

is pass@1 the same as accuracy?

Yes, pass@1 is per-problem accuracy averaged across the dataset. With n samples per problem, the unbiased estimator reduces to c/n, the fraction of samples that pass, averaged over problems. With greedy decoding and one sample it is plain accuracy. Sampling n > 1 at a fixed temperature gives a lower-variance estimate of the same quantity.

how many samples do you need to compute pass@k?

You need at least k samples per problem, and in practice several times more. The Codex paper generated 200 per problem to report k up to 100. A larger n lowers the variance of the estimate. With n equal to k, you are back to the noisy “did any of these k pass” check the estimator replaced.

can pass@k be used for tasks other than code?

Yes, pass@k works for any task with an automatic pass or fail check on each sample. Math with exact-match final answers, SQL checked against result sets, and agent tasks graded on final database state all fit. The requirement is a verifier you trust; with a fuzzy or model-based grader, false passes accumulate as k grows.

why is my pass@10 high but users still get wrong answers?

Your product most likely serves one completion, so users experience pass@1 rather than pass@10, which assumes something checks all ten candidates and surfaces a correct one. Without an execution-based filter, the number that matters is how often your selection method, greedy decoding or log-probability ranking, picks a sample that actually passes.

Sources

  1. Chen et al., Evaluating Large Language Models Trained on Code (Codex, HumanEval, pass@k estimator)
  2. OpenAI human-eval: HumanEval evaluation harness
  3. Li et al., Competition-Level Code Generation with AlphaCode (n@k)
  4. Yao et al., tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (pass^k)
  5. Yue et al., Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
  6. Miller, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
  7. Hugging Face Evaluate: code_eval metric
  8. vLLM documentation: SamplingParams
  9. MLflow Python API reference
#pass-at-k #llm-evaluation #code-generation #benchmarks #eval-metrics

Related