Pick an LLM evaluation framework by job, not popularity. lm-evaluation-harness (or HELM, now in maintenance mode) reproduces published benchmarks for model selection; promptfoo or DeepEval gates prompt changes in CI; Ragas scores RAG retrieval and generation separately; Inspect fits multi-turn agents; MLflow, Phoenix, LangSmith, Braintrust and Weave tie scores to traces.
The failure that usually starts the search is not a missing eval. It is a prompt edit that moved the score a couple of points, a deploy that went ahead on a green dashboard, and tickets two days later showing the model answering in the wrong format. Nobody can say whether the change was real, the judge behind an API alias moved, or the run replayed cached responses. So the comparison below runs on two axes: which task shape each framework fits, and whether its runs are reproducible enough to gate a deploy.
The comparison
The first axis is what each tool was built to evaluate. Benchmark harnesses score models against items with a known answer, application frameworks test a system you built, and tracing platforms attach scores to the experiment and production traces that produced them.
| Framework | Maintained by | Built for | Grading style | Best fit |
|---|---|---|---|---|
| lm-evaluation-harness | EleutherAI | Reproducing academic benchmarks across many models | Log-likelihood and exact match, defined per task | Comparing base models on MMLU, GSM8K and similar sets |
| HELM | Stanford CRFM, in maintenance mode | Multi-metric benchmarking with a public leaderboard | Accuracy plus calibration, robustness, fairness, bias, toxicity, efficiency | Reports where one accuracy number is not an acceptable summary |
| OpenAI Evals | OpenAI | A registry of small evals expressed as data plus a template | Match, includes, fuzzy and JSON match; model-graded templates | Teams on the OpenAI API wanting data-only custom evals |
| Inspect | UK AI Security Institute | Datasets, solvers and scorers in Python | Text comparison, model grading or custom scorers | Agentic, tool-using and safety evaluations |
| promptfoo | promptfoo, now part of OpenAI | Prompt and model regression testing | Declarative YAML assertions plus model-graded checks | Catching prompt regressions on every pull request |
| DeepEval | Confident AI | Application tests in a pytest idiom | Metric classes: G-Eval, DAG, RAG, agent, hallucination | Python teams who want evals to read like unit tests |
| Ragas | Vibrant Labs | RAG evaluation, plus agent, SQL and general-purpose metrics | Faithfulness, response relevancy, context precision and recall | Diagnosing retrieval and generation separately |
| MLflow 3, Phoenix, LangSmith, Braintrust, Weave | MLflow project, Arize, LangChain, Braintrust, Weights & Biases | Tracing plus evaluation | Built-in and custom LLM judges, online scoring | Linking a score to the trace that produced it |
The benchmark harnesses answer “which model should we build on”. The application frameworks answer “did today’s change break anything”. The platforms answer “which request produced this bad score”. Most teams need the second question answered on every commit and the first only when a model changes.
Gate-ability: reproducibility, CI fit, judges and traces
A run can gate a deploy only if it can be repeated, triggered per change, scored by a judge you control and traced back to the outputs behind the score. The same tools compared on that axis, leaving out HELM and OpenAI Evals, which are batch tools rather than per-commit gates:
| Framework | Licence | Reproducibility controls | CI fit | Judge control | Traces |
|---|---|---|---|---|---|
| lm-evaluation-harness | MIT | Versioned tasks, fixed seeds, sample logs | Batch, not per commit | Exact match, log-likelihood | None |
| Inspect | MIT | Python tasks, eval logs | Scheduled batch | Model-graded scorers, explicit model | Eval logs in inspect view |
| promptfoo | MIT | YAML config, response cache | GitHub Action posts to the PR | Grader set by --grader, defaultTest or per assertion | Results matrix, web viewer |
| DeepEval | Apache 2.0 | pytest test cases | deepeval test run | G-Eval, DAG; judge per metric | Via Confident AI |
| Ragas | Apache 2.0 | Metric objects | Library call | Judge per metric | None |
| MLflow 3 | Apache 2.0 | Tracked runs, datasets | mlflow.genai.evaluate in a run | Built-in judges such as Correctness and Guidelines | MLflow Tracing, built on OpenTelemetry |
| Phoenix | Elastic License 2.0 | Versioned datasets, experiments | Library call | Response and retrieval evals | OpenTelemetry via OpenInference |
| LangSmith | Hosted; self-hosting is an Enterprise add-on | Experiments per app version | SDK call | LLM-as-judge, pairwise, online | Built in |
| Braintrust | Commercial platform | Immutable experiments | Code, UI or CI/CD | Autoevals, LLM judges, online scoring | Built in |
| Weave | SDK Apache 2.0 | Evaluation runs | SDK call | Defined in your evaluation code | Built in |
The benchmark harnesses have the best reproducibility tooling here. lm-eval’s task guide asks authors to increase a task’s version number whenever the task changes, so a score can be tied to the exact task revision that produced it, and Lessons from the Trenches documents the sensitivity of models to evaluation setup that makes this necessary. Those harnesses are for choosing a model, not for the commit gate. The application frameworks are the gate. The platforms tie a score back to the trace that produced it, which no pure runner does.
Benchmark reproduction harnesses
lm-evaluation-harness is the reference implementation for standardised benchmark tasks; its README lists over 60 standard academic benchmarks with hundreds of subtasks and variants. Tasks are declared as configuration and models sit behind a common interface, so the same task file runs against a Hugging Face model, a vLLM or SGLang server, or a commercial API. Its value is consistency: two models scored by the same task file at the same shot count give a comparison that two vendor-published numbers do not. Its limitation is that it is built around items with a known correct answer, which makes it a poor fit for open-ended application output. Code benchmarks such as HumanEval report pass@k, so check which estimator a published number used before comparing it with yours.
HELM, from Stanford’s Center for Research on Foundation Models, starts from the position that a single accuracy figure is an inadequate description of a model. The original paper evaluated 30 models across 42 scenarios on seven metrics: accuracy, calibration, robustness, fairness, bias, toxicity and efficiency. The framework publishes results as a leaderboard with a web UI for inspecting individual prompts and responses, which is heavier than a task-level harness and aimed at a public, comparable report rather than a fast internal signal.
Check its status before adopting it. HELM entered maintenance mode on June 1, 2026: volunteer maintainers support it on a best-effort basis, no new features will be added, and no new evaluations will be added to the HELM leaderboards. The policy also warns that external API changes and model deprecations may break scenarios and model clients. It still produces a multi-metric report for models its clients support; a team that builds a recurring pipeline on it should budget for fixing those breaks itself.
Application-level frameworks
OpenAI Evals keeps evals in a registry. Each eval is a JSONL file of samples, each with an input and, for the basic templates, an ideal answer, plus a YAML entry under evals/registry/evals/ naming the eval class. The basic templates are Match, Includes, FuzzyMatch and JsonMatch. The model-graded template, ModelBasedClassify, wraps the completion in an evaluation prompt, with example specs for factual consistency, closed QA and head-to-head battles between two completions. An eval that follows a template needs data and YAML, not code. The repository is MIT-licensed, and its README now points to configuring and running Evals in the OpenAI Dashboard. The hosted Evals API replaces the registry file with a data source schema and testing criteria, graders such as string_check, so choose the repository for a local data-only registry and the API if the evals should live next to your OpenAI usage.
Inspect, published by the UK AI Security Institute under MIT, splits an eval into three pieces: a dataset of labelled samples with input and target columns, a solver that produces the answer, and a scorer that grades it by text comparison, model grading or a custom scheme. The solver can be a single generate() call or a full agent using tools over many turns, which is why the framework suits agentic evaluation, where the behaviour that matters happens across turns and tool calls rather than in one completion. It ships built-in agents such as ReAct, tools including bash, python, web browsing and computer use, and sandboxing for untrusted model code in Docker, Kubernetes and other back ends. Logs are designed to be read item by item in inspect view, and the repository points to a collection of over 200 pre-built evaluations.
promptfoo treats evaluation as regression testing. A YAML file lists prompts, providers and test cases with assertions attached, and the runner produces a side-by-side matrix across providers. Its model-graded assertions include llm-rubric, g-eval, factuality and RAG checks such as context-faithfulness. Because the configuration is declarative and the runner is a CLI, it drops into CI with little ceremony. Per its README, promptfoo is now part of OpenAI and remains open source under MIT.
DeepEval comes at the same job from the pytest side, so evals live next to the application’s other tests and run with deepeval test run. Its metric library covers custom judges (G-Eval and DAG), RAG metrics such as answer relevancy, faithfulness and contextual precision and recall, agent metrics such as task completion and tool correctness, and general metrics including hallucination, bias and toxicity. Results can sync to Confident AI’s hosted platform, which is where its trace view lives.
Ragas began as a RAG evaluator, and its available metrics now also cover agents and tool use (tool call accuracy, agent goal accuracy), SQL, natural-language comparison and general-purpose rubric scoring. For RAG, context precision and context recall assess retrieval, faithfulness measures whether the answer is supported by the retrieved context, and response relevancy measures whether it addresses the question. Scoring those separately is what tells a retrieval failure from a generation failure. The repository (Apache 2.0, Vibrant Labs) can also generate synthetic test sets, useful for coverage and no substitute for real inputs.
Tracing-first platforms
A pure runner reports a score. It cannot show which request, retrieval or tool call produced the failing item. The tracing platforms close that loop by storing the trace and the score together.
- MLflow 3 (Apache 2.0) runs evaluation through
mlflow.genai.evaluate, which takes a dataset, a prediction function and scorers; built-in LLM judges include Correctness and Guidelines, and feedback is attached to traces with user, timestamp and revision metadata. The repository describes its tracing as built on OpenTelemetry. - Phoenix (Arize) traces through OpenTelemetry-based OpenInference instrumentation and adds versioned datasets and experiments for tracking changes to prompts, models and retrieval. It ships under the Elastic License 2.0, not Apache or MIT.
- LangSmith records each experiment’s outputs, evaluator scores and traces per example. Pairwise evaluators compare two application versions when ranking two outputs is easier than scoring one, and online evaluators score live traffic without reference outputs.
- Braintrust calls experiments “the immutable, comparable record of your eval runs”, runnable from code, the UI or CI/CD. Scorers can be built-in autoevals, LLM judges or custom code, and online scoring evaluates production traces asynchronously.
- Weave, from Weights & Biases, logs inputs, outputs and traces and runs evaluations over them; its SDK is Apache 2.0.
Online scoring on these platforms has no ground truth, so it measures a judge’s opinion of live traffic. It is a monitoring signal, not a deploy gate.
Choosing
The practical decision tree is short.
- Are you comparing base models on published benchmarks? Use lm-evaluation-harness. Use HELM if the report must cover more than accuracy and its existing model clients still work for the models you need.
- Are you testing a RAG pipeline? Consider Ragas for a metric-oriented evaluation of retrieval and generation, or DeepEval if you want the same separation organised as pytest-style application tests.
- Are you testing an agent that uses tools over several turns? Inspect, because the solver and scorer split matches that shape and its sandboxes contain untrusted code.
- Do you mainly need a gate in CI? promptfoo if you want declarative YAML and results posted to the pull request, DeepEval if you want it inside pytest.
- Do you want a registry of small custom evals against OpenAI models? OpenAI Evals, or its hosted Evals API.
- Do you need to connect a failing score to the trace behind it? MLflow or Phoenix if you run the stack yourself; LangSmith, Braintrust or Weave if a vendor platform is acceptable.
These are not exclusive. A common arrangement runs one benchmark harness whenever a candidate model appears and one application framework on every commit, with a tracing platform behind both if production failures need to flow back into the golden set.
The metric that matters
Whichever framework runs the gate, track the paired regression delta on a pinned golden set, with a confidence interval, rather than the aggregate pass rate.
Freeze a golden set of N items and record its hash. Run baseline and candidate on the same items, scoring both on the same scale. For item i, d_i is candidate minus baseline. Report the mean of d, its standard error (standard deviation over the square root of N) and the 95% interval, mean plus or minus 1.96 standard errors. The gate: the interval must exclude a regression larger than your tolerance.
An aggregate pass rate compares two independent means, so item difficulty dominates the variance and a small real regression hides inside the interval. Pairing cancels it, because both systems answered the same question. Miller’s Adding Error Bars to Evals treats evals as experiments and gives the formulas for comparing two models and planning sample size. A difference smaller than its interval is not a result.
The second number is judge agreement with a human-labelled slice. In their MT-Bench and Chatbot Arena experiments, Zheng and co-authors reported over 80% agreement between GPT-4 judgments and human preferences, comparable to human-human agreement in those experiments. This is an empirical result for those evaluation settings, not a universal ceiling or acceptance threshold. Measure judge-human and human-human agreement on the same application-specific sample using the same rubric and agreement statistic before setting a gate; the procedure is in LLM-as-a-judge bias.
Wiring it up
The commit gate, in promptfoo: two prompt files against one provider, the grader pinned in defaultTest.options so it cannot float to promptfoo’s built-in grading provider, which is chosen from whatever credentials the runner holds, and the results posted to the pull request by the GitHub Action.
prompts:
- file://prompts/baseline.txt
- file://prompts/candidate.txt
providers:
- openai:gpt-5-mini-2025-08-07
defaultTest:
options:
provider: openai:gpt-6.1-sol
tests:
- vars:
ticket: "My invoice shows two charges for August."
assert:
- type: contains-json
- type: llm-rubric
value: "Acknowledges the duplicate charge and states one next step"
threshold: 0.8
The application model is pinned to a dated snapshot, gpt-5-mini-2025-08-07, which OpenAI’s GPT-5 mini page lists. The grader, a different model from the one under test so it is not grading its own outputs, is gpt-6.1-sol, which its model page lists as that model’s only snapshot ID. Where a grader has dated snapshots, pin the dated one; either way, record the grader ID with every run, because a model moving behind a name is the instrument change described below.
The model-selection run, in lm-eval, seeds fixed, samples logged for post-hoc diffing, candidate served through vLLM with tensor parallelism across two GPUs:
lm_eval --model vllm \
--model_args pretrained=/models/candidate,tensor_parallel_size=2,dtype=auto \
--tasks gsm8k,hellaswag --num_fewshot 5 \
--seed 0,1234,1234,1234 --batch_size auto \
--log_samples --output_path runs/candidate/
Its CLI reference marks --limit as for testing only, so a subsampled run is a smoke test, and documents that the single-command form above still works alongside the newer lm-eval run subcommand. Review the MMLU, HumanEval and GSM8K scoring protocols before interpreting the GSM8K result, so the score is tied to its task and grading method.
What you’ll see
Plot the mean paired delta per commit with its interval as a band, and a zero line. Good looks boring: the band straddles zero and narrows as the golden set grows. A real regression is a band entirely below the tolerance line, and the per-item log shows failures clustered on one category of input. An instrument change is a step on a day with no prompt diff, usually a provider moving the model behind an alias, and it vanishes when the grader is pinned to a fixed snapshot. A flat line across commits that changed the prompt is a cache replaying old responses, not stability.
What no framework provides
Every framework here ships a runner, metrics and adapters. None ships your evaluation set. The dataset decides whether the exercise is informative, and it has to come from the inputs your system actually receives and the failures you actually care about; synthetic generators help with coverage but tend to miss the hard inputs. How to build that set, including how to split development from test, is in designing an LLM evaluation that actually tells you something.
None ships a trustworthy judge either. Every framework offers model-graded metrics, and every one inherits the judge model’s position, verbosity and self-preference biases. Switching frameworks does not fix that; calibrating the judge against human labels does.
Caveats
- Every framework inherits judge bias. MT-Bench documents position, verbosity and self-enhancement bias; G-Eval reports a Spearman correlation of 0.514 with humans on summarization and a bias toward LLM-generated text.
- Caching hides drift, then dumps it. promptfoo caches successful API responses for 14 days by default, so provider drift stays invisible until entries expire, then lands as one step. Gate with
--no-cache. - Sampling budget, not framework choice, drives cost. A suite at pass@1 on 500 prompts is 500 generations; at pass@10 it is 5,000, and a model grader adds a judging call per output. HumanEval has 164 problems, so k=5 means 164 × 5 = 820 generations per run, and baseline plus candidate need 1,640 before judge calls or retries. Use the calculator to budget baseline and candidate evaluation runs separately, with each model’s token assumptions and rates, then multiply by every commit the suite runs on.
- Cardinality. Per-item logs are essential for diffing and poison for a time-series database; keep item IDs out of Prometheus labels and join samples to aggregates by run ID.
- Label leakage. Golden items drift into few-shot examples and fine-tuning sets. Keep the set out of both and disclose the prompt engineering done.
- Licence and custody. Phoenix ships under the Elastic License 2.0, not Apache or MIT. LangSmith and Braintrust are hosted platforms, so by default the golden set and every trace scored online live on the vendor’s servers; self-hosted LangSmith is an Enterprise add-on.
Use the LLM evaluation guides to connect these CI gates with dataset design, benchmark interpretation and grader checks.
None of this is monitoring. Online scoring of production traffic without references is a drift signal that belongs next to PSI and KS tests, not in the deploy gate. promptfoo’s red-team mode is an attack suite that needs its own scoping, not a regression test.
FAQ
is helm still maintained
HELM has been in maintenance mode since June 1, 2026. Volunteer maintainers keep it available on a best-effort basis, but no new features or leaderboard evaluations will be added, and external API changes may break its model clients. It still produces multi-metric reports for supported models; plan to patch breakages yourself.
deepeval vs ragas for rag evaluation
Both score retrieval and generation separately, with context precision, context recall, faithfulness and relevancy metrics. Choose Ragas for a metric-first library that can also generate synthetic test sets. Choose DeepEval if you want those checks written as pytest tests that run with deepeval test run and fail the build like any other test.
how many test cases do i need for an llm eval
Enough that the confidence interval on the paired delta is narrower than the regression you need to catch. Run baseline and candidate on the same items, compute the standard error of the per-item differences, and grow the set until the interval’s half-width is smaller than your tolerance. Miller’s error-bars paper gives the sample-size formulas.
should llm evals run on every commit
Run the application-level gate on every commit that touches prompts, retrieval or model settings, with the response cache off and the grader pinned. Run benchmark harnesses only when a candidate model appears. Budget first: generations multiply with sampling, and every model-graded check adds a judge call per output.
is promptfoo still open source now that it is part of openai
Yes. promptfoo’s README says the project is now part of OpenAI and remains open source under the MIT licence. The README also says evals run locally on your machine; the model providers and the grader your config names still receive the prompts and outputs they process, so check their data terms.