Evals ML
Isometric sci-fi assembly line moves glowing blue crystals through test stations, scales, scanners, and a secure deployment vault.
Eval Tooling

LLM Evaluation Frameworks Compared: Task Fit and CI Gates

Compare LLM evaluation tools for CI gates, reproducible runs, judge control, trace linkage and licensing, with promptfoo and lm-eval configs.

By Evals ML Editorial · ·Updated · 16 min read

Pick an LLM evaluation framework by job, not popularity. lm-evaluation-harness (or HELM, now in maintenance mode) reproduces published benchmarks for model selection; promptfoo or DeepEval gates prompt changes in CI; Ragas scores RAG retrieval and generation separately; Inspect fits multi-turn agents; MLflow, Phoenix, LangSmith, Braintrust and Weave tie scores to traces.

The failure that usually starts the search is not a missing eval. It is a prompt edit that moved the score a couple of points, a deploy that went ahead on a green dashboard, and tickets two days later showing the model answering in the wrong format. Nobody can say whether the change was real, the judge behind an API alias moved, or the run replayed cached responses. So the comparison below runs on two axes: which task shape each framework fits, and whether its runs are reproducible enough to gate a deploy.

The comparison

The first axis is what each tool was built to evaluate. Benchmark harnesses score models against items with a known answer, application frameworks test a system you built, and tracing platforms attach scores to the experiment and production traces that produced them.

FrameworkMaintained byBuilt forGrading styleBest fit
lm-evaluation-harnessEleutherAIReproducing academic benchmarks across many modelsLog-likelihood and exact match, defined per taskComparing base models on MMLU, GSM8K and similar sets
HELMStanford CRFM, in maintenance modeMulti-metric benchmarking with a public leaderboardAccuracy plus calibration, robustness, fairness, bias, toxicity, efficiencyReports where one accuracy number is not an acceptable summary
OpenAI EvalsOpenAIA registry of small evals expressed as data plus a templateMatch, includes, fuzzy and JSON match; model-graded templatesTeams on the OpenAI API wanting data-only custom evals
InspectUK AI Security InstituteDatasets, solvers and scorers in PythonText comparison, model grading or custom scorersAgentic, tool-using and safety evaluations
promptfoopromptfoo, now part of OpenAIPrompt and model regression testingDeclarative YAML assertions plus model-graded checksCatching prompt regressions on every pull request
DeepEvalConfident AIApplication tests in a pytest idiomMetric classes: G-Eval, DAG, RAG, agent, hallucinationPython teams who want evals to read like unit tests
RagasVibrant LabsRAG evaluation, plus agent, SQL and general-purpose metricsFaithfulness, response relevancy, context precision and recallDiagnosing retrieval and generation separately
MLflow 3, Phoenix, LangSmith, Braintrust, WeaveMLflow project, Arize, LangChain, Braintrust, Weights & BiasesTracing plus evaluationBuilt-in and custom LLM judges, online scoringLinking a score to the trace that produced it

The benchmark harnesses answer “which model should we build on”. The application frameworks answer “did today’s change break anything”. The platforms answer “which request produced this bad score”. Most teams need the second question answered on every commit and the first only when a model changes.

Gate-ability: reproducibility, CI fit, judges and traces

A run can gate a deploy only if it can be repeated, triggered per change, scored by a judge you control and traced back to the outputs behind the score. The same tools compared on that axis, leaving out HELM and OpenAI Evals, which are batch tools rather than per-commit gates:

FrameworkLicenceReproducibility controlsCI fitJudge controlTraces
lm-evaluation-harnessMITVersioned tasks, fixed seeds, sample logsBatch, not per commitExact match, log-likelihoodNone
InspectMITPython tasks, eval logsScheduled batchModel-graded scorers, explicit modelEval logs in inspect view
promptfooMITYAML config, response cacheGitHub Action posts to the PRGrader set by --grader, defaultTest or per assertionResults matrix, web viewer
DeepEvalApache 2.0pytest test casesdeepeval test runG-Eval, DAG; judge per metricVia Confident AI
RagasApache 2.0Metric objectsLibrary callJudge per metricNone
MLflow 3Apache 2.0Tracked runs, datasetsmlflow.genai.evaluate in a runBuilt-in judges such as Correctness and GuidelinesMLflow Tracing, built on OpenTelemetry
PhoenixElastic License 2.0Versioned datasets, experimentsLibrary callResponse and retrieval evalsOpenTelemetry via OpenInference
LangSmithHosted; self-hosting is an Enterprise add-onExperiments per app versionSDK callLLM-as-judge, pairwise, onlineBuilt in
BraintrustCommercial platformImmutable experimentsCode, UI or CI/CDAutoevals, LLM judges, online scoringBuilt in
WeaveSDK Apache 2.0Evaluation runsSDK callDefined in your evaluation codeBuilt in

The benchmark harnesses have the best reproducibility tooling here. lm-eval’s task guide asks authors to increase a task’s version number whenever the task changes, so a score can be tied to the exact task revision that produced it, and Lessons from the Trenches documents the sensitivity of models to evaluation setup that makes this necessary. Those harnesses are for choosing a model, not for the commit gate. The application frameworks are the gate. The platforms tie a score back to the trace that produced it, which no pure runner does.

Benchmark reproduction harnesses

lm-evaluation-harness is the reference implementation for standardised benchmark tasks; its README lists over 60 standard academic benchmarks with hundreds of subtasks and variants. Tasks are declared as configuration and models sit behind a common interface, so the same task file runs against a Hugging Face model, a vLLM or SGLang server, or a commercial API. Its value is consistency: two models scored by the same task file at the same shot count give a comparison that two vendor-published numbers do not. Its limitation is that it is built around items with a known correct answer, which makes it a poor fit for open-ended application output. Code benchmarks such as HumanEval report pass@k, so check which estimator a published number used before comparing it with yours.

HELM, from Stanford’s Center for Research on Foundation Models, starts from the position that a single accuracy figure is an inadequate description of a model. The original paper evaluated 30 models across 42 scenarios on seven metrics: accuracy, calibration, robustness, fairness, bias, toxicity and efficiency. The framework publishes results as a leaderboard with a web UI for inspecting individual prompts and responses, which is heavier than a task-level harness and aimed at a public, comparable report rather than a fast internal signal.

Check its status before adopting it. HELM entered maintenance mode on June 1, 2026: volunteer maintainers support it on a best-effort basis, no new features will be added, and no new evaluations will be added to the HELM leaderboards. The policy also warns that external API changes and model deprecations may break scenarios and model clients. It still produces a multi-metric report for models its clients support; a team that builds a recurring pipeline on it should budget for fixing those breaks itself.

Application-level frameworks

OpenAI Evals keeps evals in a registry. Each eval is a JSONL file of samples, each with an input and, for the basic templates, an ideal answer, plus a YAML entry under evals/registry/evals/ naming the eval class. The basic templates are Match, Includes, FuzzyMatch and JsonMatch. The model-graded template, ModelBasedClassify, wraps the completion in an evaluation prompt, with example specs for factual consistency, closed QA and head-to-head battles between two completions. An eval that follows a template needs data and YAML, not code. The repository is MIT-licensed, and its README now points to configuring and running Evals in the OpenAI Dashboard. The hosted Evals API replaces the registry file with a data source schema and testing criteria, graders such as string_check, so choose the repository for a local data-only registry and the API if the evals should live next to your OpenAI usage.

Inspect, published by the UK AI Security Institute under MIT, splits an eval into three pieces: a dataset of labelled samples with input and target columns, a solver that produces the answer, and a scorer that grades it by text comparison, model grading or a custom scheme. The solver can be a single generate() call or a full agent using tools over many turns, which is why the framework suits agentic evaluation, where the behaviour that matters happens across turns and tool calls rather than in one completion. It ships built-in agents such as ReAct, tools including bash, python, web browsing and computer use, and sandboxing for untrusted model code in Docker, Kubernetes and other back ends. Logs are designed to be read item by item in inspect view, and the repository points to a collection of over 200 pre-built evaluations.

promptfoo treats evaluation as regression testing. A YAML file lists prompts, providers and test cases with assertions attached, and the runner produces a side-by-side matrix across providers. Its model-graded assertions include llm-rubric, g-eval, factuality and RAG checks such as context-faithfulness. Because the configuration is declarative and the runner is a CLI, it drops into CI with little ceremony. Per its README, promptfoo is now part of OpenAI and remains open source under MIT.

DeepEval comes at the same job from the pytest side, so evals live next to the application’s other tests and run with deepeval test run. Its metric library covers custom judges (G-Eval and DAG), RAG metrics such as answer relevancy, faithfulness and contextual precision and recall, agent metrics such as task completion and tool correctness, and general metrics including hallucination, bias and toxicity. Results can sync to Confident AI’s hosted platform, which is where its trace view lives.

Ragas began as a RAG evaluator, and its available metrics now also cover agents and tool use (tool call accuracy, agent goal accuracy), SQL, natural-language comparison and general-purpose rubric scoring. For RAG, context precision and context recall assess retrieval, faithfulness measures whether the answer is supported by the retrieved context, and response relevancy measures whether it addresses the question. Scoring those separately is what tells a retrieval failure from a generation failure. The repository (Apache 2.0, Vibrant Labs) can also generate synthetic test sets, useful for coverage and no substitute for real inputs.

Tracing-first platforms

A pure runner reports a score. It cannot show which request, retrieval or tool call produced the failing item. The tracing platforms close that loop by storing the trace and the score together.

  • MLflow 3 (Apache 2.0) runs evaluation through mlflow.genai.evaluate, which takes a dataset, a prediction function and scorers; built-in LLM judges include Correctness and Guidelines, and feedback is attached to traces with user, timestamp and revision metadata. The repository describes its tracing as built on OpenTelemetry.
  • Phoenix (Arize) traces through OpenTelemetry-based OpenInference instrumentation and adds versioned datasets and experiments for tracking changes to prompts, models and retrieval. It ships under the Elastic License 2.0, not Apache or MIT.
  • LangSmith records each experiment’s outputs, evaluator scores and traces per example. Pairwise evaluators compare two application versions when ranking two outputs is easier than scoring one, and online evaluators score live traffic without reference outputs.
  • Braintrust calls experiments “the immutable, comparable record of your eval runs”, runnable from code, the UI or CI/CD. Scorers can be built-in autoevals, LLM judges or custom code, and online scoring evaluates production traces asynchronously.
  • Weave, from Weights & Biases, logs inputs, outputs and traces and runs evaluations over them; its SDK is Apache 2.0.

Online scoring on these platforms has no ground truth, so it measures a judge’s opinion of live traffic. It is a monitoring signal, not a deploy gate.

Choosing

The practical decision tree is short.

  1. Are you comparing base models on published benchmarks? Use lm-evaluation-harness. Use HELM if the report must cover more than accuracy and its existing model clients still work for the models you need.
  2. Are you testing a RAG pipeline? Consider Ragas for a metric-oriented evaluation of retrieval and generation, or DeepEval if you want the same separation organised as pytest-style application tests.
  3. Are you testing an agent that uses tools over several turns? Inspect, because the solver and scorer split matches that shape and its sandboxes contain untrusted code.
  4. Do you mainly need a gate in CI? promptfoo if you want declarative YAML and results posted to the pull request, DeepEval if you want it inside pytest.
  5. Do you want a registry of small custom evals against OpenAI models? OpenAI Evals, or its hosted Evals API.
  6. Do you need to connect a failing score to the trace behind it? MLflow or Phoenix if you run the stack yourself; LangSmith, Braintrust or Weave if a vendor platform is acceptable.

These are not exclusive. A common arrangement runs one benchmark harness whenever a candidate model appears and one application framework on every commit, with a tracing platform behind both if production failures need to flow back into the golden set.

The metric that matters

Whichever framework runs the gate, track the paired regression delta on a pinned golden set, with a confidence interval, rather than the aggregate pass rate.

Freeze a golden set of N items and record its hash. Run baseline and candidate on the same items, scoring both on the same scale. For item i, d_i is candidate minus baseline. Report the mean of d, its standard error (standard deviation over the square root of N) and the 95% interval, mean plus or minus 1.96 standard errors. The gate: the interval must exclude a regression larger than your tolerance.

An aggregate pass rate compares two independent means, so item difficulty dominates the variance and a small real regression hides inside the interval. Pairing cancels it, because both systems answered the same question. Miller’s Adding Error Bars to Evals treats evals as experiments and gives the formulas for comparing two models and planning sample size. A difference smaller than its interval is not a result.

The second number is judge agreement with a human-labelled slice. In their MT-Bench and Chatbot Arena experiments, Zheng and co-authors reported over 80% agreement between GPT-4 judgments and human preferences, comparable to human-human agreement in those experiments. This is an empirical result for those evaluation settings, not a universal ceiling or acceptance threshold. Measure judge-human and human-human agreement on the same application-specific sample using the same rubric and agreement statistic before setting a gate; the procedure is in LLM-as-a-judge bias.

Wiring it up

The commit gate, in promptfoo: two prompt files against one provider, the grader pinned in defaultTest.options so it cannot float to promptfoo’s built-in grading provider, which is chosen from whatever credentials the runner holds, and the results posted to the pull request by the GitHub Action.

prompts:
  - file://prompts/baseline.txt
  - file://prompts/candidate.txt
providers:
  - openai:gpt-5-mini-2025-08-07
defaultTest:
  options:
    provider: openai:gpt-6.1-sol
tests:
  - vars:
      ticket: "My invoice shows two charges for August."
    assert:
      - type: contains-json
      - type: llm-rubric
        value: "Acknowledges the duplicate charge and states one next step"
        threshold: 0.8

The application model is pinned to a dated snapshot, gpt-5-mini-2025-08-07, which OpenAI’s GPT-5 mini page lists. The grader, a different model from the one under test so it is not grading its own outputs, is gpt-6.1-sol, which its model page lists as that model’s only snapshot ID. Where a grader has dated snapshots, pin the dated one; either way, record the grader ID with every run, because a model moving behind a name is the instrument change described below.

The model-selection run, in lm-eval, seeds fixed, samples logged for post-hoc diffing, candidate served through vLLM with tensor parallelism across two GPUs:

lm_eval --model vllm \
  --model_args pretrained=/models/candidate,tensor_parallel_size=2,dtype=auto \
  --tasks gsm8k,hellaswag --num_fewshot 5 \
  --seed 0,1234,1234,1234 --batch_size auto \
  --log_samples --output_path runs/candidate/

Its CLI reference marks --limit as for testing only, so a subsampled run is a smoke test, and documents that the single-command form above still works alongside the newer lm-eval run subcommand. Review the MMLU, HumanEval and GSM8K scoring protocols before interpreting the GSM8K result, so the score is tied to its task and grading method.

What you’ll see

Plot the mean paired delta per commit with its interval as a band, and a zero line. Good looks boring: the band straddles zero and narrows as the golden set grows. A real regression is a band entirely below the tolerance line, and the per-item log shows failures clustered on one category of input. An instrument change is a step on a day with no prompt diff, usually a provider moving the model behind an alias, and it vanishes when the grader is pinned to a fixed snapshot. A flat line across commits that changed the prompt is a cache replaying old responses, not stability.

What no framework provides

Every framework here ships a runner, metrics and adapters. None ships your evaluation set. The dataset decides whether the exercise is informative, and it has to come from the inputs your system actually receives and the failures you actually care about; synthetic generators help with coverage but tend to miss the hard inputs. How to build that set, including how to split development from test, is in designing an LLM evaluation that actually tells you something.

None ships a trustworthy judge either. Every framework offers model-graded metrics, and every one inherits the judge model’s position, verbosity and self-preference biases. Switching frameworks does not fix that; calibrating the judge against human labels does.

Caveats

  • Every framework inherits judge bias. MT-Bench documents position, verbosity and self-enhancement bias; G-Eval reports a Spearman correlation of 0.514 with humans on summarization and a bias toward LLM-generated text.
  • Caching hides drift, then dumps it. promptfoo caches successful API responses for 14 days by default, so provider drift stays invisible until entries expire, then lands as one step. Gate with --no-cache.
  • Sampling budget, not framework choice, drives cost. A suite at pass@1 on 500 prompts is 500 generations; at pass@10 it is 5,000, and a model grader adds a judging call per output. HumanEval has 164 problems, so k=5 means 164 × 5 = 820 generations per run, and baseline plus candidate need 1,640 before judge calls or retries. Use the calculator to budget baseline and candidate evaluation runs separately, with each model’s token assumptions and rates, then multiply by every commit the suite runs on.
  • Cardinality. Per-item logs are essential for diffing and poison for a time-series database; keep item IDs out of Prometheus labels and join samples to aggregates by run ID.
  • Label leakage. Golden items drift into few-shot examples and fine-tuning sets. Keep the set out of both and disclose the prompt engineering done.
  • Licence and custody. Phoenix ships under the Elastic License 2.0, not Apache or MIT. LangSmith and Braintrust are hosted platforms, so by default the golden set and every trace scored online live on the vendor’s servers; self-hosted LangSmith is an Enterprise add-on.

Use the LLM evaluation guides to connect these CI gates with dataset design, benchmark interpretation and grader checks.

None of this is monitoring. Online scoring of production traffic without references is a drift signal that belongs next to PSI and KS tests, not in the deploy gate. promptfoo’s red-team mode is an attack suite that needs its own scoping, not a regression test.

FAQ

is helm still maintained

HELM has been in maintenance mode since June 1, 2026. Volunteer maintainers keep it available on a best-effort basis, but no new features or leaderboard evaluations will be added, and external API changes may break its model clients. It still produces multi-metric reports for supported models; plan to patch breakages yourself.

deepeval vs ragas for rag evaluation

Both score retrieval and generation separately, with context precision, context recall, faithfulness and relevancy metrics. Choose Ragas for a metric-first library that can also generate synthetic test sets. Choose DeepEval if you want those checks written as pytest tests that run with deepeval test run and fail the build like any other test.

how many test cases do i need for an llm eval

Enough that the confidence interval on the paired delta is narrower than the regression you need to catch. Run baseline and candidate on the same items, compute the standard error of the per-item differences, and grow the set until the interval’s half-width is smaller than your tolerance. Miller’s error-bars paper gives the sample-size formulas.

should llm evals run on every commit

Run the application-level gate on every commit that touches prompts, retrieval or model settings, with the response cache off and the grader pinned. Run benchmark harnesses only when a candidate model appears. Budget first: generations multiply with sampling, and every model-graded check adds a judge call per output.

is promptfoo still open source now that it is part of openai

Yes. promptfoo’s README says the project is now part of OpenAI and remains open source under the MIT licence. The README also says evals run locally on your machine; the model providers and the grader your config names still receive the prompts and outputs they process, so check their data terms.

Sources

  1. Miller, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
  2. Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
  3. Biderman et al., Lessons from the Trenches on Reproducible Evaluation of Language Models
  4. Liu et al., G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
  5. promptfoo documentation: GitHub Action
  6. MLflow documentation: GenAI evaluation and monitoring
  7. EleutherAI: lm-evaluation-harness
  8. lm-evaluation-harness documentation: CLI interface
  9. lm-evaluation-harness documentation: new task guide (task versioning)
  10. Stanford CRFM: HELM repository
  11. HELM documentation: Maintenance Mode Policy
  12. Liang et al., Holistic Evaluation of Language Models
  13. OpenAI: Evals framework and registry
  14. OpenAI Evals documentation: building an eval
  15. OpenAI Evals documentation: eval templates
  16. OpenAI API documentation: Evals guide
  17. UK AI Security Institute: Inspect AI
  18. Inspect AI documentation
  19. promptfoo: GitHub repository
  20. promptfoo documentation: model-graded metrics
  21. promptfoo documentation: caching
  22. Confident AI: DeepEval
  23. Vibrant Labs: Ragas
  24. Ragas documentation: available metrics
  25. MLflow: GitHub repository
  26. Arize Phoenix: GitHub repository
  27. Weights & Biases: Weave
  28. LangSmith documentation: evaluation concepts
  29. LangSmith documentation: self-hosted LangSmith
  30. Braintrust documentation: evaluate
  31. OpenAI API documentation: GPT-6.1 Sol model page
  32. OpenAI API documentation: GPT-5 mini model page
  33. OpenAI HumanEval dataset card
#llm-evaluation #eval-frameworks#ci-cd#regression-testing#llm-as-a-judge

Related