AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Making AI Evaluations Easier To Reproduce: UK AISI And EvalEval on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, which pair scores with verification, context and configuration details. The release covers five benchmarks across six frontier models, plus two cyber evaluations using a partly different model set.

The UK AI Security Institute (AISI) has begun publishing selected AI benchmark results through EvalEval’s Evaluation Cards, adding verification, context and configuration details intended to help readers understand how each result was produced. The release covers five benchmarks across six frontier models, plus two cyber evaluations that use a different, partly overlapping model set.

The release accompanies AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how model scores depend on inference-time compute and evaluation protocol. The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4.

AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones. Those use a different set of models that overlaps only partly with the main experiment, so the same model list should not be assumed to apply to them.

EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. AISI says its public reporting is available where appropriate; the announcement does not claim that every AISI evaluation or every underlying transcript is included.

The paper’s analysis of Humanity’s Last Exam illustrates why protocol details matter. The reported analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased.

At a glance
announcementWhen: announced alongside AISI’s inference-co…
The developmentUK AISI announced it is publishing evaluation results through EvalEval’s Evaluation Cards, alongside its paper on how inference-time compute affects evaluation outcomes.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Setup Details Change Score Comparisons

Benchmark scores are often cited as if they measure the same thing across models, but different evaluation protocols can produce different results. The AISI paper’s Humanity’s Last Exam analysis shows why: outcomes shifted both with the amount of inference compute used and with whether models received correctness feedback between attempts. A score reported without those conditions can leave readers unsure what performance it actually represents.

Publishing results alongside their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify when superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that treats evaluations as evidence about advanced AI capabilities.

The records do not by themselves settle which benchmark or protocol is best, but they make some of the conditions behind a result easier to see — a step toward reproducible evaluation practice in a field where repeating costly evaluations may not be feasible.

From NeurIPS Workshop to Shared Schema

The release builds on earlier collaboration between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations, and the current release applies that shared infrastructure to publicly reported AISI methods and findings.

AISI has separately worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model information into a single record.

The combined efforts address a practical reporting problem: results published across different formats and outlets may omit details needed to interpret or reproduce a run, and the field has lacked a common way to document those details alongside the scores themselves.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

Coverage and Reproduction Limits

The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. It says publicly reported methods and findings are being made available where appropriate, so the release should not be read as a complete archive of all AISI evaluation work.

The cyber evaluations use a different, partly overlapping model set, and the announcement does not enumerate that set. It also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge the current coverage and compare records consistently.

Broader Adoption of Every Eval Ever

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of the Every Eval Ever schema: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema.

Researchers in evaluation, governance and policy can already explore Evaluation Cards by benchmark or model and examine reporting practices across the collection. Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of records contributors publish. No further release date or adoption milestone was specified.

Key Questions

Which models and benchmarks are covered in the main release?

The main experiment covers five benchmarks — HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0 — across six models: Claude Opus 4, 4.5 and 4.6, and GPT-5, GPT-5.2 and GPT-5.4.

Do the cyber evaluations cover the same models?

No. The Cyber CTFs and The Last Ones evaluations use a different model set that overlaps only partly with the main experiment, and the announcement does not enumerate that set.

Does the release include all of AISI’s evaluation work?

No. AISI says publicly reported methods and findings are made available where appropriate; the announcement does not claim that every AISI evaluation or every underlying transcript is included.

Why does evaluation protocol matter for benchmark scores?

AISI’s paper found that scores on Humanity’s Last Exam changed with inference-time compute and with whether models received correctness feedback between attempts, meaning the same model can score differently under different setups.

What is Every Eval Ever (EEE)?

EEE is EvalEval’s shared schema for documenting evaluations, developed with feedback from AISI following a joint workshop at NeurIPS 2025. It organizes benchmark metadata, evaluation-run data and model metadata into a common format.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of Business Automation: AI Tools To Consider In 2026

An in-depth look at the key AI tools shaping business automation in 2026, including platforms, hardware, and frameworks, with expert insights.

Battery Re‑cell and Regeneration Technologies: Extending Battery Lifespan

Harnessing innovative re-cell and regeneration techniques can significantly extend battery lifespan—discover how these solutions revolutionize energy storage sustainability.

The Valley Of Webhooks

A widespread security vulnerability involving webhooks has been revealed, impacting thousands of organizations. Details are still emerging.

DLSS 5 Mod Shifts Load To Second GPU, Doubling Frame-Rates And Latency

A new DLSS 5 mod shifts rendering load to a second GPU, doubling frame rates but increasing latency, sparking interest among gamers and developers.