AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How UK AISI And EvalEval Are Making Benchmark Results Reproducible on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, pairing scores with verification, context and configuration details. The release covers five benchmarks across six frontier models plus two cyber evaluations, and accompanies AISI’s paper on how inference-time compute and evaluation protocols shape results.

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, a format that pairs scores with verification, evaluation context and configuration details so readers can see how the results were produced. The release covers five benchmarks across six frontier models, plus two cyber evaluations that use a different, partly overlapping model set. It accompanies AISI’s paper How Inference Compute Shapes Frontier LLM Evaluation, which examines how scores depend on inference-time compute and evaluation protocol.

The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The reported results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones, but those runs use a model set that only partly overlaps with the main experiment, so the same model list should not be assumed to apply to them.

EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. According to the announcement, publicly reported AISI methods and findings are being made available where appropriate — the release does not claim to include every AISI evaluation or every underlying transcript.

The AISI paper’s analysis of Humanity’s Last Exam illustrates why setup details matter. For that benchmark, the reported analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased. In other words, measured performance shifted with inference compute and with whether feedback was provided between attempts.

At a glance
announcementWhen: announced alongside AISI’s inference-co…
The developmentAISI has released benchmark results through EvalEval’s Evaluation Cards format, adding the setup information needed to interpret and compare scores.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Setup Details Change Score Comparisons

Benchmark scores are often cited as if they measure the same thing across models, but different evaluation protocols can produce different results. The AISI paper’s Humanity’s Last Exam analysis is a concrete example: outcomes changed depending on inference compute and on whether models received correctness feedback between attempts. A score published without those conditions leaves readers unsure what performance it actually represents.

Publishing results together with their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify cases where superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that treats evaluations as evidence about advanced AI capabilities. The records do not settle which benchmark or protocol is best, but they make some of the conditions behind a result easier to see.

From NeurIPS Workshop to Shared Schema

The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that shared infrastructure to publicly reported AISI methods and findings.

AISI has separately worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model information. Together, the efforts address a practical reporting problem: results published across different formats and outlets may omit details needed to interpret or reproduce a run, while re-running costly evaluations may not be feasible.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

Coverage and Reproduction Limits

The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. It states only that publicly reported methods and findings are being made available where appropriate, so the release should not be read as a complete archive of all AISI evaluation work.

The cyber evaluations use a different, partly overlapping model set, and the announcement does not enumerate that set. It also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge the current coverage and compare the records consistently.

Broader Adoption of Every Eval Ever

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or by model and examine reporting practices across the collection.

Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of the records contributors publish. No further release date or adoption milestone was specified.

Key Questions

Which benchmarks and models are covered in the main release?

The main experiment covers HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0 across Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. Two additional cyber evaluations — Cyber CTFs and The Last Ones — use a different, partly overlapping model set that has not been enumerated.

What is an Evaluation Card?

It is EvalEval’s format that combines evaluation results with benchmark and model information, including verification, evaluation context and configuration details, so readers can see how a score was produced.

Why does evaluation protocol affect benchmark scores?

According to AISI’s paper, scores can shift based on inference-time compute and protocol choices. In the Humanity’s Last Exam analysis, models that received correctness feedback after each attempt solved additional tasks as token use increased.

Does the release include all AISI evaluation work?

No. The announcement says publicly reported methods and findings are shared where appropriate. The number of records and transcripts available is not specified, and the release should not be read as a complete archive.

Have independent researchers reproduced the results?

That is not stated. The announcement does not say whether outside researchers have independently reproduced the published results.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Generative AI Explained: How AI Can Create Images, Music, and Text

Keen to discover how generative AI transforms creativity by crafting unique images, music, and text? The implications might surprise you.

Upcoming AI Trends In 2026: 10 Must-Watch Developments

An analysis of the confirmed and claimed AI trends to watch in 2026, highlighting their significance and what remains uncertain for the industry.

Augmented Reality Vs Virtual Reality: What’s the Difference?

Learn how augmented reality and virtual reality differ in enhancing your experiences, and discover which technology could transform your life in unexpected ways.

Isar Aerospace Launch Into Orbit [Video]

Isar Aerospace’s latest launch marks a significant milestone, with confirmed orbit insertion. Details are still emerging, but the event signals growing European space capabilities.