AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Three Ways UK AISI And EvalEval Approach Reproducible AI Benchmarks on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected results from five benchmarks across six frontier models through EvalEval’s Evaluation Cards, which pair scores with verification, context and configuration information. Two cyber evaluations are also included, using a different model set that overlaps only partly with the main group. The release is not described as a complete archive of AISI evaluations or transcripts.

The UK AI Security Institute (AISI) is publishing selected results from its evaluations of advanced AI models through EvalEval’s Evaluation Cards, adding verification, evaluation context and configuration information to help readers understand how scores were produced, as detailed in the original analysis. The release covers five benchmarks across six frontier models, as well as two cyber evaluations using a partly different model set. It accompanies AISI’s paper on how inference-time compute and evaluation protocols affect benchmark results.

The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath and Humanity’s Last Exam, along with SWE-Bench Pro and Terminal-Bench 2.0. The main results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from Cyber CTFs and The Last Ones. Those evaluations use a different model set that overlaps only partly with the main experiment, so the six-model list does not describe their coverage.

The records are linked to AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation. The paper examines how results depend on the amount of inference-time compute and the evaluation protocol. EvalEval says its cards combine verified results, evaluation context and configuration details, while its shared format organizes benchmark, evaluation-run and model information. The announcement says publicly reported methods and findings are being made available where appropriate; it does not say that every AISI evaluation or underlying transcript is included.

One example in the paper concerns Humanity’s Last Exam. Its analysis tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased. That result illustrates why the conditions behind a score matter: the reported performance can vary with both inference compute and feedback.

At a glance
reportWhen: Release announced; no record-by-record…
The developmentAISI has begun sharing selected benchmark and cyber evaluation results through EvalEval’s Evaluation Cards, alongside its paper on inference-time compute and evaluation protocols.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Evaluation Conditions Matter

Benchmark scores are often cited as if results from different models can be compared directly. But protocol choices can change measured outcomes, including how much inference compute a model receives and whether it gets feedback between attempts. A score presented without those conditions may leave readers uncertain about what it measures.

Publishing results with setup information gives researchers and practitioners more material to inspect when comparing reported runs. That can support work in AI research, model development and policy, where evaluation results may inform judgments about advanced capabilities. The cards do not establish that one benchmark or protocol is best, and they do not settle disagreements between differently configured evaluations. Their stated contribution is to make some of the circumstances behind selected results easier to see.

From Shared Schema to AISI Records

The collaboration follows earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that infrastructure to publicly reported AISI methods and findings.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring together results with information about benchmarks and models. These efforts address a reporting problem: results published in different formats may omit details needed to interpret a run, while repeating expensive evaluations may not be feasible. The announcement does not provide a full history of the projects or specify how each contributed to the records now available.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

What the Published Records Cover

The announcement does not specify how many records or transcripts are available, which setup fields appear for every benchmark, or whether external researchers have independently reproduced the results. It says material is being shared where appropriate, so the release should not be treated as a complete archive of AISI’s evaluation work.

The model set for the two cyber evaluations is not enumerated, and the announcement gives no record-by-record release dates or process for resolving disagreements between results gathered under different protocols. Those gaps limit how precisely readers can assess the collection’s coverage and compare its records. It also remains unclear how consistently contributors will fill in the available fields as more organisations use the shared format.

Broader Use of EEE

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Its proposed next step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model.

No further release date or adoption milestone was specified. Wider use could make cross-study comparisons easier, but that will depend on contributors publishing records with consistent and sufficiently complete information. The extent of future AISI releases, including whether further records or transcripts will be added, has not been announced.

Key Questions

What has AISI released?

AISI is sharing selected evaluation results through EvalEval’s Evaluation Cards. The records include verification, evaluation context and configuration information, according to the announcement.

Which models are in the main benchmark results?

The main experiment covers Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The two cyber evaluations use a different, partly overlapping model set.

Why include evaluation setup details with scores?

Results can vary with conditions such as inference-time compute and feedback. Recording those conditions helps readers interpret what a score represents and compare it with other reported runs.

Does the release include every AISI evaluation?

No such claim was made. The announcement describes selected publicly reported methods and findings being shared where appropriate, and does not say that every evaluation or transcript is included.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Some Say GLM-5.3-Flash Offers Great Value For Budget AI Projects

GLM-5.3-Flash offers a high-performance, multimodal AI model at a fraction of typical costs, making it attractive for budget-conscious AI applications.

Xanadu And Mitsubishi Chemical Advance Quantum Computing Work On EUV Lithography

Xanadu and Mitsubishi Chemical are advancing quantum computing applications to improve EUV lithography, aiming to enhance semiconductor manufacturing precision.

Inside AI’s Mind: 12 Questions That Reveal Its Secrets

A detailed exploration of how AI models like ChatGPT work, answering 12 key questions about their functioning, learning, understanding, and limitations.

Japan banks eye Anthropic’s Mythos in gearing up cybersecurity drive

Japan’s three major banks are enhancing cybersecurity defenses with Anthropic’s Mythos AI, following concerns over AI-driven vulnerabilities in financial systems.