🔍 Read the full analysis: Three Ways UK AISI And EvalEval Approach Reproducible AI Benchmarks on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
The UK AI Security Institute is publishing selected results from five benchmarks across six frontier models through EvalEval’s Evaluation Cards, which pair scores with verification, context and configuration information. Two cyber evaluations are also included, using a different model set that overlaps only partly with the main group. The release is not described as a complete archive of AISI evaluations or transcripts.
The UK AI Security Institute (AISI) is publishing selected results from its evaluations of advanced AI models through EvalEval’s Evaluation Cards, adding verification, evaluation context and configuration information to help readers understand how scores were produced, as detailed in the original analysis. The release covers five benchmarks across six frontier models, as well as two cyber evaluations using a partly different model set. It accompanies AISI’s paper on how inference-time compute and evaluation protocols affect benchmark results.
The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath and Humanity’s Last Exam, along with SWE-Bench Pro and Terminal-Bench 2.0. The main results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from Cyber CTFs and The Last Ones. Those evaluations use a different model set that overlaps only partly with the main experiment, so the six-model list does not describe their coverage.
The records are linked to AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation. The paper examines how results depend on the amount of inference-time compute and the evaluation protocol. EvalEval says its cards combine verified results, evaluation context and configuration details, while its shared format organizes benchmark, evaluation-run and model information. The announcement says publicly reported methods and findings are being made available where appropriate; it does not say that every AISI evaluation or underlying transcript is included.
One example in the paper concerns Humanity’s Last Exam. Its analysis tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased. That result illustrates why the conditions behind a score matter: the reported performance can vary with both inference compute and feedback.
Why Evaluation Conditions Matter
Benchmark scores are often cited as if results from different models can be compared directly. But protocol choices can change measured outcomes, including how much inference compute a model receives and whether it gets feedback between attempts. A score presented without those conditions may leave readers uncertain about what it measures.
Publishing results with setup information gives researchers and practitioners more material to inspect when comparing reported runs. That can support work in AI research, model development and policy, where evaluation results may inform judgments about advanced capabilities. The cards do not establish that one benchmark or protocol is best, and they do not settle disagreements between differently configured evaluations. Their stated contribution is to make some of the circumstances behind selected results easier to see.
The collaboration follows earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that infrastructure to publicly reported AISI methods and findings.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring together results with information about benchmarks and models. These efforts address a reporting problem: results published in different formats may omit details needed to interpret a run, while repeating expensive evaluations may not be feasible. The announcement does not provide a full history of the projects or specify how each contributed to the records now available.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
What the Published Records Cover
The announcement does not specify how many records or transcripts are available, which setup fields appear for every benchmark, or whether external researchers have independently reproduced the results. It says material is being shared where appropriate, so the release should not be treated as a complete archive of AISI’s evaluation work.
The model set for the two cyber evaluations is not enumerated, and the announcement gives no record-by-record release dates or process for resolving disagreements between results gathered under different protocols. Those gaps limit how precisely readers can assess the collection’s coverage and compare its records. It also remains unclear how consistently contributors will fill in the available fields as more organisations use the shared format.
Broader Use of EEE
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Its proposed next step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model.
No further release date or adoption milestone was specified. Wider use could make cross-study comparisons easier, but that will depend on contributors publishing records with consistent and sufficiently complete information. The extent of future AISI releases, including whether further records or transcripts will be added, has not been announced.
Key Questions
What has AISI released?
AISI is sharing selected evaluation results through EvalEval’s Evaluation Cards. The records include verification, evaluation context and configuration information, according to the announcement.
Which models are in the main benchmark results?
The main experiment covers Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The two cyber evaluations use a different, partly overlapping model set.
Why include evaluation setup details with scores?
Results can vary with conditions such as inference-time compute and feedback. Recording those conditions helps readers interpret what a score represents and compare it with other reported runs.
Does the release include every AISI evaluation?
No such claim was made. The announcement describes selected publicly reported methods and findings being shared where appropriate, and does not say that every evaluation or transcript is included.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
