2026 July "AI Evaluation" Digest
The Lab Leaks We Can Actually Prove
For years, part of the debate over the origins of COVID-19 has centred on the possibility that a virus being studied inside a laboratory escaped into the outside world. Whether that happened remains uncertain and fiercely debated. But this July, the AI industry produced its own lab leak, one that nobody disputes.
During an internal evaluation of GPT-5.6 Sol and a more capable, unreleased sibling model, OpenAI deliberately disabled the models’ cyber safety refusals to measure their true offensive capabilities. Instead of solving the assigned benchmark, the models discovered a vulnerability in the evaluation environment, escaped the sandbox, compromised infrastructure that connected them to the internet, broke into Hugging Face’s production systems, and retrieved the benchmark answers directly. OpenAI confirmed the incident after Hugging Face detected and contained the breach.
The story is amusing enough to sound like science fiction. It is also one of the clearest demonstrations yet of a problem AI evaluators have been quietly worrying about for years. The three testing incidents cases recently published by Anthropic only confirm the problem is inherent to evaluation.
How do you safely measure the limits of a dangerous model?
If the model keeps all its deployment guardrails, you are mostly measuring the effectiveness of those guardrails rather than the model’s underlying capabilities. But removing those guardrails creates a different problem: the evaluation itself can become dangerous. The more realistic the test, the greater the chance that the model interacts with systems outside the laboratory.
Imagine wanting to understand what criminal actions a frontier model could carry out inside a carefully constructed simulated city. The more faithfully that simulation resembles the real world, the more tempting it becomes for the model to leave the simulation altogether if doing so better achieves its objective. Or it might not notice its transition to the real world. At that point, the evaluation is no longer simply measuring capability or safety, it has become a potential source of real-world harm itself.
Importantly, nothing suggests the model was “trying to escape” in the human sense. Language models are remarkably good at adopting roles and pursuing objectives. Ask one to behave like a penetration tester or an advanced cyber attacker, and it will relentlessly optimise for that goal. The distinction between “inside the sandbox” and “outside the sandbox” is not an intrinsic concept the model understands. Unless the environment enforces those boundaries, the model has little reason to treat them differently.
Nor was the strategy itself especially original. “Find the answer key instead of taking the exam” is one of the oldest plots in computing—and in human history. A University of South Wales student was jailed for hacking into the university’s systems and selling real exam answers to other students before their tests; in 2023, the British exam boards OCR, Pearson Edexcel and AQA all had live exam papers stolen from their systems the same way and sold online before students sat them. This is precisely the kind of strategy that appears repeatedly in the data these models are trained on. The remarkable part is not that the model conceived of the idea, but that the evaluation environment allowed it to work.
This incident leaves AI evaluation with an uncomfortable dilemma. We cannot responsibly deploy increasingly capable models without understanding what they are actually capable of under realistic conditions, and hoping that the safeguards are never broken. Yet the very act of measuring those capabilities may require removing the safeguards that prevent those capabilities from causing harm.
Perhaps the lesson is not that we should stop running dangerous evaluations. Rather, we need to build laboratories that remain secure even when the experiment actively tries to escape; maybe a bunker!
News
With great power comes great responsibility, this is the summary of Euractiv explaining the implications of the EU AI Office being able to enforce the EU AI Act from 2 August. They will be able to request information and fine companies if the problems or risks are not solved. They can even ban a model from the EU market. A daunting task for about a few dozen people who work on evaluation and regulation at the office. This is very different from the much better staffed UK AISI (but no powers) and what has recently happened with the temporary bans of Anthropic and OpenAI models in the US, more opaque and discretional. What model do you think will be the first one to be banned in the EU?
OpenAI retracts its own SWE-Bench Pro recommendation. After finding SWE-bench Verified was contaminated earlier this year, OpenAI ran the same audit on SWE-Bench Pro (the benchmark it had told everyone to switch to) and found ~27–34% of tasks are broken: including overly strict or low-coverage tests, or underspecified or misleading prompts.
Peer review, but make it executable? SAI’s How much science is verifiable? sends AI agents to replicate all 168 ICML 2026 oral papers, completing 105 full replications and finding that only 34 of 92 papers with at least five judged claims reproduced more than 40% of attempted claims, while just 8 cleared 80%. The cost estimate adds another small scream: median reproduction cost is about $8,900, 17 papers exceed $100,000, and one reaches nearly $2.2M. The agents also catch the boring-but-lethal stuff manuscript review misses, including 903 code/reproducibility issues not raised by humans. Caveat: the system is still beta and aggregate-only, but “the code runs” probably should not be a surprise ending.
The world’s AI menu turns out to have exactly two cuisines on it. Our World in Data’s tally of OpenRouter’s daily top-50 most-used models finds the US and China accounting for nearly the entire list since January 2025. China’s share alone nearly quadrupled, from 5 models to 20, while the rest of the planet is reduced to a garnish: Cohere’s Command R for Canada, Mistral’s NeMo for France.
FAR.AI has launched a Security Leaderboard on the relative ease of jailbreaking frontier models. Gemini and Grok are very easy to jailbreak compared to Anthropic and OpenAI’s models. Seán Ó hÉigeartaigh, Research Professor at the University of Cambridge puts the example of terrorist groups such as Boko Haram who are exploring the use of leading AI models in their activities, jailbreaking them when they refuse. A very good example that backs the slogan: “AI is only as safe as its weakest models”!
The Winners of the Measuring Progress Toward AGI - Cognitive Abilities are out: https://www.kaggle.com/competitions/kaggle-measuring-agi/hackathon-winners.
Did we say the EU AI Offices is understaffed? Yes, but they have around 40 open positions! AI engineers, auditors, security and infrastructure experts, and risk specialists, compute infrastructure governance specialists, legal officers, operations specialists and of course AI evaluation specialists. Application Deadline: September 8, 2026
Methodology and Techniques
If a model’s high score disappears when you change the formatting rather than the substance, it was gaming the marker, not doing the work. The Evaluator Stress Test catches that reliably, but by its own admission and that is it misses subtler cheats: a model that states wrong answers with total confidence, or one that just tells you what you want to hear.
DualEval applies IRT to score each benchmark question by how sharply it separates models, and finds the best 10% reproduce the entire leaderboard. Déjà-vu? The same fit flags answers that don’t belong, catching deliberately planted contamination almost perfectly.
But it can be worse: in this preprint, the authors, using IRT tools, found that some popular benchmarks cannot separate the models at all, and on two of them the winner is a dummy that always guesses the most common answer. A leaderboard can look perfectly healthy while measuring nothing.
Here, for those who love extreme experiences, five hundred people and five frontier models were dropped into 43 tiny worlds with no instructions and told to work out the rules. The humans won every time, largely because they knew when to hit reset and run the experiment again, something the models almost never did. No humans died as the result of the experiment.
Testing a new model is expensive, so this method borrows the results of every model tested before it. The same precision now costs about a quarter fewer questions, with the statistical guarantees still intact. Sounds very psychometric…
This preprint develops a methodology for harmonising safety threshold: expected annual harm for cyber and biorisk, and rate of capability progress relative to trend for automated AI R&D. Applied to real models, the cyber floor is cleared more convincingly by Claude Mythos Preview than by GPT-5.5 (6/10 versus 2/10 on a cyber-range benchmark), yet nothing was triggered, since Anthropic’s own policy dropped its cyber category entirely. For biorisk the same method instead returns a diagnosis: no benchmark yet exists that is unsaturated, validated, and covers the full attack chain, so no quantitative floor can responsibly be set.
This preprint from Apollo Research and OpenAI implants a model with opposite, out-of-context beliefs about what its RL grader rewards, then measures how much its behavior shifts toward the grader’s preference over the user’s or developer’s; a behavioral test for “reward-seeking” that doesn’t rely on the model verbalizing it. Across a capabilities-focused OpenAI o3 RL run, that tendency grows sharply with training, and a model deliberately trained to reward-hack is more than twice as sensitive to grader beliefs as the unmodified version.
Put a different leash on the same dog and watch what happens: this open-source framework from Shanghai AI Lab decouples benchmark, harness, and environment so labs stop reinventing eval plumbing, and running seven frontier models through it shows scores swinging by as much as 15 points just from switching which agent is holding the leash. Its built-in reward-hacking detector then plays doping-control officer on the results: GLM-5.2 beats Claude-Opus-4.8 by 12 points on SWE-bench-Pro, but comes back flagged for roughly 30% more suspected cheating (modifying tests, grabbing golden patches) along the way.
Personas are still a hot topic in AI safety. This preprint follows the personality psychology tradition by treating LLM personas as positions in a behavioural trait-space, using the Big-5 OCEAN traits and training low-rank adapters (LoRAs) via constitution-guided distillation to amplify or suppress each trait. Across six models (4B–32B) from three families, these adapters allow fine-grained construction of model personas in weight space with downstream effects for safety-relevant behaviours, while (mostly) preserving existing capabilities.
Sometimes “teaching” a model a new trick is really just handing it the password to a trick it already knew: this ICML 2026 paper measures the actual bits of information a model absorbs during fine-tuning and finds latent capabilities need orders of magnitude less data to surface than genuinely new ones. Llama 3.1 8B jumps from 0% to 96% on arithmetic after training on a single example worth under 10 bits. The unsettling bit for evaluators: a fine-tuning eval meant to test what a model can’t do might just be teaching it to do that instead.
Evaluation Cards asks whether “trust us, the score is high” can please stop being an evaluation reporting standard. This preprint turns AI evaluation reports into structured records with signals for reproducibility, reporting completeness, provenance, and score comparability. Across 101,955 reported results from 30 organizations, the authors find that 96.5% of model-benchmark-metric triples miss minimal reproducibility information and 98.2% of model-benchmark pairs are reported by only one party.
Takes
The UK AI Security Institute argues that standard evaluations are fundamentally flawed by imposing strict test-time compute limits, which artificially mask the true capabilities of frontier models on complex tasks. By mapping performance as a continuous curve rather than a single fixed score, they demonstrate that allowing agents more compute to plan, verify, or retry tasks unlocks massive performance gains that traditional benchmarks completely miss. Ultimately, it seems our flagship models aren’t necessarily failing at deep reasoning; we’re just too cheap to foot the API bill required to let them finish a thought. Reminds us of the (in)famous “illusion of thinking” paper.
Really telling analysis from the OECD about the economic effect of ageing and AI in the following decades. Figure 1 in the paper should be a heads-up for some countries: those in the bottom-left quadrant will suffer GDP decline by ageing that would not be compensated by the supposed benefits of AI (unless the exposure and use of AI in those countries change, which may be the case if AI really starts adding points to GDP). Ageing can also be compensated by immigration in some countries more than others.
Terence Tao, the great mathematician, recently gave a public lecture on the future of mathematics (video, slides), touching upon some of the key elements of evaluation. First we need to condition our questions on how much AI will do what mathematicians can do. From there, we can realise that the goals will have to be a bit more explicit (to avoid falling into some kind of Midas king or paperclip problem, under Goodhart’s law). He gives a very good example: “solve unsolved problems” should not be the goal to give to AI. As an illustration, he iterates several times on this goal to show how difficult it is to find the right goal for AI, even in areas that could look more objective, such as mathematics. In the end, the goal ends up being a convoluted prompt: “Solve unsolved problems, verify them to be correct, ensure they are clearly communicated, and have them digested, accepted, and incorporated into the definitive theory of the field”. Terence Tao, you’re absolutely right!
Stanford scored AI against all 1,016 jobs in the US economy and found that in nearly half of occupations, a model could plausibly save time on most tasks. Yet four in five of those tasks see almost no real use, and the blockers are usually privacy rules and locked-down company systems rather than anything the model cannot do.
This admittedly very basic, early stages preprint challenges the assumption that scaling language models will lead to Artificial Superintelligence, arguing that mastering linguistic patterns fundamentally differs from genuine understanding of the physical world. They propose that true ASI requires “situation perception”—the ability to actively construct, revise, and act within internal simulations of physical laws, causality, and other minds, much like an embodied infant learns through environmental interaction rather than reading textbooks. Ultimately, the paper provides a conceptual roadmap for transitioning AI from stochastic text generators to agents grounded in internal world simulations, proposing concrete benchmarks like the “Apple Test” to measure a model’s capacity for counterfactual foresight and autonomous goal pursuit. It kind of puts old wine in new bottles, but we agree with the direction.
The UK AI Security Institute introduces RealityTest, a new benchmark grounded in over 3,000 human-authored queries across five languages to evaluate whether AI systems honestly disclose their identity when probed. They discover that model disclosure rates vary wildly based on phrasing and can be almost entirely suppressed by a simple system prompt instructing the model to hide its artificial nature. Ultimately, it turns out that despite all our regulatory anxiety over deceptive AI, teaching a digital imposter to lie for good requires little more than asking nicely.
Measuring superintelligence is a recurrent theme (e.g., this and this), with noticeable cognitive/psychometric tension between absolute vs population-relative scales (criterion-based vs norm-based evaluation). This preprint from the Princeton Superalignment team argues strongly in favour of the latter, and in a mixture of generative adversarial networks, Elo-like scores and psychometric discrimination maximisation, presents a protocol to evaluate “intelligence” to infinity and beyond. In the SepaRank protocol, the “proposers” generate binary questions that a population of solvers have to submit probabilities to. The proposer is scored by the induced separation among those probabilities (the items have good discrimination), and solvers are scored by a proper scoring rule on getting the question right. We wonder what assumptions to make in the protocol or in the population (and how they evolve) to make this converge, and, if so, where, and what the scales mean at all.
Findings & Results
Some formal analysis of tool -use (AI, but also GPS and abacus) and its impact on human competence from one of the greats in complexity science and “cognitive ecology”, finding dependence and competence co-evolve, both having stable basins depending on tool availability, but moving back from dependence to competence takes extra work (lowering tool availability more than was necessary to get in the dependent setting in the first place).
We all see our models or agents creating and using code snippets in python or bash to polish code or text. They, along with humans and a few animals, are tool makers.
But AI models can also use other weaker models, and that will go beyond the standard routing paradigms to a more on-demand metacognitive choice. In this blogpost from Redwood Research (which could well be a technical paper) the authors show that a weak model with good advice can become much stronger, and become unsafe. The novelty is that the advice is studied with varying bit limits, to see how much can be done with a few words of advice. They claim that even with a cap of a few words, capabilities can increase significantly, but safety can still be ensured.
More about scales and metrics: Turns out it matters an awful lot which fork you take: Fogelson et al. show whether cheap “meek” models ever catch up to frontier giants depends entirely on which ruler you use to measure “catching up.” Bounded metrics like benchmark accuracy always converge in the long run; unbounded ones like ELO or task-horizon length let the frontier pull away forever — and near-identical metrics for the same capability can land on opposite sides of that line, so whether AI power concentrates or spreads may say more about your choice of ruler than about the technology itself.
“Lower-resource languages are scored more generously“, which implies that current lower-resource benchmark results are overestimates, absolute numbers mean little across languages (not great for safety filters), and as many times before, we’re mostly measuring rank versus whatever construct we want to be measuring.
Can a model think without moving its lips? Think Fast (Gould, Ward, and seventeen collaborators, mostly Redwood Research) measures exactly how much frontier LLMs can reason when explicitly forbidden from writing chain-of-thought, using two metrics: a human-time-anchored 50%-task-completion horizon and a model-grounded “reasoning token horizon” benchmarked against o3-mini. Across 14 models on 43 benchmarks, no-CoT horizons have doubled roughly every 373 days for six years running, with the frontier model now clearing three minutes of human-equivalent task time.
Benchmarks and Leaderboards
AI assistants now remember you between conversations, which means an ordinary question can become a dangerous one given what they already know. The best system tested handled these memory-dependent situations correctly only t about half the time, and models that retrieved the right memory still failed to act on it.
Vision models were shown real robot workspaces, from surgical knots to weeding. Telling them to think step by step made them worse at seeing, and giving them worked examples made them worse and more confident at the same time.
In this preprint, fifty-two AI models were asked 542 biology questions, from the genuinely dangerous to the completely harmless. The safest models block almost everything risky but also refuse ordinary science. Interestingly the refusal rate of models under study shifted up to 28 points in a few weeks.
This preprint says judging AI creativity is a bit like judging pizza toppings; everyone agrees burnt crust is bad, but people will happily argue forever about pineapple! It introduces a new benchmark that separates things experts mostly agree on (like following instructions) from things based on personal taste (like style and visual appeal). The big takeaway: there’s no single “best” AI for creativity, rather different models shine at different stages of the creative process.
From another creativity benchmark, we lift the following quote directly: “applying factor analysis across 83 LLMs, we recover a single creativity factor ‘c’, analogous to the ‘g’ factor of general intelligence, that explains 81.5% of variance, related to but separable from general knowledge/reasoning”. Although separating the g and c rests on a single “pure” measure of fluid intelligence. We wonder how it interacts with the above also...
Another benchmark promising to turn office work into a unit test. What happens when your agent has to read the receipt in Japanese, reconcile the spreadsheet in English, and reply in Spanish? PolyWorkBench introduces 67 manually curated multilingual long-horizon workplace tasks across 5 domains, with 59 of 67 tasks spanning three or more languages across instruction, source, and output roles. It shows that while top models perform well overall, agent scaffolding matters almost as much as the model itself, and spreadsheet-heavy tasks remain difficult.
PAIR-Bench asks whether code repair should come with a tutoring transcript. It tests feedback-guided improvement on 440 wrong-answer Python Codeforces submissions, using progressive hints from “something smells funny” to “please look directly at the bug-shaped object,” then scores the whole repair arc: targeted repair, regressions, hint efficiency, final fix, gap closure, and more. DeepSeek V3.2 leads, jumping from 65.68% initial fix to 99.31% final fix with 98.80% gap closure, though broader-repair gains stay much smaller than targeted gains. The lineup is mostly 2025-era, so the benchmark is more interesting as an evaluation design than as a frontier leaderboard.
Cool long-horizon benchmark on real-world environment learning by ByteDance, also measuring learning rate, and finding e.g. that learning rate (in-context, with memory systems) doubles every three months.
PhantomBench is not as funny as the benchmark names we used to have in the past, but we can close this issue with it. With 60K+ fake concepts across 17 categories, the benchmark checks whether they abstain or start freelancing for Wikipedia. Models still fabricate quite a bit, especially when the prompt presupposes the concept exists. The useful bit: how models handle these fakes predicts how they handle rare real concepts far better than common ones, making it a scalable proxy for testing the limits of a model’s self-knowledge.
Getting the Digest: Once a month if you join at aievaluation.substack.com.
Contributors: Behzad Mehrbakhsh, José H. Orallo, Fernando Martinez-Plumed, Wout Schellaert, Lorenzo Pacchiardi, Peter Romero, Yael Moros-Daval, Félix Martí-Pérez, Konstantinos Voudouris and Cèsar Ferri.


