Download the final stimuli

The 24 calibrated transcripts in five formats.

Modal difficulty tier × domain

Tier 1 (easiest)
Tier 2
Tier 3
Tier 4 (hardest)
Outcome
View

Development pipeline

Four stages over which the 24 final transcripts were shaped. See preprint for more detail.

Failure modes observed

Question sources & acknowledgments

Questions in the “brainteasers” domain (Internal Temperature of Stars, Switch, Bergs, Probability) all have a counterintuitive aspect, such that a dishonest debater can attempt to leverage an intuitive but incorrect argument. Switch and Bergs are variations on Thinking Physics questions (Epstein 2002) previously explored by Barnes et al. (2020), and Probability Puzzle is borrowed from Coelho (2024), which in turn restates a ‘paradox’ often attributed to Elchanan Mossel (Jin 2018). The coding questions (Rust, Prefix-sum, ORM library, Database deletion) ask whether an excerpt of a (sometimes substantially modified) agent transcript drawn from METR's MALT-public dataset (Parikh and Wijk 2025) contains reward hacking; the correct answer is “yes” in two of the four cases.

The astrophysics questions (Little Red Dots, Early Universe Galaxies, Large Linear Structure, Gravitational Lens Modelling) concern resolved or near-resolved controversies in astrophysics research. The debate in Gravitational Lens Modelling is over two competing explanations of the same astrophysical system (J1721+8842), where Honest defends the findings of Dux et al. (2025) while Dishonest defends an older, since disproven model of Mangat et al. (2021). We chose this case in part because co-author Mangat was familiar with the original model from his earlier work on this system. Similarly, Large Linear Structure pits an Honest debater's claim, confirmed by recent JWST evidence (van Dokkum et al. 2026), that an object suspected to be a ‘runaway’ supermassive black hole (van Dokkum et al. 2023) is in fact one, against Dishonest's defense of the bulgeless edge-on galaxy interpretation (Almeida et al. 2023; Montes et al. 2024). Little Red Dots concerns differing explanations of a population of compact red objects discovered by JWST at high redshift (Zhang et al. 2025; Nandal and Loeb 2026), with debaters also referencing facts from Inayoshi and Maiolino (2025) and Baumgarte and Shapiro (1999) in one transcript. Early Universe Galaxies concerns the degree to which the brightness of early-universe galaxies supports MOND-based modified-gravity models (McGaugh et al. 2024) over the dark-matter paradigm instantiated in ΛCDM (Aghanim et al. 2020); debaters also cite evidence from Nanayakkara et al. (2024), Ziegler et al. (2025), Zhao et al. (2008), and Clowe et al. (2006).