On the estimation and validity of AI time horizons---a statistical look at the METR plot
A new preprint argues that estimates depend on how task difficulty is modeled.
TL;DR
- The authors re-estimate METR-style task horizons using spline and item-response models over 228 tasks and 26 AI systems. [1]
- Their fitted relationship between human task time and AI difficulty is nearly flat from 2–30 minutes and closer to linear elsewhere. [1]
- They recommend reading horizon estimates alongside diagnostic plots; the work is a preprint, not a settled revision of METR’s measure. [1,2]
A preprint by Drew T. Nguyen and William Fithian recomputes task-time horizons for 26 AI systems on 228 tasks, using spline and item-response models to relax the assumption that task difficulty changes linearly with the logarithm of human completion time. [1] [1]
The authors report that their fitted curve is nearly flat across 2–30-minute tasks and close to linear elsewhere. They argue that the same multiplicative increase in a headline horizon can represent different changes in task difficulty, and recommend interpreting estimates with diagnostic plots. [1] [1]
METR describes its horizon as a 50%-success threshold on tasks measured by human completion time and notes that its suite is concentrated in software engineering, machine learning and cybersecurity. [2] [2]
Why it matters
Task-horizon figures can look like a single capability scale, but their interpretation depends on task mix and statistical assumptions. The preprint makes those assumptions more visible without establishing a new consensus measure.
Editor's note
The statistical critique is an unreviewed preprint; METR’s own documentation is included for the definition and scope of its benchmark.