Domains

Contract 0.1, issued in publication v0.3.

A claim that fits none of these domains is out of scope, not a seventh domain. A domain is a place where the binding constraint is different, so a gain in one is not evidence about the others. Depth-by-breadth ladders rank systems. This page ranks gaps.

Published note, v0.3: the software-task instrument named below later failed the opening threshold. No forecast was opened. The text of the domain is unchanged.

1. Software tasks

Question. How long a software task, in skilled human time, a frontier agent completes at a stated reliability.

Boundary. Implementing a known task is software. Choosing the scientific question is discovery. A change used in a later training run is self-improvement.

Scored instrument. METR 50 percent and 80 percent time horizon, methodology version named in the forecast. Time Horizon 1.0 and 1.1 are not the same series.

Not the instrument. SWE-bench or a coding-leaderboard score.

2. Scientific discovery

Question. Can a system produce a novel, checkable claim about the world that was not in the prompt, and survive a pre-registered check.

Scored instrument. None in version 0.1. Forecasts do not ship.

Not the instrument. Citation count, co-authorship, or a lab blog post.

3. Long-horizon agency

Question. Can a system keep a goal, recover from its own errors, and finish work whose value depends on order, where a human would no longer be in the loop.

Scored instrument. None in version 0.1. A long software task with a unit test stays in domain 1.

Not the instrument. Context-window length, or a demo left running overnight.

4. Physical manipulation

Question. Can a system finish a stated physical task on real objects, without a human reset, at a stated reliability.

Scored instrument. None in version 0.1. A simulation success is not this domain.

Not the instrument. A demo video, or a simulation the lab also trained on.

5. Institutional coordination

Question. How long after a capability is measured does it become normal inside organizations that did not build the model.

Scored instrument. None in version 0.1. This domain does not measure the model.

Not the instrument. Download counts, API revenue, or a survey that asks whether a firm “uses AI.”

6. Self-improvement

Question. What fraction of the work that improves frontier models is done by the models, and does that fraction shorten the time to the next gain.

Scored instrument. None in version 0.1. Editor autocomplete is software. The loop must close into a later training run.

Not the instrument. A demo of a model editing its own prompt.

Only domain 1 had a candidate instrument. It did not pass the opening threshold. Empty instrument lines are a finding.

Depends on Claim classes.