A leaderboard of agency and blind-spot results, computed by a private methodology.
a measure that becomes a target ceases to be a good measure
Most benchmarks give every model one score and stop there. That single average hides the two things that matter most: whether a model spends its own compute wisely, and where it is confidently wrong.
ModelBench measures both directly, in one currency — prediction error, how surprised a model is by what actually came next.
Agency — given control over its own compute, does a model allocate it where it pays off? A higher agency score means the model's own choices beat a flat, even allocation of the same budget.
Blind spots — where does a model confidently predict the wrong thing? These aren't random errors. They cluster into patterns, and a leaderboard that only reports an average never shows them.
Scores here are facts. How they're computed is our method — and that part stays ours.