Private model evaluation

Last run

SealedBench

A private, continuously rotated evaluation of frontier model capability. Tests, scoring weights, and answer keys remain sealed to resist benchmark-specific training.

Model capability index
Composite of capability, token efficiency, speed, and cost · Higher is better

* Thinking set to highest available