Hands on the Desktop
The benchmark that matters in 2026 is not a quiz. It is whether a model can operate your computer — and the gap between labs is wide.
The quiet spec in July's release wave is operational, not cognitive. Claude Opus 5 completes 70.57% of OSWorld 2.0 — a benchmark of real tasks in a live operating system: files, browsers, settings, applications [1]. GPT-5.6 Sol posts 90.4 on BrowseComp, a web-navigation suite [1]. The labs are no longer only measuring what models know. They are measuring what models can do with a mouse.
For anyone building on these models, the practical read is simple: evaluate on your own workflows, not on leaderboards. A model that tops a math suite can still fumble a legacy CRM. The benchmarks to watch this year are OSWorld, SWE-bench Pro and their successors — the ones with hands.