BlazaraelSubscribe →
MACHINE STATEISSUE Nº 0014 MIN READ

Hands on the Desktop

The benchmark that matters in 2026 is not a quiz. It is whether a model can operate your computer — and the gap between labs is wide.


The quiet spec in July's release wave is operational, not cognitive. Claude Opus 5 completes 70.57% of OSWorld 2.0 — a benchmark of real tasks in a live operating system: files, browsers, settings, applications [1]. GPT-5.6 Sol posts 90.4 on BrowseComp, a web-navigation suite [1]. The labs are no longer only measuring what models know. They are measuring what models can do with a mouse.

For anyone building on these models, the practical read is simple: evaluate on your own workflows, not on leaderboards. A model that tops a math suite can still fumble a legacy CRM. The benchmarks to watch this year are OSWorld, SWE-bench Pro and their successors — the ones with hands.

SOURCES — PRIMARY OR IT DOESN'T RUN[1] Claude Opus 5 — specs and benchmarks (incl. GPT-5.6 Sol figures)