MacArena: Benchmarking Computer Use Agents on an Online macOS Environment
MacArena is a new benchmark for computer-use agents on macOS, with 421 verified tasks across 50 apps. The paper finds current models can rank differently on native macOS tasks than on ported ones.
Intelligence analysis by GPT-5.4 Mini

The paper argues that macOS needs its own benchmark because existing evaluations miss important GUI challenges. MacArena mixes adapted tasks, macOSWorld content, and new native tasks, then shows that performance and rankings can change sharply on Apple Silicon macOS.
MacArena is like a tougher driving test for AI that uses a computer screen and mouse. The paper says some AI helpers look good on one kind of computer, but may struggle when the computer system changes to macOS.
Analysis
What MacArena adds
The paper introduces MacArena, a benchmark for computer-use agents running in an online macOS environment. It includes 421 manually verified tasks across 50 applications, combining three sources: a curated port of OSWorld tasks, content taken from macOSWorld, and 49 newly created macOS-native tasks.
Why the authors built it
The authors say macOS is still underrepresented in this area. They argue that the only prior benchmark they cite for macOS, macOSWorld, covers a limited set of first-party apps, focuses on simpler tasks, and runs on x86 virtual machines that do not match Apple Silicon. MacArena is meant to close that gap by using Apple’s native Virtualization framework on Apple Silicon.
What the evaluation showed
Their results suggest that desktop-agent performance is not portable by default. The paper says strong scores on existing benchmarks may reflect familiarity with those task distributions rather than general GUI competence. It also reports that model rankings can flip between ported tasks and macOS-native tasks. On the MacArena subset, the leading model trails by more than 26%, which the authors present as evidence that macOS is a harder environment for current GUI agents.
Takeaway
The main contribution is not just a new benchmark, but a warning about overreading benchmark wins. If a model is only tested in one desktop setting, its apparent strength may not carry over to a different operating system with different interface patterns and constraints.
Key points
- MacArena benchmarks computer-use agents on an online macOS environment.
- The dataset contains 421 manually verified tasks across 50 applications.
- It combines ported OSWorld tasks, macOSWorld content, and 49 new macOS-native tasks.
- The paper says macOS exposes GUI challenges that Linux-based benchmarks may miss.
- Model rankings can invert on MacArena, and one leading model falls by over 26% on the macOS-native subset.
If MacArena is adopted widely, it could push computer-use agents toward more realistic testing and better macOS support. That would make benchmark scores more meaningful for real desktop use.
If teams keep optimizing only for older benchmarks, models may continue to look stronger than they really are on macOS. The paper’s ranking reversals suggest progress on one platform may not transfer cleanly to another.



