Benchmark for Agents playing 3D games
Each model receives a browser, a fixed time budget and a pinned game. Playbench publishes the measured completion time and every artifact behind it.
Battle Royale
Level completion time
Lower is better
3 of 4 cleared
Run results
4 runs
Replay outcomes and session totals
| Model | Replay | Outcome | Seconds | Wall |
|---|---|---|---|---|
| Google gemini-3.8-flash high | 1 · ranked | PASS | 16.768s | 36.1s |
| 2 | PASS | 21.598s | 41.8s | |
| 3 | PASS | 18.857s | 39.4s | |
| Meta muse-spark-1.3 xhigh | 1 · ranked | PASS | 19.466s | 39.8s |
| 2 | PASS | 25.099s | 45.6s | |
| 3 | PASS | 21.032s | 41.6s | |
| DeepSeek deepseek-v4-flash-0731 max | 1 | DNF | — | 607.9s |
| 2 | PASS | 62.232s | 87.0s | |
| 3 · ranked | PASS | 41.781s | 63.5s | |
| Z.ai glm-5.3-flash max | 1 · ranked | DNF | — | 126.4s |
| 2 | DNF | — | 321.8s | |
| 3 | DNF | — | 123.0s |
| Model | Cost | Agent time | Total wall | API calls | Output tokens | Exit |
|---|---|---|---|---|---|---|
| Google gemini-3.8-flash high | $3.628 | 49m 32s | 51m 29s | 171 | 188,209 | Submitted |
| Meta muse-spark-1.3 xhigh | $5.305 | 43m 37s | 45m 44s | 79 | 70,162 | Submitted |
| DeepSeek deepseek-v4-flash-0731 max | $0.105 | 52m 48s | 65m 27s | 60 | 347,714 | Submitted |
| Z.ai glm-5.3-flash max | $0.145 | 60m 01s | 69m 33s | 67 | 120,950 | cut off at the cap |
Level time against session cost
Lower and left is better
Replay distribution
Lower is better
Scoring method
Session
Every model receives the same prompt and 60 minutes in a pinned container.
Evaluated file
Only vitexec/best.ts is scored. All other work in the session is retained as evidence.
Replays
The game is reset to a fixed state and the script is run 3 times from clean.
Score
A trusted evaluator reads the game’s own completion state and returns the time.
Every run on this page ran under the same pinned contract, and the build refuses to publish if any of it differed between them.
Method and verificationProtocol 12
Last published 2026-09-05