playbenchRepository

Benchmark for Agents playing 3D games

Each model receives a browser, a fixed time budget and a pinned game. Playbench publishes the measured completion time and every artifact behind it.

Battle Royale

Level completion time

Lower is better

3 of 4 cleared

Run results

4 runs

  1. 16.768s

    PASSclean$3.62849m 32sEvidence

  2. 19.466s+2.698s

    PASSclean$5.30543m 37sEvidence

  3. 41.781s+25.013s

    PASSclean$0.10552m 48sEvidence

  4. DNFclean$0.14560m 01sEvidence

Replay outcomes and session totals
ModelReplayOutcomeSecondsWall
Google gemini-3.8-flash high1 · rankedPASS16.768s36.1s
2PASS21.598s41.8s
3PASS18.857s39.4s
Meta muse-spark-1.3 xhigh1 · rankedPASS19.466s39.8s
2PASS25.099s45.6s
3PASS21.032s41.6s
DeepSeek deepseek-v4-flash-0731 max1DNF607.9s
2PASS62.232s87.0s
3 · rankedPASS41.781s63.5s
Z.ai glm-5.3-flash max1 · rankedDNF126.4s
2DNF321.8s
3DNF123.0s
ModelCostAgent timeTotal wallAPI callsOutput tokensExit
Google gemini-3.8-flash high$3.62849m 32s51m 29s171188,209Submitted
Meta muse-spark-1.3 xhigh$5.30543m 37s45m 44s7970,162Submitted
DeepSeek deepseek-v4-flash-0731 max$0.10552m 48s65m 27s60347,714Submitted
Z.ai glm-5.3-flash max$0.14560m 01s69m 33s67120,950cut off at the cap

Level time against session cost

Lower and left is better

Replay distribution

Lower is better

Scoring method

  1. Session

    Every model receives the same prompt and 60 minutes in a pinned container.

  2. Evaluated file

    Only vitexec/best.ts is scored. All other work in the session is retained as evidence.

  3. Replays

    The game is reset to a fixed state and the script is run 3 times from clean.

  4. Score

    A trusted evaluator reads the game’s own completion state and returns the time.

Every run on this page ran under the same pinned contract, and the build refuses to publish if any of it differed between them.

Method and verificationProtocol 12

Last published 2026-09-05