playbenchRepository

Protocol 12

Comparing runs is only fair if they ran under the same conditions. The build re-reads every run's own record and stops rather than publish on a claim it cannot prove.

Shared conditions

Changing any of it increments the protocol version, and runs from different versions are never ranked together.

Protocol

12 · 4 runs

Agent budget

60 minutes

Temperature

0

Replays scored

3, from clean state

Score

Fastest replay that cleared — never a median, never an average

Harness

mini-swe-agent with one Bash tool

Toolchain

vitexec 0.4.0 · mini-swe-agent 2.4.6 · market 0.11.0 · node v22.23.2

Reproduction values
Viewport
1024x1024
GPU
T4
Container image
nvidia/cuda:12.4.1-runtime-ubuntu22.04
Temperature
0
Replay limit
600s each
Recording
3 × 60 fps
Model routing
OpenRouter provider.only + provider.order + allow_fallbacks=false + require_parameters=true
Scoring
fastest successful of 3 clean replays; trusted evaluator levelTimeMs
Prompt sha256
60be77d61034
Vitexec skill sha256
b32236406b66

Per-game configuration

A game brings its own seed, route and evaluator. Everything else it inherits from the contract above.

Game configuration

1 pinned

GameSeedRouteEvaluator
battle-royale@1.0.23320260903/?seed=20260903&round=1af4bcad746fb…

Per-run fields

Reasoning effort, the provider that actually served the request, and the digests that include them are recorded on each run instead.

Per-run values

4 runs

ModelRequestedServed byHarness configProtocol digest
Google gemini-3.8-flash highgoogle-ai-studioGoogle AI Studiof94fb7b5f924…3cc6f142f6aa…
Meta muse-spark-1.3 xhighmetaMetad98ffc5daa09…d67e3bf840d2…
DeepSeek deepseek-v4-flash-0731 maxbaiduBaiduc9bf3ccfe2a4…acd31be9ad6a…
Z.ai glm-5.3-flash maxz-aiZ.AIfceb710b6d16…1c974addc099…

Results from different protocol versions are never ranked together. When a later protocol lands, its results appear as a separate set rather than being merged into this one.

bench.jsonResult directoriesAll benchmarks