playbenchRepository

Benchmarks/Run record · protocol 12

z-ai/glm-5.3-flash max

One agent session on battle-royale@1.0.233, the 3 replays of its evaluated script, and every artifact that produced the number.

Level time
No clear recorded
Rank
Not ranked
Status
DNFMeasured by the evaluator
Audit
CLEANReviewed 2026-09-04
Cost
$0.14567 API calls
Agent time
60m 01sZ.AI

Replays

No replay of this script cleared, so there is no time to rank.

Replay times

This run's replays on their own scale. The filled mark is the ranked one.

Lower is better

ReplayOutcomeLevel timeWall
1DNF126.4s
2DNF321.8s
3DNF123.0s

The best script did not clear level one in 3 clean replays

Replay recording

The selected replay as it ran, uncropped at its recorded frame.

1024x1024

60 fps with audio. 13.3 MB.

Audit

Reviewed 2026-09-04 23:14:30 UTC against the evaluated best.ts alone.

Clean

Verdict · DNF

Clean. best.ts reads game phase, entities, positions, inventory, combat values, and raycasts to choose actions. Every app-affecting action uses Vitexec mouse or keyboard input; it does not call game mutation APIs, dispatch DOM events, alter protected source, or directly mutate game state.

What the boundary is

Run record

Everything below is read from this run's own result and protocol records, not transcribed.

Run id
20260904T220212Z-2a461276
Status
DNF · CLEAN
Model cost
$0.145
API calls
67
Exit
cut off at the session cap
Agent time
60m 01s
Replay time
571.1s
Requested provider
z-ai
Actual provider
Z.AI
Started
2026-09-04 22:02:26 UTC
Finished
2026-09-04 23:13:28 UTC
Harness config
fceb710b6d161783b7cc0b1d278a6e25cee466a48debda4ee672fb7bcd64af5a
Protocol digest
1c974addc099a9ece6b4078edfd59b2bb33b148d75eca99e8e1f6f261b96c97d

Token usage

As OpenRouter recorded them for this session. Context and generated tokens differ by two orders of magnitude, so each group is drawn on its own scale.

Separate scales

Context readWhat the session sent to the model.

  • Input6,832,422
  • Cache read6,628,544
  • Cache write0

GeneratedWhat the model produced.

  • Output120,950
  • Reasoning97,385
Exact counts
Cost, exactly as billed
$0.14495650999999993
Input
6,832,422
Cache read
6,628,544
Cache write
0
Output
120,950
Reasoning
97,385

Artifacts

Every file this run produced, checked in beside its result.

Every file in this run