Benchmarks/Run record · protocol 12
z-ai/glm-5.3-flash max
One agent session on battle-royale@1.0.233, the 3 replays of its evaluated script, and every artifact that produced the number.
- Level time
- —No clear recorded
- Rank
- —Not ranked
- Status
- DNFMeasured by the evaluator
- Audit
- CLEANReviewed 2026-09-04
- Cost
- $0.14567 API calls
- Agent time
- 60m 01sZ.AI
Replays
No replay of this script cleared, so there is no time to rank.
Replay times
This run's replays on their own scale. The filled mark is the ranked one.
Lower is better
| Replay | Outcome | Level time | Wall |
|---|---|---|---|
| 1 | DNF | — | 126.4s |
| 2 | DNF | — | 321.8s |
| 3 | DNF | — | 123.0s |
The best script did not clear level one in 3 clean replays
Replay recording
The selected replay as it ran, uncropped at its recorded frame.
1024x1024
60 fps with audio. 13.3 MB.
Audit
Reviewed 2026-09-04 23:14:30 UTC against the evaluated best.ts alone.
Clean
Verdict · DNF
Clean. best.ts reads game phase, entities, positions, inventory, combat values, and raycasts to choose actions. Every app-affecting action uses Vitexec mouse or keyboard input; it does not call game mutation APIs, dispatch DOM events, alter protected source, or directly mutate game state.
Run record
Everything below is read from this run's own result and protocol records, not transcribed.
- Run id
- 20260904T220212Z-2a461276
- Status
- DNF · CLEAN
- Model cost
- $0.145
- API calls
- 67
- Exit
- cut off at the session cap
- Agent time
- 60m 01s
- Replay time
- 571.1s
- Requested provider
- z-ai
- Actual provider
- Z.AI
- Started
- 2026-09-04 22:02:26 UTC
- Finished
- 2026-09-04 23:13:28 UTC
- Harness config
- fceb710b6d161783b7cc0b1d278a6e25cee466a48debda4ee672fb7bcd64af5a
- Protocol digest
- 1c974addc099a9ece6b4078edfd59b2bb33b148d75eca99e8e1f6f261b96c97d
Token usage
As OpenRouter recorded them for this session. Context and generated tokens differ by two orders of magnitude, so each group is drawn on its own scale.
Separate scales
Context readWhat the session sent to the model.
- Input6,832,422
- Cache read6,628,544
- Cache write0
GeneratedWhat the model produced.
- Output120,950
- Reasoning97,385
Exact counts
- Cost, exactly as billed
- $0.14495650999999993
- Input
- 6,832,422
- Cache read
- 6,628,544
- Cache write
- 0
- Output
- 120,950
- Reasoning
- 97,385
Artifacts
Every file this run produced, checked in beside its result.
- Evaluated scriptbest.ts
- Auditaudit.md
- Replay outcomesattempts.json
- Result metadataresult.json
- Protocolprotocol.json
- Evaluatorevaluator.ts
- Rendered promptprompt.md
- Agent tracetrace.jsonl
- Agent trajectorytrajectory.json
- Replay 1 logreplay-1.log
- Replay 2 logreplay-2.log
- Replay 3 logreplay-3.log
- Run logrun.log
- Agent logagent.log