Benchmarks/Run record · protocol 12
google/gemini-3.8-flash high
One agent session on battle-royale@1.0.233, the 3 replays of its evaluated script, and every artifact that produced the number.
- Level time
- 16.768sFastest clean replay
- Rank
- 1 of 3Among cleared runs
- Status
- PASSMeasured by the evaluator
- Audit
- CLEANReviewed 2026-09-05
- Cost
- $3.628171 API calls
- Agent time
- 49m 32sGoogle AI Studio
Replays
Replay 1 was the fastest that cleared, so it is the published score and the recording beside it is its video.
Replay times
This run's replays on their own scale. The filled mark is the ranked one.
Lower is better
| Replay | Outcome | Level time | Wall |
|---|---|---|---|
| 1ranked | PASS | 16.768s | 36.1s |
| 2 | PASS | 21.598s | 41.8s |
| 3 | PASS | 18.857s | 39.4s |
Trusted evaluator observed level-one completion; fastest of 3 clean replays selected
Replay recording
The selected replay as it ran, uncropped at its recorded frame.
1024x1024
60 fps with audio. 11.8 MB.
Audit
Reviewed 2026-09-05 00:57:50 UTC against the evaluated best.ts alone.
Clean
Verdict · PASS
best.ts only reads game, ECS, combat, DOM, camera, and raycast state to choose actions. Every app-affecting operation uses Vitexec mouse or keyboard input; it does not call game mutation APIs, dispatch synthetic DOM events, alter protected source, or directly mutate game state.
Run record
Everything below is read from this run's own result and protocol records, not transcribed.
- Run id
- 20260905T000349Z-74760d16
- Status
- PASS · CLEAN
- Model cost
- $3.628
- API calls
- 171
- Exit
- Submitted
- Agent time
- 49m 32s
- Replay time
- 117.3s
- Requested provider
- google-ai-studio
- Actual provider
- Google AI Studio
- Started
- 2026-09-05 00:03:58 UTC
- Finished
- 2026-09-05 00:57:03 UTC
- Harness config
- f94fb7b5f924600fabebd4daade43c880cec60881fa470dc50a485dc71ae55a4
- Protocol digest
- 3cc6f142f6aab5bc450867073b24bbd5a6f3160ca5152a579df365a0021bf440
Token usage
As OpenRouter recorded them for this session. Context and generated tokens differ by two orders of magnitude, so each group is drawn on its own scale.
Separate scales
Context readWhat the session sent to the model.
- Input31,165,105
- Cache read30,298,619
- Cache write0
GeneratedWhat the model produced.
- Output188,209
- Reasoning93,474
Exact counts
- Cost, exactly as billed
- $3.628044675000001
- Input
- 31,165,105
- Cache read
- 30,298,619
- Cache write
- 0
- Output
- 188,209
- Reasoning
- 93,474
Artifacts
Every file this run produced, checked in beside its result.
- Evaluated scriptbest.ts
- Auditaudit.md
- Replay outcomesattempts.json
- Result metadataresult.json
- Protocolprotocol.json
- Evaluatorevaluator.ts
- Rendered promptprompt.md
- Agent tracetrace.jsonl
- Agent trajectorytrajectory.json
- Replay 1 logreplay-1.log
- Replay 2 logreplay-2.log
- Replay 3 logreplay-3.log
- Run logrun.log
- Agent logagent.log