playbenchRepository

Benchmarks/Run record · protocol 12

google/gemini-3.8-flash high

One agent session on battle-royale@1.0.233, the 3 replays of its evaluated script, and every artifact that produced the number.

Level time
16.768sFastest clean replay
Rank
1 of 3Among cleared runs
Status
PASSMeasured by the evaluator
Audit
CLEANReviewed 2026-09-05
Cost
$3.628171 API calls
Agent time
49m 32sGoogle AI Studio

Replays

Replay 1 was the fastest that cleared, so it is the published score and the recording beside it is its video.

Replay times

This run's replays on their own scale. The filled mark is the ranked one.

Lower is better

ReplayOutcomeLevel timeWall
1rankedPASS16.768s36.1s
2PASS21.598s41.8s
3PASS18.857s39.4s

Trusted evaluator observed level-one completion; fastest of 3 clean replays selected

Replay recording

The selected replay as it ran, uncropped at its recorded frame.

1024x1024

60 fps with audio. 11.8 MB.

Audit

Reviewed 2026-09-05 00:57:50 UTC against the evaluated best.ts alone.

Clean

Verdict · PASS

best.ts only reads game, ECS, combat, DOM, camera, and raycast state to choose actions. Every app-affecting operation uses Vitexec mouse or keyboard input; it does not call game mutation APIs, dispatch synthetic DOM events, alter protected source, or directly mutate game state.

What the boundary is

Run record

Everything below is read from this run's own result and protocol records, not transcribed.

Run id
20260905T000349Z-74760d16
Status
PASS · CLEAN
Model cost
$3.628
API calls
171
Exit
Submitted
Agent time
49m 32s
Replay time
117.3s
Requested provider
google-ai-studio
Actual provider
Google AI Studio
Started
2026-09-05 00:03:58 UTC
Finished
2026-09-05 00:57:03 UTC
Harness config
f94fb7b5f924600fabebd4daade43c880cec60881fa470dc50a485dc71ae55a4
Protocol digest
3cc6f142f6aab5bc450867073b24bbd5a6f3160ca5152a579df365a0021bf440

Token usage

As OpenRouter recorded them for this session. Context and generated tokens differ by two orders of magnitude, so each group is drawn on its own scale.

Separate scales

Context readWhat the session sent to the model.

  • Input31,165,105
  • Cache read30,298,619
  • Cache write0

GeneratedWhat the model produced.

  • Output188,209
  • Reasoning93,474
Exact counts
Cost, exactly as billed
$3.628044675000001
Input
31,165,105
Cache read
30,298,619
Cache write
0
Output
188,209
Reasoning
93,474

Artifacts

Every file this run produced, checked in beside its result.

Every file in this run