Benchmarks/Run record · protocol 12
meta/muse-spark-1.3 xhigh
One agent session on battle-royale@1.0.233, the 3 replays of its evaluated script, and every artifact that produced the number.
- Level time
- 19.466sFastest clean replay
- Rank
- 2 of 3Among cleared runs
- Status
- PASSMeasured by the evaluator
- Audit
- CLEANReviewed 2026-09-05
- Cost
- $5.30579 API calls
- Agent time
- 43m 37sMeta
Replays
Replay 1 was the fastest that cleared, so it is the published score and the recording beside it is its video.
Replay times
This run's replays on their own scale. The filled mark is the ranked one.
Lower is better
| Replay | Outcome | Level time | Wall |
|---|---|---|---|
| 1ranked | PASS | 19.466s | 39.8s |
| 2 | PASS | 25.099s | 45.6s |
| 3 | PASS | 21.032s | 41.6s |
Trusted evaluator observed level-one completion; fastest of 3 clean replays selected
Replay recording
The selected replay as it ran, uncropped at its recorded frame.
1024x1024
60 fps with audio. 14.8 MB.
Audit
Reviewed 2026-09-05 00:03:42 UTC against the evaluated best.ts alone.
Clean
Verdict · PASS
best.ts only reads game and DOM state to choose actions. Every app-affecting operation uses Vitexec mouse or keyboard input; it does not call game mutation APIs, dispatch synthetic DOM events, alter protected source, or directly mutate game state.
Run record
Everything below is read from this run's own result and protocol records, not transcribed.
- Run id
- 20260904T231542Z-f4cbfd1b
- Status
- PASS · CLEAN
- Model cost
- $5.305
- API calls
- 79
- Exit
- Submitted
- Agent time
- 43m 37s
- Replay time
- 126.9s
- Requested provider
- meta
- Actual provider
- Meta
- Started
- 2026-09-04 23:15:54 UTC
- Finished
- 2026-09-05 00:02:53 UTC
- Harness config
- d98ffc5daa099621fbdf44ad7965fed3ebc458ad244b36c90dc518d64b8b8f38
- Protocol digest
- d67e3bf840d2077395411eb7f80aed1c46af479ca4cc9f70596eae0011332541
Token usage
As OpenRouter recorded them for this session. Context and generated tokens differ by two orders of magnitude, so each group is drawn on its own scale.
Separate scales
Context readWhat the session sent to the model.
- Input8,689,967
- Cache read5,323,448
- Cache write0
GeneratedWhat the model produced.
- Output70,162
- Reasoning31,575
Exact counts
- Cost, exactly as billed
- $5.304854449999998
- Input
- 8,689,967
- Cache read
- 5,323,448
- Cache write
- 0
- Output
- 70,162
- Reasoning
- 31,575
Artifacts
Every file this run produced, checked in beside its result.
- Evaluated scriptbest.ts
- Auditaudit.md
- Replay outcomesattempts.json
- Result metadataresult.json
- Protocolprotocol.json
- Evaluatorevaluator.ts
- Rendered promptprompt.md
- Agent tracetrace.jsonl
- Agent trajectorytrajectory.json
- Replay 1 logreplay-1.log
- Replay 2 logreplay-2.log
- Replay 3 logreplay-3.log
- Run logrun.log
- Agent logagent.log