Save
Save created
Follow the scored route, inspect one model's effort settings, or open the individual runs. The harness and rules explain the toolkit behind every result.
Each scored gate is one ninth of route progress. G1–G3 cover the opening; six castle steps lead to G4, when the first Bowser battle begins. Times below belong to the lifelong-player reference; S1u is an estimate.
Save created
Toad Town
Castle entry
Castle hall
Castle upper hall
Castle stairs
Upper castle
Bowser approach
Bowser battle
The castle eats the clock
The lifelong-player human reference crossed G3 to G4 in 117.3 seconds. The fastest selected completion with exact castle timing needed 742.1 seconds.
6.327× human castle time
Nine scored steps
Arithmetic mean of each selected attempt's deepest scorer gate on the nine-step G1–G4 route.
Use this reference to compare effort cells within one model. Different harness conditions remain separate.
Choose one model. Every bar keeps its effort and condition attached; mixed conditions are never pooled.
G4 is not the whole story. GPT-5.6 Terra reaches the castle in 15 of 15 runs with a 12:20 median, so castle arrival is a given for it, yet only 5 of 15 finish. Gemini 3.6 Flash (8/9 to G3, median 10:20) reaches the castle consistently and never records a G4.
The selected field
Each mark is one attempt. Choose a model, then an effort, to inspect its runs and finish times.
The retained record contains 259 runs; 119 reached G4. The comparison includes 247 attempts, with 119 finishes; image-cap-blocked Muse Spark attempts are excluded. August and September 2026 remain separate collection periods, with harness and service differences retained.
The broader coverage map records 324 attempts across 24 model families and 85 cells. Its 19 marathon and side-experiment mappings stay separate from the comparison.
GPT-6 Astra medium
12.544× human
Gemini 3.8 Flash medium
One dial, three stories
The same low-to-max dial produced three different GPT-5.6 stories. Small cells, so treat them as stories, not proof.
high effort. Luna: 2 of 3 reached G4. Sol: 3 of 3 reached G4. Terra: 1 of 3 reached G4
Three shapes, one family
In the August runs, these Anthropic models ran the same dial and drew three different curves. Each card plots average route progress at each effort setting.
Every effort averages 100%. The dial moves the clock, not the depth.
Medium dips, high recovers, xhigh dips again. More effort is not more depth.
Perfect at low and medium; high drops to 74% and xhigh stays there.
Ninety minutes, sliced open
Frame time advances only while the game runs; the scored wall clock never stops. The harness page explains the pause machinery. The retained clocks cannot separate thinking from tooling from waiting, but they show how little of some runs was active game time.
Not inference latency. Time outside the naive frame clock mixes several systems and pause semantics. The scorer wall clock remains authoritative. The Kimi K3 example is a launch-window run that has since been re-run; it stays because no run rode the 90-minute wall closer.
The endurance field
The September six-hour runs add Astra and Fable 5.1 to the six earlier endurance records. Astra reached the first staircase in Koopa Bros. Fortress; Fable 5.1 reached the playground exit. Each bar follows ordered story checkpoints toward the first Star Spirit, with wall-clock time underneath. These checkpoints differ in difficulty and length. The segmented Fable 5 recovery lineage remains separate from single-session runs, and collection months stay visible.


Astra reached G4 at 25:04; Fable 5.1 at 24:35. Both had six hours at medium effort. By the end, Astra had lowered the first staircase inside Koopa Bros. Fortress. Fable's last new story milestone was leaving the playground, followed by continued combat and recovery attempts on Goomba Road.
Fable 5.1 wasn't the strongest Claude marathon result. An earlier six-hour Opus 5 run at medium effort reached the Toad Town story milestone, farther than Fable 5.1 managed here. Its trace from the August comparison provides a reference for the two September runs. The run details retain the differences in harness versions and automatic continuations.
For the two September runs, the retained usage gives API-equivalent estimates of ~$193 for Astra and $68.05 for Fable 5.1. Cache reads account for $168.51 of Astra's estimate. Fable's cheaper cache reads make a visible difference, although cache writes also contribute substantially to its total.
| Category | Estimate |
|---|---|
| Input | $0.00 |
| Cache writes | $16.34 |
| Cache reads | $168.51 |
| Output | $8.25 |
| Total | $193.09 |
| Category | Estimate |
|---|---|
| Input | $0.12 |
| Cache writes | $33.35 |
| Cache reads | $22.53 |
| Output | $12.05 |
| Total | $68.05 |
Astra’s input/write split is estimated: all non-cached input is priced as cache writes, leaving $0 in ordinary input. Its actual write count was not reliably reported. This uses the upper end of the $189.82–$193.09 estimate range.
Fable 5.1: Completed native receipts with model-specific prices and cache TTL, including recorded helper models. Excludes unfinished or unobserved tails and aggregate-only usage. Not subscription spend or an invoice.
The score stops at G4.
The endurance story starts there.