Open-ended agents in Craftax
Craftax
A long-horizon, open-ended world combining survival, crafting, exploration, combat, and dungeon progression. We evaluate whether language-model agents can turn its achievement ladder into durable capability.

Cost vs. Performance
Average distinct achievements against average API cost per terminal Craftax rollout. Lines connect each model's measured reasoning efforts in low → medium → high order; larger markers mean more effort.
Hover a point for its cohort detail. Terra and Luna retain their Jul 9 launch-price points and add Jul 30 price-cut snapshots using the same measured token usage.
Per-Achievement Breakdown
All 66 possible achievements, ordered by Craftax progression and annotated by capability family and world region. Each cell is based on 20 terminal rollouts.
Trajectories
Select a model and rollout, then inspect the complete Craftax world replay, policy reasoning, actions, survival state, and achievement timeline.
Craftax · sealed V5
deepseek/deepseek-v4-flash-0731 / xhigh / seed 7
trajectory evidence
Achievements over policy turns
The Craftax analogue of Voyager's capability curve: each stepped line is the mean number of distinct achievements across the selected real rollouts. A policy turn equals one in-game action. Crafting icons mark where each cohort first hits a ≥50% unlock rate; one callout per achievement lands on the earliest such unlock.
survival envelope
Where each run takes damage—and gives out
One row summarizes each model’s terminal cohort. Lines are median vitals across still-active runs; the health ribbon spans the middle 50%. Red density bars aggregate damage, and bottom ticks show the distribution of death times.
“Damage event” means a negative health delta between sealed engine snapshots. The trace does not yet attribute every loss to a specific attacker, so this view does not invent one.
Achievement capability tree
Direct engine gates and tier prerequisites across all 66 Craftax achievements. Choose a model to color each node by the share of real terminal runs that unlocked it.
Seed-adjusted comparison
Paired achievement effects
Pick two policies. For every achievement, we compare whether each policy unlocked it on the exact same starting seed. This shows where changing the policy alters the capability ladder—not merely which model has the higher aggregate score.
Achievement score
+2.10 per run
95% paired interval [-0.60, 4.80]
Survival
0 percentage points
candidate minus reference survival rate
Inference cost
+$0.117
mean cost difference per rollout
collect coal
+7 candidate-only · −1 reference-only · 6 both
+30 pp
65%
35%
defeat skeleton
+7 candidate-only · −1 reference-only · 3 both
+30 pp
50%
20%
place torch
+6 candidate-only · −0 reference-only · 0 both
+30 pp
30%
0%
place plant
+1 candidate-only · −6 reference-only · 0 both
-25 pp
5%
30%
make torch
+6 candidate-only · −1 reference-only · 0 both
+25 pp
30%
5%
defeat zombie
+8 candidate-only · −4 reference-only · 5 both
+20 pp
65%
45%
collect stone
+4 candidate-only · −0 reference-only · 16 both
+20 pp
100%
80%
collect iron
+6 candidate-only · −3 reference-only · 3 both
+15 pp
45%
30%
collect sapling
+4 candidate-only · −1 reference-only · 15 both
+15 pp
95%
80%
make stone sword
+6 candidate-only · −4 reference-only · 5 both
+10 pp
55%
45%
enter dungeon
+3 candidate-only · −1 reference-only · 0 both
+10 pp
15%
5%
eat snail
+2 candidate-only · −0 reference-only · 0 both
+10 pp
10%
0%
place stone
+3 candidate-only · −5 reference-only · 2 both
-10 pp
25%
35%
collect food
+4 candidate-only · −2 reference-only · 14 both
+10 pp
90%
80%
make wood sword
+1 candidate-only · −0 reference-only · 19 both
+5 pp
100%
95%
place furnace
+1 candidate-only · −2 reference-only · 0 both
-5 pp
5%
10%
find bow
+2 candidate-only · −1 reference-only · 0 both
+5 pp
10%
5%
open chest
+2 candidate-only · −1 reference-only · 0 both
+5 pp
10%
5%
wake up
+1 candidate-only · −0 reference-only · 0 both
+5 pp
5%
0%
make arrow
+1 candidate-only · −0 reference-only · 0 both
+5 pp
5%
0%
make iron sword
+1 candidate-only · −0 reference-only · 0 both
+5 pp
5%
0%
make diamond sword
+1 candidate-only · −0 reference-only · 0 both
+5 pp
5%
0%
collect drink
+3 candidate-only · −4 reference-only · 10 both
-5 pp
65%
70%
make stone pickaxe
+2 candidate-only · −3 reference-only · 11 both
-5 pp
65%
70%
eat cow
+2 candidate-only · −2 reference-only · 14 both
0 pp
80%
80%
drink potion
+1 candidate-only · −1 reference-only · 0 both
0 pp
5%
5%
These are paired treatment effects of policy choice on achievement outcomes over the observed seeds. They are not CATEs: credible conditional effects require more shared seeds and pre-rollout world features such as biome, nearby resources, and threat density.
What sets the best rollouts apart
The top 115 rollouts by achievement score (score ≥ 11.0) against the remaining 346, under the current filters. Both groups played the same engine, so a gap here is a difference in how the policy spent its turns.
What accompanies a high-scoring run
A direct comparison of the top 115 runs with the other 346. Because the groups are defined by achievement score, score itself is intentionally omitted. These are associated outcomes, not independent causes. Orange is the high-scoring group; gray is everyone else. Values are group averages, except survival, which is the share that finished without dying.
Survived to a cap
70%vs42%
+29 percentage points
Tool types crafted
3.77vs1.51
+2.26 per run
Resources gathered
10.5vs3.77
+6.73 per run
Engine steps taken
how long the run stayed active
Resources gathered
resource events reached
Tool types crafted
distinct crafting unlocks
Enemy types beaten
distinct combat unlocks
Survived to a cap
finished without dying
What stronger runs spent
Output tokens: 88,103 vs 34,674
Inference cost: $0.12 vs $0.04
These are costs of longer, more capable runs—not capabilities by themselves.
“Tool” and “enemy” counts are distinct achievement types reached, not repeated crafts or kills; the trace records unlocks, not lifetime action counters.
The best runs
The ten highest-scoring rollouts under the current filters.
Dungeon
Craftax level 0 is the overworld; every level above it is dungeon. This section counts only what happens below ground — descents, floor gates, and the achievements that cannot be earned on the surface.
Descent rate by cohort
Share of each cohort’s terminal rollouts that entered the dungeon at least once. Cohorts are listed with their call budget because a policy with more calls has more turns in which to find a way down.
Floor gates
The dungeon spine in descent order. Each gate must be passed before the next is reachable, so the first empty bar is where every policy stops.
Below-ground achievements
All 30 achievements that cannot be earned on the surface. Unreached entries are kept visible: the shape of what is missing is the result.
Dungeon entries
Every terminal rollout that got below the overworld, with the turn it descended.
Score design
The headline is achievement-ladder score: each distinct capability milestone counts once. The Cost vs. Performance chart can isolate one family or region at a time without changing the canonical score used elsewhere.
Policy-visible observations and evaluator-only evidence are kept separate. Engine-rendered frames are derived from the Rust gold lane and every rollout links to its SHA-256-backed artifact.