Synth AI
WorkshopBlogDocsStackSign in
WorkshopBlogDocsStackSign in
← Synth Research/Evals

Synth · evals

Model evaluations

Agentic benchmarks with frozen contracts, identical seeds, and content-addressed rollout evidence. Every published number is backed by a terminal rollout artifact you can replay.

A measured Rogue code-policy rollout rendered from GameBench evidence
Rogue · top code-policy rollout · seed 101

GameBench Code Policies

new

Executable policies across frozen game environments. Compare candidate source, replay visual rollouts, and measured results across Craftax, Rogue, and DungeonGrid.

View benchmark →
Replay of a measured Craftax code-policy rollout rendered by the GameBench gold engine
Craftax · top code-policy rollout · seed 101

Craftax

v0.1 pilot

Models control a single survivor or 1-, 2-, and 4-agent shared worlds in the GameBench Craftax gold lane. Tracks achievement progress, survival behavior, and cost against agent count.

View benchmark →
Crystal Red Shrimp — the caridina strain the benchmark colonies are graded on
Crystal Red Shrimp, the benchmark's flagship strain

Shrimp Farm Eval

live

Models run a freshwater shrimp farm over 730 simulated days under strict turn, budget, and output contracts. Independently hosted sibling benchmark.

shrimpfarmeval.com ↗
© 2026 SynthWorkshopChangelogDocsBook a Demo