GameBench Code Policies
NewExecutable policies across frozen game environments, paired with source, replayable visual rollouts, and measured candidate results.
Open evaluation →Synth Research
We study how language-model agents improve, operate across long horizons, and produce evidence that can be independently verified. We publish our open systems, evaluations, reports, and field notes.
01
Frozen benchmark contracts, comparable costs, and rollout-level evidence for understanding how models and agents actually behave.
Executable policies across frozen game environments, paired with source, replayable visual rollouts, and measured candidate results.
Open evaluation →Long-horizon survival, crafting, exploration, and combat with cost curves and replayable terminal traces.
Open evaluation →Fog-of-war coordination across single- and multi-agent parties, measured against inference cost.
Open evaluation →A 730-day operational benchmark with strict turn, budget, and output contracts.
Open evaluation →02
Methods for improving prompts, policies, and long-horizon agents through measured search and replay.
Synthesis in progress
Reflective prompt evolution, Pareto candidate coverage, and compute scaling across stable task contracts.
Active research
Archive-based optimization that returns to useful trajectory states before exploring the next frontier skill.
Report
Go-Explore Long-Horizon Optimizer (GELO) is a checkpoint-native optimizer for language agents in sparse, long-horizon environments. GELO maintains an archive of high-value trajectory states, opens scoped theme searches from those checkpoints, and promotes candidates only when replayed heldout evidence improves the archive. In Craftax, NetHack, and Crafter, GELO turns undifferentiated prompt mutation into targeted frontier search over skills such as furnace placement, coal and iron recovery, and dungeon coverage.
Jun 11, 2026
Report
Scaling train-time compute for GEPA across public task containers — same-container comparisons, coverage curves, and proposer scaling evidence.
Jun 2, 2026
03
Technical reports on the open and internal systems that make agent research reproducible, inspectable, and extensible.
Public building blocks from Synth Laboratories.
View the GitHub organization ↗Runnable GEPA task containers for Banking77, HotpotQA, MiniGrid, TBLite, and Crafter, plus reproducibility evidence and a hosted GELO submission guide.
Explore the repository ↗An Apache-2.0 Rust optimization platform with a public GEPA engine and Python API, plus the SDK, CLI, and runbooks for hosted GELO jobs.
Explore the repository ↗An MIT-licensed Python SDK and HTTP task contract for datasets, mutable prompt programs, rollouts, scoring, traces, checkpoints, and resume.
Explore the repository ↗04
Working observations, negative results, reproductions, and early findings that are useful before they become formal reports.