Synth
ResearchEvalsBlogDocs

Synth Research

Research at Synth

We study how language-model agents improve, operate across long horizons, and produce evidence that can be independently verified. We publish our open systems, evaluations, reports, and field notes.

Explore model evaluations →Open-source work ↗

01

Evaluations

Frozen benchmark contracts, comparable costs, and rollout-level evidence for understanding how models and agents actually behave.

GameBench Code Policies

New

Executable policies across frozen game environments, paired with source, replayable visual rollouts, and measured candidate results.

Open evaluation →

Craftax

Live

Long-horizon survival, crafting, exploration, and combat with cost curves and replayable terminal traces.

Open evaluation →

DungeonGrid

Pilot

Fog-of-war coordination across single- and multi-agent parties, measured against inference cost.

Open evaluation →

Shrimp Farm

Live

A 730-day operational benchmark with strict turn, budget, and output contracts.

Open evaluation →
Evaluation reports and methodology notes are being prepared for publication.

02

Optimization & Learning

Methods for improving prompts, policies, and long-horizon agents through measured search and replay.

Synthesis in progress

GEPA

Reflective prompt evolution, Pareto candidate coverage, and compute scaling across stable task contracts.

Active research

GELO

Archive-based optimization that returns to useful trajectory states before exploring the next frontier skill.

Report

GELO: Archive-Based Optimization for Long-Horizon Agents

Go-Explore Long-Horizon Optimizer (GELO) is a checkpoint-native optimizer for language agents in sparse, long-horizon environments. GELO maintains an archive of high-value trajectory states, opens scoped theme searches from those checkpoints, and promotes candidates only when replayed heldout evidence improves the archive. In Craftax, NetHack, and Crafter, GELO turns undifferentiated prompt mutation into targeted frontier search over skills such as furnace placement, coal and iron recovery, and dungeon coverage.

Jun 11, 2026

Report

Scaling Train Time Compute for Gepa

Scaling train-time compute for GEPA across public task containers — same-container comparisons, coverage curves, and proposer scaling evidence.

Jun 2, 2026

03

Systems & Open Source

Technical reports on the open and internal systems that make agent research reproducible, inspectable, and extensible.

Public building blocks from Synth Laboratories.

View the GitHub organization ↗
01GitHub

Cookbooks

Runnable GEPA task containers for Banking77, HotpotQA, MiniGrid, TBLite, and Crafter, plus reproducibility evidence and a hosted GELO submission guide.

Explore the repository ↗
02GitHub

Optimizers

An Apache-2.0 Rust optimization platform with a public GEPA engine and Python API, plus the SDK, CLI, and runbooks for hosted GELO jobs.

Explore the repository ↗
03GitHub

Containers

An MIT-licensed Python SDK and HTTP task contract for datasets, mutable prompt programs, rollouts, scoring, traces, checkpoints, and resume.

Explore the repository ↗

04

Field Notes

Working observations, negative results, reproductions, and early findings that are useful before they become formal reports.

Field Notes publications are in preparation.
© 2026 SynthWorkshopChangelogDocsBook a Demo