Research Factory With Held-Out Craftax Proof
Research Factory adds typed research and maintenance cycles, immutable candidates, and benchmark-owned grading through synth-ai 0.15.0.
TL;DR
- Research Factory turns a bounded objective and budget into repeated research and typed maintenance cycles.
- Factories expose inspectable evidence and immutable candidates; the benchmark owner grades those candidates on held-out inputs.
- The release acceptance run passed all 12 FactoryBench lifecycle gates with
acceptance reward
1.0. - The immutable winner,
d218924a156cafb903cff486e1382c720b893656, received held-out mean reward1.16and benchmark score0.0303. - Python SDK and MCP support ship in
synth-ai[research]0.15.0.
What Shipped
Research Factory adds durable Efforts to Managed Research. Each Effort can launch research and maintenance run kinds, retain typed evidence and WorkProducts, preview scheduled wakes, and operate under explicit Factory budgets and active-run caps.
Candidate promotion remains outside the worker's authority. The Factory emits an immutable candidate identity; the benchmark owner checks it out and emits the held-out verdict.
pip install "synth-ai[research]==0.15.0"Evidence Boundary
The autonomous acceptance proof and the earlier local result are separate:
The local-harness delta is not an autonomous Factory result. This release does
not claim a 24/7 reliability window or Factory-to-pull-request code delivery.
The FactoryBench receipt spans about 59 minutes from Factory creation to
grading; its held-out sweep reports 0.328s. Neither evidence packet records a
total cost, so no cost claim is made.