Journal · Notes from the archive

Notes from the archive.

← All postsStudio · Oct 2026

Evals from the work itself

The best evals come from real work with a known outcome. A studio produces them every week.

More teams now build evals the obvious way: take the work they actually did, and check whether a model can do it. Split that work in two and it serves two purposes. Regression tests, the things a model must keep getting right. Breakthrough tests, the things nothing can do yet.

A game studio is unusually rich in this material. Every bug ticket has a reproduction, a fix that shipped and tests that pass; every balance change has the player data before and after. A model can be asked to fix the same bug in the same snapshot of the code, and the shipped fix grades it.

Two properties make these tasks valuable. They were never public, so no model has seen the answer. And they carry an outcome beyond the tests: whether the fix held, and what players did next.

We release studio data only with consent and redaction, with partner-confidential material left out.