Journal · Notes from the archive

Notes from the archive.

← All postsMeasurement · Oct 2026

The question worth paying for

Getting good at one game is cheap. Getting better at everything else because of it is the claim labs pay for.

The open question in reinforcement learning is generalization. A model trained on coding tasks gets better at coding. Does a model trained on a battle royale get better at anything but battle royales? Nobody selling environments can answer that in general, and anyone who says otherwise is selling.

What we can do is measure it for our own releases and publish the result. Every release is scored on held-out tasks from the same game, on tasks from our other games, and on a small set of non-game tasks that test the same skills: planning over a long horizon, acting under time pressure, working with or against other agents. Several seeds, with the spread shown.

We expect games to transfer better than most environments because they stress what agents still get wrong, and because the people in our data were not paid to perform. But expectation is not evidence. When the gain is small, the data card says so.

For buyers the practical advice is the same as for any data: keep your model and your test suite fixed, change only the data, and run it more than once.