Journal · Notes from the archive

Notes from the archive.

← All postsMeasurement · Oct 2026

Measured before you see it

The honest way to sell data is to show what it does to a model on tasks it never saw. We run that test first.

The cleanest test for a dataset or environment is now widely agreed: keep the model, the training recipe and a private test suite fixed, change only the data, and measure the difference. Open arenas run exactly that contest, and they add the warning every buyer should hear. One run is noise. A small gain from a single seed proves nothing.

We hold ourselves to it before a lab does. For each release we train a small open model on our tasks, score it on held-out tasks from the same game and on tasks from other games, and repeat across several seeds. The result ships in the data card with its spread, not just its best run.

If a release does not move held-out scores above the noise, it does not ship. If it only moves the game it came from, we say so. Transfer is the claim worth paying for, so it is the claim we measure.

Labs should still run their own test on their own suite. Our number is there so the first conversation starts from evidence.