Journal · Notes from the archive

Notes from the archive.

← All postsMarket · Oct 2026

Base data, and the layer above it

The RL market has sorted itself into layers. Underneath all of them is the real work environments are built from.

A year ago most training data was labels. Now the market has layers. Environment builders make the worlds and tasks a model practises in. Infrastructure companies give teams the tools to build, host and train on those environments. Service firms run reinforcement learning for companies that want their own models. Labs, and now a growing number of enterprises, buy from all three.

Underneath every layer sits the same raw material: base data. The codebases, logs, configurations and records of real work that an environment is built from. An environment is only as good as what it was built on. Contrived tasks teach contrived skills, and a grader written in an afternoon is the first thing a model learns to cheat.

That is why quality, not supply, is now the bottleneck. Platforms that test what vendors sell report that much of it fails: the reward climbs because the model found a shortcut, or the tasks were too artificial to teach anything that carries over. The test that matters is simple to state. Train an open model on the environment, watch the reward rise, then check that it also improved on tasks it never saw. Buyers have started testing the way an auditor would: commit the vendor’s full catalogue, draw tasks at random, and check that models solve some of them but not most.

A game studio sits in two layers at once. We hold the base data, a decade of server records, code, tickets and economy, and we build environments on it that are graded by servers that refereed millions of real matches. We sell both. We ship in open formats, so our environments run on the platforms labs already use, and we measure transfer before we ask anyone to pay for it.