FAQ · Licensing, privacy, data, delivery

What labs and lawyers ask.

What rights does a license give us?

Training, fine-tuning, evaluation and internal research by default. Commercial deployment of models trained on the data is included. Redistribution of the raw data is not. Every license is written per dataset, so the terms match what you are building.

Can we get exclusivity?

Yes, for a time window, a modality or a field of use. Most buyers start non-exclusive with a pilot and convert the parts they care about.

Who owns the data?

SuperGaming built and operates the games, so the data comes straight from the source: our servers, our players, our code. DataGame (a SuperTuned company) licenses it. There is no broker in the chain.

How do you handle player consent and privacy?

Data is collected under each game's terms of service and privacy policy. Player IDs are pseudonymized, nicknames are replaced, and personal data is removed before anything leaves our systems. Studio communications are released only with employee consent and redaction.

Are bots mixed in with human players?

Bots are flagged in every record. You can filter them out, or keep them as a control group. We quote volume in human play, so bots never inflate a deal.

What exactly is in a match record?

Everything the game server saw: positions and view direction for every player, every shot and hit with both players' positions, loot, abilities, vehicles, heals, the storm and final placement. Events are tick-exact, straight from the authoritative server.

Which games and genres are covered?

Battle royale and arena shooters, real-time strategy, arcade, racing, and Roblox titles, across a decade of live operation. Plus the studio around them: code, tickets, live-ops and performance data.

Do you have video?

We have gameplay footage for many titles, and we are building a renderer that turns stored matches into video, depth and segmentation, labeled as recorded or reconstructed. New capture with inputs and video is available for custom projects.

How should we evaluate your data?

Train on a slice and score it on tasks you hold back. Every slice ships with provenance and a data card so the comparison is clean, and we publish our own measured result, run over several seeds, before you see it. If it does not move your numbers, tell us.

Can you build for a specific failure?

Yes. Describe the failure: agents losing the thread over long tasks, planning under time pressure, coordinating a team, reading an economy, fixing real code. We show what in the archive bears on it, then cut a slice, package an environment or run a capture.

Will training on games help beyond games?

That is the open question in RL, and the one worth paying for. We measure it: every release reports held-out scores on the same game, on our other games and on non-game tasks that test the same skills, across several seeds. When the gain is small, the data card says so.

Has any of this leaked into pretraining data?

No. Server records and studio repositories were never on the public internet, and footage is a small part of what we hold. Private sets stay private, and because new matches arrive daily they can be rotated rather than retired.

Do your evals include human baselines?

Yes. Every benchmark ships with human scores from the same game, drawn from real players at different skill levels, plus a scripted-bot floor and the grader code.

Can we run agents in your games?

Yes. Our RL environments run on the production games, with the same rules players had. Flappy Bird is live today; more titles are being packaged.

Can we use your environments on the platform we already train on?

Yes. Each environment ships in an open task format (task, container, verifier, reward) that common RL frameworks and hosted training platforms can load. Tell us what you train on and we package for it.

Can we see a sample before signing?

Yes. We send a schema, a data card and a sample under a short NDA, usually within days of a first call.

Can we test tasks you did not pick?

Yes. We commit the full catalogue first, then a public random draw chooses the tasks you test, so the sample is not hand-picked. Each task comes with how often several models solve it, so you see the difficulty before you buy.

How is data delivered?

JSON and Parquet for records, MP4 and WebDataset for rendered views, hosted environments for RL. Delivered to your cloud bucket, versioned, with a data card for every release.

How is pricing set?

By dataset, volume of human play, exclusivity and refresh cadence. Pilots are priced to be easy to say yes to.

Do you work with companies training their own models?

Yes. Enterprises training models with reinforcement learning, and the firms that train for them, license our environments and evals the way labs do: a pilot first, then a licence.

Can my studio license its data through you?

Yes. We clean, pseudonymize and package your data to the same standard as ours, you keep ownership, and you share in every license.