The most complete record of how people play and how games are made.
Ten years of live games, 200M+ players and the studio that built them, catalogued for AI labs. Everything here comes from games we built and servers we run, and none of it has been on the public internet.
- Live data
- Real matches
- Server-graded
- Never public
Everything a studio makes, indexed.
Most vendors license one slice. We hold the whole chain, from the code that ran the game to the players who played it.
| No. | Holding | What it is | Scale | Status |
|---|---|---|---|---|
| H-01 | Match records | Server-side records of multiplayer matches: movement and aim, every hit, loot, vehicles, the storm, placement. | Millions of matches | Available |
| H-02 | Player histories | Pseudonymized careers: skill, progression, sessions and social play over years. | 200M+ players | Available |
| H-03 | Economy | Stores, markets, offers and currencies. What players valued, traded and bought. | Years of transactions | Available |
| H-04 | Action traces | Flappy Bird seeds and every input. Exact actions in a deterministic world. | Growing daily | Available |
| H-05 | Source and reviews | Game code, commits and pull requests with review threads, across titles we own. | A decade of repos | Available |
| H-06 | Studio work | Bug tickets, design docs and team threads, consented and redacted, linked to code. | Years of tickets | Available |
| H-07 | Live ops and performance | Balance patches, config changes, tests, crashes and frame times, with player outcomes. | Thousands of devices | Available |
| H-08 | Environments and evals | RL environments on production rules, and benchmarks with human baselines. | Growing library | Available |
Graded by the game. Measured on your evals.
A data purchase should be judged by what it does to your model on tasks we never saw. Ours is built for that test: the referee is a game server, the records were never public, and the people in them were playing for real.
| Ref. | What we offer | Graded by | Human baseline | How you test it |
|---|---|---|---|---|
| E-01 | Environments on production rules | The game server's own outcome: survival, placement, pipes cleared, win or loss. No model judges the result. | Players at every skill level, in the same situation | Train on it, then score a suite you hold back. |
| E-02 | Private evals that refresh | Server truth, tick by tick, for every player in the match. | Archive players from the same moment | New matches every day, so a set can be rotated instead of retired. |
| E-03 | Human play at scale | It is the baseline: 200M+ people who chose to play, bots labeled. | Score distributions by skill band | Compare model and human on the same seeds. |
| E-04 | Studio work with outcomes | Tests from the fix that actually shipped, plus what players did next. | The engineer who fixed it, and how long it took | Regression and breakthrough splits, from repos never made public. |
| E-05 | Targeted slices and capture | Your definition of the failure. | Matched human sessions | You name where the model breaks. We cut from the archive or capture to it. |
Built to be checked
Every environment ships with a reference solution that scores full marks, bad runs that score zero, hidden test files and its measured flake rate.
Measured before you see it
We train a small open model on each release and score held-out tasks across several seeds. No gain above the noise, no release.
Runs where you run
Open task format: a task, a container, a verifier and a reward file. Tools any agent can call. Versioned and pinned, so a rerun is a rerun.
Clean to buy
A data card with checked size, provenance, rights, personal-data handling and known limits. Exclusivity by title, modality or field of use, in writing.
How a match becomes data.
The game server is the referee. It decides what happened and writes it down. Those records, plus what players' devices report, land in the archive that every dataset, eval and environment is cut from.
How a game gets made, and what each step leaves behind.
Every turn of the studio's loop leaves a record. Linked together, a ticket, the code that fixed it and what players did next become tasks an agent can be scored on.
Inside one match record.
The tables in a single server record, as stored. Every child table hangs off the match and the players in it, so any hit, move or pickup can be joined to who did it and when.
The games, as played.
Footage from titles SuperGaming builds and runs. Each one is a source of records, players and code. Hover or tap a plate for full colour.


What the others leave out.
We read every data vendor's site in this market. These are the things none of them say, and all of them are true here.
The referee's record
Our data comes from the game server, the one source that knows where everyone was and what happened, tick by tick. Not pixels from one screen.
No middleman
We made the games, run the servers and wrote the code. One chain of custody, from player to license.
The whole studio
Gameplay linked to the code, tickets, economy and patches that shaped it. Cause and effect, not clips.
People who chose to play
Real stakes, real skill, no paid sessions. Bots are labeled, so you can drop them or use them as a control.
Every genre
Shooters, strategy, arcade, racing and Roblox, on mobile and PC, across emerging and mature markets.
Baselines from millions
Human scores from real competitive players at every skill level, for every benchmark we ship.
Nothing to leak
Server records were never published, so a private eval built from them stays private. Fresh matches every day mean it can be rotated, not retired.
Measured against people.
Each benchmark comes with a human row from the same game and a grader that reads the server record. None of it has been public, so none of it is in a pretraining set. Environments run on production rules.
| Ref. | Benchmark | Tests | Model gets | Scored by | Human baseline |
|---|---|---|---|---|---|
| T-01 | Next state | World modeling | Seconds of match state | Position and event error | Archive players |
| T-02 | Human or bot | Perception | One player's movement and fire | Server bot flag | Not applicable |
| T-03 | Zone routing | Planning | Position, loadout, storm | Survival | Players in the same spot |
| T-04 | Flappy Bird | Control · RL | Frames or state | Pipes cleared | Live players, same seeds |
| T-05 | Tower Conquest | Strategy | Board, hand, opponent | Win rate | Ranked ladder |
| T-06 | Fix the bug | Software agents | Real ticket and repo | Tests and the shipped fix | The engineer who fixed it |
Collected fairly
Under each game's terms and privacy policy. Studio data only with employee consent.
Pseudonymized
Player IDs replaced, nicknames removed, personal data stripped before delivery.
Labeled
Bots flagged. Every field marked recorded, reconstructed or inferred.
Cleared
Titles we hold rights to. Exclusive or not, per dataset, in writing.
Documented
Schema, data card and versioned releases, delivered to your cloud.
From the archive.
Notes from the archive.
Privacy policy.
Effective 10 October 2026. This policy explains what personal information datagame.ai collects, why, and what you can ask us to do with it.
Who we are
datagame.ai is operated by SuperTuned Inc, 1 Sansome St, Ste 3500, San Francisco, CA 94104, USA ("we", "us"). DataGame licenses game data, evals and RL environments to AI labs. For anything in this policy, write to hello@datagame.ai with "Privacy" in the subject.
What we collect
What you send us. When you use the contact form or email us: your name, work email, company and role, your message, and the details you choose to add (for example the data you hold or the timeline you have in mind). If an AI agent contacts us for you, we also receive what it says about who it acts for.
Technical data. Our servers and network providers process your IP address, browser type and the pages requested, to deliver the site, keep it secure and limit abuse. With each enquiry we store a one-way hash of your IP address, not the address itself.
Analytics. We use PostHog in cookieless mode to count visits, see which pages are read and where visitors come from (the referring site, and the country derived from your IP address at the time of the visit). We also use Google Analytics, which sets cookies only if you accept them in the cookie banner; if you decline, it sends anonymous, cookieless measurements. We do not record sessions and we do not use analytics for advertising. Google Search Console reports aggregated search data to us.
How we use it
To reply to you, to understand what you need, to prepare samples, NDAs and licences, to keep records of our business conversations, and to protect the site. We do not sell or rent your personal information, we do not share it for advertising, and we do not use what you send us to train AI models.
Who processes it for us
Railway (hosting), Cloudflare (DNS and network), Google Workspace (email), Resend (delivering form enquiries to our inbox) PostHog and Google (website analytics). They process data on our instructions and under their own security commitments. We may disclose information if the law requires it, or to a buyer of our business, which would be bound by this policy.
Cookies
If you accept analytics cookies, Google Analytics sets _ga and _ga_EQJN9E79J1 to tell visits apart; they last up to two years. We remember your choice in your browser’s local storage (dg_consent). You can change it any time with Cookie settings at the bottom of every page, or by clearing your browser data. We set no advertising cookies.
Where it is kept
Our providers store data in the United States and other countries. Where the law requires it, transfers rely on the providers' standard contractual safeguards.
How long we keep it
Enquiries are kept while we are in conversation with you and for up to two years after our last contact, unless you ask us to delete them sooner or we need them longer for a contract or a legal obligation. Server logs are kept for short periods by our providers.
Your choices and rights
You can ask us to tell you what we hold about you, correct it, delete it, or stop using it. Depending on where you live (for example the EU or UK under the GDPR, California under the CCPA, or India under the DPDP Act) you may have further rights, including to complain to your data-protection authority. We answer requests within 30 days and will not treat you differently for making one. We do not sell or share personal information as those terms are used in California law.
Data from our games
This policy covers the website. Player data from the games behind DataGame was collected under each game's own terms and privacy policy, and player identities are pseudonymized before any data is licensed. Questions about game data can go to the game's support channel or to hello@datagame.ai.
Children
This site is for businesses and is not directed at children. We do not knowingly collect personal information from anyone under 16 through it.
Security
The site is served only over HTTPS, enquiries go to a small number of people, and access to our systems is limited to those who need it. No system is perfectly secure; if we learn of a breach that affects you, we will tell you as the law requires.
Changes
If we change this policy we will update the date above, and for material changes we will say so on the site.
Terms of use.
Effective 10 October 2026. These terms cover your use of the datagame.ai website. Data, samples, evals and environments are provided only under a separate written agreement.
Who we are
The site is operated by SuperTuned Inc, 1 Sansome St, Ste 3500, San Francisco, CA 94104, USA. Contact: hello@datagame.ai.
Using the site
You may browse the site and share links to it. You may not use it to break the law, try to get around its security, overload it, or send spam or automated enquiries in bulk through the contact form or /api/lead.
AI agents and automated access
Agents may read the pages and the files we publish for them (llms.txt, llms-full.txt, catalog.json) and may send enquiries on behalf of a named person or organisation. An agent must say who it acts for. A person at DataGame answers every enquiry; nothing an agent receives from the site is a commitment to supply data.
Content and trademarks
The site's text, drawings, replay, footage and code belong to SuperTuned Inc and its affiliates or licensors. Browsing gives you no licence to the data, footage or games shown. DataGame, datagame.ai and the names of our games are trademarks of their owners.
No offer, no advice
Descriptions of data, volumes and results are a general guide. They are not an offer, and what we supply is defined only in a signed agreement, which governs if it differs from anything on this site.
Disclaimers and liability
The site is provided "as is", without warranties of any kind, to the extent the law allows. To the extent the law allows, SuperTuned Inc is not liable for indirect or consequential losses arising from your use of the site, and its total liability relating to the site is limited to US$100.
Law
These terms are governed by the laws of the State of California, USA, and disputes about them go to the courts in San Francisco, California, unless the law where you live says otherwise.
Changes
We may update these terms by posting a new version here with a new date.
What labs and lawyers ask.
What rights does a license give us?
Training, fine-tuning, evaluation and internal research by default. Commercial deployment of models trained on the data is included. Redistribution of the raw data is not. Every license is written per dataset, so the terms match what you are building.
Can we get exclusivity?
Yes, for a time window, a modality or a field of use. Most buyers start non-exclusive with a pilot and convert the parts they care about.
Who owns the data?
SuperGaming built and operates the games, so the data comes straight from the source: our servers, our players, our code. DataGame (a SuperTuned company) licenses it. There is no broker in the chain.
How do you handle player consent and privacy?
Data is collected under each game's terms of service and privacy policy. Player IDs are pseudonymized, nicknames are replaced, and personal data is removed before anything leaves our systems. Studio communications are released only with employee consent and redaction.
Are bots mixed in with human players?
Bots are flagged in every record. You can filter them out, or keep them as a control group. We quote volume in human play, so bots never inflate a deal.
What exactly is in a match record?
Everything the game server saw: positions and view direction for every player, every shot and hit with both players' positions, loot, abilities, vehicles, heals, the storm and final placement. Events are tick-exact, straight from the authoritative server.
Which games and genres are covered?
Battle royale and arena shooters, real-time strategy, arcade, racing, and Roblox titles, across a decade of live operation. Plus the studio around them: code, tickets, live-ops and performance data.
Do you have video?
We have gameplay footage for many titles, and we are building a renderer that turns stored matches into video, depth and segmentation, labeled as recorded or reconstructed. New capture with inputs and video is available for custom projects.
How should we evaluate your data?
Train on a slice and score it on tasks you hold back. Every slice ships with provenance and a data card so the comparison is clean, and we publish our own measured result, run over several seeds, before you see it. If it does not move your numbers, tell us.
Can you build for a specific failure?
Yes. Describe the failure: agents losing the thread over long tasks, planning under time pressure, coordinating a team, reading an economy, fixing real code. We show what in the archive bears on it, then cut a slice, package an environment or run a capture.
Will training on games help beyond games?
That is the open question in RL, and the one worth paying for. We measure it: every release reports held-out scores on the same game, on our other games and on non-game tasks that test the same skills, across several seeds. When the gain is small, the data card says so.
Has any of this leaked into pretraining data?
No. Server records and studio repositories were never on the public internet, and footage is a small part of what we hold. Private sets stay private, and because new matches arrive daily they can be rotated rather than retired.
Do your evals include human baselines?
Yes. Every benchmark ships with human scores from the same game, drawn from real players at different skill levels, plus a scripted-bot floor and the grader code.
Can we run agents in your games?
Yes. Our RL environments run on the production games, with the same rules players had. Flappy Bird is live today; more titles are being packaged.
Can we use your environments on the platform we already train on?
Yes. Each environment ships in an open task format (task, container, verifier, reward) that common RL frameworks and hosted training platforms can load. Tell us what you train on and we package for it.
Can we see a sample before signing?
Yes. We send a schema, a data card and a sample under a short NDA, usually within days of a first call.
Can we test tasks you did not pick?
Yes. We commit the full catalogue first, then a public random draw chooses the tasks you test, so the sample is not hand-picked. Each task comes with how often several models solve it, so you see the difficulty before you buy.
How is data delivered?
JSON and Parquet for records, MP4 and WebDataset for rendered views, hosted environments for RL. Delivered to your cloud bucket, versioned, with a data card for every release.
How is pricing set?
By dataset, volume of human play, exclusivity and refresh cadence. Pilots are priced to be easy to say yes to.
Do you work with companies training their own models?
Yes. Enterprises training models with reinforcement learning, and the firms that train for them, license our environments and evals the way labs do: a pilot first, then a licence.
Can my studio license its data through you?
Yes. We clean, pseudonymize and package your data to the same standard as ours, you keep ownership, and you share in every license.