Evil GeniusesMeijindex
EVIL GENIUSES×PEAK6
For AI labs

Human data from
verified CS2 players.

We're building a paid CS2 player cohort through Evil Geniuses: ranked gameplay, players' explanations and independent coach reviews, joined to specific moments. The replay pipeline and written review tools are built. The cohort will add verified rank, contributor consent and repeatable review coverage to each delivery.

The cohort we're building
  • Rank checks against Steam-linked match replays and FACEIT
  • Contributor, rank and consent records required before cohort delivery
  • Independent coach ratings, with prior exposure and disagreements recorded
  • We don't build models, so your task specs stay with you
Where things stand

What runs today, and what doesn't yet

Steam sign-in and connections to rank and replay dataBuilt
Replay parsing to per-tick state, including player inputsBuilt
Staged player and expert review with stored answersBuilt
First paid cohort of 20 to 30 CS2 playersNot yet
The deliverable

The planned pilot pack

The CS2 sample already has tick-aligned gameplay and a working annotation flow. These are the deliverables we're building toward for a paid cohort. Scope, acceptance criteria and delivery rights are agreed in the SOW before work starts.

Decision windows

Moments from the player's own ranked matches on Valve Premier or FACEIT, with per-tick state about 64 times a second: positions, view angles, inputs, utility and events.

The player's explanation

The player reviews key calls: what they noticed, what they considered and why they acted. Written recollections are collected today. Recorded interviews and transcripts are planned for the cohort.

Two coach scores

The planned cohort pairs two independent coach ratings per window, locked before comparison. The pack will include the rubric, disagreements and agreement rates. The current sample does not yet have dual-coach coverage.

Data card

Where the data came from, the schema, who was in the cohort and at what rank, the consent chain, known limitations, and a few complete sessions to review before you scale up.

Other tasks to scope with the cohort
Human baselines for agent evals, by rank band
Expert demonstrations of a specific skill
Skilled players and coaches reviewing your agent's play
Red-teaming game-playing agents

Delivery rights. Telemetry, labels, transcripts, scores and video each need an agreed scope of use. Player consent does not replace publisher permission. Rights and identifier handling must be checked before a commercial release; the current sample is not a rights-clearance guarantee. For agent evaluation, you provide the environment and agent; the proposed cohort provides play and review.

How we plan to run a pilot

Your task and the cohort

01
You set the task and the bar

Tell us the work, the minimum rank and what counts as acceptable. We turn that into a cohort spec.

02
We put together the cohort

Consenting CS2 players at the rank you need, recruited through the Evil Geniuses community and paid through EG's contractor setup. Everyone is 18 or over.

03
You get the work and its records

The data or scores, plus a record on each item of the contributor, their rank at the time and what they agreed to.

We don't build models. Your task definitions and deliverables aren't used to train anyone else's, and your results stay yours.

Why Counter-Strike

What public CS2 datasets leave out

Every ranked CS2 match produces an official replay that parses to per-tick state, and mature open-source parsers make that routine. Large datasets already exist: EgoCS-400K, and Reka's CS2-10k until Valve had it taken down. We aren't trying to beat them on hours. They're built from public match archives, so they don't carry a verified player, that player's consent, their reasons, or an expert's view of the decision. Behavioural cloning work on Counter-Strike hit the same limit: it trained on millions of scraped frames and still needed a smaller set of expert demonstrations. We make that smaller set.

A result we retracted

What our replay model can tell you

For each encounter we model which channels could have reached the player: what they had already seen, what a teammate could have called, what was audible and what was blocked. It's a versioned model, not a record of what they perceived. We used to cite a crosshair-error gap as evidence that it measured what a player knew. Re-run across eighteen demos, the gap fell from 15.3 degrees to 3.8, and to 0.7 once reveals that happened off screen were removed, so we retracted it. The channel model stands. It can't tell you what a player knew.

If you mix replay sources: in our corpus, Valve matchmaking replays mark who spotted whom about half as reliably as FACEIT replays of the same maps (0.44 against 0.78 on de_mirage). We record the source on every item.

Run a pilot

Start with one task and a small cohort

Send the task and the rank you need. We'll recruit the cohort, run the work and report against the acceptance criteria you set.