| Total winrate | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 47.4% 1138 | — 137 | 49.5% 114 | 45.7% 86 | 46.3% 111 | 45.9% 245 | 55.5% 215 | 53.1% 166 | 45.1% 106 | 38.2% 95 | |
| 50.6% 491 | 50.5% 114 | — 16 | 50.3% 35 | 50.0% 45 | 45.7% 93 | 53.4% 69 | 49.8% 63 | 54.7% 39 | 50.0% 33 | |
| 52.2% 398 | 54.3% 86 | 49.7% 35 | — 10 | 52.8% 34 | 44.6% 86 | 53.9% 60 | 53.1% 38 | 63.5% 41 | 45.9% 18 | |
| 48.4% 529 | 53.7% 111 | 50.0% 45 | 47.2% 34 | — 14 | 54.3% 106 | 49.9% 85 | 45.2% 64 | 45.9% 41 | 41.2% 43 | |
| 52.4% 989 | 54.1% 245 | 54.3% 93 | 55.4% 86 | 45.7% 106 | — 75 | 45.3% 132 | 52.5% 123 | 51.5% 96 | 60.2% 108 | |
| 48.9% 786 | 44.5% 215 | 46.6% 69 | 46.1% 60 | 50.1% 85 | 54.7% 132 | — 43 | 57.0% 113 | 40.8% 52 | 51.4% 60 | |
| 48.9% 668 | 46.9% 166 | 50.2% 63 | 46.9% 38 | 54.8% 64 | 47.5% 123 | 43.0% 113 | — 33 | 51.6% 45 | 50.0% 56 | |
| 49.1% 453 | 54.9% 106 | 45.3% 39 | 36.5% 41 | 54.1% 41 | 48.5% 96 | 59.2% 52 | 48.4% 45 | — 15 | 45.9% 33 | |
| 52.1% 446 | 61.8% 95 | 50.0% 33 | 54.1% 18 | 58.8% 43 | 39.8% 108 | 48.6% 60 | 50.0% 56 | 54.1% 33 | — 17 |
Confirmed 1v1 games of the selected mod, season and patch version only. Games shorter than 120 seconds are discarded as drops and disconnects, and so is anything involving a banned account.
Mirror matchups are counted separately, on the diagonal, and never take part in anything else. A race cannot be imbalanced against itself: every mirror game is by definition one win and one loss for the same race, so including it only drags the race total toward 50% in proportion to how popular that race is.
The rating tiers are not a fixed MMR value but a quantile of the games actually present in the current selection. The games are ordered by the lower MMR of the two players, and the top 20% or top 5% are kept, so the tier holds a usable sample on every mod, season and patch instead of collapsing to a handful of games.
Both players must be above the cut. If only one were required, the tier would fill up with strong players farming weak ones, which says nothing about balance. Games with no recorded MMR cannot be placed and are dropped while a tier is active. The resulting MMR value is printed just above.
A single player contributes at most 25 games to a single matchup. If somebody played 300 games of the same matchup, each of those games gets weight 25/300 instead of 1, so their whole streak is worth exactly 25 games. Where both players are over the cap, the smaller of the two weights is used.
Without this one dedicated player defines what a matchup looks like, and the table ends up measuring that person rather than the race.
Races are not equally strong on every map, and matchups are not played on the same mix of maps. Every game is therefore reweighted so that each matchup is judged on the same map distribution as the selection as a whole: a map that is over-represented in a matchup is weighted down, an under-represented one is weighted up. Maps with fewer than 30 games are pooled into a single bucket.
The reweighting factor is limited to the range 0.25–4.00. A map that a matchup was almost never played on carries too little information to be extrapolated into the result, so the correction is deliberately left incomplete rather than invented.
A race looks strong when strong players happen to be playing it, and a race looks weak the month its best player stops playing. To remove that, each game is measured against what the rating gap already predicts: the Elo win probability of the first side is 1 / (1 + 10^(−ΔMMR/400)).
Summing that over a matchup gives the number of wins the stronger side was supposed to take. We then search for the rating offset at which the predicted wins match the observed wins, and report that offset back as a winrate. The number therefore answers "how would this matchup look between equally skilled players", and it no longer moves when a top player picks up or abandons a race. Games without a recorded MMR are treated as an even matchup and contribute nothing to the correction.
After the cap and the map correction games no longer weigh 1 each, so the plain game count overstates how much is actually known. The effective sample is (Σw)² / Σw² — the number of full-weight games the weighted data is really worth. 400 games in which one player was capped down can be worth an effective 240.
Everything below uses this number, never the raw game count. The game count is still what is printed under each percentage; the effective sample is in the tooltip.
A matchup with 4 games and 4 wins is not a 100% matchup — it is no information at all. So the number shown is not the observed one. It is (adjusted winrate × effective sample + k/2) / (effective sample + k), which is exactly the same as adding k imaginary games split evenly between the two races before computing the percentage. Right now k = 42.
k is not a hand-picked constant. It is estimated from the data: we measure how far the matchups of this mod actually spread apart, subtract the spread that pure chance would produce at these sample sizes, and treat what is left as the real spread. If the mod genuinely has strong imbalances, k comes out small and almost nothing is pulled in. If the observed spread is no bigger than random noise would explain, k comes out large and everything is pulled hard toward 50%. The table is sceptical exactly to the degree the data justifies, and that scepticism is recalculated for every mod, season, patch and skill tier.
How to read it: a matchup with exactly k effective games ends up halfway between 50% and its raw value. With a quarter of k it keeps about a fifth of its deviation. With ten times k it is left practically untouched. Below 8 effective games nothing is shown at all — only the game count, because at that size the raw value is noise.
Each matchup also gets a 95% Wilson score interval, computed on the adjusted value and the effective sample. Wilson is used rather than the textbook normal interval because it stays sensible at small samples and near 0% and 100%.
A cell is coloured only when that entire interval lies on one side of 50% — that is, only when the deviation survives its own sample size. Everything else stays grey no matter how extreme the number looks. The intensity of the colour tracks the size of the deviation, not its significance, so a bright cell means "far from 50%" and a grey cell means "not established".
Raw values are a separate view and are coloured by the observed number alone, with no significance test behind them. That view exists to show what the data looks like before any correction — a bright cell there is not evidence of anything on its own.
Aggressive coloring drops the significance test entirely and paints every cell by its value, so even a deviation of a fraction of a percent gets a tint. It makes the table easier to scan at a glance and is how the old table behaved, but a coloured cell then means only "not exactly 50%", not "established".
The left column is the unweighted average of a race's matchups, not its overall win percentage. Averaging by games played would let the most popular matchups dominate and would quietly smuggle pick rates back into a balance number.
Mirrors are excluded, and matchups too thin to be displayed do not enter the average either. Switching the table to raw values replaces it with the plain overall win percentage across all non-mirror games.