Statistical analysis of the significance of rank differences: seeking test subjects!

CupcakeTrap·7/23/2014, 6:19:13 PM·5 votes·742 views

TL;DR: I'm gathering test match data for a statistical examination of how the various ranked tiers measure up. Click here for instructions on how to participate.

The Question
Imagine that a big match will be played. All you know about the teams is that they've been created by randomly drafting Summoners based on their rank.

  • Team A is all-Silver (SSSSS): 5 random Silver players plucked from the master list of all Silver-Tier Summoners.
  • Team B is 3 Silvers, 1 Gold, and 1 Bronze (GSSSB). In other words, their advantage is that they have a Gold in place of a Silver, and their disadvantage is that they have a Bronze in place of a Silver.

You're placing a bet. They'll run 10 matches like this (fresh draw of players each time), switching off Blue/Purple side. Team A will always be SSSSS, and Team B will always be GSSSB.

Cast your vote here. (And click here for poll results.) Then read on to see what we're doing to put this to the test.

The Experiment
How could this question be experimentally examined?

Well, if you have a large set of matches, you can statistically assess how often various team comps beat other team comps. If you're willing to make some modeling assumptions, e.g. an additive theory of skill or a multiplicative theory of skill, you could even run a regression of sorts and try to compute a "power rating" for each rank.

Meanwhile, you'll also want to take a look at variance: it's possible that rank just doesn't really correlate with game-winning power in a particularly strong way, at least not in mixed-rank games. (e.g., the "Elo hell effect": maybe Diamonds get confused in an all-Silver match and play poorly because they expect their teammates and opponents to behave in a certain way.)

Preliminary Results
Right now, we need more data. However, I've been able to do some very rough preliminary analysis using data from Factions matches. (Factions is a community game mode based on custom matches of the "Noxus vs. Ionia", "Piltover vs. Zaun", etc. variety.) This data is obviously imperfect, but it's what we have.

Click here to read the preliminary report. We still need lots more data. The more matches we have, the more instances of each team comp we have, and the better the tests we can run. The preliminary results above, for instance, resort to a hack-ish "winrate for each tier" approach to squeeze some sort of conclusion out of a rather small dataset.

Play A Couple Custom Matches For Science
Planning on playing some normals tonight? Why not make one of them a quirky mixed-rank custom match? This will help expand our dataset with some non-Factions matches.

Click here for instructions on how to participate. It's really nothing more elaborate than joining a chat and either starting up a quick custom match with others there or being on the lookout for invites from match organizers.

Why Does This Matter?
There are a few reasons I'm looking into this.

  • Curiosity. It's an interesting problem, IMO, and I'm curious to find out just how big of an impact different ranked players have on mixed-rank matches.
  • Duos. Matchmade games are special because matchmaker has access to true MMR, but still, I've often wondered, when I get matched up with or against (say) a Plat/Silver duo, just how big the gap between them probably is.
  • Factions. We score matches in Factions. This allows us to use match outcomes to influence how the storyline develops. However, it doesn't seem fair that an all-Gold Noxian team being at all-Silver Ionian team should be worth as many points as a win in an evenly matched game. To properly adjust point values, we need data on how the tiers stack up.

Anyway, I hope at least some of you will replace the occasional Normal with one of these balance testing matches. Click here for instructions on how to participate.

9 Comments

Angry Monster7/23/2014, 9:36:25 PM2 votes

So question about the pool of players. When you say random players, how random? Like people fighting over the same role random, random premade group of 5, Or people like in team builder who never met but are getting roles that they want?

Each variation changes the equation of who will win. Example a Plat jungle can lead his team in a premade through strategy and calls. But a random group may or may not listen since their is a basic trust issue inherent with working with people the first time.

Thoroniul7/23/2014, 6:36:24 PM1 votes

The line between silver and gold is really pretty thin. The only difference between gold V and silver I is that the gold V players got lucky to win 8 games worth of LP and got promoted. There is a big wall of gold elo hell players keeping silver people from getting promoted. The good ones who want to improve go higher into better tiers while the once who hit gold and don't care if they lose anymore hinder the higher silver tiers they are matched with.

The chance of having a really clean spirited with no troll, afk, dc, rage, toxic, and people who don't understand objective taking is rather low at the high silver level. It gets really annoying with a constant flux of winning and losing Lp without getting higher.

The difference between plat and silver is much higher. At the plat level you focus more on team composition, and objectives, and countering.

MrSc0tty7/25/2014, 5:48:38 PM1 votes

Problem: role and champ variance. If you have a team with a bronze and you tell him to just power farm jungle Udyr, or stick him on Leona support he's much less likely to feed than if he was up against a silver in midlane.