← All posts

Understanding Matchups with Reinforcement Learning

One of the challenges of playing any TCG is learning to play against all of the likely matchups you'll encounter. In Star Wars: Unlimited for example, even if you've settled on a specific deck you're going to play with, your resourcing decisions, sequencing, and overall strategy is going to depend on your opponent's deck and its strengths and weaknesses.

The Opponent

In order to learn how to play a matchup well, you need an opponent that is going to play the opposing deck effectively. For SwooLab, that means an AI opponent that learns to play through reinforcement learning. By playing against itself across millions of games, with millions of different randomly generated decks, the AI can learn to use all the same basic principles that a human player would learn.

The AI consists of a single neural net model that outputs

  1. a probability distribution over the legal moves at the current state (the policy), and
  2. an estimate of the probability of winning the game (the value).
The policy tells you what actions you should take, and the value tells you how good each position is for you.

Training the model

The model is trained in iterations that involve first collecting thousands of games of data, and then training on that data.

One common way to compare the skill level of two players is to use Elo ratings, which assigns a rating to each player based on their win-loss record. While there are some issues with applying Elo to a game with random events (like a shuffled deck) and non-transitivity (the rock, paper, scissors effect applied to decks), we can still use it to show how the rate of learning changes over time during the training process.

The chart below shows the Elo rating of the SwooLab model over the first 3,600 training iterations.

Self-play Elo by training checkpoint 0 100 200 300 400 500 600 700 800 10 30 100 300 1000 3000 training checkpoint (log scale) v10 — 0 Elo v20 — 114 Elo v30 — 193 Elo v40 — 263 Elo v50 — 309 Elo v60 — 350 Elo v70 — 370 Elo v80 — 406 Elo v90 — 424 Elo v100 — 437 Elo v110 — 456 Elo v120 — 468 Elo v130 — 482 Elo v140 — 488 Elo v150 — 493 Elo v160 — 504 Elo v170 — 500 Elo v180 — 505 Elo v190 — 492 Elo v200 — 514 Elo v220 — 522 Elo v240 — 530 Elo v260 — 541 Elo v280 — 542 Elo v300 — 537 Elo v320 — 550 Elo v340 — 552 Elo v360 — 546 Elo v380 — 562 Elo v400 — 572 Elo v420 — 548 Elo v440 — 547 Elo v460 — 567 Elo v480 — 572 Elo v500 — 568 Elo v520 — 585 Elo v540 — 590 Elo v560 — 592 Elo v580 — 601 Elo v600 — 599 Elo v640 — 607 Elo v680 — 602 Elo v720 — 606 Elo v760 — 608 Elo v800 — 611 Elo v840 — 613 Elo v880 — 634 Elo v920 — 635 Elo v960 — 631 Elo v1000 — 634 Elo v1040 — 632 Elo v1080 — 647 Elo v1120 — 650 Elo v1160 — 659 Elo v1200 — 665 Elo v1240 — 667 Elo v1280 — 667 Elo v1320 — 672 Elo v1360 — 680 Elo v1400 — 675 Elo v1440 — 669 Elo v1480 — 676 Elo v1520 — 667 Elo v1560 — 672 Elo v1600 — 682 Elo v1640 — 689 Elo v1680 — 685 Elo v1720 — 684 Elo v1760 — 686 Elo v1800 — 688 Elo v1900 — 683 Elo v2000 — 697 Elo v2100 — 696 Elo v2200 — 702 Elo v2300 — 715 Elo v2400 — 723 Elo v2500 — 732 Elo v2600 — 735 Elo v2700 — 729 Elo v2800 — 737 Elo v2900 — 735 Elo v3000 — 741 Elo v3100 — 741 Elo v3200 — 747 Elo v3300 — 744 Elo v3400 — 749 Elo v3500 — 748 Elo v3600 — 744 Elo 0 500 1000 1500 2000 2500 3000 3500 training checkpoint v10 — 0 Elo v20 — 114 Elo v30 — 193 Elo v40 — 263 Elo v50 — 309 Elo v60 — 350 Elo v70 — 370 Elo v80 — 406 Elo v90 — 424 Elo v100 — 437 Elo v110 — 456 Elo v120 — 468 Elo v130 — 482 Elo v140 — 488 Elo v150 — 493 Elo v160 — 504 Elo v170 — 500 Elo v180 — 505 Elo v190 — 492 Elo v200 — 514 Elo v220 — 522 Elo v240 — 530 Elo v260 — 541 Elo v280 — 542 Elo v300 — 537 Elo v320 — 550 Elo v340 — 552 Elo v360 — 546 Elo v380 — 562 Elo v400 — 572 Elo v420 — 548 Elo v440 — 547 Elo v460 — 567 Elo v480 — 572 Elo v500 — 568 Elo v520 — 585 Elo v540 — 590 Elo v560 — 592 Elo v580 — 601 Elo v600 — 599 Elo v640 — 607 Elo v680 — 602 Elo v720 — 606 Elo v760 — 608 Elo v800 — 611 Elo v840 — 613 Elo v880 — 634 Elo v920 — 635 Elo v960 — 631 Elo v1000 — 634 Elo v1040 — 632 Elo v1080 — 647 Elo v1120 — 650 Elo v1160 — 659 Elo v1200 — 665 Elo v1240 — 667 Elo v1280 — 667 Elo v1320 — 672 Elo v1360 — 680 Elo v1400 — 675 Elo v1440 — 669 Elo v1480 — 676 Elo v1520 — 667 Elo v1560 — 672 Elo v1600 — 682 Elo v1640 — 689 Elo v1680 — 685 Elo v1720 — 684 Elo v1760 — 686 Elo v1800 — 688 Elo v1900 — 683 Elo v2000 — 697 Elo v2100 — 696 Elo v2200 — 702 Elo v2300 — 715 Elo v2400 — 723 Elo v2500 — 732 Elo v2600 — 735 Elo v2700 — 729 Elo v2800 — 737 Elo v2900 — 735 Elo v3000 — 741 Elo v3100 — 741 Elo v3200 — 747 Elo v3300 — 744 Elo v3400 — 749 Elo v3500 — 748 Elo v3600 — 744 Elo Elo
Elo by training checkpoint.

It's interesting to see what the model learns early on vs later in the process. For example, early in training, moves are mostly random, and the model learns to not value upgrades or most events. An upgrade played randomly is just as likely to benefit your opponent as it is you, and an event aimed at the wrong unit could be detrimental. It's only after it learns to correctly target units with each event and upgrade that it begins to learn the value of these cards. We'll cover this in a future article on the site.

Simulating the Matchups

Practicing against the AI model is helpful preparation for playing the real thing, but we can do more than that. I ran an experiment, simulating some matchups over thousands of games and observed how the model plays it from both sides.

First thing to to was decide which matchups to simulate. SwooLab's AI uses a custom-built rules engine that currently only supports the first set, Spark of Rebellion. The other sets will be added soon, but for the purpose of demonstration, I've selected a set of 12 decks from the Set 1 meta as a rough stand-in for what the true meta once was.

The experiment consisted of

The matchup table

Each cell is the row deck's win rate against the column deck, over every game the two played in either seat, so a cell and its mirror across the diagonal always sum to 100%. Decks are ordered by overall Elo rating.

Hover any cell for that matchup's details:

61%
2000
57%
2000
66%
2000
55%
2000
62%
2000
48%
2000
63%
2000
74%
2000
53%
2000
73%
2000
86%
2000
39%
2000
57%
2000
57%
2000
54%
2000
61%
2000
51%
2000
53%
2000
70%
2000
54%
2000
75%
2000
75%
2000
43%
2000
43%
2000
30%
2000
67%
2000
59%
2000
66%
2000
49%
2000
39%
2000
78%
2000
77%
2000
80%
2000
34%
2000
43%
2000
70%
2000
52%
2000
42%
2000
43%
2000
63%
2000
66%
2000
57%
2000
75%
2000
77%
2000
45%
2000
46%
2000
33%
2000
48%
2000
49%
2000
63%
2000
64%
2000
44%
2000
72%
2000
60%
2000
74%
2000
38%
2000
39%
2000
41%
2000
58%
2000
51%
2000
38%
2000
52%
2000
67%
2000
60%
2000
71%
2000
71%
2000
52%
2000
49%
2000
34%
2000
57%
2000
37%
2000
62%
2000
48%
2000
38%
2000
59%
2000
72%
2000
70%
2000
37%
2000
47%
2000
51%
2000
37%
2000
36%
2000
48%
2000
52%
2000
47%
2000
48%
2000
71%
2000
73%
2000
26%
2000
30%
2000
61%
2000
34%
2000
56%
2000
33%
2000
62%
2000
53%
2000
50%
2000
60%
2000
76%
2000
47%
2000
46%
2000
22%
2000
43%
2000
28%
2000
40%
2000
41%
2000
52%
2000
50%
2000
72%
2000
70%
2000
27%
2000
25%
2000
23%
2000
25%
2000
40%
2000
29%
2000
28%
2000
29%
2000
40%
2000
28%
2000
50%
2000
14%
2000
25%
2000
20%
2000
23%
2000
26%
2000
29%
2000
30%
2000
27%
2000
24%
2000
30%
2000
50%
2000
row deck losesrow deck wins

A simulated metagame

A matchup table is only half the story. What matters competitively is how a deck does against the field it will actually face. And players are going to attempt to anticipate this field and plan to bring a deck that performs well against it. So it's not just a question of "which deck is best", it is about predicting which decks will actually get played, which is influenced by the win rates in the matrix above.

We can try to estimate the share of each deck in the field by treating the metagame as its own self-contained game. Each player selects a deck from the field with a certain probability, and then outcomes are determined by the estimated win rates in the table. Games like that have a Nash equilibrium: a single mix of decks that nobody can beat by switching to something else. Find the mix and you have found the field the format is pulling toward.

To find the Nash equilibrium, we find the deck mix that maximizes the worst case across every possible opponent. The solution has a property that makes it easy to verify: every deck that appears in the mix wins exactly 50% of the time against it, and every deck left out wins less than 50% of the time. Trying to pick a dark horse only leads to suboptimal results, which is what makes it an equilibrium.

The result, and what happens without Boba

I solved for the Nash equilibrium in a field consisting of only these 12 decks. The format collapses to just three viable decks of the 12, with two thirds of the field playing Boba Cunning:

Deck All 12 decks Boba removed
Boba Cunning64.3%
Boba ECL0%
Krennic ECL28.3%19.6%
Palpatine Command7.4%9.8%
Sabine ECL0%38.8%
Iden ECL0%20.4%
Han Command0%11.3%
Five others0%0%

How to read the numbers

The twelve-deck field is not a healthy format. Boba Cunning is dominant, only three decks are playable and nine are not. The best of the nine remaining decks only manages 49.3% against the three.

Now drop both Boba Fett decks and rerun it on the remaining ten. The format that emerges is much healthier: five decks in the mix, the largest at 38.8% rather than 64.3%. The intransitivity gets richer too: Palpatine crushes Iden and Krennic but folds to Sabine, Sabine loses to Han and Krennic, Han loses to Krennic.

It is also a much more robust result. Re-solving the equilibrium two thousand times over resampled matchup data, all five decks appear in at least 99% of the resamples. In the twelve-deck field, 27% of the time the meta collapses to just Boba Cunning. Krennic's place in the meta rests on a 2-point matchup edge, which is not much more than the noise from resampling. And without Krennic in the mix, Palpatine simply loses to Boba.

What this does and doesn't tell us

What this demonstrates is that reinforcement learning and simulations can be used to analyze a meta without playing a single game by hand. We can predict ahead of time whether a deck is likely to become too dominant and we can explore how banning specific cards will affect the new meta that forms in its wake.

However, these are not predictions of what people would have brought with them to a tournament on any given weekend. They are the answer to a narrower question: what would a field have to look like for no counter-pick to exist?

To answer this, in addition, I made a number of notable simplifications:

Conclusions

Despite the simplifications, we have the benefit of hindsight and can see that this approach supports the same conclusion that FFG ultimately came to (albeit only after Shadow of the Galaxy), which is that Boba Fett was overtuned. Simulations powered by reinforcement learning can reveal how a player can approach a given matchup with ease.

Questions about our findings? Email admin@swoolab.com.

Play against the model