cosmiccreature

On Banditry

This post will be a little bit of a side-quest; i had intended for this to be more comprehensive, but I'll be busy for a few weeks and didn't want to leave it until then.

Multi-Arm Banditry

You wake up, dazed and confused, the sound of metal jingling in your heavy right pocket. Reaching down, you pull out a heavy sack of about a thousand hard metal coins. The lights around you slowly flicker, the hum of machinery surrounding you. The first color you notice is oversaturated golds painted on red machines, then the ugly expanse of green and blue spill-resistant carpet. Slot Machines. Dozens… hundreds… thousands of them? The screens power on, the machines roaring for your attention, demanding those otherwise worthless tokens. Every slot machine is different, unique, and each more confusing than the last in this Backrooms-esque reality you have found yourself. No Fair Gambling Control Commission regulates or maintains these machines.

With nothing else to guide you, you try a random slot machine with one of your tokens. It beeps, whirrs, and a number spits out: "You win! +3". You look for a chute for new tokens to spill out, expecting three. But you find nothing. You ponder, and try a few more. Some you win, some you lose. But none give you a payout, only the promise of one. After exploring, you find a single, shiny golden token on a pedestal, shining in magnificent glory. Music plays in the background as you pick it up.

The Test

Using your supply of ~1,000 tokens to test the limitless machines that surround you, what strategies can you come up with to maximize your single real attempt with the golden token? No, really—pause and think about it for a few seconds. Not for me, but for you.

The Solution

You likely considered two opposite strategies: (1) Exploration, (2) Exploitation. Exploration spreads out your coins wide: you try 1,000 slot machines with your 1,000 tokens, remembering which singular slot machine gave the best reward. Then you go back to it with your golden token, and hope you can maximize your single-shot payout. Exploring broad works fantastic if there's only a few consistently-high expected value machines. In the other extreme, you find two machines and test them as exhaustively as you can, splitting your coins equally to try each 500 times. Now you should know which machine of the two is better and by how much; but you lack any idea of how good those machines are relative to the average machine in the casino 1.

Intuitively, there must be some midpoint between these two strategies that balances sampling enough to estimate the quality of a handful of machines, but enough variety to not find one that is rigged decisively in your favor. Maybe you came up with a strategy like that: something like sampling 20 machines 50 times. It should also be reasonable to assume that the optimal sampling strategy changes depending on exactly what your token budget is. Let me offer another solution:

  • Allocate your coins into 4 buckets: the first 3 have 300 tokens apiece, with the fourth bucket only having 100.
  • Test 100 different machines, 3 times apiece (300 coins).
  • Eliminate the "unlucky" bottom 75%, leaving you with 25 machines. Test each of them 12 times (300 coins). Now you've tested 25 candidate machines 15 times, with 225 tokens "wasted" on exploration.
  • Eliminate the "unlucky" bottom 75%, leaving you with 8 machines. Test each of them 37 times (296 coins). Now you've tested 8 candidate machines 52 times, with 480 tokens "wasted" on exploration.
  • Eliminate the "unlucky" bottom 75%, leaving you with 3 machines. Test each of them 34 times (102 coins). Now you've tested 3 candidate machines 86 times, with 740 tokens "wasted" on exploration.
  • Throw your remaining final 2 test tokens on your best machine. Now you've tested your pick 88 times, with 912 "wasted" on exploration.

This strategy of exploring wide-then-refine is a great strategy to adapting to any given sampling budget and sample space. Our first pass of 100 machines identifies both consistently-high-value machines and high-but-somewhat-infrequent payoffs, while rapidly filtering out low-value machines that seem rigged against us. The second pass eliminates most of the spuriously lucky machines that happened to high-roll during our first pass despite being overall low-EV. Our final pass lets us categorize our candidates into "inconsistent but high value" and "consistent positive value". With 86 attempts on each of our best three machines, we have a decent 2 idea of each machine's luck profile, and we feel reasonably happy with our final guess.

This strategy will miss a super-high-value-machine-that-happened-to-be-extremely-unlucky-3-times-in-a-row in the first pass that a more naive solution like "sample 20 machines 50 times", we are also 5x more likely to actually explore that super-high-value machine at all. Without any information about how rigged (or not) this casino is, we simply can't estimate the odds of that occurring. There's lots of ways that this strategy can be further improved.


  1. Strictly speaking, this task is impossible because this strange casino is a Cauchy distribution in disguise. While samples from Cauchys have observable moments, they can't perfectly estimate the hidden distribution with those samples. ↩

  2. is it decent? how do we know this? well i had to go refresh my stats: for an original 1899 style slot machine with 3 reels and 5 symbols there's 125 positions. Once you can prove it is fair, then you estimate your payoff by the sum of the paylines divided by 125 minus 1 token for entry. This needs about 620 rolls (124 degrees of freedom*5) to test if the machine is entirely fair—so, nothing like "if reels 1&2 roll apple, avoid apple on reel 3". If you're looking only to see if each reel is individually fair (12 dof but you can't detect the previous example of cheating, but you can see if reel 1 consistently avoids lucky 7, reducing the odds of your best payline) you only need 19-33 attempts; this info could be used to refine our sampling strategy. Modern machines have 5 reels with dozens of adaptive symbols, so yeah, we can't test those this way. Still, part of the point of the metaphor is that these machines aren't fair: this competition isn't about randomly selecting from only the best from your fair options; it's about exploiting powerful stuff that reliably wins much better than average and not playing the weak stuff that just loses. So is measuring whether or not a machine is "fair" or not even what you're trying to achieve in this scenario? No, you want to find the unfair machine that's in your favor. So, i'm just going to be content with my colloquial use of "decent" here. ↩