Starting a Fresh Group

When forming a new group to play RPGs with, there are quite a few considerations to take into account. So many, in fact, that I don’t think I can hit them all. I’m also certain I can’t even think of them all. However, I have a list below that covers the major areas. Also, this […] Over the last few weeks we’ve been looking at the the analysis done for the Kickstarter for “Honest Dice | Precision Machined Metal Dice You Can Trust” as a case study. We’ve gone over the general idea of how one would go about testing dice and presenting results. We’ve also gone over the analysis that was done and what was wrong with it. This week we’re looking at how it might have been done and what conclusions we can actually draw from the testing that was done and the data that was collected. Using the template established earlier an ideal analysis would look something like this: The intent of our testing is to show that Honest Dice are fairer, that is closer to the ideal distribution than several other die options. Since we’re testing dice against each other, the appropriate test is the Chi Square Test of Homogeneity. This test is specifically designed to test if sets of dice or other phenomenon have the same distribution or not. Specifically this test checks the hypothesis H0: all dice have the same distribution. Statistical significance in this test will tell us that at least one of the dice has a different distribution. Since the base test only determines if at least one die is different, if a difference is detected in the D20s or the D4s, follow up tests of homogeneity will need to be performed to determine where the differences lie. While there are 6 possible pairings of dice that might be different for the d20s and 3 possible pairings for the d4s, each additional test increases our potential family wise error rate and we’re really only interested in differences between the Honest Dice and the other dice being tested if a difference is detected in the D20s, we’ll do three follow up tests: the Honest Dice D20 vs the three other D20 options. For the D4s, if a difference is detected we’ll do two follow ups: the Honest D4 vs each of the other D4 options. Finally, with these follow up tests, significance is still only showing a difference in the two dice, not which is better. In those cases, it’s finally time to do goodness of fit tests to test the two dice against another. In this final set of tests we don’t need to worry about thresholds or family wise error rates because we’re (finally! just looking to compare the two P-values. This comparison of P-values is only valid if the preceding tests showed significance.) For these tests, we’re going to use a .05 threshold for significance. This is a common middle of the road threshold. Given the probabilities involved in D20s (.05 chance of rolling any given side, a .1 threshold seems unreasonably high. .05 frankly seems high too, but because of limited data (see below) going with .01 seems unlikely to give a fair chance for finding significance. If I were designing this test from scratch, I would use a .01 threshold and simply increase sample size but that’s unfortunately not an option. For the follow up tests, since we have Family Wise Error Rate concerns, we need to pick an adjustment to our significance threshold to account for it. Since all of our follow up tests are going to be the Honest Die option vs another die, it’s reasonable to assume that if the Honest Die is the one with the different distribution then the follow up tests are not independent of one another. Thus for a raw threshold number, we’ll use the Bonferroni adjustment, which is to just to divide the intended threshold by the number of tests. Thus for the D20 follow ups with three tests, our threshold will be .05/3=.017 and for the D4 follow ups with two tests, our threshold will be .025. We’ll also use the Holm’s step down procedure which is similar to the Bonferroni adjustment and which makes fewer assumptions about distribution than alternative step tests. We don’t need to use two options, and in fact two options may give us conflicting results, but I’m interested in using both these techniques as I don’t have experience with them, and I don’t want to use both options and only report one. For our D20 tests, our sample size will be 2000, for our D6 tests we’ll use a sample size of 1000 and for our D4 tests we’ll use a sample size of 500. We’re using these sample sizes because that’s the size of the data set we have available, not because those are ideal minimum sample sizes. To determine ideal minimum sample sizes we need to know Effect size (.1), significance threshold (.05, .017, or .025 as appropriate), desired power (.05) and degrees of freedom ( faces-1*dice-1) for each test to feed them into G*Power. Thus for each test the ideal sample sizes I’d like to have for both .05 and .01 base threshold are: Test Threshold Degrees of Freedom Sample Size .05 Sample Size .01 D20 Test .05 (20-1)*(4-1)=57 4533 5647 D20 Follow ups .017 (20-1)*(2-1)=19 3571 4354 D6 Test .05 (6-1)*(2-1)=5 1979 2577 D4 Test .05 (4-1)*(3-1)=6 2086 2705 D4 Follow ups .025 (4-1)*(2-1)=3 1962 2491 Performing all these tests requires a data set. Others doing follow up analysis like this is precisely why providing your data set is a best practice. However, despite the fact that a data set was not shared, one can be reverse engineered from the graphs provided using the following method: Take the image of the graph provided and find the y coordinate of the pixel that forms the 0 line on the graph. For each bar, find the y coordinate of the pixel that

Categories: Blog
Videogames Heaven Blog
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.