Hypothesis Testing: The Two Kinds of Error and the Trade-off Between Significance and Power

Sayantan R. Majumdar, Ritabrata K. Ganguly

Abstract


A statistical hypothesis test decides, from data, whether to reject a claim, such as that a treatment has no effect, but because the data are random the decision can be wrong in two distinct ways: the test may reject the claim when it is in fact true, or fail to reject it when it is in fact false; these are the two kinds of error, and they cannot both be made small at once for a fixed amount of data. This study examines hypothesis testing, the two kinds of error, and the trade-off between significance and power. The chances of the two errors were examined as the threshold for rejecting the claim was varied, the distributions of the test result under the claim and under a real effect overlapping, so that the trade-off between the two errors could be seen. The two errors traded off against each other: setting a strict threshold, so that the claim is rejected only on strong evidence, made it unlikely to reject a true claim, a small chance of the first error, the significance level, but more likely to miss a real effect, a large chance of thesecond error and so low power; setting a lenient threshold did the reverse,catching real effects more often but rejecting true claims more often too. The two errors could not both be made small at once for a fixed sample, because the distributions of the test result under the claim and under an effect overlapped, and moving the threshold to reduce one error enlarged the other. The only way to reduce both together was to gather more data, which separated the two distributions and so shrank the overlap. The threshold was therefore chosen to fix the chance of the first error at a small, conventional level, accepting the resulting power, with the sample made large enough that the power was adequate. The study shows that the two errors of a test trade off through the rejection threshold and are reduced together only by more data, and it discusses the implications for the design and interpretation of statistical tests. KEYWORDS: Hypothesis testing, Type I error, Type II error, Significancelevel, Statistical power, Null Hypothesis, Sample size, Statistical inference,Trade-off

Full Text:

PDF 9-16

Refbacks

  • There are currently no refbacks.