Transcription of Post Hoc Power: Tables and Commentary
1 post Hoc Power: Tables and CommentaryRussell V. LenthJuly, 2007 The University of IowaDepartment of Statistics and Actuarial ScienceTechnical Report No. 378 AbstractPost hoc power is the retrospective power of an observed effect based on thesample size and parameter estimates derived from a given data set. Many scientistsrecommend using post hoc power as a follow-up analysis, especially if a finding isnonsignificant. This article presents Tables of post hoc power for commontandFtests. These Tables make it explicitly clear that for a given significance level, post hocpower depends only on thePvalue and the degrees of freedom. It is hoped that thisarticle will lead to greater understanding of what post hoc power is and is not. Wealso present a grand unified formula for post hoc power based on a reformulationof the problem, and a discussion of alternative words: post hoc power, Observed power,Pvalue, Grand unified formula1 IntroductionPower analysis has received an increasing amount of attention in the social-scienceliterature ( , Cohen, 1988; Bausell and Li, 2002; Murphy and Myors, 2004).
2 Usedprospectively, it is used to determine an adequate sample size for a planned study (see,for example, Kraemer and Thiemann, 1987); for a stated effect size and significance levelfor a statistical test, one finds the sample size for which the power of the test will achievea specified studies are not planned with such a prospective power calculation, however;and there is substantial evidence ( , Mone et al., 1996; Maxwell, 2004) that manypublished studies in the social sciences are under-powered. Perhaps in response to this,some researchers ( , Fagley, 1985; Hallahan and Rosenthal, 1996; Onwuegbuzie andLeech, 2004) recommend that power be computed retrospectively. There are differingapproaches to retrospective power, but the one of interest in this article is a powercalculation based on the observed value of the effect size, as well as other auxiliaryquantities such as the error standard deviation, while the significance level of the test isheld at a specified value.
3 We will refer to such power calculations as post hoc power (PHP). Advocates of PHP recommend its use especially when a statisticallynonsignificant result is obtained. The thinking here is that such a lack of significancecould be due either to low power or to a truly small effect; if the post hoc power is foundto be high, then the argument is made that the nonsignificance must then be due to asmall effect is substantial literature, much of it outside of the social sciences ( , Goodmanand Berlin, 1994; Zumbo and Hubley, 1998; Levine and Ensom, 2001; Hoenig andHeisey, 2001), that takes an opposing view to PHP practices. Lenth (2001) points out thatPHP is simply a function of thePvalue of the test, and thus adds no new and Maxwell (2005) show that PHP does not necessarily provide an accurateestimate of true power. Hoenig and Heisey (2001) discuss several misconceptionsconnected with retrospective power.
4 Among other things, they demonstrate that when atest is nonsignificant, then the higher the PHP, the more evidence there isagainstthe nullhypothesis. They also point out that, in lieu of PHP, a correct and effective way toestablish that an effect is small is to use an equivalence test (Schuirmann, 1987).In this article, we derive and present new Tables that directly give exact PHP for allstandard scenarios involvingttests (Section 2) andFtests (Section 3). (The PHP ofcertainztests and 2tests can also be obtained as limiting cases.) All that is needed toobtain PHP in these settings is the significance level, thePvalue of the test, and thedegrees of freedom. If one desires a PHP calculation, this is obviously a convenientresource for obtaining exact power with very little effort; however, the broader goal is todemonstrate explicitly what PHP is, and what it is not.
5 In Section 4, we present a slightreformulation of the PHP problem that leads to a grand unified formula for post hocpower that is universal to all tests and is a simple head calculation. The results arediscussed in Section 5, along with possible alternative practices regarding 1 may be used to obtain the post hoc power (PHP) for most common one-andtwo-tailedttests, when the significance level is =.05. The only required information(beyond ) is thePvalue of the test and the degrees of freedom. Computational detailsare provided later in this section; for now, here is an illustration based on an example inHallahan and Rosenthal (1996). They discuss the results of a hypothetical study where anew treatment is tested to see if it improves cognitive functioning of stroke are 20 patients in the control group and 20 in the treatment group, and theobserved difference between the groups is.
6 4 standard deviations somewhat short of a medium effect on the scale proposed by Cohen (1988) with aPvalue of .225(two-sample pooledttest, two-tailed). In this case, we have =38 degrees of to the bottom half of Table 1 (for 2-sided tests) and linearly interpolating, weobtain a post hoc power of about .234 (the exact value, using the algorithm used toproduce Table 1, is .2251.) This agrees with the value of .23 reported in the briefly discuss some patterns in these Tables . First, PHP is a decreasing function of2 Table 1: post hoc power of attest when the significance level is =.05. It depends on thePvalue, the degrees of freedom , and whether it is one- or two-tailed. post hoc power ofaztest may be obtained using the entries for = .Pvalue of testAlternative , for any number of degrees of freedom and alternative.
7 In general, except forvery small degrees of freedom, the power of a marginally significant test (P= =.05) isaround one half, with the two-tailed powers generally higher than the one-tailed the test is significant, the power is higher than .5; and when the test is nonsignificant,the power is usually less than .5. Thus, it is an empty question whether the PHP is highwhen significance is not Derivation of the tablesConsider the null hypothesisH0: = 0, where is some parameter and 0is a specifiednull value (often zero). We have available an estimator , and thetstatistic has the formt= 0se( )(1)wherese( )is an estimate of the standard error of whenH0is true. Assume that:1. For all , is normally distributed with mean ; its standard deviation will bedenoted .32. For all , se( )2/ 2has a 2distribution with degrees of freedom. The value of is andse( )are conditions hold for most commont-test settings, such as a one-sample test of amean, pooled or paired comparisons of two means, and tests of regression coefficientsunder standard homogeneity us re-write (1) in the formt=[( )/ ] + [( 0)/ ]se( )/ =Z+ Q/ (2)where = ( 0)/.
8 According to the stated assumptions,ZandQare independent,Zis standard normal, andQis 2with degrees of freedom. This characterizes thenoncentraltdistribution with degrees of freedom and noncentrality parameter . (See,for example, Hogg et al., 2005, page 442). The power of the test is then defined asP(t RH1, ), whereRH1, is the set oftvalues for whichH0is rejected, based on thestated alternativeH1and significance level .Notice that the form of = ( 0)/ is exactly that of thetstatistic, with populationvalues substituted in place of andse( ). In calculating PHP, we substitute the observedvalues of and the observed error standard deviation (and thus the observedse( )) fortheir population counterparts; thus, the noncentrality parameter used in PHP is =t,the observedtstatistic itself. If one is given only thePvalue and the degrees of freedom,the inverse of thetdistribution may be used to obtain the observedtstatistic (or itsabsolute value, in the case of the two-tailed test), hence the noncentrality parameter ,hence the post hoc power.
9 Table 1 is computed using this process. Computations wereperformed in the R statistical package (R Development Core Team, 2006), using itsbuilt-in functionsqtandpt(percentiles and cumulative probabilities of the central ornoncentraltdistribution). post hoc power of certainztests can be obtained from the limiting case when .This can be verified by noting that thezstatistic has the same form as (1) withse( )set toits known value . Then the denominator in (2) reduces to 1. However, keep in mind theunderlying condition in our derivation that the standard error of is regardless of thetrue value of ; this condition doesnothold inztests involving proportions, because thestandard error of a proportion depends on the value of the proportion 2 provides PHP values for a variety of fixed-effectFtests such as those obtained inthe analysis of linear models with homogeneous-variance assumptions.
10 Given asignificance level of =.05 (the only case covered in the Tables ), the only otherinformation needed to obtain PHP is thePvalue and the numerator and denominatordegrees of freedom ( 1and 2respectively). For example, suppose that we have datafrom an experiment where scores were measured on 40 children randomly assigned to 54 Table 2: post hoc power of a fixed-effectsFtest when the significance level is =. depends on thePvalue of the test and the degrees of freedom for the numerator ( 1)and the denominator ( 2). post hoc power of a 2test with 1degrees of freedom may beobtained using the entries for 2= . post hoc power for 1=1 may be obtained fromthe two-tailedt-test results in Table 1, with = of test 1