Transcription of STATISTICAL GRAND ROUNDS Equivalence and …
1 STATISTICAL GRAND ROUNDSE quivalence and noninferiority testing in RegressionModels and Repeated-Measures DesignsEdward J. Mascha, PhD,* and Daniel I. Sessler, MD Equivalence and noninferiority designs are useful when the superiority of one intervention overanother is neither expected nor required. Equivalence trials test whether a difference betweengroups falls within a prespecified Equivalence region, whereas noninferiority trials test whethera preferred intervention is either better or at least not worse than the comparator, withworsebeing defined a priori. Special designs and analyses are needed because neither of theseconclusions can be reached from a nonsignificant test for superiority. Using the data from acompanion article, we demonstrate analyses of basic Equivalence and noninferiority designs,along with more complex model -based methods.
2 We first give an overview of methods for designand analysis of data from superiority, Equivalence , and noninferiority trials, including how toanalyze each type of design using linear regression models. We then show how the analogoushypotheses can be tested in a repeated-measures setting in which there are multiple outcomesper subject. We especially address interactions between the repeated factor, usually time, andtreatment. Although we focus on the analysis of continuous outcomes, extensions to other datatypes as well as sample size consideration are discussed. (Anesth Analg 2011;112:678 87) Equivalence and noninferiority designs are gainingpopularity in medical research, and for good an era when multiple successful treatments areavailable for many conditions and diseases, investigatorsoften compare a new treatment to an existing one and askwhether the new treatment is at least as ,2 Forexample, it is now rarely considered ethical to compare anovel treatment with placebo when effective treatments arealready available.
3 However, it is often of considerableinterest to evaluate whether a new treatment is at least aseffective as an existing one, especially if the novel treatmentis less expensive, easier to use, or causes fewer side these comparative efficacy or active-comparator trials, the hypothesis is that 2 treatments, perhaps 2 anes-thetics, have comparable of comparability or Equivalence are not justifiedfrom a nonsignificant test for superiority, because thenegative result may simply result from a lack of power(Type II error) in the presence of a truly nontrivial popu-lation effect. Rather, an Equivalence design is needed, withthe null hypothesis being that the difference betweenmeans or proportions is outside of an a priori specifiedequivalence ,4If the observed confidence interval(CI) lies within the a priori region, the null hypothesis isrejected, and Equivalence claimed.
4 In addition to CIs, STATISTICAL tests can be used to assess whether the truedifference lies within the Equivalence designs are useful when the goal is todemonstrate that a preferred treatment is at least as goodas or not worse than a competitor or standard ,6 For example, if the preferred treatment is lessexpensive or safer, it would suffice to show it was at leastnot worse (and perhaps better) than a comparator on theprimary measure of efficacy. Also, cost effectiveness mightbe assessed, for example, by simply requiring noninferior-ity on either cost or effectiveness, and superiority on was the approach taken in the design andanalysis of the companion paper in this issue of the journalby Ruetzler et al.
5 ,7in which researchers tested the hypoth-esis that intraoperative distal esophageal (core) tempera-tures are not! C lower (a priori specified noninferiority !) during elective open abdominal surgery under generalanesthesia in patients warmed with a warm water sleeve onone arm than with an upper body forced air cover. Patientswere randomly assigned to intraoperative warming witheither a circulating water sleeve (n"37) or forced air (n"34); intraoperative core temperature was measured every15 minutes, beginning 15 minutes after intubation (Fig. 1).Because temperatures were recorded over time, the Ruet-zler et al. trial was a repeated-measures design. We usethese data to illustrate various approaches to noninferiorityas well as to Equivalence and superiority this article, we refer to Ruetzler et al.
6 As thecompanion paper. Figure 2 depicts sample CIs and the appropriate infer-ence for the 3 types of designs that we discuss. In asuperiority trial the null hypothesis of no difference is onlyrejected if the observed CI for the treatment difference doesnot overlap zero. Thus, result A in Figure 2 can claimsuperiority of test treatment T to standard S, but result Bcannot. In a noninferiority trial, one treatment is deemed not worse than the other only if the CI for the differencelies above a prespecified noninferiority !(thus, result C canclaim noninferiority , but result D cannot). Finally, in anequivalence trial, 2 treatments are deemed equivalent From the *Department of Quantitative Health Sciences and Department ofOutcomes Research, Cleveland Clinic, Cleveland, for publication November 9, : No authors declare no conflict of digital content is available for this article.
7 Direct URL citationsappear in the printed text and are provided in the HTML and PDF versionsof this article on the journal s Web site ( ).Reprints will not be available from the correspondence to Edward J. Mascha, PhD, Department of Quan-titative Health Sciences, Cleveland Clinic, 9500 Euclid Avenue, JJN3 Cleve-land, OH 44195. Address e-mail to 2011 International Anesthesia Research SocietyDOI: 2011 Volume 112 Number 3only if the CI falls within the prespecified equivalenceregion (result E can claim Equivalence , but result F cannot).Our goal is to review STATISTICAL approaches for equiva-lence and noninferiority trials in various clinical settings;for comparison, we also briefly present analysis of conven-tional superiority trials.
8 We first review basic methods fordesign and analysis of each type of 12We thendemonstrate how to analyze these designs in a linear regres-sion model , including the repeated-measures setting in whichthere are multiple outcomes per subject, as in the companionpaper. We give particular attention to assessing the interactionbetween the repeated factor, which is usually elapsed time inperioperative studies, and the intervention. Although wefocus on continuous outcomes, we briefly review noninferi-ority and Equivalence testing methods for some additionaloutcome types. Sample size considerations for these designsare also discussed. Throughout this article, we illustrate ourexamples with data from the companion , Equivalence , ANDNONINFERIORITY DESIGNS THE BASICSS uperiority TestingFor a study designed to assess superiority of one interven-tion over another for a continuous outcome, the null andalternative hypotheses areH0:"E#"S$0 and H1:"E#"S%0,(1)where"Eand"Sare the population means for the respec-tive experimental (E) and standard (S) interventions.
9 As-suming a normal distribution for the outcomes in eachgroup and equal variances, we use the Studentttest toassess superiority of E to S (or S to E). The test statistic isTsup$" E#" S!sp2#1/nE&1/nS$,(2)where" Eand" Sare the observed means (and" E%" Sestimates the treatment effect),SPis the pooled estimateof the common SD across groups, whereSP$"#nE#1$sE2&#nS#1$sS2nE&nS#2#1/ 2,sE2andsS2are the observedvari-ances ( , squared standard deviations) andnEandnSare thesample sizes for the E and S interventions, respectively. Thedenominator of Equation (2) is the estimated SD of" E#" S,alsocalledtheestimatedSEofthedifferenc e,orSE " E#" a 2-sided test of superiority we compare the absolutevalue ofTsupto atdistribution with nE%nS%2 degrees offreedom (df) at the'/2 level, where'( alpha ) is the apriori specified significance level or type I error for thestudy, typically The 2-sided superiorityPvalue istwicethe probability of observing a value greater than$Tsup$if the null hypothesis were true.
10 The null hypothesis isrejected if thePvalue is smaller than the designated'.Correspondingly, the null hypothesis is rejected if the100(1 ')% CI does not overlap primary outcome in the companion paper was coretemperature during surgery. The study was designed as anoninferiority trial; but as an example, we first apply a testof superiority to the dataset. Because the study had arepeated-measures design, with temperature measured ev-ery 15 minutes intraoperatively, the dataset has 1 row persubject per repeated measurement, with variables site_id(1"Cleveland Clinic, 2"Medical University of Vienna),pt_id"patient ID, sleeve (1"warming sleeve, 0"forcedair), time_m (minutes after induction), esophtemp (esoph-ageal temperature at specified time) and preoperative full dataset is included in Appendix 1, which consistsof raw data from the companion paper7(see SupplementalDigital Content 1, )