Transcription of STAT331 Logrank Test Introduction - Stanford University
1 STAT331 Logrank TestIntroduction:The Logrank test is the most commonly-used statistical testfor comparing the survival distributions of two or more groups (such as dif-ferent treatment groups in a clinical trial). The purpose of this unit is tointroduce the Logrank test from a heuristic perspective and to discuss popu-lar extensions. Formal investigation of the properties of the Logrank test willbe covered in later that we have 2 groups of individuals, say group 0 and group 1. Ingroupj, there underlying survival times with common de-notedFj( ), for j=0,1.
2 The corresponding hazard and survival functions forgroupjare denotedhj( ) andSj( ), usual, we assume that the observations are subject to noninformative rightcensoring: within each group, theTiandCiare want a nonparametric test ofH0:F0( ) =F1( );or equivalently, ofS0( ) =S1( );orh0( ) =h1( ).If we knewF0andF1were in the same parametric family ( ,Sj(t) =e jt),thenH0is expressible as a point/region in a Euclidean parameter space. How-ever, we instead want a nonparametric test; that is, a test whose validity doesnot depend on any parametric the following picture shows, there are many ways in whichS0( ) andS1( )can differ:1 SSSSSSttt101010t0transient differenceparallel after tdiverging0 SSSStt1010late emergingnot stochasticallydifferenceorderedIt is intuitively clear that a UMP (Uniformly Most Powerful) test cannotexist forH0:S0( ) =S1( )vsH1: notH0 Two options in this case are to select a directional test or an omnibus test.
3 (a)directional test:These are oriented to a specific type of difference; ,S1(t) = [S0(t)] for some . As a result, they might (and often do) havepoor power against certain other alternatives.(b)omnibus test:These tests attempt to have some power againstmost or all types of differences. As a result, they sometimes have substantiallylower power than a directional test for certain alternatives. For example, atest might be based on | S1(t) S2(t)|dtover some time is difficult to make the choice between directional tests , or between direc-tional vs omnibus tests , in the abstract.
4 It involves several factors, includingprior expectations of the likely differences, properties of various tests for avariety of settings, and practical consequences of a false negative Test:Early work (1960s) in this area fell along 2 lines:(a) Modify rank tests to allow censoring (Gehan, 1965).(b) Adapt methods used for analyzing 2 2 contingency tables to accom-modate censoring (Mantel, 1966).We introduce the Logrank test from the latter perspective as it easily includestests developed from the former and provides good insight into the propertiesof the Logrank Test Construction:Denote the distincttimes of observed failuresas 1< 2< < k, and defineYi( j) = # persons in groupiwho are at risk at j(i= 0;1;j= 1;2; : : : ; k)Y( j) =Y0( j) +Y1( j) = # at risk at j(both groups)dij= # in groupiwho fail (uncensored) at j(i= 0;1;j= 1;2.)
5 K)dj=d0j+d1j= total # failures at jThe information at time jcan be summarized in the following 2x2 table:observed toat riskfail at jat jgroup 0d0jY0( j) d0jY0( j)group 1d1jY1( j) d1jY1( j)djY( j) djY( j)Note:d0j=Y0( j) can be viewed as an estimator ofh0( j).SupposeH0:F0( ) =F1( ) holds. conditional on the 4 marginal totals, asingle element (sayd1j) defines the table. Furthermore, with this condition-ing and assumingH0,d1jhas the hypergeometric distribution; that is:3P[d1j=d] =(djd)(Y( j) djY1( j) d)/(Y( j)Y1( j))ford=max(0; dj Y0( j)); ; min(dj; Y1( j)):The mean and variance ofd1junderH0are thusEj=(Y1( j)Y( j))djVj=Y( j) Y1( j)Y( j) 1 Y1( j)(djY( j))(1 djY( j))=Y0( j)Y1( j)dj(Y( j) dj)Y( j)2(Y( j) 1)DefineOj=d1j.
6 Fisher s test would tell us to consider extreme values ofd1jas evidence , defineO= kj=1Oj= total # failures in group 1E= kj=1 EjV= kj=1 Vjand letZ=O E V= j(Oj Ej) jVj:Then underH0, it is argued thatZapx N(0;1)(or thatZ2apx 21)This approximation can be used to obtain an approximate test forH0bycomparing the observed value of Z (orZ2) to the tail area of the standardnormal (chi-square) :Group 0 : 3:1;6:8+;9;9;11:3+;16:2 Group 1 : 8:7;9;10:1+;12:1+;18:7;23:1+Thenk= 5 and 1; : : : ; 5= 3:1;8:7;9;16:2;18:7 Group 0 Group 1 1= 3:11560661 11 12 2= 8:70441561 9 10 3= 92241453 6 9 4= 16:21010221 2 3 5= 18:70001121 1 2Oj=01101Ej=1=26=1015=92=31Vj=1=46=255=9 2=90O= 3; E= 3:44; V= 1:26; Z= :39 (2-sidedP=:70)Comments: Note that the test statistics is only affected by ranks of the observedtimes (both censored or failure).
7 WhileEjmay be a conditional expectation for eachj, it is not clear thatEhas such an interpretation. Also, the creation ofZand its approx-imation as aN(0;1) suggests that the contributions from each jare independent. Is this true/accurate? Then, isZL N(0;1) underH0? Note the similarity of the Logrank test to techniques for combining 2 2tables across strata ( , cities). Note that the sequencesY0( 1); Y0( 2); Y0( 3); : : :andY1( 1); Y1( 2); Y1( 3); : : :are nonincreasing, and as soon as one reaches 0 [ ,Y0( 5) = 0 at5 5= 18:7], it mustfollow thatOj=EjandVj= 0 at and beyond thistime.
8 Thus, we would get the same answer ( ,Z) if the constructionstopped at the last time when bothY0( j) andY1( j) are>0. Although it is not obvious from the construction, the Logrank test is adirectional test oriented towards alternatives whereS1(t) = (S0(t)) , orequivalently, whenh1(t)=h0(t) = . We will see later that the logrankstatistic arises as a score test from a partial likelihood function for Cox sproportional hazards model. While the heuristic arguments leading to the approximation of the nulldistribution of the Logrank test seem reasonable, is the result correct?
9 In addition, how does the test behave as a function of the amount ofcensoring or the hazard functions in the two treatment groups? Wereturn to these important practical questions in later Some Extensions of the Logrank Test:Strati ed Logrank test: Suppose that we have two groups (say, 2 treat-ments), as before, but that we want to control (adjust) for a categoricalcovariate ( , gender). Then there are 4=2x2 types of individuals. Forexample, their respective survivor functions might be as shown below. Ifwe still want to compare treatment groups, but also adjust for gender, astratified Logrank testcould be used.
10 SupposeS(l)j( ) denotes the survivalfunction for groupjin stratuml, and considerH0:S(l)0( ) =S(l)1( ); l=1; ; }}FemalesMalesS(t)ttr. 1tr. 07 The stratified Logrank test is useful when the distibution of the stratum vari-able in the two groups is not the same, but the distribution of the relevantcovariates in each stratum is the same in both groups (within each stratum,the groups have a comparable prognosis). The stratified Logrank test can alsobe useful to gain data intoLgroups, whereL= # levels of the categorical co-variates on which you want to stratify ( ,L= 2 when stratifying bygender) ; E; V(say,O(l); E(l); V(l)) within each group, just as withthe ordinary l=1(O(l) E(l))vuutL l=1V(l)apx N(0;1) underH0 Note 1: Intuitively, it should be clear how this statistic attempts to adjustfor the stratification variable and, assuming the [O(l); E(l); V(l)] are approxi-mately uncorrelated, that the statistic will be approximatelyN(0;1).