Transcription of Propensity Analysis in Stata Revision: 1
1 Propensity Analysis in StataRevision: LuntOctober 14, 2014 Contents1. The data used in this tutorial .. Additional programs required ..42. Checking Balance53. Calculating Propensity Using Logistic Regression .. Diagnostics for the Propensity score .. Using the Propensity score .. Stratification .. Weighting .. matching ..154. Rechecking Stratification .. Weighting .. matching ..185. Assessing the effect of Na vely .. Stratification .. Weighting .. matching ..226. Trimming247. Alternative Analyses2618.
2 Conclusions27A. Complete do file for tutorial28B. Do file used to generate example dataset292 List of of Propensity score .. of Log Odds of Propensity score .. of observed and predicted associations between confounders andtreatment .. of case-control differences after na ve matching .. of case-control differences after matching with a caliper of .17 List of and standard deviation of confounders in treated and untreated balance of confounders between treated and untreated .. test on initial Propensity model .. interactions in Propensity model .. of fit of improved Propensity model.
3 Test for non-linearity .. into quintiles of Propensity score .. into deciles of Propensity score .. balance of confounders between treated and untreated after stratify-ing into quintiles .. balance of confounders between treated and untreated after stratify-ing into deciles .. balance of confounders between treated and untreated after weighting balance of confounders between treated and untreated after na vematching .. balance of confounders between treated and untreated after match-ing within caliper .. ve Analysis of the effect of treatment .. of the effect of treatment, stratifying by Propensity score in 5 strata.
4 Of the effect of treatment, stratifying by Propensity score in 10 of the effect of treatment, using weighting .. of the effect of treatment, using weighting, restricted to commonsupport .. of the effect of treatment, using na ve matching .. of the effect of treatment, using matching with a caliper .. Analysis of the effect of treatment, using matching with caliper .. of the effect of treatment, using weighting, trimmed at the fifth centile of the effect of treatment, using weighting, trimmed at the fifth centile 2631. IntroductionPropensity scores can be very useful in the Analysis of observational studies.
5 They enableus to balance a large number of covariates between two groups (referred to as exposed andunexposed in this tutorial) by balancing a single variable, the Propensity score . There are threeways to use the Propensity score to do this balancing: matching , stratification and will explore all three ways in this models depend on the potential outcomes model popularized by Don Rubin[1].In this model, we assume every subject has two potential outcomes: one if they were treated,the other if they are not treated. The aim is to compare treated subjects to untreated subjectswith the same potential outcomes: this ensures that the difference between treated and un-treated subjects is due to the treatment, since the outcomes in both groups would have beenthe same had the treated subjects not received treatment.
6 Rosenbaum and Rubin [2] haveshown that subjects with the same Propensity score have, on average, the same potential out-comes, so comparing treated and untreated subjects with the same Propensity score gives anunbiased estimate of the effect of The data used in this tutorialWe will use simulated data for this tutorial, since that way we can know what the correctanswer is, and compare the results we get with different methods with the correct answer. Theoutcome we are interested in is the variabley, which is normally distributed. The treatmentvariable,t, has the effect of reducingyby 1.
7 However, there are three confounding variables,x1,x2andx3: an increase in any of these variables increases the probability of receivingtreatment, and also increases the outcomey. So you can think ofyas being a measure ofdisease severity, with those with the highest disease severity being more likely to receivetreatment. This data can be loaded into Stata with the commandsglobal datadir "$ " Additional programs requiredI have written some ado-files which make Analysis with Propensity scores a little easier, andwhich we will use throughout this tutorial. They can be downloaded by entering the followingcommand in Stata :net from clicking on Propensity and finally clicking on click here to install.
8 Wewill also use thepbalchkcommand which can be installed in the same Checking BalanceBefore we start analysing the data, it will be useful to see how big a problem we have. We willtherefore compare all of the confounders between the treated and untreated. One way to dothis is with thetabstatcommand: Listing 1 shows how we can get the mean and standarddeviation for each variable in the treated and 1 Mean and standard deviation of confounders in treated and untreated subjects. tabstat x*, by(t) statistics(mean sd) columns(statistics)Summary for variables: x1 x2 x3by categories of: tt | mean sd---------+--------------------0 |.
9 8412321| | +--------------------1 | .4777602 .943132| | +--------------------Total | .0087739 .9968548| | can see that there is a difference of about 1 inx1, 2 inx2and 6 inx3. However,since we don t know the units in which these variables are measured in, we don t know if thedifference inx3is more important than the difference inx1or not. We could do a significancetest, but that is very sample-size dependent, and does not tell us how big any differencesbetween treated and untreated are. We are better looking at standardised differences : thedifference in terms of standard can get this data easily from thepbalchkprogram, the syntax of which ispbalchktreatvar testvarswheretreatvaris the treatment variable (in our caset) andtestvarsare the potentialconfounders, in our casex1,x2andx3.
10 The result frompbalchkis shown in Listing 2 Checking balance of confounders between treated and untreated. pbalchk t x1 x2 x3 Mean in treated Mean in Untreated Standardised | | | : Significant imbalance exists in the following variables:x1 x2 x35 The output shows us that the treated and untreated differ by about 1 SD inx1andx2, andby SD inx3. So the treated and untreated are more similar inx3than they are Calculating Propensity Using Logistic RegressionWe use logistic regression to calculate the Propensity scores.