Example: tourism industry

Econ 582 Fixed Effects Estimation of Panel Data

Econ 582 Fixed Effects Estimation of Panel DataEric ZivotMay 28, 2012 Panel Data Framework =x0 + =1 (individuals); =1 (time periods)y 1=X ( ) ( 1)+ Main question: Isx uncorrelated with ?1. If yes, then we have a SUR type model with common If no, then we have a multi-equation system with common coefficients andendogenous regressors. We need to use an Estimation procedure to deal withthe endogeneity. In the Panel set-up, under certain assumptions, we can dealwith the endogeneity without using instruments using the so-calledfixed effects(FE) Components Assumption = + =unobservedfixed effect [x ]=0 [ 0 ]= Thefixed effect component (which is actually an unobserved random vari-able) captures unobserved heterogeneity across individuals that isfixed the error components assumption, the RE and FE models are defined asfollows:RE model: [x ]=0FE model: [x ]6=0 Example: Panel wage equation 69 = + 69 + + 69 + + 69 80 = + 80 + + 80 + + 80 [ ]6=0 Here captures unobserved ability that is correlated with.

Remark: With panel data, as we saw in the last lecture, the endogeneity due to unobserved heterogeneity (i.e., [x ] 6=0 ) can be eliminated without the use of instruments. To see this, consider the difference in log-wages over time:

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Econ 582 Fixed Effects Estimation of Panel Data

1 Econ 582 Fixed Effects Estimation of Panel DataEric ZivotMay 28, 2012 Panel Data Framework =x0 + =1 (individuals); =1 (time periods)y 1=X ( ) ( 1)+ Main question: Isx uncorrelated with ?1. If yes, then we have a SUR type model with common If no, then we have a multi-equation system with common coefficients andendogenous regressors. We need to use an Estimation procedure to deal withthe endogeneity. In the Panel set-up, under certain assumptions, we can dealwith the endogeneity without using instruments using the so-calledfixed effects(FE) Components Assumption = + =unobservedfixed effect [x ]=0 [ 0 ]= Thefixed effect component (which is actually an unobserved random vari-able) captures unobserved heterogeneity across individuals that isfixed the error components assumption, the RE and FE models are defined asfollows:RE model: [x ]=0FE model: [x ]6=0 Example: Panel wage equation 69 = + 69 + + 69 + + 69 80 = + 80 + + 80 + + 80 [ ]6=0 Here captures unobserved ability that is correlated with.

2 Itisassumedthat ability does not vary over : The existence of guarantees across time error correlation: [ 69 80]= [ + 69 + 80 ]= [ 2 ]+ [ 80]+ [ 69]+ [ 69 80]6=0 Remark: With Panel data, as we saw in the last lecture, the endogeneity due tounobserved heterogeneity ( , [x ]6=0) can be eliminated without theuse of instruments. To see this, consider the difference in log-wages over time: 80 69 =( )+ ( 80 69 )+ ( 80 69 )+( )+ 80 69 = ( 80 69 )+ ( 80 69 )+ 80 69 Hence, we can consistently estimate and by using thefirst differenced data! Fixed Effects EstimationKey insight: With Panel data, can be consistently estimated without are 3 equivalent approaches1. Within group estimator2. Least squares dummy variable estimator3. First difference estimatorWithin group estimatorTo illustrate the within group estimator consider the simplified Panel regressionwith a single regressor = + + [ ]6=0 [ ]=0 Trick to removefixed effect :First, for each average over time = + + =1 X =1 =1 X =1 =1 X =1 Second, form the transformed regression = ( )+( )+ or = + Now, stack by observation for =1 giving the giant regression y = x + or y 1= x 1+ 1 The within-group FE estimator is pooled OLS on the transformed regression(stacked by observation) =( x0 x) 1 x0 y= X =1 x0 x 1 X =1 x0 y Remarks1.

3 Ifx does not vary with ( =x )then x =0and we cannotestimate 2. Must be careful computing the degrees of freedom for the FE are total observations and parameters in so it appreas that thereare degrees of freedom. However, you lose 1 degree of freedom for eachfixed effect eliminated. So the actual degrees of freedom are = ( 1) Matrix Algebra Derivation of Within Group Fixed Effects EstimatorConsider the general model (assume all variables vary with and ) =x0 + + Stack the observations for =1 givingy 1=X ( ) ( 1)+ (1 1)1 ( 1)+ 1 DefineQ =I 1 (10 1 ) 110 =I P P =1 (10 1 ) 110 = 11 10 NoteP 1 =1 Q 1 =0P y =1 (10 1 ) 110 y =1 Q y =(I P )y =y 1 = y The transformed error components model is thenQ y =Q X + Q 1 +Q y = X + The giant regression (stacked by observation) is y = X + or y 1= X( ) ( 1)+ 1 Note: Unless [ 0]= 2 I = is not FE estimator is again pooled OLS on the transformed system = X0 X 1 X0 y= X =1 X0 X 1 X =1 X0 y = X =1(Q X )0Q X 1 X =1(Q X )0Q y = X =1X0 Q X 1 X =1X0 Q y sinceQ is Squares Dummy Variable ModelConsider the general model =x0 + + Stack the observations over givingy 1=X + 1 + Now create the giant regression = + 1 + ory 1=X( ) ( 1)+D( ) ( 1)+ ( 1)=X +(I 1 ) + D=I 1 Aside: Partitioned RegressionConsider the partitioned regression equationy 1=X1 1 1 1 1+X2 2 2 2 2+ The LS estimators for 1and 2can be expressed as 1=(X01Q2X1) 1X01Q2y Q2=I PX2 2=(X02Q1X2) 1X02Q1y Q1=I PX1wherePX1=X1(X01X1) 1X01 PX2=X2(X02X2) 1X02 Result.

4 The FE estimator is the partitioned OLS estimator of in the giantregression =(X0Q X) 1X0Q yQ =I P P =D(D0D) 1D0 D=I 1 Now,P =(I 1 )h(I 1 )0(I 1 )i 1(I 1 )0=(I 1 )[I 10 1 ] 1(I 1 )0=(I 1 )[I 10 1 1] I 10 =I P Therefore,Q =I P =I (I P )=I Q As a result =(X0Q X) 1X0Q y=(X0(I Q )X) 1X0(I Q )y= X =1 X0 X 1 X =1 X0 y which is exactly the result we got when we transformed the model by subtract-ing offgroup means. That is, the LSDV estimator of in the FE model isnumerically identical to the Within estimator of Recovering Estimates of With the LSDV approach, it is straightforward to deduce estimates for Byexamining the normal equations for in the Giant Regression, one can deducethat = x0 =1 X =1 x =1 X =1 Comments: For short panels (small ) is inconsistent ( Fixed and )FE as a First Difference EstimatorResults: When =2 pooled OLS on thefirst differenced model is numericallyidentical to the LSDV and Within estimators of When 2 pooled OLS on thefirst differenced model is not numericallythe same as the LSDV and Within estimators of It is consistent, butgenerally less efficient that the LSDV and Within estimators.

5 When 2and [ 0 ]= 2 I (no serial correlation), then pooledGLS on thefirst differenced model is numerically the same as the of First Differenced (FD) Estimator and LSDV and WithinEstimatorsConsider the error components model =x0 + + [ 0 ]= 2 I Note: The spherical error assumption is made so that GLS Estimation is 1st differences over : 2= 2 1= x0 2 + = ( 1)= x0 + Notice that differencing removes thefixed matrix form the model isC0y ( 1) 1=C0X +C0 or y = X + whereC0( 1) = 11000 110000 11 Note that [ 0 ]= [C0 0 C]=C0 [ 0 ]C= 2 C0 CThe transformed Giant Regression is y = X + or y( 1) 1= X(( 1) ) ( 1)+ ( 1) 1 Notice that [ 0]( 1) ( 1)= 2 C0C0 00 2 C0C 2 C0C = 2 hI (C0C)iPooled GLS on the transformed regression gives =h X0 I (C0C) 1 Xi 1 X0 I (C0C) 1 y= X =1X 0C(C0C) 1C0X 1 X =1X 0C(C0C) 1C0y = X =1X 0P X 1 X =1X 0P y = X =1X 0Q X 1 X =1X 0Q y = X =1 X0 X 1 X =1 X0 y Now.

6 X0 I (C0C) 1 X=h X01 X0 i (C0C) (C0C) 1 X = X =1 X0 (C0C) 1 X = X =1(C0X )0(C0C) 1C0X = X =1X 0C(C0C) 1C0X = X =1X 0P X P =C(C0C) 1C0 The resultP =C(C0C) 1C0=Q =I 1 (10 1 ) 110 follows from the result thatC01 = 11000 110000 11 = That is,P =C(C0C) 1C0is an idempotent matrix satisfyingP 1 =0 It projects onto the space orthogonal to1 andthisisexactlywhatQ does:Q 1 =0 Hence, when [ 0 ]= 2 I pooled GLS on the FD model is numericallyequivalent to the LSDV and Within estimators of Panel Robust Statistical Inference =x0 + + =1 ; =1 =X + =1 2. = 1 + (error components)3.{y X }is (over but not )4. [ ]6=0(Endogeneity)5. [ ]=0for =1 2 Comments It is reasonable to assume independence over ( , cross-sectional inde-pendence due to random sampling at given ) Errors are generally serially correlated over for a given ( is autocor-related) and heteroskedastic over (cross-sectional heteroskedasticity) OLS standard errors are typically downward biased due to serial correlation Panel robust standard errors correct for serial correlation and Panel Data Estimators in Common NotationThe Within (LSDV) and FD estimators of in =x0 + + =1.

7 =1 are pooled OLS estimators in a transformed model = x0 + Within Estimator: = x =x x FD Estimator: = 1 x =x x 1 1In matrix notation, we have y 1= X + =1 and the giant regression is y 1= X 1+ Pooled OLS on the giant regression is = X0 X 1 X0 y= X =1 X0 X 1 X =1 X0 y = Asymptotic Distribution Theory(Advanced)Usingy = X + we have = X =1 X0 X 1 X =1 X0 X + = X =1 X0 X 1 X =1 X0 X + X =1 X0 X 1 X =1 X0 = + X =1 X0 X 1 X =1 X0 It follows that = 1 X =1 X0 X 11 X =1 X0 The Law of Large Numbers (LLN) gives1 X =1 X0 X [ X0 X ] 1and the Central Limit Theorem (CLT) gives1 X =1 X0 (0 [ X0 0 X ])Hence, by Slutsky s theorem = 1 X =1 X0 X 11 X =1 X0 [ X0 X ] 1 (0 [ X0 0 X ]) (0 avar( ))whereavar( )= [ X0 X ] 1 [ X0 0 X ] [ X0 X ] 1 Therefore, ( 1davar( ))Remark.

8 The sandwich form ofavar( )suggests that it is an inefficientestimator in Robust InferenceUnder serial correlation and heteroskedasticity, it can be shown (Econ 583)1 X =1 X0 X [ X0 X ] 1 X =1 X0 0 X [ X0 0 X ] = y X Then a consistent estimate foravar( )isdavar( )= 1 X =1 X0 X 1 1 X =1 X0 0 X 1 1 X =1 X0 X 1 Remarks The formula fordavar( )is the Panel robust covariance estimate. Itis not the same as the usual White correction for heteroskedasticity in apooled OLS regression. The White correction does not account for serialcorrelation. It is also not the Newey-West correction for heteroskedasticityand autocorrelation. The formula fordavar( )is also known as the cluster robust covarianceestimate when the clustering variable is (each individual is a cluster)Special Case: When is Spherical [ 0 ]= 2 I For the Within estimator, we have =Q so that [ 0 ]= [Q 0 Q ]= 2 Q This implies that there is no serial correlation (correlation across time) and nocross-sectional heteroskedasticity.

9 [ X0 0 X ]= [ X0 [ 0 ]] X ]= 2 [ X0 X ]andavar( )= 2 [ X0 X ] 1 Then a consistent estimate foravar( )isdavar( )= 2 1 X =1 X0 X 1 2 =1 X =1 0 =( y X )Remark: In this case the Within estimator is an efficient estimator.


Related search queries