Example: air traffic controller

Lecture 5 Multiple Choice Models Part I –MNL, Nested Logit

RS Lecture 171 Lecture 5 Multiple Choice ModelsPart I MNL, Nested LogitDCM: Different Models Popular Models :1. Probit model 2. Binary Logit Model3. Multinomial Logit Model4. Nested Logit model5. Ordered Logit model Relevant literature:- Train (2003): Discrete Choice Methods with Simulation- Franses and Paap (2001): Quantitative Models in Market Research- Hensher, Rose and Greene (2005): Applied Choice AnalysisRS Lecture 17 Multinomial Logit (MNL) model In many of the situations, discrete responses are more complex than the binary case:- Single Choice out of more than two alternatives: Electoral choices and interest in explaining the vote for a particular party. - Multiple choices: Travel to work in rush hour, and travel to work out of rush hour, as well as the Choice of bus or car. The distinction should not be exaggerated: we could always enumerate travel-time, travel-mode Choice combinations and then treat the problem as making a single decision.

1 Lecture 5 Multiple Choice Models Part I –MNL, Nested Logit DCM: Different Models •Popular Models: 1. ProbitModel 2. Binary LogitModel ... 1 for engineer, 2 for lawyer, etc. (categories)-Opinions are usually coded with scales, where 1 stands for ... 2.016 (21.33) BL β1

Tags:

  Lecture, Model, Multiple, Part, Choice, Lecture 5 multiple choice models part, 1 lecture 5 multiple choice models part

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Lecture 5 Multiple Choice Models Part I –MNL, Nested Logit

1 RS Lecture 171 Lecture 5 Multiple Choice ModelsPart I MNL, Nested LogitDCM: Different Models Popular Models :1. Probit model 2. Binary Logit Model3. Multinomial Logit Model4. Nested Logit model5. Ordered Logit model Relevant literature:- Train (2003): Discrete Choice Methods with Simulation- Franses and Paap (2001): Quantitative Models in Market Research- Hensher, Rose and Greene (2005): Applied Choice AnalysisRS Lecture 17 Multinomial Logit (MNL) model In many of the situations, discrete responses are more complex than the binary case:- Single Choice out of more than two alternatives: Electoral choices and interest in explaining the vote for a particular party. - Multiple choices: Travel to work in rush hour, and travel to work out of rush hour, as well as the Choice of bus or car. The distinction should not be exaggerated: we could always enumerate travel-time, travel-mode Choice combinations and then treat the problem as making a single decision.

2 In a few cases, the values associated with the choices will themselves be meaningful, for example, number of patents: y = 0; 1,2,.. (count data). In most cases, the values are Logit (MNL) model In most cases, the value of the dependent variable is merely a coding for some qualitative outcome:- Labor force participation: we code yes" as 1 and no" as 0(qualitative choices)- Occupational field: 0 for economist, 1 for engineer, 2 for lawyer, etc. (categories)- Opinions are usually coded with scales, where 1 stands for strongly disagree", 2 for disagree", 3 for neutral", etc. Nothing conceptually difficult about moving from a binary to a multi-response framework, but numerical difficulties can be big. A simple model to generalized: The Logit Lecture 17 Multinomial Logit (MNL) model Now, we have a Choice between J (greater than 2) categories Dependent variable yn= 1, 2, 3.

3 J Explanatory variables zn, different across individuals, not across choices (standard MNLmodel). The MLN specifies for Choice j = 1,2,.., J: xn, different across (individuals and) choices (conditional MNLmodel). The conditional Logit model specifies for Choice j: Both Models are easy to estimate. +===lljnnzzzzjyP)'exp(1)'exp()|( ==ljnjnnnxxxjyP)'exp()'exp()|( Multinomial Logit (MNL) model The MNL can be viewed as a special case of the conditional logitmodel. Suppose we have a vector of individual characteristics Ziof dimension K, and J vectors of coefficients j, each of dimension K. Then define, We are back in the conditional Logit model . RS Lecture 17 MNL Link with Utility Maximization The modeling approach (McFadden s) is similar to the binary case. - Random Utility for individual n,associated with Choice j:Un1= Vnj+ nj= j+z n j+w n j+ nj- utility from decision j- same parameters for all , if yn= jif (Unj- Uni) > 0(n selects jover i.)

4 - Like in the binary case, we get:- Specify distribution for f( ) => Logit independence across utility functions- identical variances (means absorbed in constants)nnninjnninjnjninnidfijVVIijVVj ijyP < = < === )()()(Prob],|[Prob If we add a constant to a parameter ( i+c), given the Logistic distribution, exp(c) will cancel out. Cannot distinguish between ( i+c) and i. Need a normalization select a reference category, say i, and set coefficients equal to 0 , i=0. (Typically, i=J.) Conditional MNL model (xn: different across (individuals and) choices) ==lnlnjnnxxxjyP)'exp()'exp()|( ijxxxjyPxxiyPilnlnjnnilnlnn +==+== )'exp(1)'exp()|()'exp(11)|( ==lnlnjnnxxxjyP)'exp()'exp()|( MNL model - IdentificationRS Lecture 17 The interpretation of parameters is based on partial effects: Derivative (marginal effect) Elasticity (proportional changes)Note: The elasticity is the same for all choices j.

5 A change in the cost of air travel has the same effect on all other forms of travel. (This result is called independecne from irrelevant alternatives (IIA). Not a realistic property. Many experiments reject it.)knjnjnknnPPxxjyP = = )1()|(knjnkknjnjnjnknknjPxPPPxxP = = )1()1(loglogMNL model Interpretation & Effects Interpretation of parameters Probability-ratio Does not depend on the other alternatives! A change in attribute xnkdoes not affect the log-odds ratio between choices jand i. This result is called independence from irrelevant alternatives (IIA). Implication of MNL Models pointed out by Luce (1959).Note: The log-odds ratio of each response follow a linear model . A regression can be used for the comparison of two choices at a time.)(')|()|(ln)'exp()'exp()|()|(ninjnn nnninjnnnnxxxiyPxjyPxxxiyPxjyP = ===== MNL model Interpretation & EffectsRS Lecture 17 Estimation ML estimation))'exp(ln())'(()))'exp(ln()'ex p((ln()'exp()'exp(ln)()|(ln)()|()( = = = == == knjnjnjnjknjnjnjnjnjknjnjnjnjnnnjnjDnnxx DxxDxxDLogLxjyPDLogLxjyPLnjwhere Dnj=1 if jis selected, 0 otherwise)MNL model Estimation Estimation- ML: A lot of , with a lot of unknowns (parameters).

6 Each covariate has J-1 coefficients. We use numerical procedures, G-N or N-R often work well. - Alternative estimation proceduresSimulation-assisted estimation (Train, )Bayesian estimation (Train, ) MNL model EstimationRS Lecture 17 Example (from Bucklin and Gupta (1992)): Ui= constant for brand-size i BLhi= loyalty of household h to brand of brandsize i LBPhit= 1 if i was last brand purchased, 0 otherwise SLhi= loyalty of household h to size of brandsize i LSPhit= 1 if i was last size purchased, 0 otherwise Priceit= actual shelf price of brand-size i at time t Promoit= promotional status of brand-size i at time t itithithihithiihitjhjthithtLSPSLLBPBLuUU UinciPPromoPrice)exp()exp()|(654321 + + + + + +== MNL model Application - PIM Data scanner panel data 117 weeks: 65 for initialization, 52 for estimation 565 households.

7 300 selected randomly for estimation, remaining hh = holdout sample for validation Data set for estimation: 30,966 shopping trips, 2,275 purchases in the category (liquid laundry detergent) Estimation limited to the 7 top-selling brands (80% of category purchases), representing 28 brand-size combinations (= level of analysis for the Choice model )MNL model Application - PIMRS Lecture modelFull modelBICU (pseudo R )LL# Goodness-of-FitMNL model Application - ( ).548 ( ) ( ).512 ( ) ( ) ( )BL 1 LBP 2SL 3 LSP 4 Price 5 Promo 6 Coefficients (t-statistic)Parameter Estimation ResultsMNL model Application Travel Mode Data: 4 Travel Modes: Air, Bus, Train, Car. N=210----------------------------------- ------------------------Discrete Choice (multinomial Logit ) modelDependent variable ChoiceLog likelihood function based on N = 210, K = 7 Information Criteria: Normalization=1/NNormalized UnnormalizedAIC IC Quinn * Log-L fncn R-sqrd R2 AdjConstants only.

8 0951 .0850 Chi-squared[ 4] = [ chi squared > value ] = .00000 Response data are given as ind. choicesNumber of 210, skipped 0 obs--------+---------------------------- ----------------------Variable| Coefficient Standard Error P[|Z|>z]--------+----------------------- ---------------------------GC| .03711** .01484 .0124 INVC| ** .01668 .0010 INVT| ** .00215 .0000 HINCA| .02922** .00931 .0017A_AIR| ** .69281 .0064A_TRAIN| .69364** .25010 .0055A_BUS| .24817 .4132--------+-------------------------- ------------------------RS Lecture 17 CLOGIT Fit Measure: Based on the log likelihood Based on the model predictions+---------------------------- --------------------------+| Cross tabulation of actual vs.

9 Predicted choices. || Row indicator is actual, column is predicted. || Predicted total is F(k,j,i)=Sum(i=1,..,N) P(k,j,i). || Column totals may be subject to rounding error. |+-------------------------------------- ----------------+Matrix Crosstab has 5 rows and 5 TRAIN BUS CAR Total+---------------------------------- ------------------------------------AIR | (16) | (19) | (4) | (17) | in parentheses below show the number of correct predictions by a model with only Choice specific likelihood function only .0951 .0850 Chi-squared[ 4] = model Application Travel Mode Scale parameter Variance of the extreme value distribution Var[ ] = /6- If true utility is U*nj= * xnj+ *njwith Var( *nj)= ( /6), the estimated representative utility Vnj= xnjinvolves a rescaling of * => = * / * and can not be estimated separately Take into account that the estimated coefficients indicate the variable s effect relative tothe variance of unobserved factors Include scale parameters if subsamples in a pooled estimation (may) have different error variancesMNL model ScalingRS Lecture 17 Scale parameter in the case of pooled estimation of subsamples with different error variance For each subsamples, multiply utility by s, which is estimated simultaneously with Normalization: set s equal to 1 for 1 subs.

10 Values of s reflect diff s in error variation s>1 : error variance is smaller in s than in the reference subsample s<1 : error variance is larger in s than in the reference subsampleMNL model Scaling Example(from Breugelmans et al (2005), based on Andrews and Currim (2002); Swait and Louvi re (1993)): Data from online experiment, 2 product categories Three different assortments, assigned to different respondent groups Assortment 1: small assortment Assortment 2 = extended with addirional brands Assortment 3 = extended with add types Explanatory variables are the same (hh char s, MM), with exception of the constants A scale factor is introduced for assortment 2 and 3 (assortment 1 is reference with scale factor =1)MNL model ApplicationRS Lecture 17 Table 1: Descriptive stats for each assortment (margarine and cereals)MARGARINEA ttributeAssortment 1 (limited)Assortment 2 (add new flavors of existing brands)Assortment 3 (add new brands of existing flavors)BrandCommon aCommonCommonAdd new brandsFlavorCommonCommonCommonAdd new flavors# alternatives111917# respondents105116100# purchase occasions275279278# screens needed< 1 > 1 > 1 CEREALSA ttributeAssortment 1 (limited)Assortment 2 (add new flavors of existing brands)Assortment 3 (add new brands of existing flavors)


Related search queries