Transcription of Thinking fast or slow? A reinforcement learning approach
1 P < and detailsConclusionCost-benefit arbitration between multiple RL systemsFlexible adapation based on reward-advantageIn progress: w s relationship with outgroup bias and psychiatric symptomsExperiment 4. New paradigm with stakes1x50%5x50%Prediction:w5x > w1xExperiment Parameter r pExp 1. w1x .54 < .001 w5x .32 < .001 Exp. 2 w1x .92 w5x .81 Stakes manipulationResultsCorrelations1x5x1x5xE xp. 1 Exp. 2** = 94 Accuracy-demand tradeoff in novel 2-step rateExperiment simulationsDoes w predict reward?r = **n = 184 Novel paradigmRL modelQnet = w QMB + (1 - w) QMFQMF( ) = QMF( ) + RPEStay probabilityw = 0 SameDifferentPrevious start statew = 1 SameDifferentPrevious start statePrevious outcomeWinLossQMB( ) = Q( ) = QMB( )model-freemodel-basedBehavioral predictionsTaskDrift rate = 2chance of winning pieces of space treasure:Range = [ ],No accuracy-demand tradeoff in Daw 2-step taskRL rateExperiment 2 Does w predict reward?
2 R = = rateExperiment 1. Stakes manipulationStakes manipulation1x50%5x50%Prediction:If model-based planning is costly, participants should plan more when stakes are highw5x > = 98 Daw et al. (2011) 2-step taskBehavioral predictionsRL modelQnet = w QMB + (1 - w) QMFw indicates model-basedness (degree of system 2)Task70%70%chance of winning space treasure ( )(changing slowly)Win: Loss:Bounds = [ ]Drift rate ( ) = distribution050100150200 Trial numberReward ( ) = QMF( ) + RPEQMB( ) = max ( ) + max ( ) w = 1 WinLosePrevious outcomePrev. transitionCommonRarew = outcomePrev. transitionCommonRareStay probabilityw = 0 WinLosePrevious outcomePrev. transitionCommonRaremodel-freedatamodel- basedDual process theory and reinforcement learningTheories of judgement and decision making posit existence of two systems:Kahneman (2003)Accuracy-efficiency tradeoff between System 1 and System 2?
3 Often assumed that systems engage in a cost-benefit trade-off,but direct evidence for this has been 2 HabitualAutomaticComputationally cheapSystem 1 Goal-directedDeliberativeComputationally expensiveRecent advances in computer science and reinforcement learning :Daw et al. (2011)Model-free ) Department of Psychology, Harvard Universityb) Center for Brain Science, Harvard UniversityWouter Koola, Samuel J. Gershmana,b, & Fiery A. CushmanaThinking fast or slow ? A reinforcement learning approach