Example: stock market

Empirical Asset Pricing via Machine Learning

[16:04 6/4/2020 ]Page: 2223 2223 2274 Empirical Asset Pricing via MachineLearning Shihao GuBooth School of Business, University of ChicagoBryan KellyYale University, AQR Capital Management, and NBERD acheng XiuBooth School of Business, University of ChicagoWe perform a comparative analysis of Machine Learning methods for the canonical problemof Empirical Asset Pricing : measuring Asset risk premiums. We demonstrate large economicgains to investors using Machine Learning forecasts, in some cases doubling the performanceof leading regression-based strategies from the literature. We identify the best-performingmethods (trees and neural networks) and trace their predictive gains to allowing nonlinearpredictor interactions missed by other methods. All methods agree on the same set ofdominant predictive signals, a set that includes variations on momentum, liquidity, andvolatility.

benchmarks for the predictive accuracy of machine learning methods in measuring risk premiums of the aggregate market and individual stocks. This accuracy is summarized two ways. The first is a high out-of-sample predictive R2 relative to preceding literature that is robust across a variety of machine learning specifications.

Tags:

  Machine, Market, Learning, Machine learning

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Empirical Asset Pricing via Machine Learning

1 [16:04 6/4/2020 ]Page: 2223 2223 2274 Empirical Asset Pricing via MachineLearning Shihao GuBooth School of Business, University of ChicagoBryan KellyYale University, AQR Capital Management, and NBERD acheng XiuBooth School of Business, University of ChicagoWe perform a comparative analysis of Machine Learning methods for the canonical problemof Empirical Asset Pricing : measuring Asset risk premiums. We demonstrate large economicgains to investors using Machine Learning forecasts, in some cases doubling the performanceof leading regression-based strategies from the literature. We identify the best-performingmethods (trees and neural networks) and trace their predictive gains to allowing nonlinearpredictor interactions missed by other methods. All methods agree on the same set ofdominant predictive signals, a set that includes variations on momentum, liquidity, andvolatility.

2 (JELC52, C55, C58, G0, G1, G17)Received September 4, 2018; editorial decision September 22, 2019 by Editor AndrewKarolyi. Authors have furnished an Internet Appendix, which is available on the OxfordUniversity Press Web site next to the link to the final published paper online. We benefitted from discussions with Joseph Babcock, Si Chen (discussant), Rob Engle, Andrea Frazzini,Amit Goyal (discussant), Lasse Pedersen, Lin Peng (discussant), Alberto Rossi (discussant), and Guofu Zhou(discussant) and seminar and conference participants at Erasmus School of Economics, NYU, Northwestern,Imperial College, National University of Singapore, UIBE, Nanjing University, Tsinghua PBC School of Finance,Fannie Mae, Securities and Exchange Commission, City University of Hong Kong, Shenzhen FinanceInstitute at CUHK, NBER Summer Institute, New Methods for the Cross Section of Returns Conference,Chicago Quantitative Alliance Conference, Norwegian Financial Research Conference, EFA, China InternationalConference in Finance, 10th World Congress of the Bachelier Finance Society, Financial Engineering and RiskManagement International Symposium, Toulouse Financial Econometrics Conference, Chicago Conference onNew Aspects of Statistics, Financial Econometrics, and Data Science, Tsinghua Workshop on Big Data andInternet Economics, Q group, IQ-KAP Research Prize Symposium.

3 Wolfe Research, INQUIRE UK, AustralasianFinance and Banking Conference, Goldman Sachs Global Alternative Risk Premia Conference, AFA, and SwissFinance Institute. We gratefully acknowledge the computing support from the Research Computing Center at theUniversity of Chicago. The views and opinions expressed are those of the authors and do not necessarily reflectthe views of AQR Capital Management, its affiliates, or its employees; do not constitute an offer, solicitationof an offer, or any advice or recommendation, to purchase any securities or other financial instruments, andmay not be construed as such. Supplementary data can be found onThe Review of Financial Studiesweb correspondence to Shihao Gu, University of Chicago, Booth School of Business, 5807 S. Woodlawn Ave.,Chicago, IL 60637; telephone: +1(310)869-0675. E-mail: Review of Financial Studies33 (2020) 2223 2273 The Authors 2020.

4 Published by Oxford University Press. This is an Open Access article distributed underthe terms of the Creative Commons Attribution Non-Commercial NoDerivs License( ), which permits non-commercial reproduction anddistribution of the work, in any medium, provided the original work is not altered or transformed in any way,and that the work is properly cited. For commercial re-use, please contact Access publication February 26, 2020 Downloaded from by University of Chicago Libraries user on 18 April 2020[16:04 6/4/2020 ]Page: 2224 2223 2274 The Review of Financial Studies/v 33 n 5 2020In this article, we conduct a comparative analysis of Machine learningmethods for finance. We do so in the context of perhaps the most widely studiedproblem in finance, that of measuring equity risk primary contributions are twofold. First, we provide a new set ofbenchmarks for the predictive accuracy of Machine Learning methods inmeasuring risk premiums of the aggregate market and individual stocks.

5 Thisaccuracy is summarized two ways. The first is a high out-of-sample predictiveR2relative to preceding literature that is robust across a variety of machinelearning specifications. Second, and more importantly, we demonstrate largeeconomic gains to investors using Machine Learning forecasts. A portfoliostrategy that times the S&P 500 with neural network forecasts enjoys anannualized out-of-sample Sharpe ratio of versus the Sharpe ratioof a buy-and-hold investor. And a value-weighted long-short decile spreadstrategy that takes positions based on stock-level neural network forecastsearns an annualized out-of-sample Sharpe ratio of , more than doublingthe performance of a leading regression-based strategy from the prediction is economically meaningful. The fundamental goal of assetpricing is to understand the behavior of risk expected returnswere perfectly observed, we would still need theories to explain their behaviorand Empirical analysis to test those theories.

6 But risk premiums are notoriouslydifficulttomeasure:marketeffi ciencyforcesreturnvariationtobedominated byunforecastable news that obscures risk premiums. Our research highlights gainsthat can be achieved in prediction and identifies the most informative predictorvariables. This helps resolve the problem of risk premium measurement, whichthen facilitates more reliable investigation into economic mechanisms of , we synthesize the Empirical Asset Pricing literature with the fieldof Machine Learning . Relative to traditional Empirical methods in Asset Pricing , Machine Learning accommodates a far more expansive list of potential predictorvariables and richer specifications of functional form. It is this flexibility thatallows us to push the frontier of risk premium measurement. Interest in machinelearning methods for finance has grown tremendously in both academia andindustry. This article provides a comparative overview of Machine learningmethods applied to the two canonical problems of Empirical Asset Pricing :predicting returns in the cross-section and time series.

7 Our view is that the bestway for researchers to understand the usefulness of Machine Learning in the1 Our focus is on measuring conditional expected stock returns in excess of the risk-free rate. Academic financetraditionally refers to this quantity as the risk premium because of its close connection with equilibriumcompensation for bearing equity investment risk. We use the terms expected return and risk premium interchangeably. One may be interested in potentially distinguishing between different components of expectedreturns, such as those due to systematic risk compensation, idiosyncratic risk compensation, or even due tomispricing. For Machine Learning approaches to this problem, see Gu, Kelly, and Xiu (2019) and Kelly, Pruitt,and Su (2019).2224 Downloaded from by University of Chicago Libraries user on 18 April 2020[16:04 6/4/2020 ]Page: 2225 2223 2274 Empirical Asset Pricing via Machine Learningfield of Asset Pricing is to apply and compare the performance of each of itsmethods in familiar Empirical definition of Machine Learning is inchoate and is often context use the term to describe (a) a diverse collection of high-dimensional modelsfor statistical prediction, combined with (b) so-called regularization methodsfor model selection and mitigation of overfit, and (c) efficient algorithms forsearching among a vast number of potential model high-dimensional nature of Machine Learning methods (element (a)of this definition) enhances their flexibility relative to more traditionaleconometric prediction techniques.

8 This flexibility brings hope of betterapproximating the unknown and likely complex data generating processunderlying equity risk premiums. With enhanced flexibility, however, comes ahigher propensity of overfitting the data. Element (b) of our Machine learningdefinition describes refinements in implementation that emphasize stable out-of-sample performance to explicitly guard against overfit. Finally, with manypredictors it becomes infeasible to exhaustively traverse and compare all modelpermutations. Element (c) describes clever Machine Learning tools designed toapproximate an optimal specification with manageable computational for analysis with Machine Learning , two main research agendas have monopolized modern Empirical assetpricing research. The first seeks to describe and understand differences equity risk premium. Measurement of an Asset s risk premium isfundamentally a problem of prediction the risk premium is the ,whosemethodsare largely specialized for prediction tasks, is thus ideally suited to the problemof risk premium , the collection of candidate conditioning variables for the riskpremium is large.

9 The profession has accumulated a staggering list of predictorsthat various researchers have argued possess forecasting power for number of stock-level predictive characteristics reported in the literaturenumbers in the hundreds and macroeconomic predictors of the aggregatemarket number in the , predictors are often close cousinsand highly correlated. Traditional prediction methods break down when thepredictor count approaches the observation count or predictors are highlycorrelated. With an emphasis on variable selection and dimension reductiontechniques, Machine Learning is well suited for such challenging prediction2 Green et al. (2013) count 330 stock-level predictive signals in published or circulated drafts. Harvey, Liu, andZhu (2016) study 316 factors, which include firm characteristics and common factors, for describing stockreturn behavior. They note that this is only a subset of those studied in the literature.

10 Welch and Goyal (2008)analyze nearly twenty predictors for the aggregate market return. In both stock and aggregate return predictions,there presumably exists a much larger set of predictors that were tested but failed to predict returns and werethus never from by University of Chicago Libraries user on 18 April 2020[16:04 6/4/2020 ]Page: 2226 2223 2274 The Review of Financial Studies/v 33 n 5 2020problems by reducing degrees of freedom and condensing redundant variationamong , further complicating the problem is ambiguity about the functionalforms through which the high-dimensional predictor sets enter into riskpremiums. Should they enter linearly? If nonlinearities are needed, whichform should they take? Must we consider interactions among predictors? Suchquestions rapidly proliferate the set of potential model specifications. Thetheoretical literature offers little guidance for winnowing the list of conditioningvariables and functional forms.


Related search queries