Transcription of The Principles of Deep Learning Theory arXiv:2106.10165v2 ...
1 The Principles of deep Learning TheoryAn Effective Theory Approach to Understanding Neural NetworksDaniel A. Roberts and Sho Yaidabased on research in col laboration withBoris [ ] 24 Aug 2021iiContentsPrefacevii0 An Effective Theory Approach .. The Theoretical Minimum ..41 Gaussian Integrals .. Probability, Correlation and Statistics, and All That .. Nearly-Gaussian Distributions .. 282 Neural Function Approximation .. Activation Functions .. Ensembles .. 473 Effective Theory of deep Linear Networks at deep Linear Networks .. Criticality .. Fluctuations .. Chaos .. 654 RG Flow of First Layer: Good-Old Gaussian .. Second Layer: Genesis of Non-Gaussianity .. Deeper Layers: Accumulation of Non-Gaussianity .. Marginalization Rules.
2 Subleading Corrections .. RG Flow and RG Flow .. 1035 Effective Theory of Preactivations at Criticality Analysis of the Kernel .. Criticality for Scale-Invariant Activations .. Universality beyond Scale-Invariant Activations .. Strategy .. Criticality: sigmoid, softplus, nonlinear monomials, etc.. 0 Universality Class: tanh, sin, etc.. Universality Classes: SWISH, etc. and GELU, etc.. Fluctuations .. for the Scale-Invariant Universality Class .. for theK?= 0 Universality Class .. Finite-Angle Analysis for the Scale-Invariant Universality Class .. 1456 Bayesian Bayesian Probability .. Bayesian Inference and Neural Networks .. Model Fitting .. Model Comparison .. Bayesian Inference at Infinite Width .. Evidence for Criticality .. s Not Wire Together .. of Representation Learning .
3 Bayesian Inference at Finite Width .. Learning , Inc.. s Wire Together .. of Representation Learning .. 1857 Gradient-Based Supervised Learning .. Gradient Descent and Function Approximation .. 1948 RG Flow of the Neural Tangent Forward Equation for the NTK .. First Layer: Deterministic NTK .. Second Layer: Fluctuating NTK .. Deeper Layers: Accumulation of NTK Fluctuations .. :Interlayer Correlations .. Mean .. Cross Correlations .. Variance .. 2219 Effective Theory of the NTK at Criticality Analysis of the NTK .. Scale-Invariant Universality Class .. 0 Universality Class .. Criticality, Exploding and Vanishing Problems, and None of That .. 241iv10 Kernel A Small Step .. No Wiring .. No Representation Learning .. A Giant Leap .. Newton s Method.
4 Algorithm Independence .. : Cross-Entropy Loss .. Kernel Prediction .. Generalization .. Bias-Variance Tradeoff and Criticality .. Interpolation and Extrapolation .. Linear Models and Kernel Methods .. Linear Models .. Kernel Methods .. Infinite-Width Networks as Linear Models .. 28711 Representation Differential of the Neural Tangent Kernel .. RG Flow of the dNTK .. Forward Equation for the dNTK .. First Layer: Zero dNTK .. Second Layer: Nonzero dNTK .. Deeper Layers: Growing dNTK .. Effective Theory of the dNTK at Initialization .. Scale-Invariant Universality Class .. 0 Universality Class .. Nonlinear Models and Nearly-Kernel Methods .. Nonlinear Models .. Nearly-Kernel Methods .. Finite-Width Networks as Nonlinear Models .. 329 The End of Training335.
5 1 Two More Differentials .. 337 .2 Training at Finite Width .. 347 . A Small Step Following a Giant Leap .. 352 . Many Many Steps of Gradient Descent .. 357 . Prediction at Finite Width .. 374 .3 RG Flow of the ddNTKs: The Full Expressions .. 385 Epilogue: Model Complexity from the Macroscopic Perspective391vA Information in deep Entropy and Mutual Information .. Information at Infinite Width: Criticality .. Information at Finite Width: Optimal Aspect Ratio .. 412B Residual Residual Multilayer Perceptrons .. Residual Infinite Width: Criticality Analysis .. Residual Finite Width: Optimal Aspect Ratio .. Residual Building Blocks .. 436 References439 Index447viPrefaceThis has necessitated a complete break from the historical line of development, but thisbreak is an advantage through enabling the approach to the new ideas to be made asdirect as A.
6 M. Dirac in the 1930 preface ofThe Principles of Quantum Mechanics[1].This is a research monograph in the style of a textbook about the Theory of deep this book might look a little different from the other deep Learning books thatyou ve seen before, we assure you that it is appropriate for everyone with knowledgeof linear algebra, multivariable calculus, and informal probability Theory , and with ahealthy interest in neural networks. Practitioner and theorist alike, we want all of youto enjoy this book. Now, let us tell you some and foremost, in this book we ve strived for pedagogy in every choice we vemade, placing intuition above formality. This doesn t mean that calculations are incom-plete or sloppy; quite the opposite, we ve tried to provide full details of every calculation of which there are certainly very many and place a particular emphasis on the toolsneeded to carry out related calculations of interest.
7 In fact, understanding how the calcu-lations are done is as important as knowing their results, and thus often our pedagogicalfocus is on the details , while we present the details of all our calculations, we ve kept the experi-mental confirmations to the privacy of our own computerized notebooks. Our reasonfor this is simple: while there s much to learn from explaining a derivation, there s notmuch more to learn from printing a verification plot that shows two curves lying on topof each other. Given the simplicity of modern deep - Learning codes and the availabilityof compute, it s easy to verify any formula on your own; we certainly have thoroughlychecked them all this way, so if knowledge of the existence of such plots are comfortingto you, know at least that they do exist on our personal and cloud-based hard , our main focus is on realistic models that are used by the deep learningcommunity in practice: we want to studydeepneural networks.
8 In particular, thismeans that(i)a number of special results on single-hidden-layer networks will not bediscussed and(ii)theinfinite-width limitof a neural network which corresponds to azero-hidden-layer network will be introduced only as a starting point. All such idealizedmodels will eventually beperturbeduntil they correspond to a real model. We certainlyacknowledge that there s a vibrant community of deep - Learning theorists devoted toviiexploring different kinds of idealized theoretical limits. However, our interests are fixedfirmly on providing explanations for the tools and approaches used by practitioners, inan effort to shed light on what makes them work so , a large part of the book is focused on deep multilayer perceptrons. We madethis choice in order to pedagogically illustrate the power of the effective Theory framework not due to any technical obstruction and along the way we give pointers for how thisformalism can be extended to other architectures of interest.
9 In fact, we expect thatmany of our results have a broad applicability, and we ve tried to focus on aspects thatwe expect to have lasting and universal value to the deep Learning , while much of the material is novel and appears for the first time in thisbook, and while much of our framing, notation, language, and emphasis breaks withthe historical line of development, we re also very much indebted to the deep learningcommunity. With that in mind, throughout the book we will try to reference importantprior contributions, with an emphasis on recent seminal deep - Learning results rather thanon being completely comprehensive. Additional references for those interested can easilybe found within the work that we , this book initially grew out of a research project in collaboration with BorisHanin. To account for his effort and then support, we ve accordingly commemorated himon the cover.
10 More broadly, we ve variously appreciated the artwork, discussions, en-couragement, epigraphs, feedback, management, refereeing, reintroduction, and supportfrom Rafael Araujo, L eon Bottou, Paul Dirac, Ethan Dyer, John Frank, Ross Girshick,Vince Higgs, Yoni Kahn, Yann LeCun, Kyle Mahowald, Eric Mintun, Xiaoliang Qi,Mike Rabbat, David Schwab, Stephen Shenker, Eva Silverstein, PJ Steiner, DJ Strouse,and Jesse Thaler. Organizationally, we re grateful to FAIR and Facebook, Diffeo andSalesforce, MIT and IAIFI, and Cambridge University Press and the , given intense (and variously uncertain) spacetime and energy-momentumcommitment that writing this book entailed, Dan is grateful to Aya, Lumi, and LisaYaida; from the dual sample-space perspective, Sho is grateful to Adrienne Rothschildsand would be retroactively grateful to any hypothetical future Mark or Emily that wouldhave otherwise been thanked in this , we hope that this book spreads our optimism that itispossible to havea general Theory of deep Learning , one that s both derived from first Principles and atthe same time focused on describing how realistic models actually work: nearly-simplephenomena in practice should correspond to nearly-simple effective theories.