{"title": "Data-dependent Sample Complexity of Deep Neural Networks via Lipschitz Augmentation", "book": "Advances in Neural Information Processing Systems", "page_first": 9725, "page_last": 9736, "abstract": "Existing Rademacher complexity bounds for neural networks rely only on norm control of the weight matrices and depend exponentially on depth via a product of the matrix norms. Lower bounds show that this exponential dependence on depth is unavoidable when no additional properties of the training data are considered. We suspect that this conundrum comes from the fact that these bounds depend on the training data only through the margin. In practice, many data-dependent techniques such as Batchnorm improve the generalization performance. For feedforward neural nets as well as RNNs, we obtain tighter Rademacher complexity bounds by considering additional data-dependent properties of the network: the norms of the hidden layers of the network, and the norms of the Jacobians of each layer with respect to all previous layers. Our bounds scale polynomially in depth when these empirical quantities are small, as is usually the case in practice. To obtain these bounds, we develop general tools for augmenting a sequence of functions to make their composition Lipschitz and then covering the augmented functions. Inspired by our theory, we directly regularize the network\u2019s Jacobians during training and empirically demonstrate that this improves test performance.", "full_text": "Data-dependentSampleComplexityofDeepNeuralNetworksviaLipschitzAugmentationColinWeiComputerScienceDepartmentStanfordUniversitycolinwei@stanford.eduTengyuMaComputerScienceDepartmentStanfordUniversitytengyuma@stanford.eduAbstractExistingRademachercomplexityboundsforneuralnetworksrelyonlyonnormcontroloftheweightmatricesanddependexponentiallyondepthviaaproductofthematrixnorms.Lowerboundsshowthatthisexponentialdependenceondepthisunavoidablewhennoadditionalpropertiesofthetrainingdataareconsidered.Wesuspectthatthisconundrumcomesfromthefactthattheseboundsdependonthetrainingdataonlythroughthemargin.Inpractice,manydata-dependenttechniquessuchasBatchnormimprovethegeneralizationperformance.ForfeedforwardneuralnetsaswellasRNNs,weobtaintighterRademachercomplexityboundsbyconsideringadditionaldata-dependentpropertiesofthenetwork:thenormsofthehiddenlayersofthenetwork,andthenormsoftheJacobiansofeachlayerwithrespecttoallpreviouslayers.Ourboundsscalepolynomiallyindepthwhentheseempiricalquantitiesaresmall,asisusuallythecaseinpractice.Toobtainthesebounds,wedevelopgeneraltoolsforaugmentingasequenceoffunctionstomaketheircompositionLipschitzandthencoveringtheaugmentedfunctions.Inspiredbyourtheory,wedirectlyregularizethenetwork\u2019sJacobiansduringtrainingandempiricallydemonstratethatthisimprovestestperformance.1IntroductionDeepnetworkstrainedinpracticetypicallyusemanymoreparametersthantrainingexamples,andthereforehavethecapacitytoover\ufb01ttothetrainingset[Zhangetal.,2016].Fortunately,therearealsomanyknown(andunknown)sourcesofregularizationduringtraining:modelcapacityregularizationsuchassimpleweightdecay,implicitoralgorithmicregularization[Gunasekaretal.,2017,2018b,Soudryetal.,2018,Lietal.,2018],and\ufb01nallyregularizationthatdependsonthetrainingdatasuchasBatchnorm[IoffeandSzegedy,2015],layernormalization[Baetal.,2016],groupnormalization[WuandHe,2018],pathnormalization[Neyshaburetal.,2015a],dropout[Srivastavaetal.,2014,Wageretal.,2013],andregularizingthevarianceofactivations[LittwinandWolf,2018].Inmanycases,itremainsunclearwhydata-dependentregularizationcanimprovethe\ufb01naltesterror\u2014forexample,whyBatchnormempiricallyimprovesthegeneralizationperformanceinpractice[IoffeandSzegedy,2015,Zhangetal.,2019].Wedonothavemanytoolsforanalyzingdata-dependentregularizationintheliterature;withtheexceptionofDziugaiteandRoy[2018],[Aroraetal.,2018]and[NagarajanandKolter,2019](withwhichwecomparelaterinmoredetail),existingboundstypicallyconsiderpropertiesoftheweightsofthelearnedmodelbutlittleabouttheirinteractionswiththetrainingset.Formally,de\ufb01neadata-dependentpropertyasanyfunctionofthelearnedmodelandthetrainingdata.Inthiswork,weprovetightergeneralizationboundsbyconsideringadditionaldata-dependentpropertiesofthenetwork.Optimizingtheseboundsleadstodata-dependentregularizationtechniquesthatempiricallyimproveperformance.33rdConferenceonNeuralInformationProcessingSystems(NeurIPS2019),Vancouver,Canada.\fOnewell-understoodandimportantdata-dependentpropertyisthetrainingmargin:Bartlettetal.[2017]showthatnetworkswithlargernormalizedmarginshavebettergeneralizationguarantees.However,neuralnetsarecomplex,sothereremainmanyotherdata-dependentpropertieswhichcouldpotentiallyleadtobettergeneralization.WeextendtheboundsandtechniquesofBartlettetal.[2017]byconsideringadditionalproperties:thehiddenlayernormsandinterlayerJacobiannorms.Our\ufb01nalgeneralizationbound(Theorem5.1)isapolynomialinthehiddenlayernormsandLipschitzconstantsonthetrainingdata.Wegiveasimpli\ufb01edversionbelowforexpositionalpurposes.LetFdenoteaneuralnetworkwithsmoothactivation\u03c6parameterizedbyweightmatrices{W(i)}ri=1thatperfectlyclassi\ufb01esthetrainingdatawithmargin\u03b3>0.Lettdenotethemaximum\u20182normofanyhiddenlayerortrainingdatapoint,and\u03c3themaximumoperatornormofanyinterlayerJacobian,wherebothquantitiesareevaluatedonlyonthetrainingdata.Theorem1.1(Simpli\ufb01edversionofTheorem5.1).Suppose\u03c3,t\u22651.Withprobability1\u2212\u03b4overthetrainingdata,wecanboundthetesterrorofFbyL0-1(F)\u2264eO\uf8eb\uf8ec\uf8ed(\u03c3\u03b3+r3\u03c32)t(cid:16)1+PikW(i)>k2/32,1(cid:17)3/2+r2\u03c3(cid:16)1+PikW(i)k2/31,1(cid:17)3/2\u221an+rslog(1\u03b4)n\uf8f6\uf8f7\uf8f8Thenotation\u02dcOhideslogarithmicfactorsind,r,\u03c3,tandthematrixnorms.Thek\u00b7k2,1normisformallyde\ufb01nedinSection3.Thedegreeofthedependencieson\u03c3maylookunconventional\u2014thisismostlyduetothedramaticsimpli\ufb01cationfromourfullTheorem5.1,whichobtainsamorenaturalboundthatconsidersallinterlayerJacobiannormsinsteadofonlythemaximum.Ourboundispolynomialint,\u03c3,andnetworkdepth,butindependentofwidth.Inpractice,tand\u03c3havebeenobservedtobemuchsmallerthantheproductofmatrixnorms[Aroraetal.,2018,NagarajanandKolter,2019].Weremarkthatourboundisnothomogeneousbecausethesmoothactivationsarenothomogeneousandcancauseasecondordereffectonthenetworkoutputs.Incontrast,theboundsofNeyshaburetal.[2015b],Bartlettetal.[2017],Neyshaburetal.[2017a],Golowichetal.[2017]alldependonaproductofnormsofweightmatriceswhichscalesexponentiallyinthenetworkdepth,andwhichcanbethoughtofasaworstcaseLipschitzconstantofthenetwork.Infact,lowerboundsshowthatwithonlynorm-basedconstraintsonthehypothesisclass,thisproductofnormsisunavoidableforRademachercomplexity-basedapproaches(seeforexampleTheorem3.4of[Bartlettetal.,2017]andTheorem7of[Golowichetal.,2017]).Wecircumventtheselowerboundsbyadditionallyconsideringthemodel\u2019sJacobiannorms\u2013empiricalLipschitzconstantswhicharemuchsmallerthantheproductofnormsbecausetheyareonlycomputedonthetrainingdata.TheboundofAroraetal.[2018]dependsonsimilarquantitiesrelatedtonoisestabilitybutonlyholdsforacompressednetworkandnottheoriginal.TheboundofNagarajanandKolter[2019]alsodependspolynomiallyontheJacobiannormsratherthanexponentiallyindepth;howevertheseboundsalsorequirethattheinputstotheactivationlayersareboundedawayfrom0,anassumptionthatdoesnotholdinpractice[NagarajanandKolter,2019].Wedonotrequirethisassumptionbecauseweconsidernetworkswithsmoothactivations,whereastheboundofNagarajanandKolter[2019]appliestorelunets.InSectionG,weadditionallypresentageneralizationboundforrecurrentneuralnetsthatscalespolynomiallyinthesamequantitiesasourboundforstandardneuralnets.PriorgeneralizationboundsforRNNseitherrequireparametercounting[KoiranandSontag,1997]ordependexponentiallyondepth[Zhangetal.,2018,Chenetal.,2019].InFigure1,weplotthedistributionoverthesumofproductsofJacobianandhiddenlayernorms(whichistheleadingtermoftheboundinourfullTheorem5.1)foraWideResNet[ZagoruykoandKomodakis,2016]trainedwithandwithoutBatchnorm.Figure1showsthatthissumblowsupfornetworkstrainedwithoutBatchnorm,indicatingthatthetermsinourboundareempiricallyrelevantforexplainingdata-dependentregularization.AnimmediatebottleneckinprovingTheorem1.1isthatstandardtoolsrequire\ufb01xingthehypothesisclassbeforelookingattrainingdata,whereasconditioningondata-dependentpropertiesmakesthehypothesisclassarandomobjectdependingonthedata.Anaturalattemptistoaugmenttheloss2\fFigure1:Leth1,h2,h3denotethe1st,2nd,and3rdblocksofa16-layerWideResNetandJitheJacobianoftheoutputw.r.tlayeri.Inlog-scaleweplotahistogramofthe100largestvaluesonthetrainingsetofP3i=1khikkJik/\u03b3foraWideRes-NettrainedwithandwithoutBatchnormonCI-FAR10,where\u03b3istheexample\u2019smargin.withindicatorsontheintendeddata-dependentquantities{\u03b3i},withdesiredbounds{\u03bai}asfollows:laug=(lold\u22121)Yproperties\u03b3i1(\u03b3i\u2264\u03bai)+1Thisaugmentedlossupperboundstheoriginallosslold\u2208[0,1],withequalitywhenallpropertiesholdforthetrainingdata.Theaugmentationletsusreasonaboutahypothesisclassthatisindependentofthedatabydirectlyconditioningondata-dependentpropertiesintheloss.Themainchallengeswiththisapproacharetwofold:1)designingthecorrectsetofpropertiesand2)provinggeneralizationofthe\ufb01nallosslaug,acomplicatedfunctionofthenetwork.Ourmaintooliscoveringnumbers:Lemma4.1showsthatacompositionoffunctions(i.e,aneuralnetwork)haslowcoveringnumberiftheoutputisworst-caseLipschitzateachlevelofthecompositionandinternallayersareboundedinnorm.Unfortunately,thestandardneuralnetlosssatis\ufb01esneitheroftheseproperties(withoutexponentialdependenciesondepth).However,byaugmentingwithproperties\u03b3,wecanguaranteetheyhold.Onetechnicalchallengeisthataugmentingthelossmakesithardertoreasonaboutcovering,astheindicatorscanintroducecomplicateddependenciesbetweenlayers.Ourmaintechnicalcontributionsare:1)WedemonstratehowtoaugmentacompositionoffunctionstomakeitLipschitzatalllayers,andthuseasytocover.Beforethisaugmentation,theLipschitzconstantcouldscaleexponentiallyindepth(Theorem4.4).2)Wereducecoveringacomplicatedsequenceofoperationstocoveringtheindividualoperations(Theorem4.3).3)Bycombining1and2,itfollowscleanlythatouraugmentedlossonneuralnetworkshaslowcoveringnumberandthereforehasgoodgeneralization.Ourboundscalespolynomially,notexponentially,inthedepthofthenetworkwhenthenetworkhasgoodLipschitzconstantsonthetrainingdata(Theorem5.1).Asacomplementtothemaintheoreticalresultsinthispaper,weshowempiricallyinSection6thatdirectlyregularizingourcomplexitymeasurecanresultinimprovedtestperformance.2RelatedWorkZhangetal.[2016]andNeyshaburetal.[2017b]showthatgeneralizatonindeeplearningoftendisobeysconventionalstatisticalwisdom.Oneoftheapproachesadoptedtorwardsexplaininggeneralizationisimplicitregularization;numerousrecentworkshaveshownthatthetrainingmethodprefersminimumnormormaximummarginsolutions[Soudryetal.,2018,Lietal.,2018,JiandTelgarsky,2018,Gunasekaretal.,2017,2018a,b,Weietal.,2018].Withtheexceptionof[Weietal.,2018],thesepapersanalyzesimpli\ufb01edsettingsanddonotapplytolargerneuralnetworks.ThispapermorecloselyfollowsalineofworkrelatedtoRademachercomplexityboundsforneuralnetworks[Neyshaburetal.,2015b,2018,Bartlettetal.,2017,Golowichetal.,2017].Foracomparison,seetheintroduction.TherehasalsobeenworkonderivingPAC-Bayesianboundsforgeneralization[Neyshaburetal.,2017b,a,NagarajanandKolter,2019].DziugaiteandRoy[2017a]optimizeaboundtocomputenon-vacuousboundsforgeneralizationerror.Anotherlineofworkanalyzesneuralnetsviatheirbehavioronnoisyinputs.Neyshaburetal.[2017b]provePAC-Bayesiangeneralizationboundsforrandomnetworksunderassumptionsonthenetwork\u2019sempiricalnoisestability.Aroraetal.[2018]developanotionofnoisestabilitythatallowsforcompressionofanetworkunderanappropriatenoisedistribution.Theyadditionallyprovethatthecompressednetworkgeneralizeswell.Incomparison,ourLipschitznessconstructionalsorelatestonoisestability,butourboundsholdfortheoriginalnetworkanddonotrelyontheparticularnoisedistribution.3\fNagarajanandKolter[2019]usePAC-BayesboundstoproveasimilarresultasoursforgeneralizationofanetworkwithboundedhiddenlayerandJacobiannorms.Themaindifferenceisthattheirboundsdependontheinverserelupreactivations,whicharefoundtobelargeinpractice[NagarajanandKolter,2019];ourboundsapplytosmoothactivationsandavoidthisdependenceatthecostofanadditionalfactorintheJacobiannorm(showntobeempiricallysmall).Wenotethatthechoiceofsmoothactivationsisempiricallyjusti\ufb01ed[Clevertetal.,2015,Klambaueretal.,2017].WealsoworkwithRademachercomplexityandcoveringnumbersinsteadofthePAC-Bayesframework.ItisrelativelysimpletoadaptourtechniquestorelunetworkstoproduceasimilarresulttothatofNagarajanandKolter[2019],byconditioningonlargepre-activationvaluesinourLipschitzaugmentationstep(seeSection4.2).InSectionH,weprovideasketchofthisargumentandobtainaboundforrelunetworksthatispolynomialinhiddenlayerandJacobiannormsandinversepreactivations.However,itisnotobvioushowtoadapttheargumentofNagarajanandKolter[2019]toactivationfunctionswhosederivativesarenotpiecewise-constant.DziugaiteandRoy[2018,2017b]developPAC-Bayesboundsfordata-dependentpriorsobtainedviasomedifferentiallyprivatemechanism.Theirboundsareforarandomizedclassi\ufb01ersampledfromtheprior,whereasweanalyzeadeterministic,\ufb01xedmodel.Novaketal.[2018]empiricallydemonstratethatthesensitivityofaneuralnettoinputnoisecorrelateswithgeneralization.Sokoli\u00b4cetal.[2017],KruegerandMemisevic[2015]proposestability-basedregularizersforneuralnets.Hardtetal.[2015]showthatmodelswhichtrainfastertendtogeneralizebetter.Keskaretal.[2016],Hofferetal.[2017]studytheeffectofbatchsizeongeneralization.Brutzkusetal.[2017]analyzeaneuralnetworktrainedonhingelossandlinearlyseparabledataandshowthatgradientdescentrecoverstheexactseparatinghyperplane.3NotationLet1(E)betheindicatorfunctionofeventE.Letl0-1denotethestandard0-1loss.For\u03ba\u22650,Let1\u2264\u03ba(\u00b7)bethesoftenedindicatorfunctionde\ufb01nedas1\u2264\u03ba(t)=(1ift\u2264\u03ba2\u2212t/\u03baif\u03ba\u2264t\u22642\u03ba0if2\u03ba\u2264tNotethat1\u2264\u03bais\u03ba\u22121-Lipschitz.De\ufb01nethenormk\u00b7kp,qbykAkp,q,(cid:16)Pj(cid:0)PiApi,j(cid:1)q/p(cid:17)1/q.LetPnbeauniformdistributionovernpoints{x1,...,xn}\u2282Dx.LetfbeafunctionthatmapsDxtosomeoutputspaceDf,andassumebothspacesareequippedwithsomenorms|||\u00b7|||(thesenormscanbedifferentbutweusethesamenotationsforthem).ThentheL2(Pn,|||\u00b7|||)normofthefunctionfisde\ufb01nedaskfkL2(Pn,|||\u00b7|||),(cid:16)1nPi|||f(xi)|||2(cid:17)1/2.WeuseDtodenotetotalderivativeoperator,andthusDf(x)representstheJacobianoffatx.SupposeFisafamilyoffunctionsfromDxtoDf.LetC(\u0001,F,\u03c1)bethecoveringnumberofthefunctionclassFw.r.t.metric\u03c1withcoversize\u0001.Inmanycases,thecoveringnumberdependsontheexamplesthroughthenormsoftheexamples,andinthispaperweonlyworkwiththesecases.Thus,weletN(\u0001,F,s)bethemaximumcoveringnumberforanypossiblendatapointswithnormnotlargerthans.Precisely,ifwede\ufb01nePn,stobethesetofallpossibleuniformdistributionssupportedonndatapointswithnormsnotlargerthans,thenN(\u0001,F,s),supPn\u2208Pn,sC(\u0001,F,L2(Pn,|||\u00b7|||)).SupposeFcontainsfunctionswithminputsthatmapfromatensorproductmEuclideanspacetoEuclideanspace,thenwede\ufb01neN(\u0001,F,(s1,...,sm)),supP:\u2200(x1,...,xm)\u2208supp(P)kxik\u2264siC(\u0001,F,L2(P)).4OverviewofMainResultsandProofTechniquesInthissection,wegiveageneraloverviewofthemaintechnicalresultsandoutlinehowtoprovethemwithminimalnotation.Wewillpointtolatersectionswheremanystatementsareformalized.Tosimplifythecoremathematicalreasoning,weabstractfeed-forwardneuralnetworks(includingresidualnetworks)ascompositionsofoperations.LetF1,...,Fkbeasequenceoffamiliesoffunctions(correspondingtofamiliesofsinglelayerneuralnetsinthedeeplearningsetting)and\u2018be4\faLipschitzlossfunctiontakingvaluesin[0,1].Westudythecompositionsof\u2018andfunctionsinFi\u2019s:L,\u2018\u25e6Fk\u25e6Fk\u22121\u00b7\u00b7\u00b7\u25e6F1={\u2018\u25e6fk\u25e6fk\u22121\u25e6\u00b7\u00b7\u00b7\u25e6f1:\u2200i,fi\u2208Fi}(1)Textbookresults[BartlettandMendelson,2002]boundthegeneralizationerrorbytheRademachercomplexity(formallyde\ufb01nedinSectionC)ofthefamilyoflossesL,whichinturnisboundedbythecoveringnumberofLthroughDudley\u2019sentropyintegraltheorem[Dudley,1967].Modulominornuances,thekeyremainingquestionistogiveatightcoveringnumberboundforthefamilyLforeverytargetcoversize\u0001inacertainrange(often,considering\u0001\u2208[1/nO(1),1]suf\ufb01ces).Asalludedtointheintroduction,generalizationerrorboundsobtainedthroughthismachineryonlydependonthe(training)datathroughthemargininthelossfunction,andouraimistoutilizemoredata-dependentproperties.Towardsunderstandingwhichdata-dependentpropertiesareusefultoregularize,itishelpfultorevisitthedata-independentcoveringtechniqueof[Bartlettetal.,2017],theskeletonofwhichissummarizedbelow.RecallthatN(\u0001,F,s)denotesthecoveringnumberforarbitraryndatapointswithnormlessthans.Thefollowinglemmasaysthatiftheintermediatevariable(orthehiddenlayer)fi\u25e6\u00b7\u00b7\u00b7\u25e6f1(x)isbounded,andthecompositionoftherestofthefunctionsl\u25e6fk\u25e6\u00b7\u00b7\u00b7\u25e6fi+1(x)isLipschitz,thensmallcoveringnumberoflocalfunctionsimplysmallcoveringnumberforthecompositionoffunctions.Lemma4.1.[abstractionoftechniquesin[Bartlettetal.,2017]]Inthecontextabove,assume:1.foranyx\u2208supp(Pn),|||fi\u25e6\u00b7\u00b7\u00b7\u25e6f1(x)|||\u2264si.2.\u2018\u25e6fk\u25e6\u00b7\u00b7\u00b7\u25e6fi+1is\u03bai-Lipschitzforalli.Then,wehavethefollowingcoveringnumberboundforL(foranychoiceof\u00011,...,\u0001k>0):logN(Pki=1\u03bai\u0001i,L,s0)\u2264Pki=1logN(\u0001i,Fi,si\u22121).ThelemmasaysthatthelogcoveringnumberandthecoversizescalelinearlyiftheLipschitznessparametersandnormsremainconstant.However,thesetwoquantities,intheworstcase,caneasilyscaleexponentiallyinthenumberoflayers,andtheyarethemainsourcesofthedependencyofproductofspectral/Frobeniusnormsoflayersin[Golowichetal.,2017,Bartlettetal.,2017,Neyshaburetal.,2017a,2015b]Moreprecisely,theworst-caseLipschitznessoverallpossibledatapointscanbeexponentiallybiggerthantheaverage/typicalLipschitznessforexamplesrandomlydrawnfromthetrainingortestdistribution.WeaimtobridgethisgapbyderivingageneralizationerrorboundthatonlydependsontheLipschitznessandboundednessonthetrainingexamples.Ourgeneralapproach,partiallyinspiredbymargintheory,istoaugmentthelossfunctionbysoftindicatorsofLipschitznessandboundedness.Lethibeshorthandnotationforfi\u25e6\u00b7\u00b7\u00b7\u25e6f1,thei-thintermediatevalue,andletz(x),\u2018(hk(x))betheoriginalloss.Our\ufb01rstattemptconsidered:\u02dcz0(x),1+(z(x)\u22121)\u00b7kYi=11\u2264si(khi(x)k)\u00b7kYi=11\u2264\u03bai(k\u2202z/\u2202hikop)(2)Sinceztakesvaluesin[0,1],theaugmentedloss\u02dcz0isanupperboundontheoriginallosszwithequalitywhenalltheindicatorsaresatis\ufb01edwithvalue1.Thehopewasthattheindicatorswould\ufb02attenthoseregionswherehiisnotboundedandwherezisnotLipschitzinhi.However,therearetwoimmediateissues.First,thesoftindicatorsfunctionsarethemselvesfunctionsofhi.It\u2019sunclearwhethertheaugmentedfunctioncanbeLipschitzwithasmallconstantw.r.thi,andthuswecannotapplyLemma4.1.1Second,theaugmentedlossfunctionbecomescomplicatedanddoesn\u2019tfallintothesequentialcomputationformofLemma4.1,andthereforeevenifLipschitznessisnotanissue,weneednewcoveringtechniquesbeyondLemma4.1.Weaddressthe\ufb01rstissuebyrecursivelyaugmentingthelossfunctionbymultiplyingmoresoftindicatorsthatboundtheJacobianofthecurrentfunction.The\ufb01nalloss\u02dczreads:2\u02dcz(x),1+(z(x)\u22121)\u00b7kYi=11\u2264si(khi(x)k)\u00b7Y1\u2264i\u2264j\u2264k1\u2264\u03baj\u2190i(kDfj\u25e6\u00b7\u00b7\u00b7\u25e6fi[hi\u22121]kop)(3)1Apriori,it\u2019salsounclearwhat\u201cLipschitzinhi\u201dmeanssincethe\u00afz0doesnotonlydependonxthroughhi.Wewillformalizethisinlatersectionafterde\ufb01ningproperlanguageaboutdependenciesbetweenvariables.2Unlikeinequation(2),wedon\u2019taugmenttheJacobianofthelossw.r.tthelayers.Thisallowsustodealwithnon-differentiablelossfunctionssuchasramploss.5\fwhere\u03baj\u2190i\u2019sareuser-de\ufb01nedparameters.Forourapplicationtoneuralnets,weinstantiatesiasthemaximumnormoflayeriand\u03baj\u2190iasthemaximumnormoftheJacobianbetweenlayerjandiacrossthetrainingdataset.Apolynomialin\u03ba,scanbeshowntoboundtheworst-caseLipschitznessofthefunctionw.r.t.theintermediatevariablesintheformulaabove.3Byourchoiceof\u03ba,s,a)thetraininglossisunaffectedbytheaugmentationandb)theworst-caseLipschitznessofthelossiscontrolledbyapolynomialoftheLipschitznessonthetrainingexamples.WeprovideaninformaloverviewofouraugmentationprocedureinSection4.2andformallystatede\ufb01nitionsandguaranteesinSectionB.ThedownsideoftheLipschitzaugmentationisthatitfurthercomplicatesthelossfunction.Towardscoveringthelossfunction(assumingLipschitzproperties)ef\ufb01ciently,weextendLemma4.1,whichworksforsequentialcompositionsoffunctions,togeneralfamiliesofformulas,orcomputationalgraphs.WeinformallyoverviewthisextensioninSection4.1usingaminimalsetofnotations,andinSectionA,wegiveaformalpresentationoftheseresults.CombiningtheLipschitzaugmentationandgraphscoveringresults,weobtainacoveringnumberboundofaugmentedloss.ThetheorembelowisformallystatedinTheoremB.3ofSectionB.Theorem4.2.Let\u02dcLbethefamilyofaugmentedlossesde\ufb01nedin(3).Forcoverresolutions\u0001iandvalues\u02dc\u03baithatarepolynomialintheparameterssi,\u03baj\u2190i,weobtainthefollowingcoveringnumberboundfor\u02dcL:logN(Xi\u0001i\u02dc\u03bai,\u02dcL,s0)\u2264XilogN(\u0001i,Fi,si\u22121)+XilogN(\u0001i,DFi,si\u22121)whereDFidenotesthefunctionclassobtainedfromapplyingthetotalderivativeoperatortoallfunctionsinFi.Now,followingthestandardtechniqueofboundingRademachercomplexityviacoveringnumbers,wecanobtaingeneralizationerrorboundsforaugmentedloss.Forthedemonstrationofourtechnique,supposethatthefollowingsimpli\ufb01cationholds:logN(\u0001i,DFi,si\u22121)=logN(\u0001i,Fi,si\u22121)=s2i\u22121/\u00012i.Thenafterminimizingthecoveringnumberboundin\u0001iviastandardtechniques,weobtainthebelowgeneralizationerrorboundontheoriginallossforparameters\u02dc\u03baialludedtoinTheorem4.2andformallyde\ufb01nedinTheoremB.2.Whenthetrainingexamplessatisfytheaugmentedindicators,Etrain[\u02dcz]=Etrain[z],andbecause\u02dczboundszfromabove,wehaveEtest[z]\u2212Etrain[z]\u2264Etest[\u02dcz]\u2212Etrain[\u02dcz]\u2264eO (cid:16)Pi\u02dc\u03ba2/3is2/3i\u22121(cid:17)3/2\u221an+rlog(1/\u03b4)n!(4)4.1OverviewofComputationalGraphCoveringToobtaintheaugmented\u02dczde\ufb01nedin(3),weneededtoconditionondata-dependentpropertieswhichintroduceddependenciesbetweenthevariouslayers.Becauseofthis,Lemma4.1isnolongersuf\ufb01cienttocover\u02dcz.Inthissection,weinformallyoverviewhowtoextendLemma4.1tocovermoregeneralfunctionsviathenotionofcomputationalgraphs.Forspaceconstraints,thissectionisadramaticallyabbreviatedandinformalversionofSectionA.AcomputationalgraphG(V,E,{RV})isanacyclicdirectedgraphwiththreecomponents:thesetofnodesVcorrespondstovariables,thesetofedgesEdescribesdependenciesbetweenthesevariables,and{RV}containsalistofcompositionrulesindexedbythevariablesV\u2019s,representingtheprocessofcomputingVfromitsdirectpredecessors.Forsimplicity,weassumethegraphcontainsauniquesink,denotedbyOG,andwecallitthe\u201coutputnode\u201d.WealsooverloadthenotationOGtodenotethefunctionthatthecomputationalgraphG\ufb01nallycomputes.LetIG={I1,...,Ip}bethesubsetofnodeswithnopredecessors,whichwecallthe\u201cinputnodes\u201dofthegraph.Thenotionofafamilyofcomputationalgraphsgeneralizesthesequentialfamilyoffunctioncom-positionsin(1).LetG={G(V,E,{RV})}beafamilyofcomputationalgraphswithsharednodes,edges,outputnode,andinputnodes(denotedbyI).LetRVbethecollectionofallpossiblecompo-sitionrulesusedfornodeVbythegraphsinthefamilyG.ThisfamilyGde\ufb01nesasetoffunctionsOG,{OG:G\u2208G}.3Asmentionedinfootnote1,wewillformalizetheprecisemeaningofLipschitznesslater.6\fThetheorembelowextendsLemma4.1.Inthecomputationalgraphinterpretation,Lemma4.1appliestoasequentialfamilyofcomputationalgraphswithkinternalnodesV1,...,Vk,whereeachVicomputesthefunctionfi,andtheoutputcomputesthecompositionOG=\u2018\u25e6fk\u00b7\u00b7\u00b7\u25e6f1=z.However,theaugmentedloss\u02dcznolongerhasthissequentialstructure,requiringthebelowtheoremforcoveringgenericfamiliesofcomputationalgraphs.Weshowthatcoveringageneralfamilyofcomputationalgraphscanbereducedtocoveringallthelocalcompositionrules.Theorem4.3(InformalandweakerversionofTheoremA.3).Supposethatthereisanordering(V1,...,Vm)ofthenodes,sothataftercuttingoutnodesV1,...,Vi\u22121,thenodeVibecomesaleafnodeandtheoutputOGis\u03baVi-Lipschitzw.r.ttoViforallG\u2208G.Inaddition,assumethatforallG\u2208G,thenodeV\u2019svaluehasnormatmostsV.Letpr(V)beallthepredecessorsofVandspr(V)bethelistofnormupperboundsofthepredecessorsofV.Then,smallcoveringnumbersforallofthelocalcompositionrulesofVwithresolution\u0001VwouldimplysmallcoveringnumberforthefamilyofcomputationalgraphswithresolutionPV\u0001V\u03baV:logN(XV\u2208V\\I\u222a{O}\u03baV\u0001V+\u0001O,OG,sI)\u2264XV\u2208V\\IlogN(\u0001V,RV,spr(V))(5)InSectionAweformalizethenotionof\u201ccutting\u201dnodesfromthegraph.TheconditionthatnodeV\u2019svaluehasnormatmostsVisasimpli\ufb01cationmadeforexpositionalpurposes;ourfullTheo-remA.3alsoappliesifOGcollapsestoaconstantwhenevernodeV\u2019svaluehasnormgreaterthansV.Thisallowsforthesoftenedindicators1\u2264si(khi(x)k)usedin(3).4.2LipschitzAugmentationofComputationalGraphsThecoveringnumberboundofTheorem4.3reliesonLipschitznessw.r.tinternalnodesofthegraphunderaworst-casechoiceofinputs.Fordeepnetworks,thiscanscaleexponentiallyindepthviatheproductofweightnormsandeasilybelargerthantheaverageLipschitz-nessovertypicalinputs.Inthissection,weexplainageneraloperationtoaugmentsequentialgraphs(suchasneuralnets)intographswithbetterworst-caseLipschitzconstants,sotoolssuchasTheorem4.3canbeapplied.Thissectionisheavilysimpli\ufb01edforspaceconstraints.Formalde\ufb01nitionsandtheoremstatementsareinSectionB.Theaugmentationreliesonintroducingtermssuchasthesoftindicatorsinequation(2)and(3)whichconditionondata-dependentproperties.AsoutlinedinSection4,theywilltranslatetothedata-dependentpropertiesinthegeneralizationbounds.Wealsorequiretheaugmentedfunctiontoupperboundtheoriginal.Wewillpresentagenericapproachtoaugmentfunctioncompositionssuchasz,\u2018\u25e6fk\u25e6...\u25e6f1,whoseLipschitzconstantsarepotentiallyexponentialindepth,withonlypropertiesinvolvingthenormsoftheinter-layerJacobians.Wewillproduce\u02dcz,whoseworst-caseLipschitznessw.r.t.internalnodescanbepolynomialindepth.InformalexplanationofLipschitzaugmentation:InthesamesettingofSection4,recallthatin(2),our\ufb01rstunsuccessfulattempttosmoothoutthefunctionwasbymultiplyingindicatorsonthenormsofthederivativesoftheoutput:Qki=11\u2264\u03bai(k\u2202z/\u2202hikop).Thedif\ufb01cultyliesincontrollingtheLipschitznessofthenewtermsk\u2202z/\u2202hikopthatweintroduce:bythechainrule,wehavetheexpansion\u2202z\u2202hi=\u2202z\u2202hk\u2202hk\u2202hk\u22121\u00b7\u00b7\u00b7\u2202hi+1\u2202hi,whereeachhj0isitselfafunctionofhjforj0>j.Thismeans\u2202z\u2202hiisacomplicatedfunctionintheintermediatevariableshjfor1\u2264j\u2264k.BoundingtheLipschitznessof\u2202z\u2202hirequiresaccountingfortheLipschitznessofeveryterminitsexpansion,whichischallengingandcreatescomplicateddependenciesbetweenvariables.Ourkeyinsightisthatbyconsideringamorecomplicatedaugmentationwhichconditionsonthederivativesbetweenallintermediatevariables,wecanstillcontrolLipschitznessofthesystem,leadingtothemoreinvolvedaugmentationpresentedin(3).OurmaintechnicalcontributionisTheorem4.4,whichweinformallystatebelow.Theorem4.4(InformalversionofTheoremB.2).Thefunctions\u02dcz(de\ufb01nedin(3))canbecomputedbyafamilyofcomputationalgraphseGillustratedinFigure2.ThisfamilyhasinternalnodesViandJicomputinghiandDfi[hi\u22121],respectively,andcomputesamodi\ufb01edoutputrulethataugmentsthe7\foriginalwithsoftindicators.ThesesoftindicatorsconditionthatthenormsoftheJacobiansandhiareboundedbyparameters\u03baj\u2190i,si.Importantly,theoutputO\u02dcGis\u02dc\u03baVi,\u02dc\u03baJi-Lipschitzw.r.t.Vi,Ji,respectively,aftercuttingnodesV1,J1,...,Vi\u22121,Ji\u22121,forparameters\u02dc\u03baVi,\u02dc\u03baJithatarepolynomialsin\u03baj\u2190i,si.Inaddition,theaugmentedfunction\u02dczwillupperboundtheoriginalwithequalitywhenalltheindicatorsaresatis\ufb01ed.Thecruxoftheproofisleveragingthechainruletodecompose\u2202z\u2202hiintoaproductandthenapplyingatelescopingargumenttoboundthedifferenceintheproductbydifferencesinindividualterms.InSectionBwepresentaformalversionofthisresultandalsoapplyTheorem4.3toproduceacoveringnumberboundforeG.5ApplicationtoNeuralNetworksFigure2:Lipschitzaugmentation(informallyde\ufb01ned).Inthissectionweprovideourgeneralizationboundforneuralnets,whichwasobtainedusingmachineryfromSection4.1.De\ufb01neaneuralnetworkFparameterizedbyrweightmatri-ces{W(i)}byF(x)=W(r)\u03c6(\u00b7\u00b7\u00b7\u03c6(W(1)(x))\u00b7\u00b7\u00b7).Weusetheconventionthatactivationsandmatrixmultiplicationsaretreatedasdistinctlayersindexedwithasubscript,withoddlay-ersapplyingamatrixmultiplicationandevenlayersapplying\u03c6(seeExampleA.1foravisualization).AdditionalnotationdetailsandtheproofareinSectionC.ThebelowresultfollowsfrommodelingtheneuralnetlossasasequentialcomputationalgraphandusingouraugmentationproceduretomakeitLipschitzinitsnodeswithparameters\u03bahidden,(i),\u03bajacobian,(i).ThenwecovertheaugmentedlosstobounditsRademachercomplexity.Theorem5.1.Assumethattheactivation\u03c6is1-Lipschitzwitha\u00af\u03c3\u03c6-Lipschitzderivative.Fixreferencematrices{A(i)},{B(i)}.Withprobability1\u2212\u03b4overtherandomdrawsofthedataPn,allneuralnetworksFwithparameters{W(i)}andpositivemargin\u03b3satisfy:E(x,y)\u223cP[l0-1(F(x),y)]\u2264\u02dcO\uf8eb\uf8ec\uf8ed(cid:16)Pi(\u03bahidden,(i)a(i)t(i\u22121))2/3+(\u03bajacobian,(i)b(i))2/3(cid:17)3/2\u221an+rrlog(1/\u03b4)n\uf8f6\uf8f7\uf8f8where\u03bajacobian,(i),P1\u2264j\u22642i\u22121\u2264j0\u22642r\u22121\u03c3j0\u21902i\u03c32i\u22122\u2190j\u03c3j0\u2190j,and\u03bahidden,(i),\u03be+\u03c32r\u22121\u21902i\u03b3+Pi\u2264i0<r\u03c32i0\u21902it(i0)+P1\u2264j\u2264j0\u22642r\u22121Pj0j00=max{2i,j},j00even\u00af\u03c3\u03c6\u03c3j0\u2190j00+1\u03c3j00\u22121\u21902i\u03c3j00\u22121\u2190j\u03c3j0\u2190j.Intheseexpressions,wede\ufb01ne\u03c3j\u22121\u2190j=1,\u03be=poly(r)\u22121,and:a(i),kW(i)>\u2212A(i)>k2,1+\u03be,b(i),kW(i)\u2212B(i)k1,1+\u03bet(0),maxx\u2208Pnkxk+\u03be,t(i),maxx\u2208PnkF2i\u21901(x)k+\u03be\u03c3j0\u2190j,maxx\u2208PnkQj0\u2190j(x)kop+\u03be,and\u03b3,min(x,y)\u2208Pn[F(x)]y\u2212maxy06=y[F(x)]y0>0whereQj0\u2190jcomputestheJacobianoflayerj0w.r.t.layerj.Notethatthetrainingerrorhereis0becauseoftheexistenceofpositivemargin\u03b3.Wenotethatourboundhasnoexplicitdependenceonwidthandinsteaddependsonthek\u00b7k2,1,k\u00b7k1,1normsoftheweightsoffsetbyreferencematrices{A(i)},{B(i)}.Thesenormscanavoidscalingwiththewidthofthenetworkifthedifferencebetweentheweightsandreferencematricesissparse.Thereferencematrices{A(i)},{B(i)}areusefulifthereissomepriorbeliefbeforetrainingaboutwhatweightmatricesarelearned,andtheyalsoappearintheboundsofBartlettetal.[2017].InSectionG,wealsoshowthatourtechniquescaneasilybeextendedtoprovidegeneralizationboundsforRNNsscalingpolynomiallyindepthviathesamequantitiest(i),\u03c3j0\u2190j.8\fTable1:TesterrorforamodeltrainedonCIFAR10invarioussettings.SettingNormalizationJacobianRegTestErrorBaselineBatchNorm\u00d74.43%Lowlearningrate(0.01)BatchNorm\u00d75.98%X5.46%NodataaugmentationBatchNorm\u00d710.44%X8.25%NoBatchNormNone\u00d76.65%LayerNorm[Baetal.,2016]\u00d76.20%X5.57%6ExperimentsThoughthemainpurposeofthepaperistostudythedata-dependentgeneralizationboundsfromatheoreticalperspective,weprovidepreliminaryexperimentsdemonstratingthattheproposedcomplexitymeasureandgeneralizationboundsareempiricallyrelevant.Weshowthatregularizingthecomplexitymeasureleadstobettertestaccuracy.InspiredbyTheorem5.1,wedirectlyregularizetheJacobianoftheclassi\ufb01cationmarginw.r.toutputsofnormalizationlayersandafterresidualblocks.Ourreasoningisthatnormalizationlayerscontrolthehiddenlayernorms,soadditionallyregularizingtheJacobiansresultsinregularizationoftheproduct,whichappearsinourbound.We\ufb01ndthatthisiseffectiveforimprovingtestaccuracyinavarietyofsettings.WenotethatSokoli\u00b4cetal.[2017]showpositiveexperimentalresultsforasimilarregularizationtechniqueindata-limitedsettings.Supposethatm(F(x),y)=[F(x)]y\u2212maxj6=y[t]jdenotesthemarginofthenetworkforex-ample(x,y).Lettingh(i)denotesomehiddenlayerofthenetwork,wede\ufb01nethenotationJ(i),\u2202\u2202h(i)m(F(x),y)andusetrainingobjective\u02c6Lreg[F],E(x,y)\u223cPn\"l(x,y)+\u03bb Xi1(kJ(i)(x)k2F\u2265\u03c3)kJ(i)(x)k2F!#whereldenotesthestandardcrossentropyloss,and\u03bb,\u03c3arehyperparameters.NotetheJacobianistakenwithrespecttoascalaroutputandthereforeisavector,soitiseasytocompute.ForaWideResNet16[ZagoruykoandKomodakis,2016]architecture,wetrainusingtheaboveobjective.ThethresholdontheFrobeniusnormintheregularizationisinspiredbythetruncationsinouraugmentedloss(inallourexperiments,wechoose\u03c3=0.1).Wetunethecoef\ufb01cient\u03bbasahyperparameter.Inourexperiments,wetooktheregularizedindicesitobelastlayersineachresidualblockaswellaslayersinresidualblocksfollowingaBatchNorminthestandardWideResNet16architecture.IntheLayerNormsetting,wesimplyreplacedBatchNormlayerswithLayerNorm.TheremaininghyperparametersettingsarestandardforWideResNet;foradditionaldetailsseeSectionI.1.Figure1showstheresultsformodelstrainedandtestedonCIFAR10inlowlearningrateandnodataaugmentationsettings,whicharesettingswheregeneralizationtypicallysuffers.WealsoexperimentwithreplacingBatchNormlayerswithLayerNormandadditionallyregularizingtheJacobian.Weobserveimprovementsintesterrorforallthesesettings.InSectionI.2,weempiricallydemonstratethatourcomplexitymeasureindeedavoidstheexponentialscalingindepthforaWideResNetmodeltrainedonCIFAR10.7ConclusionInthispaper,wetacklethequestionofhowdata-dependentpropertiesaffectgeneralization.WeprovetightergeneralizationboundsthatdependpolynomiallyonthehiddenlayernormsandnormsoftheinterlayerJacobians.Toprovethesebounds,weworkwiththeabstractionofcomputationalgraphsanddevelopgeneraltoolstoaugmentanysequentialfamilyofcomputationalgraphsintoaLipschitzfamilyandthencoverthisLipschitzfamily.Thisaugmentationandcoveringprocedureappliestoanysequenceoffunctioncompositions.Aninterestingdirectionforfutureworkistogeneralizeourtechniquestoarbitrarycomputationalgraphstructures.Additionally,encouragedbyourpromisingpreliminaryresults,webelievethereistheexcitingempiricaldirectionofapplyingtheseboundstodevelopbetterdata-dependentregularization.9\fAcknowledgmentsCWwassupportedbyaNSFGraduateResearchFellowship.ToyotaResearchInstitute(TRI)providedfundstoassisttheauthorswiththeirresearchbutthisarticlesolelyre\ufb02ectstheopinionsandconclusionsofitsauthorsandnotTRIoranyotherToyotaentity.ReferencesSanjeevArora,RongGe,BehnamNeyshabur,andYiZhang.Strongergeneralizationboundsfordeepnetsviaacompressionapproach.arXivpreprintarXiv:1802.05296,2018.JimmyLeiBa,JamieRyanKiros,andGeoffreyEHinton.Layernormalization.arXivpreprintarXiv:1607.06450,2016.PeterLBartlettandShaharMendelson.Rademacherandgaussiancomplexities:Riskboundsandstructuralresults.JournalofMachineLearningResearch,3(Nov):463\u2013482,2002.PeterLBartlett,DylanJFoster,andMatusJTelgarsky.Spectrally-normalizedmarginboundsforneuralnetworks.InAdvancesinNeuralInformationProcessingSystems,pages6240\u20136249,2017.FriedrichLBauer.Computationalgraphsandroundingerror.SIAMJournalonNumericalAnalysis,11(1):87\u201396,1974.AlonBrutzkus,AmirGloberson,EranMalach,andShaiShalev-Shwartz.Sgdlearnsover-parameterizednetworksthatprovablygeneralizeonlinearlyseparabledata.arXivpreprintarXiv:1710.10174,2017.MinshuoChen,XingguoLi,andTuoZhao.Ongeneralizationboundsofafamilyofrecurrentneuralnetworks,2019.URLhttps://openreview.net/forum?id=Skf-oo0qt7.Djork-Arn\u00e9Clevert,ThomasUnterthiner,andSeppHochreiter.Fastandaccuratedeepnetworklearningbyexponentiallinearunits(elus).arXivpreprintarXiv:1511.07289,2015.RMDudley.Thesizesofcompactsubsetsofhilbertspaceandcontinuityofgaussianprocesses.JournalofFunctionalAnalysis,1(3):290\u2013330,1967.GintareKarolinaDziugaiteandDanielMRoy.Computingnonvacuousgeneralizationboundsfordeep(stochastic)neuralnetworkswithmanymoreparametersthantrainingdata.arXivpreprintarXiv:1703.11008,2017a.GintareKarolinaDziugaiteandDanielMRoy.Entropy-sgdoptimizesthepriorofapac-bayesbound:Generalizationpropertiesofentropy-sgdanddata-dependentpriors.arXivpreprintarXiv:1712.09376,2017b.GintareKarolinaDziugaiteandDanielMRoy.Data-dependentpac-bayespriorsviadifferentialprivacy.InAdvancesinNeuralInformationProcessingSystems,pages8430\u20138441,2018.NoahGolowich,AlexanderRakhlin,andOhadShamir.Size-independentsamplecomplexityofneuralnetworks.arXivpreprintarXiv:1712.06541,2017.SuriyaGunasekar,BlakeEWoodworth,SrinadhBhojanapalli,BehnamNeyshabur,andNatiSrebro.Implicitregularizationinmatrixfactorization.InAdvancesinNeuralInformationProcessingSystems,pages6151\u20136159,2017.SuriyaGunasekar,JasonLee,DanielSoudry,andNathanSrebro.Characterizingimplicitbiasintermsofoptimizationgeometry.arXivpreprintarXiv:1802.08246,2018a.SuriyaGunasekar,JasonLee,DanielSoudry,andNathanSrebro.Implicitbiasofgradientdescentonlinearconvolutionalnetworks.arXivpreprintarXiv:1806.00468,2018b.MoritzHardt,BenjaminRecht,andYoramSinger.Trainfaster,generalizebetter:Stabilityofstochasticgradientdescent.arXivpreprintarXiv:1509.01240,2015.10\fEladHoffer,ItayHubara,andDanielSoudry.Trainlonger,generalizebetter:closingthegeneraliza-tiongapinlargebatchtrainingofneuralnetworks.InAdvancesinNeuralInformationProcessingSystems,pages1731\u20131741,2017.SergeyIoffeandChristianSzegedy.Batchnormalization:Acceleratingdeepnetworktrainingbyreducinginternalcovariateshift.arXivpreprintarXiv:1502.03167,2015.ZiweiJiandMatusTelgarsky.Riskandparameterconvergenceoflogisticregression.arXivpreprintarXiv:1803.07300,2018.NitishShirishKeskar,DheevatsaMudigere,JorgeNocedal,MikhailSmelyanskiy,andPingTakPeterTang.Onlarge-batchtrainingfordeeplearning:Generalizationgapandsharpminima.arXivpreprintarXiv:1609.04836,2016.G\u00fcnterKlambauer,ThomasUnterthiner,AndreasMayr,andSeppHochreiter.Self-normalizingneuralnetworks.InAdvancesinneuralinformationprocessingsystems,pages971\u2013980,2017.PascalKoiranandEduardoDSontag.Vapnik-chervonenkisdimensionofrecurrentneuralnetworks.InEuropeanConferenceonComputationalLearningTheory,pages223\u2013237.Springer,1997.DavidKruegerandRolandMemisevic.Regularizingrnnsbystabilizingactivations.arXivpreprintarXiv:1511.08400,2015.YuanzhiLi,TengyuMa,andHongyangZhang.Algorithmicregularizationinover-parameterizedmatrixsensingandneuralnetworkswithquadraticactivations.InConferenceOnLearningTheory,pages2\u201347,2018.EtaiLittwinandLiorWolf.Regularizingbythevarianceoftheactivations\u2019sample-variances.InAdvancesinNeuralInformationProcessingSystems,pages2115\u20132125,2018.VaishnavhNagarajanandZicoKolter.DeterministicPAC-bayesiangeneralizationboundsfordeepnetworksviageneralizingnoise-resilience.InInternationalConferenceonLearningRepresenta-tions,2019.URLhttps://openreview.net/forum?id=Hygn2o0qKX.BehnamNeyshabur,RyotaTomioka,RuslanSalakhutdinov,andNathanSrebro.Data-dependentpathnormalizationinneuralnetworks.arXivpreprintarXiv:1511.06747,2015a.BehnamNeyshabur,RyotaTomioka,andNathanSrebro.Norm-basedcapacitycontrolinneuralnetworks.InConferenceonLearningTheory,pages1376\u20131401,2015b.BehnamNeyshabur,SrinadhBhojanapalli,DavidMcAllester,andNathanSrebro.Apac-bayesianapproachtospectrally-normalizedmarginboundsforneuralnetworks.arXivpreprintarXiv:1707.09564,2017a.BehnamNeyshabur,SrinadhBhojanapalli,DavidMcAllester,andNatiSrebro.Exploringgeneraliza-tionindeeplearning.InAdvancesinNeuralInformationProcessingSystems,pages5947\u20135956,2017b.BehnamNeyshabur,ZhiyuanLi,SrinadhBhojanapalli,YannLeCun,andNathanSrebro.Towardsunderstandingtheroleofover-parametrizationingeneralizationofneuralnetworks.arXivpreprintarXiv:1805.12076,2018.RomanNovak,YasamanBahri,DanielAAbola\ufb01a,JeffreyPennington,andJaschaSohl-Dickstein.Sensitivityandgeneralizationinneuralnetworks:anempiricalstudy.arXivpreprintarXiv:1802.08760,2018.JureSokoli\u00b4c,RajaGiryes,GuillermoSapiro,andMiguelRDRodrigues.Robustlargemargindeepneuralnetworks.IEEETransactionsonSignalProcessing,65(16):4265\u20134280,2017.DanielSoudry,EladHoffer,MorShpigelNacson,SuriyaGunasekar,andNathanSrebro.Theimplicitbiasofgradientdescentonseparabledata.TheJournalofMachineLearningResearch,19(1):2822\u20132878,2018.11\fNitishSrivastava,GeoffreyHinton,AlexKrizhevsky,IlyaSutskever,andRuslanSalakhutdinov.Dropout:asimplewaytopreventneuralnetworksfromover\ufb01tting.TheJournalofMachineLearningResearch,15(1):1929\u20131958,2014.StefanWager,SidaWang,andPercySLiang.Dropouttrainingasadaptiveregularization.InAdvancesinneuralinformationprocessingsystems,pages351\u2013359,2013.ColinWei,JasonDLee,QiangLiu,andTengyuMa.Onthemargintheoryoffeedforwardneuralnetworks.arXivpreprintarXiv:1810.05369,2018.Wikipediacontributors.Chainrule\u2014Wikipedia,thefreeencyclopedia,2019.YuxinWuandKaimingHe.Groupnormalization.arXivpreprintarXiv:1803.08494,2018.SergeyZagoruykoandNikosKomodakis.Wideresidualnetworks.arXivpreprintarXiv:1605.07146,2016.ChiyuanZhang,SamyBengio,MoritzHardt,BenjaminRecht,andOriolVinyals.Understandingdeeplearningrequiresrethinkinggeneralization.arXivpreprintarXiv:1611.03530,2016.HongyiZhang,YannN.Dauphin,andTengyuMa.Residuallearningwithoutnormalizationviabetterinitialization.InInternationalConferenceonLearningRepresentations,2019.URLhttps://openreview.net/forum?id=H1gsz30cKX.JiongZhang,QiLei,andInderjitSDhillon.Stabilizinggradientsfordeepneuralnetworksviaef\ufb01cientsvdparameterization.arXivpreprintarXiv:1803.09327,2018.12\f", "award": [], "sourceid": 5136, "authors": [{"given_name": "Colin", "family_name": "Wei", "institution": "Stanford University"}, {"given_name": "Tengyu", "family_name": "Ma", "institution": "Stanford University"}]}