{"title": "Local Smoothness in Variance Reduced Optimization", "book": "Advances in Neural Information Processing Systems", "page_first": 2179, "page_last": 2187, "abstract": "Abstract We propose a family of non-uniform sampling strategies to provably speed up a class of stochastic optimization algorithms with linear convergence including Stochastic Variance Reduced Gradient (SVRG) and Stochastic Dual Coordinate Ascent (SDCA). For a large family of penalized empirical risk minimization problems, our methods exploit data dependent local smoothness of the loss functions near the optimum, while maintaining convergence guarantees. Our bounds are the first to quantify the advantage gained from local smoothness which are significant for some problems significantly better. Empirically, we provide thorough numerical results to back up our theory. Additionally we present algorithms exploiting local smoothness in more aggressive ways, which perform even better in practice.", "full_text": "LocalSmoothnessinVarianceReducedOptimizationDanielVainsencher,HanLiuTongZhangDept.ofOperationsResearch&FinancialEngineeringDept.ofStatisticsPrincetonUniversityRutgersUniversityPrinceton,NJ08544Piscataway,NJ,08854{daniel.vainsencher,han.liu}@princeton.edutzhang@stat.rutgers.eduAbstractWeproposeafamilyofnon-uniformsamplingstrategiestoprovablyspeedupaclassofstochasticoptimizationalgorithmswithlinearconvergenceincludingStochasticVarianceReducedGradient(SVRG)andStochasticDualCoordinateAscent(SDCA).Foralargefamilyofpenalizedempiricalriskminimizationprob-lems,ourmethodsexploitdatadependentlocalsmoothnessofthelossfunctionsneartheoptimum,whilemaintainingconvergenceguarantees.Ourboundsarethe\ufb01rsttoquantifytheadvantagegainedfromlocalsmoothnesswhicharesigni\ufb01cantforsomeproblemssigni\ufb01cantlybetter.Empirically,weprovidethoroughnumer-icalresultstobackupourtheory.Additionallywepresentalgorithmsexploitinglocalsmoothnessinmoreaggressiveways,whichperformevenbetterinpractice.1IntroductionWeconsiderminimizationoffunctionsofformP(w)=n\u22121nXi=1\u03c6i(cid:0)x\u22a4iw(cid:1)+R(w)wheretheconvex\u03c6icorrespondstoalossofwonsomedataxi,RisaconvexregularizerandPis\u00b5stronglyconvex,sothatP(w\u2032)\u2265P(w)+hw\u2032\u2212w,\u25bdP(w)i+\u00b52kw\u2032\u2212wk2.Inaddition,weassumeeach\u03c6iissmoothingeneralandnear\ufb02atinsomeregion;examplesincludeSVM,regressionwiththeabsoluteerroror\u03b5insensitiveloss,smoothapproximationsofthose,andalsologisticregression.Stochasticoptimizationalgorithmsconsideroneloss\u03c6iatatime,chosenatrandomaccordingtoadistributionptwhichmaychangeovertime.Recentalgorithmscombine\u03c6iwithinformationaboutpreviouslyseenlossestoacceleratetheprocess,achievinglinearconvergencerate,includingStochasticVarianceReducedGradient(SVRG)[2],StochasticAveragedGradient(SAG)[4],andStochasticDualCoordinateAscent(SDCA)[6].TheexpectednumberofiterationsrequiredbythesealgorithmsisofformO(cid:0)(n+L/\u00b5)log(cid:0)\u03b5\u22121(cid:1)(cid:1)whereLisaLipschitzconstantofalllossgradients\u25bd\u03c6i,measuringtheirsmoothness.Dif\ufb01cultproblems,havingaconditionnumberL/\u00b5muchlargerthann,arecalledillconditioned,andhavemotivatedthedevelopmentofacceleratedalgorithms[5,8,3].Someofthesealgorithmshavebeenadaptedtoallowimportancesamplingwhereptisnonuniform;theeffectonconvergenceboundsistoreplacetheuniformboundLdescribedabovebyLavg,theaverageoverLi,lossspeci\ufb01cLipschitzbounds.Inpractice,foranimportantclassofproblems,alargeproportionof\u03c6ineedtobesampledonlyveryfewtimes,andothersinde\ufb01nitely.AsanexamplewetakeaninstanceofsmoothSVM,with\u00b5=n\u22121andL\u224830,solvedviastandardSDCA.InFigure1weobservethedecayofanupperboundontheupdatespossiblefordifferentsamples,wherechoosingasamplethatiswhiteproducesnoupdate.Thelargemajorityofthe\ufb01gureiswhite,indicatingwastedeffort.For95%oflosses,thealgorithmcapturedallrelevantinformationafterjust3visits.Sincethenonwhitezoneisnearlyconstantovertime,detectingandfocusingonthefewimportantlossesshouldbepossible.This1\frepresentsbothasuccessofSDCAandasigni\ufb01cantroomforimprovement,asfocusingjusthalftheeffortontheactivelosseswouldincreaseeffectivenessbyafactorof10.SimilarphenomenaoccurundertheSVRGandSAGalgorithmsaswell.Butisthephenomenonspeci\ufb01ctoasingleproblem,orgeneral?forwhatproblemscanweexpectthesetofusefullossestobesmallandnearconstant?Figure1:SDCAonsmoothedSVM.DualresidualsupperboundtheSDCAupdatesize;whiteindicateszerohencewastedeffort.Thedualresidualsquicklybecomesparse;thesupportisstable.Allowingpttochangeovertime,thephenomenondescribedindeedcanbeexploited;Figure2showssigni\ufb01cantspeedupsobtainedbyourvariantsofSVRGandSDCA.ComparisonsonotherdatasetsaregiveninSection4.Themechanismbywhichspeedupisobtainedisspeci\ufb01ctoeachal-gorithm,buttheunderlyingphenomenonweexploitisthesame:manyproblemsaremuchsmootherlocallythanglobally.Firstconsiderasinglesmoothedhingeloss\u03c6i,asusedinsmoothedSVMwithsmoothingparameter\u03b3.Thenon-smoothnessofthehingelossisspreadin\u03c6ioveranintervaloflength\u03b3,asillustratedinFigure3andgivenby\u03c6i(a)=\uf8f1\uf8f2\uf8f30a>11\u2212a\u2212\u03b3/2a<1\u2212\u03b3(a\u22121)2/(2\u03b3)otherwise.TheLipschitzconstantofdda\u03c6i(a)is\u03b3\u22121,henceitentersintotheglobalestimateofconditionnum-berLavgasLi=kxik/\u03b3;henceapproximatingthehingelossmoreprecisely,withasmaller\u03b3,makestheproblemsstrictlymoreillconditioned.Butoutsidethatintervaloflength\u03b3,\u03c6icanbelocallyapproximatedasaf\ufb01ne,havingaconstantgradient;intoacorrectexpressionoflocalcondi-tioning,sayonintervalBinthe\ufb01gure,itshouldcontributenothing.Sosmaller\u03b3cansometimesmaketheproblem(locally)betterconditioned.AsetIoflosseshavingconstantgradientsoverasubsetofthehypothesisspacecanbesummarizedforpurposesofoptimizationbyasingleaf\ufb01ne05001000150020002500300035004000Effective passes over data10-1310-1210-1110-1010-910-810-710-610-510-410-310-210-1100Duality gap/suboptimalitySVRG solving smoothed hinge loss SVM on MNIST 0/1. Loss gradient is 3.33e+01 Lip. smooth. 6.77e-05 strong convexity.Uniform sampling ([2])Global smoothness sampling ([7])Local SVRG (Alg. 1)Empirical Affinity SVRG (Alg. 4)050100150200250300350Effective passes over data10-1310-1210-1110-1010-910-810-710-610-510-410-310-210-1100Duality gap/suboptimalitySDCA solving smoothed hinge loss SVM on MNIST 0/1. Loss gradient is 3.33e+01 Lip. smooth. 6.77e-05 strong convexity.Uniform sampling ([6])Global smoothness sampling ([10])Affine-SDCA (Alg. 2)Empirical \u00a2 SDCA (Alg. 3)Figure2:OntheleftweseevariantsofSVRGwith\u03b7=1/(8L),ontherightvariantsofSDCA.2\fFigure3:Aloss\u03c6ithatisnear\ufb02at(Hessianvanishes,nearconstantgradient)ona\u201cball\u201dB\u2282R.Bwithradius2rkxikisinducedbythe(Euclidean)ballofhypothesesB(wt,r),thatweproveincludesw\u2217.Thentheloss\u03c6idoesnotcontributetocurvatureintheregionofinterest,andanaf\ufb01nemodelofthesumofsuch\u03c6ionBcanreplacesamplingfromthem.We\ufb01ndrinalgorithmsbycombiningstrongconvexitywithquantitiessuchasdualitygaporgradientnorm.function,sosamplingfromIshouldnotbenecessary.ItsohappensthatSAG,SVRGandSDCAnaturallydosuchmodeling,henceneedonlylightmodi\ufb01cationstorealizesigni\ufb01cantgains.WeprovidethedetailsforSVRGinSection2(theSAGcaseissimilar)andforSDCAinSection3.Otherlosses,whilenowhereaf\ufb01ne,arelocallysmooth:thelogisticregressionlosshasgradientswithlocalLipschitzconstantsthatdecayexponentiallywithdistancefromahyperplanedependentonxi.Forsuchlosseswecannotforgosamplingany\u03c6ipermanently,butwecanstillobtainboundsbene\ufb01ttingfromlocalsmoothnessforanSVRGvariant.Nextwede\ufb01neformallytherelevantgeometricpropertiesoftheoptimizationproblemandrelatethemtoprovableconvergenceimprovementsoverexistinggenericbounds;wegivedetailedboundsinthesequel.ThroughoutB(c,r)isaEuclideanballofradiusraroundc.De\ufb01nition1.WeshalldenoteLi,r=maxw\u2208B(w\u2217,r)(cid:13)(cid:13)\u25bd2\u03c6i(cid:0)x\u22a4i\u00b7(cid:1)(cid:13)(cid:13)2whichisalsotheuniformLipschitzcoef\ufb01cientof\u25bd\u03c6ithatholdatdistanceatmostrfromw\u2217.Remark2.Algorithmswillusesimilarquantitiesnotdependentonknowingw\u2217suchas\u02dcLi,raroundaknown\u02dcw.De\ufb01nition3.Wede\ufb01netheaverageballsmoothnessfunctionS:R\u2192Rofaproblemby:S(r)=nXi=1Li,\u221e/nXi=1Li,r.InTheorem5weseethatAlgorithm1requiresfewerstochasticgradientsamplestoreducelosssub-optimalitybyaconstantfactorthanSVRGwithimportancesamplingaccordingtoglobalsmooth-ness.Onceithascerti\ufb01edthattheoptimumw\u2217iswithinrofthecurrentiteratew0itusesS(2r)timeslessstochasticgradientsteps.Thenextmeasuresimilarlyincreaseswhenmanylossesareaf\ufb01neonaballaroundtheoptimum.De\ufb01nition4.Wede\ufb01netheballaf\ufb01nityfunctionS:R\u2192[0,n]ofaproblemby:A(r)= n\u22121nXi=11{Li,r>0}!\u22121.InTheorem10weseesimilarlythatAlgorithm2requiresfeweraccessesof\u03c6itoreducethedualitygaptoany\u03b5>0thanSDCAwithimportancesamplingaccordingtoglobalsmoothness.Onceithascerti\ufb01edthattheoptimumiswithindistancerofthecurrentprimaliteratew=w(cid:0)\u03b10(cid:1)itaccessesA(2r)timesfewer\u03c6i.Inbothcases,localsmoothnessandaf\ufb01nityenableustofocusaconstantportionofsamplingeffortonthefewerlossesstillchallengingneartheoptimum;whenthesearefew,theratios(andhence3\falgorithmicadvantage)arelarge.Weobtaintheseprovablespeedupsoveralreadyfastalgorithmsbyusingthatlocalsmoothnesswhichwecancertify.FornonsmoothlossessuchasSVMandandabsolutelossregression,wecansimilarlyignoreirrelevantlosses,leadingtosigni\ufb01cantpracticalimprovements;thecurrenttheoryforsuchlossesisinsuf\ufb01cienttoquantifythespeedupsaswedoforsmoothlosses.Weobtainalgorithmsthataresimplerandsometimesmuchfasterbyusingthemorequalitativeobservationthatasiteratestendtoanoptimum,thesetofrelevantlossesisgenerallystableandshrinking.Thenalgorithmscanestimatethesetofrelevantlossesdirectlyfromquantitiesobservedinperformingstochasticiterations,sidesteppingtheloosenessofestimatingr.Therearetwopreviousworksinthisgeneraldirection.The\ufb01rstpaperworkcombiningnon-uniformsamplingandempiricalestimationoflosssmoothnessis[4].TheynoteexcellentempiricalperformanceonavariantofSAG,butwithouttheoryensuringconvergence.Weprovidesimilarlyfast(andboundfree)variantsofSDCA(Section3.2)andSVRG(Section2.2).AdynamicimportancesamplingvariantofSDCAwasreportedin[1]withoutrelationtolocalsmoothness;wediscusstheconnectioninSection3.2LocalsmoothnessandgradientdescentalgorithmsInthissectionwedescribehowSVRG,incontrasttotheclassicalstochasticgradientdescent(SGD),naturallyexposeslocalsmoothnessinlosses.ThenwepresenttwovariantsofSVRGthatrealizethesegains.WebeginbyconsideringasinglelosswhenclosetotheoptimumandforsimplicityassumeR\u22610.AssumeasmallballB=B(w,r)aroundourcurrentestimatewincludesaroundtheoptimumw\u2217,andBiscontainedina\ufb02atregionof\u03c6i,andthisholdsforalargeproportionofthenlosses.SGDanditsdescendentSVRG(withimportancesampling)useupdatesofformwt+1=wt\u2212\u03b7vti/(pin),whereEi\u223cpvti/(pin)=\u25bdF(wt)isanunbiasedestimatorofthefullgradientofthelosstermF(w)=n\u22121Pni=1\u03c6i(cid:0)x\u22a4iw(cid:1).SVRGusesvti=(cid:0)\u25bd\u03c6i(cid:0)x\u22a4iwt(cid:1)\u2212\u25bd\u03c6i(cid:0)x\u22a4i\u02dcw(cid:1)(cid:1)/(pin)+\u25bdF(\u02dcw)where\u02dcwissomereferencepoint,withtheadvantagethatvtihasvariancethatvanishesaswt,\u02dcw\u2192w\u2217.Wepointoutinadditionthatwhen\u02dcw,wt\u2208Band\u25bd\u03c6i(cid:0)x\u22a4i\u00b7(cid:1)isconstantonBtheeffectsofsampling\u03c6icancelsoutandvti=\u25bdF(\u02dcw).Inparticular,wecansetpti=0withnolossofinfor-mation.Moregenerallywhen\u25bd\u03c6i(cid:0)x\u22a4i\u00b7(cid:1)isnearconstantonB(smallLi,r)thedifferencebetweenthesampledvaluesof\u25bd\u03c6iinvtiisverysmallandpticanbesimilarlysmall.Weformalizethisinthenextsection,wherewelocalizeexistingtheorythatappliedimportancesamplingtoadaptSVRGstaticallytolosseswithvariedglobalsmoothness.2.1TheLocalSVRGalgorithmHalvingthesuboptimalityofasolutionusingSVRGhastwoparts:computinganexactgradientatareferencepoint,andperformingmanystochasticgradientdescentsteps.Thesamplingdistribution,stepsizeandnumberofiterationsinthelatteraredeterminedbysmoothnessofthelosses.Algorithm1,Local-SVRG,replacestheglobalboundsongradientchangeLiwithlocalonesLi,r,madevalidbyrestrictingiterationstoasmallballcerti\ufb01edtocontaintheoptimum.Thisallowsustoleveragepreviousalgorithmsandanalysis,maintainingpreviousguaranteesandimprovingonthemwhenS(r)islarge.ForthissectionweassumeP=F;asintheinitialversionofSVRG[2],wemayincorporateasmoothregularizer(thoughinadifferentway,explainedlater).Thisallowsustoapplytheex-istingProx-SVRGalgorithm[7]anditstheory;insteadofusingtheproximaloperatorfor\ufb01xedregularization,weuseittolocalize(byprojections)thestochasticdescenttoaballBaroundtheref-erencepoint\u02dcwseeAlgorithm1.ThenthetheorydevelopedaroundimportancesamplingandglobalsmoothnessappliestosharperlocalsmoothnessestimatesthatholdonB(ignoring\u03c6iwhichareaf\ufb01neonBisaspecialcase).Thisallowsforfewerstochasticiterationsandusingalargerstepsize,obtainingspeedupsthatareproblemdependentbutoftenlargeinlatestages;seeFigure2.Thisisformalizedinthefollowingtheorem.4\fAlgorithm1LocalSVRGisanapplicationofProxSVRGwith\u02dcwdependentregularization.Thisportionreducessuboptimalitybyaconstantfactor,applyiterativelytominimizeloss.1.Compute\u02dcv=\u25bdF(\u02dcw)2.De\ufb01ner=2\u00b5k\u02dcvk,R(w)=iB(\u02dcw,r)=(cid:26)0w\u2208B(\u02dcw,r)\u221eotherwise(by\u00b5strongconvexity,w\u2217\u2208B(\u02dcw,r))3.Foreachi,compute\u02dcLi,r=maxw\u2208B(\u02dcw,r)\u25bd2\u03c6i(cid:0)x\u22a4iw(cid:1)4.De\ufb01neaprobabilitydistribution:pi\u221d\u02dcLi,r,weightedLipschitzconstant\u02dcLp=maxi\u02dcLi,r/(npi)andstepsize\u03b7=116\u02dcLp.5.ApplytheinnerloopofProx-SVRG:(a)Setw0=\u02dcw(b)Fort\u2208{1,...,m}:i.Chooseit\u223cpii.Computevt=(cid:0)\u25bd\u03c6it(cid:0)wt\u22121(cid:1)\u2212\u25bd\u03c6it(\u02dcx)(cid:1)/(npit)+\u02dcviii.wt=prox\u03b7R(cid:0)wt\u22121\u2212\u03b7vt(cid:1)(c)Return\u02c6w=m\u22121Pt\u2208[m]wtTheorem5.Let\u02dcwbeaninitialsolutionsuchthat\u25bdF(\u02dcw)certi\ufb01esthatw\u2217\u2208B=B(\u02dcw,r).Algorithm1\ufb01nds\u02c6wwithEF(\u02c6w)\u2212F(w\u2217)\u2264(F(\u02dcw)\u2212F(w\u2217))/2usingO(d(n+m))time,wherem=128\u00b5n\u22121Pni=1Li,2r+3.Remark6.Inthedif\ufb01cultcasethatisillconditionedevenlocallysothat128n\u22121Pni=1Li,2r\u226bn\u00b5,thetermnisnegligibleandtheratiobetweencomplexitiesofAlgorithm1andanSVRGusingglobalsmoothnessapproachesS(2r).Proof.Intheinitialpassonthedata,compute\u25bdF(\u02dcw),rand\u02dcLi,r\u2264Li,2r.WethenapplyasingleroundofAlgorithmProx-SVRGof[7],withtheregularizerR(x)=\u03c7B(\u02dcw,r)localizingaroundthereferencepoint.ThenwemayapplyTheorem1of[7]withlocal\u02dcLi,rinsteadoftheglobalLirequiredthereforgeneralproximaloperators.Thisallowsustousethecorrespondinglargerstepsize\u03b7=116Lp=116n\u22121Pni=1\u02dcLi,r.Remark7.Theuseofprojections(hencetherestrictiontosmoothregularization)isnecessarybe-causethelocalsmoothnessisrestrictedtoB,andventuringoutsideBwithalargestepsizemaycompromiseconvergenceentirely.WhileexcursionsoutsideBaredif\ufb01culttocontrolintheory,inpracticeskippingtheprojectionentirelydoesnotseemtohurtconvergence.Informally,steppingfarfromBrequiresmovingconsistentlyagainst\u25bdF,whichisanunlikelyevent.Remark8.Thetheoryrequiresmstochasticstepsperexactgradienttoguaranteeanyimprovementatall,butforillconditionedproblemsthisisoftenverypessimistic.Inpractice,the\ufb01rstO(n)stochasticstepsafteranexactgradientprovidemostofthebene\ufb01t.Inthisheuristicscenario,thecomputationalbene\ufb01tofTheorem5isthroughthesamplingdistributionandthelargerstepsize.Enlargingthestepsizewithoutaccompanyingtheoryoftengainsacorrespondingspeeduptoacertainprecisionbuttheriskofnonconvergencematerializesfrequently.While[2]incorporatedasmoothRbyaddingittoeverylossfunction,thiscouldreducethesmooth-ness(increase\u02dcLi,r)inherentinthelosseshencereducingthebene\ufb01tsofourapproach.Weinsteadproposetoaddasinglelossfunctionde\ufb01nedasnR;thatthisisnotofform\u03c6i(cid:0)x\u22a4iw(cid:1)posesnorealdif\ufb01cultybecauseLocal-SVRGdependsonlossesonlythroughtheirgradientsandsmoothness.Themaindif\ufb01cultywiththeapproachofthissectionisthatinearlystagesrislarge,inpartbecause\u00b5isoftenverysmall(\u00b5=n\u2212\u03b1for\u03b1\u2208{0.5,1}arecommonchoices),leadingtoloosebounds5\fon\u02dcLi,r.Insomecasesthespeedupisonlyobtainedwhentheprecisionisalreadysatisfactory;weconsideralessconservativeschemeinthenextsection.2.2TheEmpiricalAf\ufb01nitySVRGalgorithmLocal-SVRGreliesonlocalsmoothnesstocertifythatsome\u2206ti=(cid:13)(cid:13)\u25bd\u03c6i(cid:0)x\u22a4iwt(cid:1)\u2212\u25bd\u03c6i(cid:0)x\u22a4i\u02dcw(cid:1)(cid:13)(cid:13)aresmall.Incontrast,EmpiricalAf\ufb01nitySVRG(Algorithm4)takes\u2206ti>ttobeevidencethatalossisactive;when\u2206ti=0severaltimes,thatisevidenceoflocalaf\ufb01nityoftheloss,henceitcanbesampledlessoften.Thisstrategydeemphasizeslocallyaf\ufb01nelossesevenwhenristoolargetocertifyit,therebyfocusesworkontherelevantlossesmuchearlier.HalfofthetimewesampleproportionaltotheglobalboundsLiwhichkeepsestimatesof\u2206ticurrent,andalsoboundsthevariancewhensome\u2206tiincreasesfromzerotopositive.Abene\ufb01tofusing\u2206tiisthatitisobservedateverysampleofiwithoutadditionalwork.PseudocodefortheslightlylongAlgorithm4isinthesupplementarymaterialforspacereasons.3StochasticDualCoordinateAscent(SDCA)TheSDCAalgorithmsolvesPthroughthedualproblemD(\u03b1)=\u2212n\u22121nXi=1\u03c6\u2217i(\u2212\u03b1i)+R\u2217(w(\u03b1))wherew(\u03b1)=\u25bdR\u2217(cid:0)1\u03bbnPni=1xi\u03b1i(cid:1).Ateachiteration,SDCAchoosesiatrandomaccordingtopt,andupdatesthe\u03b1icorrespondingtotheloss\u03c6itoincreaseD.Thisschemehasbeenusedforparticularlossesbefore,andwasanalyzedin[6]obtaininglinearratesforgeneralsmoothlosses,uniformsamplingandl2regularization,andrecentlygeneralizedin[10]tootherregularizersandgeneralsamplingdistributions.Inparticular,[10]showimprovedboundsandperformancebystati-callyadaptingtotheglobalsmoothnesspropertiesoflosses;usingadistributionpi\u221d1+Li(n\u00b5)\u22121,itsuf\ufb01cestoperformO(cid:16)(cid:16)n+Lavg\u00b5(cid:17)log(cid:16)(cid:16)n+Lavg\u00b5(cid:17)\u03b5\u22121(cid:17)(cid:17)iterationstoobtainanexpecteddu-alitygapofatmost\u03b5.WhileSDCAisverydifferentfromgradientdescentmethods,itsharesthepropertythatwhenthecurrentstateofthealgorithm(intheformof\u03b1i)alreadymatchesthederiva-tiveinformationfor\u03c6i,theupdatedoesnotrequire\u03c6iandcanbeskipped.Aswe\u2019veseeninFigure1,manylossesconverge\u03b1i\u2192\u03b1\u2217iveryquickly;wewillshowthatlocalaf\ufb01nityisasuf\ufb01cientcondition.3.1TheAf\ufb01ne-SDCAalgorithmThealgorithmicapproachforexploitinglocallyaf\ufb01nelossesinSDCAisverydifferentfromthatforgradientdescentstylealgorithms;forsomeaf\ufb01nelosseswecertifyearlythatsome\u03b1iareintheir\ufb01nalform(seeLemma9)andhenceforthignorethem.Thisappliesonlytolocallyaf\ufb01ne(notjustsmooth)losses,butunlikeLocal-SVRG,doesnotrequiremodifyingthealgorithmforexplicitlocalization.Weuseareductiontoobtainimprovedrateswhilereusingthetheoryof[9]fortheremainingpoints.TheseresultsarestatedforsquaredEuclideanregularization,butholdforstronglyconvexRasin[10].Lemma9.Letwt=w(\u03b1t)\u2208B(w\u2217,r),andlet{gi}=Sw\u2208B(wt,r)\u03c6\u2032i(cid:0)x\u22a4iw(cid:1);inotherwords,\u03c6i(cid:0)x\u22a4i\u00b7(cid:1)isaf\ufb01neonB(wt,r)whichincludesw\u2217.Thenwecancomputetheoptimalvalue\u03b1\u2217i=\u2212gi.Proof.AsstatedinSection7of[6],foreachi,wehave\u2212\u03b1\u2217i=\u03c6\u2032i(cid:0)x\u22a4iw\u2217(cid:1).Thenif\u03c6\u2032i(cid:0)x\u22a4iw(cid:1)isaconstantsingletononB(wt,r)containingw\u2217,theninparticularthatis\u2212\u03b1\u2217i.ThelemmaenablesAlgorithm2toignoreagrowingproportionoflosses.Theoverallconvergencethisenablesisgivenbythefollowing.6\fAlgorithm2Af\ufb01ne-SDCA:adaptingtolocallyaf\ufb01ne\u03c6i,withspeedupapproximatelyA(r).1.\u03b10=0\u2208Rn,I0=\u2205.2.For\u03c4\u2208{1,...}:(a)\u02dcw\u03c4=w(cid:0)\u03b1(\u03c4\u22121)m(cid:1);Computer\u03c4=q2(cid:0)P(\u02dcw\u03c4)\u2212D(cid:0)\u03b1(\u03c4\u22121)m(cid:1)(cid:1)/\u00b5(b)ComputeI\u03c4=ni:(cid:12)(cid:12)(cid:12)Sw\u2208B(w\u03c4,r)\u03c6\u2032i(cid:0)x\u22a4i\u02dcw\u03c4(cid:1)(cid:12)(cid:12)(cid:12)=1o(c)Fori\u2208I\u03c4\\I\u03c4\u22121:\u03b1(\u03c4\u22121)ni=\u2212\u03c6\u2032i(cid:0)x\u22a4i\u02dcw\u03c4(cid:1)(d)p\u03c4i\u221d(0i\u2208I\u03c41+Li(n\u00b5)\u22121otherwise,si=(cid:26)0i\u2208I\u03c4s/p\u03c4iotherwise(e)Fort\u2208[(\u03c4\u22121)m+1,\u03c4m]:i.Chooseit\u223cp\u03c4ii.Compute\u2206\u03b1tit=sit\u00b7(cid:0)\u03c6\u2032it(cid:0)x\u22a4itw(\u03b1t)(cid:1)\u2212\u03b1t\u22121it(cid:1)iii.\u03b1tj=(\u03b1t\u22121j+\u2206\u03b1tjj=it\u03b1t\u22121jotherwiseTheorem10.Ifatepoch\u03c4Algorithm2isatdualitygap\u03b5\u03c4,itwillachieveexpecteddualitygap\u03b5inatmost(cid:16)n\u2032+A\u22121(2r)L\u2032avg\u00b5(cid:17)log(cid:16)(cid:16)n\u2032+A\u22121(2r)L\u2032avg\u00b5(cid:17)\u03b5\u03c4\u03b5(cid:17)iterations,wheren\u2032=n\u2212|I\u03c4|andL\u2032avg=n\u2032\u22121Pi\u2208[n]\\I\u03c4Li\u00b5.Remark11.AssumingLi=Lforsimplicity,andrecallingA(2r)\u2264n/n\u2032,we\ufb01ndthenumberofiterationsisreducedbyafactorofatleastA(2r),comparedtousingpi\u221d1+Li(n\u00b5)\u22121.Incontrast,thecostofthesteps2ato2daddedbyAlgorithm2isatmostafactorofO((m+n)/m),whichmaybedriventowardsonebythechoiceofm.Recentwork[1]modi\ufb01edSDCAfordynamicimportancesamplingdependentonthesocalleddualresidual:\u03bai=\u03b1i+\u03c6\u2032i(cid:0)x\u22a4iw(\u03b1)(cid:1)(whereby\u03c6\u2032i(w)werefertothederivativeof\u03c6iatw)whichis0at\u03b1\u2217.Theyexhibitpracticalimprovementinconvergence,especiallyforsmoothSVM,andtheoreticalspeedupswhen\u03baissparse(foranimpracticalversionofthealgorithm),but[1]doesnottelluswhenthispre-conditionholds,northemagnitudeoftheexpectedbene\ufb01tintermsofpropertiesoftheproblem(asopposedtoalgorithmstatesuchas\u03ba).Inthecontextoflocally\ufb02atlossessuchassmoothSVM,weanswerthesequestionsthroughlocalsmoothness:Lemma9shows\u03baitendstozeroforlossesthatarelocallyaf\ufb01neonaballaroundtheoptimum,andthepracticalAlgorithm2realizesthebene\ufb01twhenthiscerti\ufb01cationcomesintoplay,asquanti\ufb01edintermsofA(r).3.2TheEmpirical\u2206SDCAalgorithmAlgorithm2useslocalaf\ufb01nityandasmalldualitygaptocertifytheoptimalityofsome\u03b1i,avoidingcalculating\u2206\u03b1ithatarezerooruseless;naturallyrissmallenoughonlylateintheprocess.Algo-rithm3insteaddedicateshalfofsamplesinproportiontothemagnitudeofrecent\u2206\u03b1i(theotherhalfchosenuniformly).AsFigure2illustrates,thisapproachleadstosigni\ufb01cantspeedupmuchearlierthantheapproachbasedondualitygapcerti\ufb01cationoflocalaf\ufb01nity.WhileweitisnotclearthatwecanproveforAlgorithm3aboundthatstrictlyimprovesonAlgorithm2,itisworthnotingthatexceptfor(probablyrare)updatestoi\u2208I\u03c4,andafactorof2,theempiricalalgorithmshouldquicklydetectalllocallyaf\ufb01nelosseshenceobtainatleastthespeedupofthecertifyingalgorithm.Inaddition,itnaturallyadaptstotheexpectedsmallupdatesoflocallysmoothlosses.Notethat\u2206\u03b1iiscloselyrelatedto(andmightbereplacableby)\u03ba,butthecurrentalgorithmdifferssigni\ufb01cantlyfromthosein[1]inhowthesequantitiesareusedtoguidesampling.7\fAlgorithm3Empirical\u2206SDCA1.\u03b10=0\u2208Rn,Ati=0.2.For\u03c4\u2208{1,...}:(a)p\u03c4=0.5p\u03c4,1+0.5p2wherep\u03c4,1i\u221dA(\u03c4\u22121)miandp2i=n\u22121(b)Fort\u2208[(\u03c4\u22121)m+1,\u03c4m]:i.Chooseit\u223cp\u03c4ii.Compute\u2206\u03b1tit=sit\u00b7(cid:0)\u03c6\u2032it(cid:0)x\u22a4itw(\u03b1t)(cid:1)\u2212\u03b1t\u22121it(cid:1)iii.Atj=(0.5At\u22121j+0.5(cid:12)(cid:12)\u2206\u03b1tj(cid:12)(cid:12)j=itAt\u22121jotherwiseiv.\u03b1tj=(\u03b1t\u22121j+\u2206\u03b1tjj=it\u03b1t\u22121jotherwise4EmpiricalevaluationWeappliedthesamealgorithmswithalmost1thesameparametersto4additionalclassi\ufb01cationdatasetstodemonstratetheimpactofouralgorithmvariantsmorewidely.TheresultsforSDCAareinFigure4,thoseforSVRGinFigure5inSection7inthesupplementarymaterialforlackofspace.0100200300400Effective passes over data10-1310-1210-1110-1010-910-810-710-610-510-410-310-210-1100Duality gap/suboptimalitySDCA solving smoothed hinge loss SVM on Mushroom. Loss gradient is 3.33e+01 Lip. smooth. 1.23e-04 strong convexity.Uniform sampling ([6])Global smoothness sampling ([10])Affine-SDCA (Alg. 2)Empirical \u00a2 SDCA (Alg. 3)020406080Effective passes over data10-1310-1210-1110-1010-910-810-710-610-510-410-310-210-1100Duality gap/suboptimalitySDCA solving smoothed hinge loss SVM on w8a. Loss gradient is 3.33e+01 Lip. smooth. 2.01e-05 strong convexity.Uniform sampling ([6])Global smoothness sampling ([10])Affine-SDCA (Alg. 2)Empirical \u00a2 SDCA (Alg. 3)051015202530Effective passes over data10-1310-1210-1110-1010-910-810-710-610-510-410-310-210-1100Duality gap/suboptimalitySDCA solving smoothed hinge loss SVM on Dorothea. Loss gradient is 3.33e+01 Lip. smooth. 1.25e-03 strong convexity.Uniform sampling ([6])Global smoothness sampling ([10])Affine-SDCA (Alg. 2)Empirical \u00a2 SDCA (Alg. 3)0100200300400500Effective passes over data10-1310-1210-1110-1010-910-810-710-610-510-410-310-210-1100Duality gap/suboptimalitySDCA solving smoothed hinge loss SVM on ijcnn1. Loss gradient is 3.33e+01 Lip. smooth. 5.22e-06 strong convexity.Uniform sampling ([6])Global smoothness sampling ([10])Affine-SDCA (Alg. 2)Empirical \u00a2 SDCA (Alg. 3)Figure4:SDCAvariantresultsonfouradditionaldatasets.Theadvantagesofusinglocalsmoothnessaresigni\ufb01cantontheharderdatasets.References[1]DominikCsiba,ZhengQu,andPeterRicht\u00b4arik.Stochasticdualcoordinateascentwithadaptiveprobabilities.arXivpreprintarXiv:1502.08053,2015.[2]RieJohnsonandTongZhang.Acceleratingstochasticgradientdescentusingpredictivevari-ancereduction.InAdvancesinNeuralInformationProcessingSystems,pages315\u2013323,2013.1Ononeofthenewdatasets,SVRGwitharatioofstep-sizetoLavgmoreaggressivethantheorysuggestsstoppedconverging;hencewechangedallrunstousethepermissible1/8.Nootherparameterswerechangedadaptedtothedataset.8\f[3]QihangLin,ZhaosongLu,andLinXiao.Anacceleratedproximalcoordinategradientmethodanditsapplicationtoregularizedempiricalriskminimization.arXivpreprintarXiv:1407.1296,2014.[4]MarkSchmidt,NicolasLeRoux,andFrancisBach.Minimizing\ufb01nitesumswiththestochasticaveragegradient.arXivpreprintarXiv:1309.2388,2013.[5]ShaiShalev-ShwartzandTongZhang.Acceleratedproximalstochasticdualcoordinateascentforregularizedlossminimization.MathematicalProgramming,pages1\u201341,2013.[6]ShaiShalev-ShwartzandTongZhang.Stochasticdualcoordinateascentmethodsforregular-izedloss.TheJournalofMachineLearningResearch,14(1):567\u2013599,2013.[7]LinXiaoandTongZhang.Aproximalstochasticgradientmethodwithprogressivevariancereduction.SIAMJournalonOptimization,24(4):2057\u20132075,2014.[8]YuchenZhangandLinXiao.Stochasticprimal-dualcoordinatemethodforregularizedempir-icalriskminimization.arXivpreprintarXiv:1409.3257,2014.[9]PeilinZhaoandTongZhang.Stochasticoptimizationwithimportancesampling.arXivpreprintarXiv:1401.2753,2014.[10]PeilinZhaoandTongZhang.Stochasticoptimizationwithimportancesamplingforregularizedlossminimization.ProceedingsofThe32ndInternationalConferenceonMachineLearning,2015.9\f", "award": [], "sourceid": 1297, "authors": [{"given_name": "Daniel", "family_name": "Vainsencher", "institution": "Princeton University"}, {"given_name": "Han", "family_name": "Liu", "institution": "Princeton University"}, {"given_name": "Tong", "family_name": "Zhang", "institution": "Rutgers"}]}