{"title": "Mirror Descent Meets Fixed Share (and feels no regret)", "book": "Advances in Neural Information Processing Systems", "page_first": 980, "page_last": 988, "abstract": "Mirror descent with an entropic regularizer is known to achieve shifting regret bounds that are logarithmic in the dimension. This is done using either a carefully designed projection or by a weight sharing technique. Via a novel unified analysis, we show that these two approaches deliver essentially equivalent bounds on a notion of regret generalizing shifting, adaptive, discounted, and other related regrets. Our analysis also captures and extends the generalized weight sharing technique of Bousquet and Warmuth, and can be refined in several ways, including improvements for small losses and adaptive tuning of parameters.", "full_text": "Mirror Descent Meets Fixed Share\n\n(and feels no regret)\n\nNicol\u00f2 Cesa-Bianchi\n\nUniversit\u00e0 degli Studi di Milano\n\nnicolo.cesa-bianchi@unimi.it\n\nG\u00e1bor Lugosi\n\nICREA & Universitat Pompeu Fabra, Barcelona\n\ngabor.lugosi@upf.edu\n\nPierre Gaillard\n\nEcole Normale Sup\u00e9rieure\u2217, Paris\npierre.gaillard@ens.fr\n\nGilles Stoltz\n\nEcole Normale Sup\u00e9rieure\u2217, Paris &\nHEC Paris, Jouy-en-Josas, France\n\ngilles.stoltz@ens.fr\n\nAbstract\n\nMirror descent with an entropic regularizer is known to achieve shifting regret\nbounds that are logarithmic in the dimension. This is done using either a carefully\ndesigned projection or by a weight sharing technique. Via a novel uni\ufb01ed analysis,\nwe show that these two approaches deliver essentially equivalent bounds on a no-\ntion of regret generalizing shifting, adaptive, discounted, and other related regrets.\nOur analysis also captures and extends the generalized weight sharing technique\nof Bousquet and Warmuth, and can be re\ufb01ned in several ways, including improve-\nments for small losses and adaptive tuning of parameters.\n\n1\n\nIntroduction\n\nOnline convex optimization is a sequential prediction paradigm in which, at each time step, the\nlearner chooses an element from a \ufb01xed convex set S and then is given access to a convex loss\nfunction de\ufb01ned on the same set. The value of the function on the chosen element is the learner\u2019s\nloss. Many problems such as prediction with expert advice, sequential investment, and online re-\ngression/classi\ufb01cation can be viewed as special cases of this general framework. Online learning\nalgorithms are designed to minimize the regret. The standard notion of regret is the difference\nbetween the learner\u2019s cumulative loss and the cumulative loss of the single best element in S. A\nmuch harder criterion to minimize is shifting regret, which is de\ufb01ned as the difference between the\nlearner\u2019s cumulative loss and the cumulative loss of an arbitrary sequence of elements in S. Shifting\nregret bounds are typically expressed in terms of the shift, a notion of regularity measuring the length\nof the trajectory in S described by the comparison sequence (i.e., the sequence of elements against\nwhich the regret is evaluated). In online convex optimization, shifting regret bounds for convex sub-\nsets S \u2286 Rd are obtained for the projected online mirror descent (or follow-the-regularized-leader)\nalgorithm. In this case the shift is typically computed in terms of the p-norm of the difference of\nconsecutive elements in the comparison sequence \u2014see [1, 2] and [3].\nWe focus on the important special case when S is the simplex. In [1] shifting bounds are shown for\nprojected mirror descent with entropic regularizers using a 1-norm to measure the shift.1 When the\ncomparison sequence is restricted to the corners of the simplex (which is the setting of prediction\nwith expert advice), then the shift is naturally de\ufb01ned to be the number of times the trajectory moves\n\n\u2217Ecole Normale Sup\u00e9rieure, Paris \u2013 CNRS \u2013 INRIA, within the project-team CLASSIC\n1Similar 1-norm shifting bounds can also be proven using the analysis of [2]. However, without using\nentropic regularizers it is not clear how to achieve a logarithmic dependence on the dimension, which is one of\nthe advantages of working in the simplex.\n\n1\n\n\fto a different corner. This problem is often called \u201ctracking the best expert\u201d \u2014see, e.g., [4, 5, 1, 6, 7],\nand it is well known that exponential weights with weight sharing, which corresponds to the \ufb01xed-\nshare algorithm of [4], achieves a good shifting bound in this setting. In [6] the authors introduce a\ngeneralization of the \ufb01xed-share algorithm, and prove various shifting bounds for any trajectory in\nthe simplex. However, their bounds are expressed using a quantity that corresponds to a proper shift\nonly for trajectories on the simplex corners.\nIn this paper we offer a uni\ufb01ed analysis of mirror descent, \ufb01xed share, and the generalized \ufb01xed\nshare of [6] for the setting of online convex optimization in the simplex. Our bounds are expressed\nin terms of a notion of shift based on the total variation distance. Our analysis relies on a generalized\nnotion of shifting regret which includes, as special cases, related notions of regret such as adaptive\nregret, discounted regret, and regret with time-selection functions. Perhaps surprisingly, we show\nthat projected mirror descent and \ufb01xed share achieve essentially the same generalized regret bound.\nFinally, we show that widespread techniques in online learning, such as improvements for small\nlosses and adaptive tuning of parameters, are all easily captured by our analysis.\n\n2 Preliminaries\n\n(cid:80)T\n\nFor simplicity, we derive our results in the setting of online linear optimization. As we show in the\nsupplementary material, these results can be easily extended to the more general setting of online\nconvex optimization through a standard linearization step.\nOnline linear optimization may be cast as a repeated game between the forecaster and the environ-\n\nment as follows. We use \u2206d to denote the simplex(cid:8)q \u2208 [0, 1]d : (cid:107)q(cid:107)1 = 1(cid:9).\n1. Forecaster chooses(cid:98)pt = ((cid:98)p1,t, . . . ,(cid:98)pd,t) \u2208 \u2206d\n3. Forecaster suffers loss(cid:98)p\nThe goal of the forecaster is to minimize the accumulated loss, e.g.,(cid:98)LT =(cid:80)T\nt=1(cid:98)p\n\n2. Environment chooses a loss vector (cid:96)t = ((cid:96)1,t, . . . , (cid:96)d,t) \u2208 [0, 1]d\n\nOnline linear optimization in the simplex. For each round t = 1, . . . , T ,\n\n(cid:62)\nt (cid:96)t .\n\n(cid:62)\nt (cid:96)t. In the now\nclassical problem of prediction with expert advice, the goal of the forecaster is to compete with the\nbest \ufb01xed component (often called \u201cexpert\u201d) chosen in hindsight, that is, with mini=1,...,T\nt=1 (cid:96)i,t;\nor even to compete with a richer class of sequences of components. In Section 3 we state more\nspeci\ufb01cally the goals considered in this paper.\nWe start by introducing our main algorithmic tool, described in Figure 1, a share algorithm whose\nformulation generalizes the seemingly unrelated formulations of the algorithms studied in [4, 1, 6]. It\nis parameterized by the \u201cmixing functions\u201d \u03c8t : [0, 1]td \u2192 \u2206d for t (cid:62) 2 that assign probabilities to\npast \u201cpre-weights\u201d as de\ufb01ned below. In all examples discussed in this paper, these mixing functions\nare quite simple, but working with such a general model makes the main ideas more transparent. We\nthen provide a simple lemma that serves as the starting point2 for analyzing different instances of\nthis generalized share algorithm.\nLemma 1. For all t (cid:62) 1 and for all qt \u2208 \u2206d, Algorithm 1 satis\ufb01es\nvi,t+1(cid:98)pi,t\n(cid:98)pj,t e\u2212\u03b7 (cid:96)j,t\nBy de\ufb01nition of vi,t+1, for all i = 1, . . . , d we then have(cid:80)d\n\n\uf8f6\uf8f8 +\nj=1(cid:98)pj,t e\u2212\u03b7 (cid:96)j,t = (cid:98)pi,t e\u2212\u03b7 (cid:96)i,t/vi,t+1,\nt (cid:96)t (cid:54) (cid:96)i,t + (1/\u03b7) ln(vi,t+1/(cid:98)pi,t) + \u03b7/8. The proof is concluded by taking a\n\n(cid:0)(cid:98)pt \u2212 qt\nd(cid:88)\n\nProof. By Hoeffding\u2019s inequality (see, e.g., [3, Section A.1.1]),\n\nd(cid:88)\n\uf8eb\uf8ed d(cid:88)\n\n(cid:98)pj,t (cid:96)j,t (cid:54) \u2212 1\n\n(cid:96)t (cid:54) 1\n\u03b7\n\n(cid:1)(cid:62)\n\n(cid:98)p\n\n(cid:62)\n\n+\n\n.\n\n\u03b7\n8\n\nqi,t ln\n\ni=1\n\nwhich implies\nconvex aggregation with respect to qt.\n\nj=1\n\nln\n\n\u03b7\n\nj=1\n\n\u03b7\n8\n\n.\n\n(1)\n\n2We only deal with linear losses in this paper. However, it is straightforward that for sequences of \u03b7\u2013exp-\n\nconcave loss functions, the additional term \u03b7/8 in the bound is no longer needed.\n\n2\n\n\fParameters: learning rate \u03b7 > 0 and mixing functions \u03c8t for t (cid:62) 2\n\nInitialization:(cid:98)p1 = v1 = (1/d, . . . , 1/d)\n\nFor each round t = 1, . . . , T ,\n\n2. Observe loss (cid:96)t \u2208 [0, 1]d ;\n3. [loss update] For each j = 1, . . . , d de\ufb01ne\n\n1. Predict(cid:98)pt ;\n(cid:98)pj,t e\u2212\u03b7 (cid:96)j,t\n(cid:80)d\ni=1(cid:98)pi,t e\u2212\u03b7 (cid:96)i,t\nVt+1 =(cid:2)vi,s\n(cid:3)\n4. [shared update] De\ufb01ne(cid:98)pt+1 = \u03c8t+1\n\n(cid:0)Vt+1\n\n1(cid:54)i(cid:54)d, 1(cid:54)s(cid:54)t+1\n\nvj,t+1 =\n\n(cid:1).\n\nthe current pre-weights,\n\nand vt+1 = (v1,t+1, . . . , vd,t+1);\nthe d \u00d7 (t + 1) matrix of all past and current pre-weights;\n\nAlgorithm 1: The generalized share algorithm.\n\n3 A generalized shifting regret for the simplex\n\nWe now introduce a generalized notion of shifting regret which uni\ufb01es and generalizes the notions of\ndiscounted regret (see [3, Section 2.11]), adaptive regret (see [8]), and shifting regret (see [2]). For\na \ufb01xed horizon T , a sequence of discount factors \u03b2t,T (cid:62) 0 for t = 1, . . . , T assigns varying weights\nto the instantaneous losses suffered at each round. We compare the total loss of the forecaster with\nthe loss of an arbitrary sequence of vectors q1, . . . , qT in the simplex \u2206d. Our goal is to bound the\nregret\n\nT(cid:88)\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:62)\n\n\u03b2t,T(cid:98)p\n\nt=1\n\nt=1\n\n\u03b2t,T q(cid:62)\nt (cid:96)t\n\nin terms of the \u201cregularity\u201d of the comparison sequence q1, . . . , qT and of the variations of the\ndiscounting weights \u03b2t,T . By setting ut = \u03b2t,T q(cid:62)\n\n+, we can rephrase the above regret as\n\nt \u2208 Rd\n\nu(cid:62)\nt (cid:96)t .\n\n(2)\n\nT(cid:88)\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:62)\n\n(cid:107)ut(cid:107)1(cid:98)p\n\nt=1\n\nt=1\n\nT(cid:88)\n\nt=2\n\nm(uT\n\n1 ) =\n\nIn the literature on tracking the best expert [4, 5, 1, 6], the regularity of the sequence u1, . . . , uT is\nmeasured as the number of times ut (cid:54)= ut+1. We introduce the following regularity measure\n\nDTV(ut, ut\u22121)\n\n+, we de\ufb01ne DTV(x, y) =(cid:80)\n\n(3)\n\n(xi \u2212 yi).\nwhere for x = (x1, . . . , xd), y = (y1, . . . , yd) \u2208 Rd\nNote that when x, y \u2208 \u2206d, we recover the total variation distance DTV(x, y) = 1\n2 (cid:107)x \u2212 y(cid:107)1, while\nfor general x, y \u2208 Rd\n+, the quantity DTV(x, y) is not necessarily symmetric and is always bounded\nby (cid:107)x \u2212 y(cid:107)1. The traditional shifting regret of [4, 5, 1, 6] is obtained from (2) when all ut are such\nthat (cid:107)ut(cid:107)1 = 1.\n\nxi(cid:62)yi\n\n4 Projected update\n\nThe shifting variant of the EG algorithm analyzed in [1] is a special case of the generalized share\nalgorithm in which the function \u03c8t+1 performs a projection of the pre-weights on the convex set\nd = [\u03b1/d, 1]d \u2229 \u2206d. Here \u03b1 \u2208 (0, 1) is a \ufb01xed parameter. We can prove (using techniques similar\n\u2206\u03b1\nto the ones shown in the next section\u2014see the supplementary material) the following bound which\ngeneralizes [1, Theorem 16].\n\n3\n\n\fTheorem 1. For all T (cid:62) 1, for all sequences (cid:96)1, . . . , (cid:96)t \u2208 [0, 1]d of loss vectors, and for all\nu1, . . . , uT \u2208 Rd\n\n+, if Algorithm 1 is run with the above update, then\n\nt (cid:96)t (cid:54) (cid:107)u1(cid:107)1 ln d\nu(cid:62)\n\n\u03b7\n\n+\n\nm(uT\n1 )\n\n\u03b7\n\nln\n\nd\n\u03b1\n\n+\n\n+ \u03b1\n\n(cid:107)ut(cid:107)1 .\n\n(4)\n\n(cid:16) \u03b7\n\n8\n\n(cid:17) T(cid:88)\n\nt=1\n\nT(cid:88)\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:62)\n\n(cid:107)ut(cid:107)1(cid:98)p\n\nt=1\n\nt=1\n\nThis bound can be optimized by a proper tuning of \u03b1 and \u03b7 parameters. We show a similarly tuned\n(and slightly better) bound in Corollary 1.\n\n5 Fixed-share update\n\nNext, we consider a different instance of the generalized share algorithm corresponding to the update\n\n(cid:98)pj,t+1 =\n\nd(cid:88)\n\n(cid:16) \u03b1\n\nd\n\ni=1\n\n(cid:17)\n\n+ (1 \u2212 \u03b1)1i=j\n\nvi,t+1 =\n\n\u03b1\nd\n\n+ (1 \u2212 \u03b1)vj,t+1 ,\n\n0 (cid:54) \u03b1 (cid:54) 1\n\n(5)\n\nDespite seemingly different statements, this update in Algorithm 1 can be seen to lead exactly to the\n\ufb01xed-share algorithm of [4] for prediction with expert advice. We now show that this update delivers\na bound on the regret almost equivalent to (though slightly better than) that achieved by projection\non the subset \u2206\u03b1\nTheorem 2. With the above update, for all T (cid:62) 1, for all sequences (cid:96)1, . . . , (cid:96)T of loss vectors\n(cid:96)t \u2208 [0, 1]d, and for all u1, . . . , uT \u2208 Rd\n+,\n\nd of the simplex.\n\nT(cid:88)\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:62)\n\n(cid:107)ut(cid:107)1 (cid:98)p\n\nt=1\n\nt=1\n\nt (cid:96)t (cid:54) (cid:107)u1(cid:107)1 ln d\nu(cid:62)\n\n\u03b7\n\n+\n\nT(cid:88)\n\n+\n\nt=1\nm(uT\n1 )\n\n\u03b7\n8\n\n\u03b7\n\n(cid:107)ut(cid:107)1\n\nln\n\nd\n\u03b1\n\n+\n\n(cid:80)T\nt=2 (cid:107)ut(cid:107)1 \u2212 m(uT\n1 )\n\n\u03b7\n\nln\n\n1\n\n1 \u2212 \u03b1\n\n.\n\n1 ) and(cid:80)T\n\nNote that if we only consider vectors of the form ut = qt = (0, . . . , 0, 1, 0, . . . , 0) then m(qT\n1 )\ncorresponds to the number of times qt+1 (cid:54)= qt in the sequence qT\n1 . We thus recover [4, Theorem 1]\nand [6, Lemma 6] from the much more general Theorem 2.\nThe \ufb01xed-share forecaster does not need to \u201cknow\u201d anything in advance about the sequence of\nthe norms (cid:107)ut(cid:107) for the bound above to be valid. Of course, in order to minimize the obtained\nupper bound, the tuning parameters \u03b1, \u03b7 need to be optimized and their values will depend on the\nt=1 (cid:107)ut(cid:107)1 for the sequences one wishes to compete against. This\nmaximal values of m(uT\nis illustrated in the following corollary, whose proof is omitted. Therein, h(x) = \u2212x ln x \u2212 (1 \u2212\nx) ln(1 \u2212 x) denotes the binary entropy function for x \u2208 [0, 1]. We recall3 that h(x) (cid:54) x ln(e/x)\nfor x \u2208 [0, 1].\nCorollary 1. Suppose Algorithm 1 is run with the update (5). Let m0 > 0 and U0 > 0. For all T (cid:62)\n1, for all sequences (cid:96)1, . . . , (cid:96)T of loss vectors (cid:96)t \u2208 [0, 1]d, and for all sequences u1, . . . , uT \u2208 Rd\nwith (cid:107)u1(cid:107)1 + m(uT\nT(cid:88)\n(cid:107)ut(cid:107)1 (cid:98)p\n\n1 ) (cid:54) m0 and(cid:80)T\n(cid:118)(cid:117)(cid:117)(cid:116) U0\n\n(cid:118)(cid:117)(cid:117)(cid:116) U0 m0\n\nt (cid:96)t\u2212 T(cid:88)\n\n(cid:32)\nt=1 (cid:107)ut(cid:107)1\n\n(cid:18) e U0\n\n(cid:18) m0\n\nm0 ln d + U0 h\n\n(cid:19)(cid:33)\n\n+\n\n(cid:19)(cid:33)\n\nu(cid:62)\nt (cid:96)t (cid:54)\n\n(cid:54) U0,\n\nln d + ln\n\n(cid:32)\n\n(cid:54)\n\n(cid:62)\n\nt=1\n\nt=1\nwhenever \u03b7 and \u03b1 are optimally chosen in terms of m0 and U0.\nProof of Theorem 2. Applying Lemma 1 with qt = ut/(cid:107)ut(cid:107)1, and multiplying by (cid:107)ut(cid:107)1, we get\nfor all t (cid:62) 1 and ut \u2208 Rd\n\nm0\n\nU0\n\n2\n\n2\n\n+\n\n(cid:107)ut(cid:107)1 (cid:98)p\n\n3As can be seen by noting that ln(cid:0)1/(1 \u2212 x)(cid:1) < x/(1 \u2212 x)\n\ni=1\n\nui,t ln\n\n(cid:62)\nt (cid:96)t \u2212 u(cid:62)\n\nt (cid:96)t (cid:54) 1\n\u03b7\n\nvi,t+1(cid:98)pi,t\n\n+\n\n\u03b7\n8\n\n(cid:107)ut(cid:107)1 .\n\n(6)\n\nd(cid:88)\n\n4\n\n\fd(cid:88)\n\ni=1\n\nWe now examine\n\nui,t ln\n\nvi,t+1(cid:98)pi,t\n\n(cid:18)\n\nd(cid:88)\n\ni=1\n\n=\n\nui,t ln\n\n(cid:19)\n\n+\n\n\u2212 ui,t\u22121 ln\n\n1\nvi,t\n\n(cid:18)\n\nd(cid:88)\n\ni=1\n\n(cid:19)\n\nui,t\u22121 ln\n\n1\nvi,t\n\n\u2212 ui,t ln\n\n1\n\nvi,t+1\n\n1(cid:98)pi,t\n(cid:19)\n\n(cid:88)\n(cid:88)\n\ni : ui,t(cid:62)ui,t\u22121\n\n+\n\n(cid:18)\n(cid:18)\n\n(cid:124)\n\n(ui,t \u2212 ui,t\u22121) ln\n\n+ ui,t\u22121 ln\n\n(ui,t \u2212 ui,t\u22121) ln\n\n+ui,t ln\n\n1(cid:98)pi,t\n\n1\nvi,t\n\n(cid:125)\n\n(cid:123)(cid:122)\n\n.\n\n(7)\n\n(8)\n\n(cid:19)\nvi,t(cid:98)pi,t\n(cid:19)\nvi,t(cid:98)pi,t\n\n.\n\nFor the \ufb01rst term on the right-hand side, we have\n\n(cid:18)\n\nd(cid:88)\n\ni=1\n\nui,t ln\n\n1(cid:98)pi,t\n\n\u2212 ui,t\u22121 ln\n\n1\nvi,t\n\n=\n\ni : ui,t<ui,t\u22121\n\nIn view of the update (5), we have 1/(cid:98)pi,t (cid:54) d/\u03b1 and vi,t/(cid:98)pi,t (cid:54) 1/(1 \u2212 \u03b1). Substituting in (8), we\nd(cid:88)\n\n(cid:19)\n\nget\n\n(cid:54)0\n\n(cid:18)\n1(cid:98)pi,t\n(cid:54) (cid:88)\n\nui,t ln\n\ni=1\n\ni : ui,t(cid:62)ui,t\u22121\n\n\u2212 ui,t\u22121 ln\n\n1\nvi,t\n\n(ui,t \u2212 ui,t\u22121) ln\n\n+\n\nd\n\u03b1\n\n\uf8eb\uf8ed (cid:88)\nui,t \u2212 (cid:88)\n(cid:123)(cid:122)\n\ni : ui,t(cid:62)ui,t\u22121\n\ni: ui,t(cid:62)ui,t\u22121\n\n\uf8eb\uf8ed d(cid:88)\n(cid:124)\n\ni=1\n\n= DTV(ut, ut\u22121) ln\n\nd\n\u03b1\n\n+\n\n\uf8f6\uf8f8 ln\n\nui,t\n\n1\n\n1 \u2212 \u03b1\n\nui,t\u22121 +\n\ni: ui,t<ui,t\u22121\n\n(ui,t \u2212 ui,t\u22121)\n\nln\n\n1\n\n1 \u2212 \u03b1\n\n.\n\n(cid:88)\n\uf8f6\uf8f8\n(cid:125)\n\nThe sum of the second term in (7) telescopes. Substituting the obtained bounds in the \ufb01rst sum of\nthe right-hand side in (7), and summing over t = 2, . . . , T , leads to\n\nT(cid:88)\n\nd(cid:88)\n\nt=2\n\ni=1\n\nui,t ln\n\nvi,t+1(cid:98)pi,t\n\n(cid:54) m(uT\n\n1 ) ln\n\nd\n\u03b1\n\n+\n\n=(cid:107)ut(cid:107)1\u2212DTV(ut,ut\u22121)\n\n(cid:32) T(cid:88)\n\nt=2\n\n(cid:33)\n(cid:107)ut(cid:107)1 \u2212 m(uT\n1 )\nd(cid:88)\n\nln\n\n1\n\n1 \u2212 \u03b1\n\nWe hence get from (6), which we use in particular for t = 1,\n\nT(cid:88)\n\nt=1\n\n(cid:107)ut(cid:107)1(cid:98)p\n\n(cid:62)\nt (cid:96)t \u2212 u(cid:62)\n\nt (cid:96)t (cid:54) 1\n\u03b7\n\nd(cid:88)\n\ni=1\n\nui,1 ln\n\n1(cid:98)pi,1\n\n+\n\n\u03b7\n8\n\nT(cid:88)\n\nt=1\n\n6 Applications\n\n+\n\nm(uT\n1 )\n\n\u03b7\n\nln\n\nd\n\u03b1\n\n+\n\n+\n\nui,1 ln\n\ni=1\n\n1\nvi,2\n\n1\n\n(cid:124)\n\n\u2212 ui,T ln\nvi,T +1\n(cid:54)0\n\n(cid:123)(cid:122)\n\n(cid:125)\n\n.\n\n(cid:107)ut(cid:107)1\n(cid:80)T\nt=2 (cid:107)ut(cid:107)1 m(uT\n1 )\n\n\u03b7\n\nln\n\n1\n\n1 \u2212 \u03b1\n\n.\n\nWe now show how our regret bounds can be specialized to obtain bounds on adaptive and discounted\nregret, and on regret with time-selection functions. We show regret bounds only for the speci\ufb01c\ninstance of the generalized share algorithm using update (5); but the discussion below also holds up\nto minor modi\ufb01cations for the forecaster studied in Theorem 1.\n\nAdaptive regret was introduced by [8] and can be viewed as a variant of discounted regret where\nthe monotonicity assumption is dropped. For \u03c40 \u2208 {1, . . . , T}, the \u03c40-adaptive regret of a forecaster\nis de\ufb01ned by\n\n(cid:41)\n\n(cid:98)p\n(cid:62)\nt (cid:96)t \u2212 min\nq\u2208\u2206d\n\ns(cid:88)\n\nt=r\n\nq(cid:62)(cid:96)t\n\n.\n\n(9)\n\nR\u03c40\u2212adapt\n\nT\n\n=\n\nmax\n\n[r, s] \u2282 [1, T ]\ns + 1 \u2212 r (cid:54) \u03c40\n\n(cid:40) s(cid:88)\n\nt=r\n\n5\n\n\fThe fact that this is a special case of (2) clearly emerges from the proof of Corollary 2 below here.\nAdaptive regret is an alternative way to measure the performance of a forecaster against a changing\nenvironment. It is a straightforward observation that adaptive regret bounds also lead to shifting\nregret bounds (in terms of hard shifts). In this paper we note that these two notions of regret share\nan even tighter connection, as they can be both viewed as instances of the same alma mater notion\nof regret, i.e., the generalized shifting regret introduced in Section 3. The work [8] essentially\nconsidered the case of online convex optimization with exp-concave loss function; in case of general\nconvex functions, they also mentioned that the greedy projection forecaster of [2] enjoys adaptive\nregret guarantees. This is obtained in much the same way as we obtain an adaptive regret bound for\nthe \ufb01xed-share forecaster in the next result.\nCorollary 2. Suppose that Algorithm 1 is run with the shared update (5). Then for all T (cid:62) 1, for\nall sequences (cid:96)1, . . . , (cid:96)T of loss vectors (cid:96)t \u2208 [0, 1]d, and for all \u03c40 \u2208 {1, . . . , T},\n\nR\u03c40\u2212adapt\n\nT\n\n(cid:54)\n\n(cid:115)\n\n(cid:18)\n\n\u03c40\n2\n\n(cid:18) 1\n\n(cid:19)\n\n\u03c40\n\n(cid:19)\n\n(cid:54)\n\n(cid:114) \u03c40\n\n2\n\n\u03c40 h\n\n+ ln d\n\nln(ed\u03c40)\n\nwhenever \u03b7 and \u03b1 are chosen optimally (depending on \u03c40 and T ).\n\nAs mentioned in [8], standard lower bounds on the regret show that the obtained bound is optimal\nup to the logarithmic factors.\nProof. For 1 (cid:54) r (cid:54) s (cid:54) T and q \u2208 \u2206d, the regret in the right-hand side of (9) equals the\n1 de\ufb01ned as ut = q for t = r, . . . , s and\nregret considered in Theorem 2 against the sequence uT\n0 = (0, . . . , 0) for the remaining t. When r (cid:62) 2, this sequence is such that DTV(ur, ur\u22121) =\n1 ) = 1, while (cid:107)u1(cid:107)1 = 0.\nDTV(q, 0) = 1 and DTV(us+1, us) = DTV(0, q) = 0 so that m(uT\nWhen r = 1, we have (cid:107)u1(cid:107)1 = 1 and m(uT\n1 ) + (cid:107)u1(cid:107)1 = 1, that\nis, m0 = 1. Specializing the bound of Theorem 2 with the additional choice U0 = \u03c40 gives the\nresult.\n\n1 ) = 0. In all cases, m(uT\n\nDiscounted regret was introduced in [3, Section 2.11] and is de\ufb01ned by\n\nT(cid:88)\n\nt=1\n\nmax\nq\u2208\u2206d\n\n(cid:0)(cid:98)p\n\n\u03b2t,T\n\n(cid:62)\nt (cid:96)t \u2212 q(cid:62)(cid:96)t\n\n(cid:1) .\n\n(10)\n\nThe discount factors \u03b2t,T measure the relative importance of more recent losses to older losses. For\ninstance, for a given horizon T , the discounts \u03b2t,T may be larger as t is closer to T . On the contrary,\nin a game-theoretic setting, the earlier losses may matter more then the more recent ones (because of\ninterest rates), in which case \u03b2t,T would be smaller as t gets closer to T . We mostly consider below\nmonotonic sequences of discounts (both non-decreasing and non-increasing). Up to a normalization,\nwe assume that all discounts \u03b2t,T are in [0, 1]. As shown in [3], a minimal requirement to get non-\n\ntrivial bounds is that the sum of the discounts satis\ufb01es UT =(cid:80)\n\nt(cid:54)T \u03b2t,T \u2192 \u221e as T \u2192 \u221e.\n\n\u221a\n\nA natural objective is to show that the quantity in (10) is o(UT ), for instance, by bounding it by\nUT . We claim that Corollary 1 does so, at least whenever the sequences\nsomething of the order of\n(\u03b2t,T ) are monotonic for all T . To support this claim, we only need to show that m0 = 1 is a suitable\nvalue to deal with (10). Indeed, for all T (cid:62) 1 and for all q \u2208 \u2206d, the measure of regularity involved\nin the corollary satis\ufb01es\n\nT(cid:88)\n\n(cid:0)\u03b2t,T \u2212 \u03b2t\u22121,T\n\n(cid:1)\n\n+ = max(cid:8)\u03b21,T , \u03b2T,T\n\n(cid:9) (cid:54) 1 ,\n\n(cid:107)\u03b21,T q(cid:107)1 + m(cid:0)(\u03b2t,T q)t(cid:54)T\n\n(cid:1) = \u03b21,T +\n\nt=2\n\nwhere the second equality follows from the monotonicity assumption on the discounts.\nThe values of the discounts for all t and T are usually known in advance. However, the horizon T\nis not. Hence, a calibration issue may arise. The online tuning of the parameters \u03b1 and \u03b7 shown\nUT for all\nin Section 7.3 entails a forecaster that can get discounted regret bounds of the order\nT . The fundamental reason for this is that the discounts only come in the de\ufb01nition of the \ufb01xed-\nshare forecaster via their sums. In contrast, the forecaster discussed in [3, Section 2.11] weighs each\ninstance t directly with \u03b2t,T (i.e., in the very de\ufb01nition of the forecaster) and enjoys therefore no\nregret guarantees for horizons other than T (neither before T nor after T ). Therein, the knowledge\n\n\u221a\n\n6\n\n\f(cid:113)(cid:80)\n\nof the horizon T is so crucial that it cannot be dealt with easily, not even with online calibration of\nthe parameters or with a doubling trick. We insist that for the \ufb01xed-share forecaster, much \ufb02exibility\nis gained as some of the discounts \u03b2t,T can change in a drastic manner for a round T to values\n\u03b2t,T +1 for the next round. However we must admit that the bound of [3, Section 2.11] is smaller\nthan the one obtained above, as it of the order of\nt(cid:54)T \u03b2t,T\nbound. Again, this improvement was made possible because of the knowledge of the time horizon.\nAs for the comparison to the setting of discounted losses of [9], we note that the latter can be cast as\na special case of our setting (since the discounting weights take the special form \u03b2t,T = \u03b3t . . . \u03b3T\u22121\ntherein, for some sequence \u03b3s of positive numbers). In particular, the \ufb01xed-share forecaster can\nsatisfy the bound stated in [9, Theorem 2], for instance, by using the online tuning techniques of\nSection 7.3. A \ufb01nal reference to mention is the setting of time-selection functions of [10, Section 6],\nwhich basically corresponds to knowing in advance the weights (cid:107)ut(cid:107)1 of the comparison sequence\nu1, . . . , uT the forecaster will be evaluated against. We thus generalize their results as well.\n\nt,T , in contrast to our\n\n(cid:113)(cid:80)\n\nt(cid:54)T \u03b22\n\n7 Re\ufb01nements and extensions\n\nWe now show that techniques for re\ufb01ning the standard online analysis can be easily applied to our\nframework. We focus on the following: improvement for small losses, sparse target sequences, and\ndynamic tuning of parameters. Not all of them where within reach of previous analyses.\n\n7.1\n\nImprovement for small losses\n\nThe regret bounds of the \ufb01xed-share forecaster can be signi\ufb01cantly improved when the cumulative\nloss of the best sequence of experts is small. The next result improves on Corollary 1 whenever\nL0 (cid:28) U0. For concreteness, we focus on the \ufb01xed-share update (5).\nCorollary 3. Suppose Algorithm 1 is run with the update (5). Let m0 > 0, U0 > 0, and L0 > 0.\nFor all T (cid:62) 1, for all sequences (cid:96)1, . . . , (cid:96)T of loss vectors (cid:96)t \u2208 [0, 1]d, and for all sequences\nu1, . . . , uT \u2208 Rd\n\n+ with (cid:107)u1(cid:107)1 + m(uT\n\nt=1 (cid:107)ut(cid:107)1\n\nt (cid:96)t (cid:54) L0,\n\nt=1 u(cid:62)\n\n(cid:54) U0, and(cid:80)T\n(cid:19)(cid:33)\n(cid:18) e U0\n\nm0\n\n(cid:18) e U0\n\n(cid:19)\n\nm0\n\nln d + ln\n\n+ ln d + ln\n\nT(cid:88)\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:62)\n\n(cid:107)ut(cid:107)1 (cid:98)p\n\nt=1\n\nt=1\n\nu(cid:62)\nt (cid:96)t (cid:54)\n\n1 ) (cid:54) m0,(cid:80)T\n(cid:118)(cid:117)(cid:117)(cid:116)L0 m0\n(cid:32)\n\nwhenever \u03b7 and \u03b1 are optimally chosen in terms of m0, U0, and L0.\n\nHere again, the parameters \u03b1 and \u03b7 may be tuned online using the techniques shown in Section 7.3.\nThe above re\ufb01nement is obtained by mimicking the analysis of Hedge forecasters for small losses\n(see, e.g., [3, Section 2.4]). In particular, one should substitute Lemma 1 with the following lemma\nin the analysis carried out in Section 5; its proof follows from the mere replacement of Hoeffding\u2019s\ninequality by [3, Lemma A.3], which states that for all \u03b7 \u2208 R and for all random variable X taking\nvalues in [0, 1], one has ln E[e\u2212\u03b7X ] (cid:54) (e\u2212\u03b7 \u2212 1)EX.\nLemma 2. Algorithm 1 satis\ufb01es 1 \u2212 e\u2212\u03b7\n(cid:62)\nt (cid:96)t \u2212 q(cid:62)\n\nfor all qt \u2208 \u2206d.\n\nd(cid:88)\n\n(cid:19)\n\nqi,t ln\n\n(cid:98)p\n\n(cid:18) vi,t(cid:98)pi,t+1\n\nt (cid:96)t (cid:54) 1\n\u03b7\n\ni=1\n\n\u03b7\n\n7.2 Sparse target sequences\n\nThe work [6] introduced forecasters that are able to ef\ufb01ciently compete with the best sequence of\nexperts among all those sequences that only switch a bounded number of times and also take a\nsmall number of different values. Such \u201csparse\u201d sequences of experts appear naturally in many\napplications. In this section we show that their algorithms in fact work very well in comparison with\na much larger class of sequences u1, . . . , uT that are \u201cregular\u201d\u2014that is, m(uT\n1 ), de\ufb01ned in (3) is\nsmall\u2014and \u201csparse\u201d in the sense that the quantity n(uT\ni=1 maxt=1,...,T ui,t is small. Note\nthat when qt \u2208 \u2206d for all t, then two interesting upper bounds can be provided. First, denoting\nthe union of the supports of these convex combinations by S \u2286 {1, . . . , d}, we have n(qT\n1 ) (cid:54) |S|,\nthe cardinality of S. Also, n(qT\ncombinations. Thus, n(uT\n\nt = 1, . . . , T}(cid:12)(cid:12), the cardinality of the pool of convex\n\n1 ) (cid:54) (cid:12)(cid:12){qt,\n\n1 ) generalizes the notion of sparsity of [6].\n\n1 ) =(cid:80)d\n\n7\n\n\fHere we consider a family of shared updates of the form\n\n(cid:98)pj,t = (1 \u2212 \u03b1)vj,t + \u03b1\n\nwj,t\nZt\n\n,\n\n(cid:80)d\n\n0 (cid:54) \u03b1 (cid:54) 1 ,\n\n(11)\n\nwhere the wj,t are nonnegative weights that may depend on past and current pre-weights and Zt =\ni=1 wi,t is a normalization constant. Shared updates of this form were proposed by [6, Sections 3\nand 5.2]. Apart from generalizing the regret bounds of [6], we believe that the analysis given below\nis signi\ufb01cantly simpler and more transparent. We are also able to slightly improve their original\nbounds.\nWe focus on choices of the weights wj,t that satisfy the following conditions: there exists a constant\nC (cid:62) 1 such that for all j = 1, . . . , d and t = 1, . . . , T ,\n\nand\n\nThe next result improves on Theorem 2 when T (cid:28) d and n(uT\ndimension (or number of experts) d is large but the sequence uT\nthe supplementary material; it is a variation on the proof of Theorem 2.\nTheorem 3. Suppose Algorithm 1 is run with the shared update (11) with weights satisfying the\nconditions (12). Then for all T (cid:62) 1, for all sequences (cid:96)1, . . . , (cid:96)T of loss vectors (cid:96)t \u2208 [0, 1]d, and\nfor all sequences u1, . . . , uT \u2208 Rd\n+,\n\n(12)\n1 ), that is, when the\n1 is sparse. Its proof can be found in\n\n1 ) (cid:28) m(uT\n\nC wj,t+1 (cid:62) wj,t .\n\nvj,t (cid:54) wj,t (cid:54) 1\n\nt (cid:96)t (cid:54) n(uT\nu(cid:62)\n1 ) ln d\n\u03b7\n\n+\n\nn(uT\n\n1 ) T ln C\n\n\u03b7\n\n+\n\nm(uT\n1 )\n\n\u03b7\n\nln\n\nmaxt(cid:54)T Zt\n\n\u03b1\n\n+\n\nT(cid:88)\n\u03b7\n(cid:80)T\n8\nt=2 (cid:107)ut(cid:107)1 \u2212 m(uT\n1 )\n\n(cid:107)ut(cid:107)1\n\nt=1\n\n+\n\n\u03b7\n\nln\n\n1\n\n1 \u2212 \u03b1\n\n.\n\nT(cid:88)\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:62)\n\n(cid:107)ut(cid:107)1 (cid:98)p\n\nt=1\n\nt=1\n\nCorollaries 8 and 9 of [6] can now be generalized (and even improved); we do so\u2014in the supple-\nmentary material\u2014by showing two speci\ufb01c instances of the generic update (11) that satisfy (12).\n\n7.3 Online tuning of the parameters\n\nvj,t+1 =(cid:98)p\n\n(cid:44) d(cid:88)\n\ni=1\n\n(cid:98)p\n\nThe forecasters studied above need their parameters \u03b7 and \u03b1 to be tuned according to various quan-\ntities, including the time horizon T . We show here how the trick of [11] of having these parameters\nvary over time can be extended to our setting. For the sake of concreteness we focus on the \ufb01xed-\nshare update, i.e., Algorithm 1 run with the update (5). We respectively replace steps 3 and 4 of its\ndescription by the loss and shared updates\n\n\u03b7t\n\n\u03b7t\u22121\nj,t\n\ne\u2212\u03b7t(cid:96)j,t\n\n\u03b7t\n\n\u03b7t\u22121\ni,t\n\ne\u2212\u03b7t(cid:96)i,t\n\n+ (1 \u2212 \u03b1t) vj,t+1 ,\n\n\u03b1t\nd\n\nand\n\npj,t+1 =\n\n(13)\nfor all t (cid:62) 1 and all j \u2208 {1, . . . , d}, where (\u03b7\u03c4 ) and (\u03b1\u03c4 ) are two sequences of positive numbers,\nindexed by \u03c4 (cid:62) 1. We also conventionally de\ufb01ne \u03b70 = \u03b71. Theorem 2 is then adapted in the\nfollowing way (when \u03b7t \u2261 \u03b7 and \u03b1t \u2261 \u03b1, Theorem 2 is exactly recovered).\nTheorem 4. The forecaster based on the updates (13) is such that whenever \u03b7t (cid:54) \u03b7t\u22121 and \u03b1t (cid:54)\n\u03b1t\u22121 for all t (cid:62) 1, the following performance bound is achieved. For all T (cid:62) 1, for all sequences\n(cid:96)1, . . . , (cid:96)T of loss vectors (cid:96)t \u2208 [0, 1]d, and for all u1, . . . , uT \u2208 Rd\n+,\n\u2212 1\n\u03b7t\u22121\n(cid:107)ut(cid:107)1\n\u03b7t\u22121\n\n(cid:18) 1\nT(cid:88)\n\n(cid:107)ut(cid:107)1\nd(1 \u2212 \u03b1T )\n\nt (cid:96)t \u2212 T(cid:88)\n\n(cid:32)(cid:107)ut(cid:107)1\n\n(cid:107)ut(cid:107)1 (cid:98)p\n\nm(uT\n1 )\n\u03b7T\n\n(cid:19)(cid:33)\n\nu(cid:62)\nt (cid:96)t (cid:54)\n\nT(cid:88)\n\nT(cid:88)\n\nT(cid:88)\n\n1\n\n1 \u2212 \u03b1t\n\n\u03b7t\u22121\n\n+\n\nln\n\nln d\n\n\u03b7t\n\n+\n\nln\n\n+\n\n\u03b1T\n\n\u03b71\n\nt=1\n\nt=2\n\nt=1\n\n+\n\n(cid:62)\n\n(cid:107)ut(cid:107)1 .\n\n8\n\nt=1\n\nt=2\n\nDue to space constraints, we provide an illustration of this bound only in the supplementary material.\n\nAcknowledgments\n\nThe authors acknowledge support from the French National Research Agency (ANR) under grant\nEXPLO/RA (\u201cExploration\u2013exploitation for ef\ufb01cient resource allocation\u201d) and by the PASCAL2\nNetwork of Excellence under EC grant no. 506778.\n\n8\n\n\fReferences\n[1] M. Herbster and M. Warmuth. Tracking the best linear predictor. Journal of Machine Learning\n\nResearch, 1:281\u2013309, 2001.\n\n[2] M. Zinkevich. Online convex programming and generalized in\ufb01nitesimal gradient ascent. In\n\nProceedings of the 20th International Conference on Machine Learning, ICML 2003, 2003.\n\n[3] N. Cesa-Bianchi and G. Lugosi. Prediction, learning, and games. Cambridge University Press,\n\n2006.\n\n[4] M. Herbster and M. Warmuth. Tracking the best expert. Machine Learning, 32:151\u2013178, 1998.\n[5] V. Vovk. Derandomizing stochastic prediction strategies. Machine Learning, 35(3):247\u2013282,\n\nJun. 1999.\n\n[6] O. Bousquet and M.K. Warmuth. Tracking a small set of experts by mixing past posteriors.\n\nJournal of Machine Learning Research, 3:363\u2013396, 2002.\n\n[7] A. Gy\u00f6rgy, T. Linder, and G. Lugosi. Tracking the best of many experts. In Proceedings of the\n18th Annual Conference on Learning Theory (COLT), pages 204\u2013216, Bertinoro, Italy, Jun.\n2005. Springer.\n\n[8] E. Hazan and C. Seshadhri. Ef\ufb01cient learning algorithms for changing environments. Proceed-\n\nings of the 26th International Conference of Machine Learning (ICML), 2009.\n\n[9] A. Chernov and F. Zhdanov. Prediction with expert advice under discounted loss. In Proceed-\nings of the 21st International Conference on Algorithmic Learning Theory, ALT 2010, pages\n255\u2013269. Springer, 2008.\n\n[10] A. Blum and Y. Mansour. From extermal to internal regret. Journal of Machine Learning\n\nResearch, 8:1307\u20131324, 2007.\n\n[11] P. Auer, N. Cesa-Bianchi, and C. Gentile. Adaptive and self-con\ufb01dent on-line learning algo-\n\nrithms. Journal of Computer and System Sciences, 64:48\u201375, 2002.\n\n9\n\n\f", "award": [], "sourceid": 471, "authors": [{"given_name": "Nicol\u00f2", "family_name": "Cesa-bianchi", "institution": null}, {"given_name": "Pierre", "family_name": "Gaillard", "institution": null}, {"given_name": "Gabor", "family_name": "Lugosi", "institution": null}, {"given_name": "Gilles", "family_name": "Stoltz", "institution": null}]}