{"title": "Large Margin Classifiers: Convex Loss, Low Noise, and Convergence Rates", "book": "Advances in Neural Information Processing Systems", "page_first": 1173, "page_last": 1180, "abstract": "", "full_text": "Large margin classi\ufb01ers: convex loss, low noise,\n\nand convergence rates\n\nPeter L. Bartlett, Michael I. Jordan and Jon D. McAuliffe\nDivision of Computer Science and Department of Statistics\n\nUniversity of California, Berkeley\n\nBerkeley, CA 94720\n\nfbartlett,jordan,jong@stat.berkeley.edu\n\nAbstract\n\nMany classi\ufb01cation algorithms, including the support vector machine,\nboosting and logistic regression, can be viewed as minimum contrast\nmethods that minimize a convex surrogate of the 0-1 loss function. We\ncharacterize the statistical consequences of using such a surrogate by pro-\nviding a general quantitative relationship between the risk as assessed us-\ning the 0-1 loss and the risk as assessed using any nonnegative surrogate\nloss function. We show that this relationship gives nontrivial bounds un-\nder the weakest possible condition on the loss function\u2014that it satisfy a\npointwise form of Fisher consistency for classi\ufb01cation. The relationship\nis based on a variational transformation of the loss function that is easy\nto compute in many applications. We also present a re\ufb01ned version of\nthis result in the case of low noise. Finally, we present applications of\nour results to the estimation of convergence rates in the general setting of\nfunction classes that are scaled hulls of a \ufb01nite-dimensional base class.\n\n1 Introduction\n\nConvexity has played an increasingly important role in machine learning in recent years,\nechoing its growing prominence throughout applied mathematics (Boyd and Vandenberghe,\n2003). In particular, a wide variety of two-class classi\ufb01cation methods choose a real-valued\nclassi\ufb01er f based on the minimization of a convex surrogate (cid:30)(yf (x)) in the place of an\nintractable loss function 1(sign(f (x)) 6= y). Examples of this tactic include the support\nvector machine, AdaBoost, and logistic regression, which are based on the exponential\nloss, the hinge loss and the logistic loss, respectively.\nWhat are the statistical consequences of choosing models and estimation procedures so as\nto exploit the computational advantages of convexity? In the setting of 0-1 loss, some basic\nanswers have begun to emerge. In particular, it is possible to demonstrate the Bayes-risk\nconsistency of methods based on minimizing convex surrogates for 0-1 loss, with appropri-\nate regularization. Lugosi and Vayatis (2003) have provided such a result for any differen-\ntiable, monotone, strictly convex loss function (cid:30) that satis\ufb01es (cid:30)(0) = 1. This handles many\ncommon cases although it does not handle the SVM. Steinwart (2002) has demonstrated\nconsistency for the SVM as well, where F is a reproducing kernel Hilbert space and (cid:30) is\ncontinuous. Other results on Bayes-risk consistency have been presented by Jiang (2003),\nZhang (2003), and Mannor et al. (2002).\n\n\fTo carry this agenda further, it is necessary to \ufb01nd general quantitative relationships be-\ntween the approximation and estimation errors associated with (cid:30), and those associated with\n0-1 loss. This point has been emphasized by Zhang (2003), who has presented several ex-\namples of such relationships. We simplify and extend Zhang\u2019s results, developing a general\nmethodology for \ufb01nding quantitative relationships between the risk associated with (cid:30) and\nthe risk associated with 0-1 loss. In particular, let R(f ) denote the risk based on 0-1 loss and\nlet R(cid:3) = inf f R(f ) denote the Bayes risk. Similarly, let us refer to R(cid:30)(f ) = E(cid:30)(Y f (X))\nas the \u201c(cid:30)-risk,\u201d and let R(cid:3)\n(cid:30) = inf f R(cid:30)(f ) denote the \u201coptimal (cid:30)-risk.\u201d We show that, for\nall measurable f,\n\n (R(f ) (cid:0) R(cid:3)) (cid:20) R(cid:30)(f ) (cid:0) R(cid:3)\n(cid:30);\n\n(1)\nfor a nondecreasing function : [0; 1] ! [0;1), and that no better bound is possible.\nMoreover, we present a general variational representation of in terms of (cid:30), and show\nhow this representation allows us to infer various properties of .\nThis result suggests that if is well-behaved then minimization of R(cid:30)(f ) may provide a\nreasonable surrogate for minimization of R(f ). Moreover, the result provides a quantitative\n(cid:30) into\nway to transfer assessments of statistical error in terms of \u201cexcess (cid:30)-risk\u201d R(cid:30)(f )(cid:0) R(cid:3)\nassessments of error in terms of \u201cexcess risk\u201d R(f ) (cid:0) R(cid:3).\nAlthough our principal goal is to understand the implications of convexity in classi\ufb01cation,\nwe do not impose a convexity assumption on (cid:30) at the outset. Indeed, while conditions\nsuch as convexity, continuity, and differentiability of (cid:30) are easy to verify and have natural\nrelationships to optimization procedures, it is not immediately obvious how to relate such\nconditions to their statistical consequences. Thus, in Section 2 we consider the weakest\npossible condition on (cid:30)\u2014that it is \u201cclassi\ufb01cation-calibrated,\u201d which is essentially a point-\nwise form of Fisher consistency for classi\ufb01cation. We show that minimizing (cid:30)-risk leads\nto minimal risk precisely when (cid:30) is classi\ufb01cation-calibrated.\nBuilding on (1), in Section 3 we study the low noise setting, in which the posterior prob-\nability (cid:17)(X) is not too close to 1=2. We show that in this setting we are able to obtain an\nimprovement in the relationship between excess (cid:30)-risk and excess risk.\nSection 4 turns to the estimation of convergence rates for empirical (cid:30)-risk minimization\nin the low noise setting. We \ufb01nd that for convex (cid:30) satisfying a certain uniform convexity\ncondition, empirical (cid:30)-risk minimization yields convergence of misclassi\ufb01cation risk to\nthat of the best-performing classi\ufb01er in F, and the rate of convergence can be strictly faster\nthan the classical parametric rate of n(cid:0)1=2.\n\n2 Relating excess risk to excess (cid:30)-risk\n\nThere are three sources of error to be considered in a statistical analysis of classi\ufb01cation\nproblems: the classical estimation error due to \ufb01nite sample size, the classical approxima-\ntion error due to the size of the function space F, and an additional source of approximation\nerror due to the use of a surrogate in place of the 0-1 loss function. It is this last source of\nerror that is our focus in this section. We give estimates for this error that are valid for any\nmeasurable function. Since the error is de\ufb01ned in terms of the probability distribution, we\nwork with population expectations in this section.\nFix an input space X and let (X; Y ); (X1; Y1); : : : ; (Xn; Yn) 2 X (cid:2) f(cid:6)1g be i.i.d., with\ndistribution P . De\ufb01ne (cid:17) : X ! [0; 1] as (cid:17)(x) = P (Y = 1jX = x).\nDe\ufb01ne the f0; 1g-risk, or just risk, of f as R(f ) = P (sign(f (X)) 6= Y ), where sign((cid:11)) =\n1 for (cid:11) > 0 and (cid:0)1 otherwise. Based on the sample Dn = ((X1; Y1); : : : ; (Xn; Yn)), we\nwant to choose a function fn with small risk. De\ufb01ne the Bayes risk R(cid:3) = inf f R(f ), where\nthe in\ufb01mum is over all measurable f. Then any f satisfying sign(f (X)) = sign((cid:17)(X) (cid:0)\n1=2) a.s. on f(cid:17)(X) 6= 1=2g has R(f ) = R(cid:3).\n\n\fFix a function (cid:30) :  ! [0;1). De\ufb01ne the (cid:30)-risk of f as R(cid:30)(f ) = E(cid:30)(Y f (X)). We can\nview (cid:30) as specifying a contrast function that is minimized in determining a discriminant f.\nDe\ufb01ne C(cid:17)((cid:11)) = (cid:17)(cid:30)((cid:11)) + (1 (cid:0) (cid:17))(cid:30)((cid:0)(cid:11)), so that the conditional (cid:30)-risk at x 2 X is\nE((cid:30)(Y f (X))jX = x) = C(cid:17)(x)(f (x)) = (cid:17)(x)(cid:30)(f (x)) + (1 (cid:0) (cid:17)(x))(cid:30)((cid:0)f (x)):\nAs a useful illustration for the de\ufb01nitions that follow, consider a singleton domain X =\nfx0g. Minimizing (cid:30)-risk corresponds to choosing f (x0) to minimize C(cid:17)(x0)(f (x0)).\nFor (cid:17) 2 [0; 1], de\ufb01ne the optimal conditional (cid:30)-risk\n\nH((cid:17)) = inf\n(cid:11)2\n\nC(cid:17)((cid:11)) = inf\n(cid:11)2\n\n((cid:17)(cid:30)((cid:11)) + (1 (cid:0) (cid:17))(cid:30)((cid:0)(cid:11))):\n\n(cid:30) := inf f R(cid:30)(f ) = EH((cid:17)(X)), where the in\ufb01mum is\n\nThen the optimal (cid:30)-risk satis\ufb01es R(cid:3)\nover measurable functions. For (cid:17) 2 [0; 1], de\ufb01ne\ninf\n\nH (cid:0)((cid:17)) =\n\nC(cid:17)((cid:11)) =\n\ninf\n\n(cid:11):(cid:11)(2(cid:17)(cid:0)1)(cid:20)0\n\n(cid:11):(cid:11)(2(cid:17)(cid:0)1)(cid:20)0\n\n((cid:17)(cid:30)((cid:11)) + (1 (cid:0) (cid:17))(cid:30)((cid:0)(cid:11))):\n\nThis is the optimal value of the conditional (cid:30)-risk, under the constraint that the sign of the\nargument (cid:11) disagrees with that of 2(cid:17) (cid:0) 1.\nWe now turn to the basic condition we impose on (cid:30). This condition generalizes the re-\nquirement that the minimizer of C(cid:17)((cid:11)) (if it exists) has the correct sign. This is a minimal\ncondition that can be viewed as a form of Fisher consistency for classi\ufb01cation (Lin, 2001).\nDe\ufb01nition 1. We say that (cid:30) is classi\ufb01cation-calibrated if, for any (cid:17) 6= 1=2,\n\nH (cid:0)((cid:17)) > H((cid:17)):\n\nThe following functional transform of the loss function will be useful in our main result.\nDe\ufb01nition 2. We de\ufb01ne the -transform of a loss function as follows. Given (cid:30) :  !\n[0;1), de\ufb01ne the function : [0; 1] ! [0;1) by = ~ (cid:3)(cid:3), where\n2 (cid:19) (cid:0) H(cid:18) 1 + (cid:18)\n2 (cid:19) ;\n\n~ ((cid:18)) = H (cid:0)(cid:18) 1 + (cid:18)\n\nis the Fenchel-Legendre biconjugate of g : [0; 1] ! \n\n. Equivalently,\nand g(cid:3)(cid:3) : [0; 1] ! \nthe epigraph of g(cid:3)(cid:3) is the closure of the convex hull of the epigraph of g. (Recall that the\nepigraph of a function g is the set f(x; t) : x 2 [0; 1]; g(x) (cid:20) tg.)\nIt is immediate from the de\ufb01nitions that ~ and are nonnegative and that they are also con-\ntinuous on [0; 1]. We calculate the -transform for exponential loss, logistic loss, quadratic\nloss and truncated quadratic loss, tabulating the results in Table 1. All of these loss func-\ntions can be veri\ufb01ed to be classi\ufb01cation-calibrated. (The other parameters listed in the table\nwill be referred to later.)\nThe importance of the -transform is shown by the following theorem.\n\nexponential\nlogistic\nquadratic\ntruncated quadratic\n\n(cid:30)((cid:11))\n\ne(cid:0)(cid:11)\n\nln(1 + e(cid:0)2(cid:11))\n\n(1 (cid:0) (cid:11))2\n\n(maxf0; 1 (cid:0) (cid:11)g)2\n\n ((cid:18))\n\n1 (cid:0) p1 (cid:0) (cid:18)2\n\n(cid:18)\n(cid:18)2\n(cid:18)2\n\nLB\n\neB\n2\n\n(cid:14)((cid:15))\n\ne(cid:0)B(cid:15)2=8\ne(cid:0)2B(cid:15)2=4\n\n2(B + 1)\n2(B + 1)\n\n(cid:15)2=4\n(cid:15)2=4\n\nTable 1: Four convex loss functions and the corresponding -transform. On the interval\n[(cid:0)B; B], each loss function has the indicated Lipschitz constant LB and modulus of con-\nvexity (cid:14)((cid:15)) with respect to d(cid:30). All have a quadratic modulus of convexity.\n\n\u0001\n\u0001\n\fTheorem 3.\n\n1. For any nonnegative loss function (cid:30), any measurable f : X ! \n\nand any probability distribution on X (cid:2) f(cid:6)1g,\n\n (R(f ) (cid:0) R(cid:3)) (cid:20) R(cid:30)(f ) (cid:0) R(cid:3)\n(cid:30):\n\n2. Suppose jXj (cid:21) 2. For any nonnegative loss function (cid:30), any (cid:15) > 0 and any (cid:18) 2\n[0; 1], there is a probability distribution on X (cid:2) f(cid:6)1g and a function f : X ! \nsuch that R(f ) (cid:0) R(cid:3) = (cid:18) and ((cid:18)) (cid:20) R(cid:30)(f ) (cid:0) R(cid:3)\n\n(cid:30) (cid:20) ((cid:18)) + (cid:15).\n\n3. The following conditions are equivalent.\n\n(a) (cid:30) is classi\ufb01cation-calibrated.\n(b) For any sequence ((cid:18)i) in [0; 1], ((cid:18)i) ! 0 if and only if (cid:18)i ! 0.\n(c) For every sequence of measurable functions fi : X ! \n\nbility distribution on X (cid:2) f(cid:6)1g, R(cid:30)(fi) ! R(cid:3)\n\n(cid:30) implies R(fi) ! R(cid:3).\n\nand every proba-\n\nRemark: It can be shown that classi\ufb01cation-calibration implies is invertible on [0; 1], in\n(cid:30)).\nwhich case it is meaningful to write the upper bound on excess risk as (cid:0)1(R(cid:30)(f ) (cid:0) R(cid:3)\nRemark: Zhang (2003) has given a comparison theorem like Part 1, for convex (cid:30) that\nsatisfy certain conditions. Lugosi and Vayatis (2003) and Steinwart (2002) have shown\nlimiting results like Part 3c under other conditions on (cid:30). All of these conditions are stronger\nthan the ones we assume here.\nThe following lemma summarizes various useful properties of H, H (cid:0) and .\nLemma 4. The functions H, H (cid:0) and have the following properties, for all (cid:17) 2 [0; 1]:\n1. H and H (cid:0) are symmetric about 1=2: H((cid:17)) = H(1 (cid:0) (cid:17)), H (cid:0)((cid:17)) = H (cid:0)(1 (cid:0) (cid:17)).\n2. H is concave and satis\ufb01es H((cid:17)) (cid:20) H(1=2) = H (cid:0)(1=2).\n3. If (cid:30) is classi\ufb01cation-calibrated, then H((cid:17)) < H(1=2) for (cid:17) 6= 1=2.\n4. H (cid:0) is concave on [0; 1=2] and [1=2; 1], and satis\ufb01es H (cid:0)((cid:17)) (cid:21) H((cid:17)).\n5. H, H (cid:0) and ~ are continuous on [0; 1].\n6. is continuous on [0; 1], is nonnegative and minimal at 0, and (0) = 0.\n7. (cid:30) is classi\ufb01cation-calibrated iff ((cid:18)) > 0 for all (cid:18) 2 (0; 1].\nProof. (Of Theorem 3). For Part 1, it is straightforward to show that\n\nR(f ) (cid:0) R(cid:3) = E (1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)]j2(cid:17)(X) (cid:0) 1j) ;\n\nwhere 1 [(cid:8)] is 1 if the predicate (cid:8) is true and 0 otherwise. From the de\ufb01nition, is convex,\nso we can apply Jensen\u2019s inequality, the fact that (0) = 0 (Lemma 4, part 6) and the fact\nthat ((cid:18)) (cid:20) ~ ((cid:18)), to show that\n (R(f ) (cid:0) R(cid:3))\n(cid:20) E (1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)]j2(cid:17)(X) (cid:0) 1j)\n= E (1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)] (j2(cid:17)(X) (cid:0) 1j))\n(cid:20) E(cid:16)1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)] ~ (j2(cid:17)(X) (cid:0) 1j)(cid:17)\n= E(cid:0)1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)](cid:0)H (cid:0)((cid:17)(X)) (cid:0) H((cid:17)(X))(cid:1)(cid:1)\n= E(cid:18)1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)](cid:18)\n(cid:20) E(cid:0)C(cid:17)(X)(f (X)) (cid:0) H((cid:17)(X))(cid:1)\n= R(cid:30)(f ) (cid:0) R(cid:3)\n(cid:30);\n\nC(cid:17)(X)((cid:11)) (cid:0) H((cid:17)(X))(cid:19)(cid:19)\n\n(cid:11):(cid:11)(2(cid:17)(X)(cid:0)1)(cid:20)0\n\ninf\n\n\fwhere the last inequality used the fact that for any x, and in particular when sign(f (x)) =\nsign((cid:17)(x) (cid:0) 1=2), we have C(cid:17)(x)(f (x)) (cid:21) H((cid:17)(x)).\nFor Part 2, the \ufb01rst inequality is from Part 1. For the second, \ufb01x (cid:15) > 0 and (cid:18) 2 [0; 1]. From\nthe de\ufb01nition of , we can choose (cid:13); (cid:11)1; (cid:11)2 2 [0; 1] for which (cid:18) = (cid:13)(cid:11)1 + (1 (cid:0) (cid:13))(cid:11)2 and\n ((cid:18)) (cid:21) (cid:13) ~ ((cid:11)1) + (1 (cid:0) (cid:13)) ~ ((cid:11)2) (cid:0) (cid:15)=2. Choose distinct x1; x2 2 X , and choose PX such\nthat PXfx1g = (cid:13), PXfx2g = 1 (cid:0) (cid:13), (cid:17)(x1) = (1 + (cid:11)1)=2, and (cid:17)(x2) = (1 + (cid:11)2)=2.\nsuch that f (x1) (cid:20) 0, f (x2) (cid:20) 0,\nFrom the de\ufb01nition of H (cid:0), we can choose f : X ! \nC(cid:17)(x1)(f (x1)) (cid:20) H (cid:0)((cid:17)(x1)) + (cid:15)=2 and C(cid:17)(x2)(f (x2)) (cid:20) H (cid:0)((cid:17)(x2)) + (cid:15)=2. Then it is\n(cid:30) (cid:20) (cid:13) ~ ((cid:11)1) + (1(cid:0) (cid:13)) ~ ((cid:11)2) + (cid:15)=2 (cid:20) ((cid:18)) + (cid:15). Furthermore,\neasy to verify that R(cid:30)(f )(cid:0) R(cid:3)\nsince sign(f (xi)) = (cid:0)1 but (cid:17)(xi) (cid:21) 1=2, we have R(f ) (cid:0) R(cid:3) = Ej2(cid:17)(X) (cid:0) 1j = (cid:18).\nFor Part 3, \ufb01rst note that, for any (cid:30), is continuous on [0; 1] and (0) = 0 by Lemma 4,\npart 6, and hence (cid:18)i ! 0 implies ((cid:18)i) ! 0. Thus, we can replace condition (3b) by\n\n(3b\u2019) For any sequence ((cid:18)i) in [0; 1], ((cid:18)i) ! 0 implies (cid:18)i ! 0 .\n\nTo see that (3a) implies (3b\u2019), let (cid:30) be classi\ufb01cation-calibrated, and let ((cid:18)i) be a se-\nquence that does not converge to 0. De\ufb01ne c = lim sup (cid:18)i > 0, and pass to a sub-\nsequence with lim (cid:18)i = c. Then lim ((cid:18)i) = (c) by continuity, and (c) > 0 by\nclassi\ufb01cation-calibration (Lemma 4, part 7). Thus, for the original sequence ((cid:18)i), we see\nlim sup ((cid:18)i) > 0, so we cannot have ((cid:18)i) ! 0.\nPart 1 implies that (3b\u2019) implies (3c). The proof that (3c) implies (3a) is straightforward;\nsee Bartlett et al. (2003).\n\nThe following observation is easy to verify. It shows that if (cid:30) is convex, the classi\ufb01cation-\ncalibration condition is easy to verify and the transform is a little easier to compute.\nLemma 5. Suppose (cid:30) is convex. Then we have\n\n1. (cid:30) is classi\ufb01cation-calibrated if and only if it is differentiable at 0 and (cid:30)0(0) < 0.\n2. If (cid:30) is classi\ufb01cation-calibrated, then ~ is convex, hence = ~ .\n\nAll of the classi\ufb01cation procedures mentioned in earlier sections utilize surrogate loss func-\ntions which are either upper bounds on 0-1 loss or can be transformed into upper bounds\nvia a positive scaling factor. It is easy to verify that this is necessary.\nLemma 6. If (cid:30) :  ! [0;1) is classi\ufb01cation-calibrated, then there is a (cid:13) > 0 such that\n(cid:13)(cid:30)((cid:11)) (cid:21) 1 [(cid:11) (cid:20) 0] for all (cid:11) 2 \n3 Tighter bounds under low noise conditions\n\n.\n\nIn a study of the convergence rate of empirical risk minimization, Tsybakov (2001) pro-\nvided a useful condition on the behavior of the posterior probability near the optimal deci-\nsion boundary fx : (cid:17)(x) = 1=2g. Tsybakov\u2019s condition is useful in our setting as well; as\nwe show in this section, it allows us to obtain a re\ufb01nement of Theorem 3.\nRecall that\n\nR(f ) (cid:0) R(cid:3) = E (1 [sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)]j2(cid:17)(X) (cid:0) 1j)\n\n(cid:20) PX (sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)) ;\n\n(2)\nwith equality provided that (cid:17)(X) is almost surely either 1 or 0. We say that P has noise\nexponent (cid:11) (cid:21) 0 if there is a c > 0 such that every measurable f : X !  has\nPX (sign(f (X)) 6= sign((cid:17)(X) (cid:0) 1=2)) (cid:20) c (R(f ) (cid:0) R(cid:3))(cid:11) :\n\n(3)\nNotice that we must have (cid:11) (cid:20) 1, in view of (2). If (cid:11) = 0, this imposes no constraint on the\nnoise: take c = 1 to see that every probability measure P satis\ufb01es (3). On the other hand,\nit is easy to verify that (cid:11) = 1 if and only if j2(cid:17)(X) (cid:0) 1j (cid:21) 1=c a.s. [PX].\n\n\fTheorem 7. Suppose P has noise exponent 0 < (cid:11) (cid:20) 1, and (cid:30) is classi\ufb01cation-calibrated.\nThen there is a c > 0 such that for any f : X ! \n\n,\n\nc (R(f ) (cid:0) R(cid:3))(cid:11) (R(f ) (cid:0) R(cid:3))1(cid:0)(cid:11)\n\n2c\n\n! (cid:20) R(cid:30)(f ) (cid:0) R(cid:3)\n\n(cid:30):\n\nFurthermore, this never gives a worse rate than the result of Theorem 3, since\n\n(R(f ) (cid:0) R(cid:3))(cid:11) (R(f ) (cid:0) R(cid:3))1(cid:0)(cid:11)\n\n2c\n\n! (cid:21) (cid:18) R(f ) (cid:0) R(cid:3)\n\n2c\n\n(cid:19) :\n\nThe proof follows closely that of Theorem 3(1), with the modi\ufb01cation that we approximate\nthe error integral separately over subsets of the input space with low and high noise.\n\n4 Estimation rates\nLarge margin algorithms choose ^f from a class F to minimize empirical (cid:30)-risk,\n\nn\n\n1\nn\n\n^R(cid:30)(f ) = ^E(cid:30)(Y f (X)) =\n\nXi=1\nWe have seen how the excess risk depends on the excess (cid:30)-risk. In this section, we examine\nthe convergence of ^f\u2019s excess (cid:30)-risk, R(cid:30)( ^f ) (cid:0) R(cid:3)\n(cid:30). We can split this excess risk into an\nestimation error term and an approximation error term:\n\n(cid:30)(Yif (Xi)):\n\nR(cid:30)( ^f ) (cid:0) R(cid:3)\n\n(cid:30) = (R(cid:30)( ^f ) (cid:0) inf\n\nf 2F\n\nR(cid:30)(f )) + ( inf\nf 2F\n\nR(cid:30)(f ) (cid:0) R(cid:3)\n(cid:30)):\n\nWe focus on the \ufb01rst term, the estimation error term. For simplicity, we assume throughout\nthat some f (cid:3) 2 F achieves the in\ufb01mum, R(cid:30)(f (cid:3)) = inf f 2F R(cid:30)(f ).\nThe simplest way to bound R(cid:30)( ^f ) (cid:0) R(cid:30)(f (cid:3)) is to show that ^R(cid:30)(f ) and R(cid:30)(f ) are close,\nuniformly over F. This approach can give the wrong rate. For example, for a nontrivial\nclass F, the resulting estimation error bound can decrease no faster than 1=pn. However, if\nF is a small class (for instance, a VC-class) and R(cid:30)(f (cid:3)) = 0, then R(cid:30)( ^f ) should decrease\nas log n=n. Lee et al. (1996) showed that fast rates are also possible for the quadratic\nloss (cid:30)((cid:11)) = (1 (cid:0) (cid:11))2 if F is convex, even if R(cid:30)(f (cid:3)) > 0. In particular, because the\nquadratic loss function is strictly convex, it is possible to bound the variance of the excess\nloss (difference between the loss of a function f and that of the optimal f (cid:3)) in terms of its\nexpectation. Since the variance decreases as we approach the optimal f (cid:3), the risk of the\nempirical minimizer converges more quickly to the optimal risk than the simple uniform\nconvergence results would suggest. Mendelson (2002) improved this result, and extended\nit from prediction in L2(PX ) to prediction in Lp(PX) for other values of p. The proof\nused the idea of the modulus of convexity of a norm. This idea can be used to give a\nsimpler proof of a more general bound when the loss function satis\ufb01es a strict convexity\ncondition, and we obtain risk bounds. The modulus of convexity of an arbitrary strictly\nconvex function (rather than a norm) is a key notion in formulating our results.\nDe\ufb01nition 8 (Modulus of convexity). Given a pseudometric d de\ufb01ned on a vector space\n, the modulus of convexity of f with respect to d is the\nS, and a convex function f : S ! \nfunction (cid:14) : [0;1) ! [0;1] satisfying\n(cid:14)((cid:15)) = inf(cid:26) f (x1) + f (x2)\n\n(cid:19) : x1; x2 2 S; d(x1; x2) (cid:21) (cid:15)(cid:27) :\n\n(cid:0) f(cid:18) x1 + x2\n\n2\n\n2\n\nIf (cid:14)((cid:15)) > 0 for all (cid:15) > 0, we say that f is strictly convex with respect to d.\n\n\fWe consider loss functions (cid:30) that also satisfy a Lipschitz condition with respect to a pseu-\n: we say that (cid:30) :  ! \ndometric d on \nis Lipschitz with respect to d, with constant L,\n, j(cid:30)(a) (cid:0) (cid:30)(b)j (cid:20) L (cid:1) d(a; b): (Note that if d is a metric and (cid:30) is convex,\nif for all a; b 2 \nthen (cid:30) necessarily satis\ufb01es a Lipschitz condition on any compact subset of \nWe consider four loss functions that satisfy these conditions: the exponential loss function\nused in AdaBoost, the deviance function for logistic regression, the quadratic loss function,\nand the truncated quadratic loss function; see Table 1. We use the pseudometric\n\n.)\n\nd(cid:30)(a; b) = inf fja (cid:0) (cid:11)j + j(cid:12) (cid:0) bj : (cid:30) constant on (minf(cid:11); (cid:12)g; maxf(cid:11); (cid:12)g)g :\nFor all except the truncated quadratic loss function, this corresponds to the standard metric\non \n, d(cid:30)(a; b) = ja(cid:0) bj. In all cases, d(cid:30)(a; b) (cid:20) ja(cid:0) bj, but for the truncated quadratic, d(cid:30)\nignores differences to the right of 1. It is easy to calculate the Lipschitz constant and mod-\nulus of convexity for each of these loss functions. These parameters are given in Table 1.\nIn the following result, we consider the function class used by algorithms such as\nAdaBoost: the class of linear combinations of classi\ufb01ers from a \ufb01xed base class. We as-\nsume that this base class has \ufb01nite Vapnik-Chervonenkis dimension, and we constrain the\nsize of the class by restricting the \u20181 norm of the linear parameters. If G is the VC-class,\nwe write F = B absconv(G), for some constant B, where\n\n; gi 2 G; k(cid:11)k1 = B) :\n\nB absconv(G) =( m\nXi=1\n(cid:11)igi : m 2 \nTheorem 9. Let (cid:30) :  ! \nbe a convex loss function. Suppose that, on the interval\n[(cid:0)B; B], (cid:30) is Lipschitz with constant LB and has modulus of convexity (cid:14)((cid:15)) = aB(cid:15)2 (both\nwith respect to the pseudometric d).\nFor any probability distribution P on X (cid:2) Y that has noise exponent (cid:11) = 1, there is\na constant c0 for which the following is true. For i.i.d. data (X1; Y1); : : : ; (Xn; Yn), let\n^f 2 F be the minimizer of the empirical (cid:30)-risk, R(cid:30)(f ) = ^E(cid:30)(Y f (X)). Suppose that\nF = B absconv(G), where G (cid:18) f(cid:6)1gX has dV C (G) = d, and\n\n; (cid:11)i 2 \n\n(cid:15)(cid:3) (cid:21) BLB max((cid:18) LBaB\n\nB (cid:19)1=(d+1)\n\n; 1) n(cid:0)(d+2)=(2d+2)\n\nThen with probability at least 1 (cid:0) e(cid:0)x,\n\nR( ^f ) (cid:20) R(cid:3) + c0(cid:18)(cid:15)(cid:3) +\n\nLB(LB=aB + B)x\n\nn\n\n+ inf\nf 2F\n\n(cid:30)(cid:19) :\nR(cid:30)(f ) (cid:0) R(cid:3)\n\nNotice that the rate obtained here is strictly faster than the classical n(cid:0)1=2 parametric rate,\neven though the class is in\ufb01nite dimensional and the optimal element of F can have risk\nlarger than the Bayes risk. The key idea in the proof is similar to ideas from Lee et al.\n(1996), Mendelson (2002), but simpler. Let f (cid:3) be the minimizer of (cid:30)-risk in a function\nclass F. If the class F is convex and the loss function (cid:30) is strictly convex and Lipschitz,\nthen the variance of the excess loss, gf (x; y) = (cid:30)(yf (x)) (cid:0) (cid:30)(yf (cid:3)(x)), decreases with\nits expectation. Thus, as a function f 2 F approaches the optimum, f (cid:3), the two losses\n(cid:30)(Y ^f (X)) and (cid:30)(Y f (cid:3)(X)) become strongly correlated. This leads to the faster rates.\nMore formally, suppose that (cid:30) is L-Lipschitz and has modulus of convexity (cid:14)((cid:15)) (cid:21) c(cid:15)r\nf (cid:20) L2 (Egf =(2c))2=r. For the\nwith r (cid:20) 2. Then it is straightforward to show that Eg2\ndetails, see Bartlett et al. (2003).\n\n5 Conclusions\n\nWe have studied the relationship between properties of a nonnegative margin-based loss\nfunction (cid:30) and the statistical performance of the classi\ufb01er which, based on an i.i.d. training\n\n\fset, minimizes empirical (cid:30)-risk over a class of functions. We \ufb01rst derived a universal upper\nbound on the population misclassi\ufb01cation risk of any thresholded measurable classi\ufb01er in\nterms of its corresponding population (cid:30)-risk. The bound is governed by the -transform, a\nconvexi\ufb01ed variational transform of (cid:30). It is the tightest possible upper bound uniform over\nall probability distributions and measurable functions in this setting.\nUsing this upper bound, we characterized the class of loss functions which guarantee that\nevery (cid:30)-risk consistent classi\ufb01er sequence is also Bayes-risk consistent, under any popu-\nlation distribution. Here (cid:30)-risk consistency denotes sequential convergence of population\n(cid:30)-risks to the smallest possible (cid:30)-risk of any measurable classi\ufb01er. The characteristic prop-\nerty of such a (cid:30), which we term classi\ufb01cation-calibration, is a kind of pointwise Fisher con-\nsistency for the conditional (cid:30)-risk at each x 2 X . The necessity of classi\ufb01cation-calibration\nis apparent; the suf\ufb01ciency underscores its fundamental importance in elaborating the sta-\ntistical behavior of large-margin classi\ufb01ers.\nUnder the low noise assumption of Tsybakov (2001), we sharpened our original upper\nbound and studied the Bayes-risk consistency of ^f, the minimizer of empirical (cid:30)-risk over a\nconvex, bounded class of functions F which is not too complex. We found that, for convex\n(cid:30) satisfying a certain uniform strict convexity condition, empirical (cid:30)-risk minimization\nyields convergence of misclassi\ufb01cation risk to that of the best-performing classi\ufb01er in F,\nas the sample size grows. Furthermore, the rate of convergence can be strictly faster than\nthe classical n(cid:0)1=2, depending on the strictness of convexity of (cid:30) and the complexity of F.\nAcknowledgments\n\nWe would like to thank Gilles Blanchard, Olivier Bousquet, Pascal Massart, Ron Meir,\nShahar Mendelson, Martin Wainwright and Bin Yu for helpful discussions.\n\nReferences\nBartlett, P. L., Jordan, M. I., and McAuliffe, J. M. (2003). Convexity, classi\ufb01cation and risk bounds.\n\nTechnical Report 638, Dept. of Statistics, UC Berkeley. [www.stat.berkeley.edu/tech-reports].\n\nBoyd, S. and Vandenberghe, L. (2003). Convex Optimization. [www.stanford.edu/(cid:24)boyd].\nJiang, W. (2003). Process consistency for Adaboost. Annals of Statistics, in press.\nLee, W. S., Bartlett, P. L., and Williamson, R. C. (1996). Ef\ufb01cient agnostic learning of neural net-\n\nworks with bounded fan-in. IEEE Transactions on Information Theory, 42(6):2118\u20132132.\n\nLin, Y. (2001). A note on margin-based loss functions in classi\ufb01cation. Technical Report 1044r,\n\nDepartment of Statistics, University of Wisconsin.\n\nLugosi, G. and Vayatis, N. (2003). On the Bayes risk consistency of regularized boosting methods.\n\nAnnals of Statistics, in press.\n\nMannor, S., Meir, R., and Zhang, T. (2002). The consistency of greedy algorithms for classi\ufb01cation.\n\nIn Proceedings of the Annual Conference on Computational Learning Theory, pages 319\u2013333.\n\nMendelson, S. (2002). Improving the sample complexity using global data. IEEE Transactions on\n\nInformation Theory, 48(7):1977\u20131991.\n\nSteinwart, I. (2002). Consistency of support vector machines and other regularized classi\ufb01ers. Tech-\n\nnical Report 02-03, University of Jena, Department of Mathematics and Computer Science.\n\nTsybakov, A. (2001). Optimal aggregation of classi\ufb01ers in statistical learning. Technical Report\n\nPMA-682, Universit\u00b4e Paris VI.\n\nZhang, T. (2003). Statistical behavior and consistency of classi\ufb01cation methods based on convex risk\n\nminimization. Annals of Statistics, in press.\n\n\f", "award": [], "sourceid": 2416, "authors": [{"given_name": "Peter", "family_name": "Bartlett", "institution": null}, {"given_name": "Michael", "family_name": "Jordan", "institution": null}, {"given_name": "Jon", "family_name": "Mcauliffe", "institution": null}]}