{"title": "Empirical Localization of Homogeneous Divergences on Discrete Sample Spaces", "book": "Advances in Neural Information Processing Systems", "page_first": 820, "page_last": 828, "abstract": "In this paper, we propose a novel parameter estimator for probabilistic models on discrete space. The proposed estimator is derived from minimization of homogeneous divergence and can be constructed without calculation of the normalization constant, which is frequently infeasible for models in the discrete space. We investigate statistical properties of the proposed estimator such as consistency and asymptotic normality, and reveal a relationship with the alpha-divergence. Small experiments show that the proposed estimator attains comparable performance to the MLE with drastically lower computational cost.", "full_text": "Empirical Localization of Homogeneous Divergences\n\non Discrete Sample Spaces\n\nTakashi Takenouchi\n\nDepartment of Complex and Intelligent Systems\n\nFuture University Hakodate\n\n116-2 Kamedanakano, Hakodate, Hokkaido, 040-8655, Japan\n\nttakashi@fun.ac.jp\n\nDepartment of Computer Science and Mathematical Informatics\n\nTakafumi Kanamori\n\nNagoya University\n\nFurocho, Chikusaku, Nagoya 464-8601, Japan\n\nkanamori@is.nagoya-u.ac.jp\n\nAbstract\n\nIn this paper, we propose a novel parameter estimator for probabilistic models on\ndiscrete space. The proposed estimator is derived from minimization of homoge-\nneous divergence and can be constructed without calculation of the normalization\nconstant, which is frequently infeasible for models in the discrete space. We in-\nvestigate statistical properties of the proposed estimator such as consistency and\nasymptotic normality, and reveal a relationship with the information geometry.\nSome experiments show that the proposed estimator attains comparable perfor-\nmance to the maximum likelihood estimator with drastically lower computational\ncost.\n\n1 Introduction\n\nParameter estimation of probabilistic models on discrete space is a popular and important issue in\nthe \ufb01elds of machine learning and pattern recognition. For example, the Boltzmann machine (with\nhidden variables) [1] [2] [3] is a very popular probabilistic model to represent binary variables, and\nattracts increasing attention in the context of Deep learning [4]. A training of the Boltzmann ma-\nchine, i.e., estimation of parameters is usually done by the maximum likelihood estimation (MLE).\nThe MLE for the Boltzmann machine cannot be explicitly solved and the gradient-based optimiza-\ntion is frequently used. A dif\ufb01culty of the gradient-based optimization is that the calculation of the\ngradient requires calculation of a normalization constant or a partition function in each step of the op-\ntimization and its computational cost is sometimes exponential order. The problem of computational\ncost is common to the other probabilistic models on discrete spaces and various kinds of approxi-\nmation methods have been proposed to solve the dif\ufb01culty. One approach tries to approximate the\nprobabilistic model by a tractable model by the mean-\ufb01eld approximation, which considers a model\nassuming independence of variables [5]. Another approach such as the contrastive divergence [6]\navoids the exponential time calculation by the Markov Chain Monte Carlo (MCMC) sampling.\nIn the literature of parameters estimation of probabilistic model for continuous variables, [7] em-\nploys a score function which is a gradient of log-density with respect to the data vector rather than\nparameters. This approach makes it possible to estimate parameters without calculating the normal-\nization term by focusing on the shape of the density function. [8] extended the method to discrete\nvariables, which de\ufb01nes information of \u201cneighbor\u201d by contrasting probability with that of a \ufb02ipped\n\n1\n\n\fvariable. [9] proposed a generalized local scoring rules on discrete sample spaces and [10] proposed\nan approximated estimator with the Bregman divergence.\nIn this paper, we propose a novel parameter estimator for models on discrete space, which does not\nrequire calculation of the normalization constant. The proposed estimator is de\ufb01ned by minimization\nof a risk function derived by an unnormalized model and the homogeneous divergence having a weak\ncoincidence axiom. The derived risk function is convex for various kind of models including higher\norder Boltzmann machine. We investigate statistical properties of the proposed estimator such as the\nconsistency and reveal a relationship between the proposed estimator and the (cid:11)-divergence [11].\n\n2 Settings\nLet X be a d-dimensional vector of random variables in a discrete space X (typically f+1;(cid:0)1gd)\nx2X f (x). Let M and\nand a bracket \u27e8f\u27e9 be summation of a function f (x) on X , i.e., \u27e8f\u27e9 =\nP be a space of all non-negative \ufb01nite measures on X and a subspace consisting of all probability\nmeasures on X , respectively.\n\n\u2211\n\nM = ff (x)j\u27e8f\u27e9 < 1; f (x) (cid:21) 0g ; P = ff (x)j\u27e8f\u27e9 = 1; f (x) (cid:21) 0g :\n\nIn this paper, we focus on parameter estimation of a probabilistic model (cid:22)q(cid:18)(x) on X , written as\n\nq(cid:18)(x)\n\n(cid:22)q(cid:18)(x) =\n\n(1)\nwhere (cid:18) is an m-dimensional vector of parameters, q(cid:18)(x) is an unnormalized model in M and\nZ(cid:18) = \u27e8q(cid:18)\u27e9 is a normalization constant. A computation of the normalization constant Z(cid:18) sometimes\nrequires calculation of exponential order and is sometimes dif\ufb01cult for models on the discrete space.\nNote that the unnormalized model q(cid:18)(x) is not normalized and \u27e8q(cid:18)\u27e9 =\nx2X q(cid:18)(x) = 1 does not\nnecessarily hold. Let (cid:18)(x) be a function on X and throughout the paper, we assume without loss\nof generality that the unnormalized model q(cid:18)(x) can be written as\n\n\u2211\n\nZ(cid:18)\n\nq(cid:18)(x) = exp( (cid:18)(x)):\n\n(2)\nRemark 1. By setting (cid:18)(x) as (cid:18)(x) (cid:0) log Z(cid:18), the normalized model (1) can be written as (2).\nExample 1. The Bernoulli distribution on X = f+1;(cid:0)1g is a simplest example of the probabilistic\nmodel (1) with the function (cid:18)(x) = (cid:18)x.\nExample 2. With a function (cid:18);k(x) = (x1; : : : ; xd; x1x2; : : : ; xd(cid:0)1xd; x1x2x3; : : :)(cid:18), we can de-\n\ufb01ne a k-th order Boltzmann machine [1, 12].\nExample 3. Let xo 2 f+1;(cid:0)1gd1 and xh 2 f+1;(cid:0)1gd2 be an observed vector and hidden vector,\nrespectively, and x =\ncatenated vector. A function h;(cid:18)(xo) for the Boltzmann machine with hidden variables is written\nas\n\n) 2 f+1;(cid:0)1gd1+d2 where T indicates the transpose, be a con-\n\no ; xT\nh\n\n(\n\nxT\n\n\u2211\n\n h;(cid:18)(xo) = log\n\nexp( (cid:18);2(x));\n\nxh\n\n(3)\n\nxh\n\nis the summation with respect to the hidden variable xh.\n\nwhere\nLet us assume that a dataset D = fxign\ni=1 generated by an underlying distribution p(x), is given and\nZ be a set of all patterns which appear in the dataset D. An empirical distribution ~p(x) associated\nwith the dataset D is de\ufb01ned as\n\n{\n\n~p(x) =\n\nnx\nn\n0\n\nx 2 Z;\notherwise;\n\n\u2211\n\n\u2211\n\nn\n\ni=1 I(xi = x) is a number of pattern x appeared in the dataset D.\n\nwhere nx =\nDe\ufb01nition 1. For the unnormalized model (2) and distributions p(x) and ~p(x) in P , probability\nfunctions r(cid:11);(cid:18)(x) and ~r(cid:11);(cid:18)(x) on X are de\ufb01ned by\n\n\u27e8\n\n\u27e9\n\np(x)(cid:11)q(cid:18)(x)1(cid:0)(cid:11)\n\np(cid:11)q1(cid:0)(cid:11)\n\n(cid:18)\n\nr(cid:11);(cid:18)(x) =\n\n\u27e8\n\n~p(x)(cid:11)q(cid:18)(x)1(cid:0)(cid:11)\n\n~p(cid:11)q1(cid:0)(cid:11)\n\n(cid:18)\n\n:\n\n\u27e9\n\n; ~r(cid:11);(cid:18)(x) =\n\n2\n\n\f\u2211\n\nThe distribution r(cid:11);(cid:18) (~r(cid:11);(cid:18)) is an e-mixture model of the unnormalized model (2) and p(x) (~p(x))\nwith ratio (cid:11) [11].\nRemark 2. We observe that r0;(cid:18)(x) = ~r0;(cid:18)(x) = (cid:22)q(cid:18)(x), r1;(cid:18)(x) = p(x), ~r1;(cid:18)(x) = ~p(x). Also if\np(x) = (cid:22)q(cid:18)0 (x), r(cid:11);(cid:18)0 (x) = (cid:22)q(cid:18)0(x) holds for an arbitrary (cid:11).\n\nn\n\nTo estimate the parameter (cid:18) of probabilistic model (cid:22)q(cid:18), the MLE de\ufb01ned by ^(cid:18)mle = argmax(cid:18) L((cid:18)) is\nfrequently employed, where L((cid:18)) =\ni=1 log (cid:22)q(cid:18)(xi) is the log-likelihood of the parameter (cid:18) with\nthe model (cid:22)q(cid:18). Though the MLE is asymptotically consistent and ef\ufb01cient estimator, a main drawback\nof the MLE is that computational cost for probabilistic models on the discrete space sometimes\nbecomes exponential. Unfortunately the MLE does not have an explicit solution in general, the\n\u27e9(cid:0)\nestimation of the parameter can be done by the gradient based optimization with a gradient \u27e8~p \n\u27e8(cid:22)q(cid:18) \n@(cid:18) . While the \ufb01rst term can be easily calculated, the second\nterm includes calculation of the normalization term Z(cid:18), which requires 2d times summation for\nX = f+1;(cid:0)1gd and is not feasible when d is large.\n\n\u27e9 of log-likelihood, where \n\n\u2032\n(cid:18) = @ (cid:18)\n\n\u2032\n(cid:18)\n\n\u2032\n(cid:18)\n\n3 Homogeneous Divergences for Statistical Inference\n\nDivergences are an extension of the squared distance and are often used in statistical inference. A\nformal de\ufb01nition of the divergence D(f; g) is a non-negative valued function on M(cid:2)M or on P(cid:2)P\nsuch that D(f; f ) = 0 holds for arbitrary f. Many popular divergences such as the Kullback-Leilber\n(KL) divergence de\ufb01ned on P (cid:2) P enjoy the coincidence axiom, i.e., D(f; g) = 0 leads to f = g.\nThe parameter in the statistical model (cid:22)q(cid:18) is estimated by minimizing the divergence D(~p; (cid:22)q(cid:18)), with\nrespect to (cid:18).\nIn the statistical inference using unnormalized models, the coincidence axiom of the divergence is\nnot suitable, since the probability and the unnormalized model do not exactly match in general. Our\npurpose is to estimate the underlying distribution up to a constant factor using unnormalized models.\nHence, divergences having the property of the weak coincidence axiom, i.e., D(f; g) = 0 if and only\nif g = cf for some c > 0, are good candidate. As a class of divergences with the weak coincidence\naxiom, we focus on homogeneous divergences that satisfy the equality D(f; g) = D(f; cg) for any\nf; g 2 M and any c > 0.\nA representative of homogeneous divergences is the pseudo-spherical (PS) divergence [13], or in\nother words, (cid:13)-divergence [14], that is de\ufb01ned from the H\u00a8older inequality. Assume that (cid:13) is a\npositive constant. For all non-negative functions f; g in M, the H\u00a8older inequality\n\n\u27e8\ng(cid:13)+1\nholds. The inequality becomes an equality if and only if f and g are linearly dependent. The PS-\ndivergence D(cid:13)(f; g) for f; g 2 M is de\ufb01ned by\n\n\u27e9 (cid:13)\n(cid:13)+1 (cid:0) \u27e8f g(cid:13)\u27e9 (cid:21) 0\n\n\u27e9 1\n\nf (cid:13)+1\n\n\u27e8\n\n(cid:13)+1\n\n\u27e8\n\n\u27e9\n\n\u27e8\n\n\u27e9 (cid:0) log \u27e8f g(cid:13)\u27e9 ;\n\nD(cid:13)(f; g) =\n\n1\n\n1 + (cid:13)\n\nlog\n\nf (cid:13)+1\n\n+\n\n(cid:13)\n\n1 + (cid:13)\n\nlog\n\ng(cid:13)+1\n\n(cid:13) > 0:\n\n(4)\n\nThe PS divergence is homogeneous, and the H\u00a8older inequality ensures the non-negativity and the\nweak coincidence axiom of the PS-divergence. One can con\ufb01rm that the scaled PS-divergence,\n(cid:0)1D(cid:13), converges to the extended KL-divergence de\ufb01ned on M(cid:2)M, as (cid:13) ! 0. The PS-divergence\n(cid:13)\nis used to obtain a robust estimator [14].\nAs shown in (4), the standard PS-divergence from the empirical distribution ~p to the unnormalized\n\u27e9, that may be infeasible in our setup. To circumvent\nmodel q(cid:18) requires the computation of \u27e8q(cid:13)+1\nsuch an expensive computation, we employ a trick and substitute a model ~pq(cid:18) localized by the\nempirical distribution for q(cid:18), which makes it possible to replace the total sum in \u27e8q(cid:13)+1\n\u27e9 with the\nempirical mean. More precisely, let us consider the PS-divergence from f = (p(cid:11)q1(cid:0)(cid:11))\n1+(cid:13) to\n1+(cid:13) for the probability distribution p 2 P and the unnormalized model q 2 M,\nq1(cid:0)(cid:11)\ng = (p(cid:11)\n\u2032 are two distinct real numbers. Then, the divergence vanishes if and only if p(cid:11)q1(cid:0)(cid:11) /\nwhere (cid:11); (cid:11)\n, i.e., q / p. We de\ufb01ne the localized PS-divergence S(cid:11);(cid:11)\u2032;(cid:13)(p; q) by\nq1(cid:0)(cid:11)\np(cid:11)\nq1(cid:0)(cid:11)\nS(cid:11);(cid:11)\u2032;(cid:13)(p; q) = D(cid:13)((p(cid:11)q1(cid:0)(cid:11))1=(1+(cid:13)); (p(cid:11)\n(cid:13)\n\n\u27e8\n\n\u27e9\n\n\u27e9\n\n\u27e8\n\n1\n\n)\n\n(cid:18)\n\n(cid:18)\n\n\u2032\n\n\u2032\n\n\u2032\n\n1\n\n\u2032\n\n\u2032\n\n\u2032\n\n1\n\n\u2032\u27e9 (cid:0) log\n\np(cid:12)q1(cid:0)(cid:12)\n\n;\n\n(5)\n\n)1=(1+(cid:13)))\nlog\u27e8p(cid:11)\nq1(cid:0)(cid:11)\n\n\u2032\n\np(cid:11)q1(cid:0)(cid:11)\n\n+\n\nlog\n\n=\n\n1 + (cid:13)\n\n1 + (cid:13)\n\n3\n\n\f\u27e8\n\n\u27e9\n\n\u2211\n\n(\n\n)(cid:11)\n\n\u2032\n\n)=(1 + (cid:13)). Substituting the empirical distribution ~p into p, the total sum over\nwhere (cid:12) = ((cid:11) + (cid:13)(cid:11)\nX is replaced with a variant of the empirical mean such as\nq1(cid:0)(cid:11)(x)\n\u2032\nfor a non-zero real number (cid:11). Since S(cid:11);(cid:11)\u2032;(cid:13)(p; q) = S(cid:11)\u2032;(cid:11);1=(cid:13)(p; q) holds, we can assume (cid:11) > (cid:11)\nwithout loss of generality. In summary, the conditions of the real parameters (cid:11); (cid:11)\n; (cid:13) are given by\n\n~p(cid:11)q1(cid:0)(cid:11)\n\nx2Z\n\u2032\n\nnx\nn\n\n=\n\n\u2032\n\n; (cid:11) \u0338= 0; (cid:11)\n\n\u2032 \u0338= 0; (cid:11) + (cid:13)(cid:11)\n\n\u2032 \u0338= 0;\n\n(cid:13) > 0; (cid:11) > (cid:11)\nwhere the last condition denotes (cid:12) \u0338= 0.\nLet us consider another aspect of the computational issue about the localized PS-divergence. For the\nprobability distribution p and the unnormalized exponential model q(cid:18), we show that the localized\n\u2032 and (cid:13) are properly chosen.\nPS-divergence S(cid:11);(cid:11)\u2032;(cid:13)(p; q(cid:18)) is convex in (cid:18), when the parameters (cid:11); (cid:11)\nTheorem 1. Let p 2 P be any probability distribution, and let q(cid:18) be the unnormalized exponential\nmodel q(cid:18)(x) = exp((cid:18)T \u03d5(x)), where \u03d5(x) is any vector-valued function corresponding to the suf\ufb01-\ncient statistic in the (normalized) exponential model (cid:22)q(cid:18). For a given (cid:12), the localized PS-divergence\nS(cid:11);(cid:11)\u2032;(cid:13)(p; q(cid:18)) is convex in (cid:18) for any (cid:11); (cid:11)\n)=(1 + (cid:13)) if and only if (cid:12) = 1.\nProof. After some calculation, we have @2 log\u27e8p(cid:11)q1(cid:0)(cid:11)\n= (1 (cid:0) (cid:11))2Vr(cid:11);(cid:18) [\u03d5], where Vr(cid:11);(cid:18) [\u03d5] is the\ncovariance matrix of \u03d5(x) under the probability r(cid:11);(cid:18)(x). Thus, the Hessian matrix of S(cid:11);(cid:11)\u2032;(cid:13)(p; q(cid:18))\nis written as\n@2\n\n; (cid:13) satisfying (cid:12) = ((cid:11) + (cid:13)(cid:11)\n\n(cid:18)\n@(cid:18)@(cid:18)T\n\n)2\n\n\u27e9\n\n\u2032\n\n\u2032\n\n\u2032\n\nVr(cid:11)\u2032;(cid:18) [\u03d5] (cid:0) (1 (cid:0) (cid:12))2Vr(cid:12);(cid:18) [\u03d5]:\n\n@(cid:18)@(cid:18)T S(cid:11);(cid:11)\u2032;(cid:13)(p; q(cid:18)) =\n\n(1 (cid:0) (cid:11))2\n1 + (cid:13)\n\nVr(cid:11);(cid:18) [\u03d5] +\n\n(cid:13)(1 (cid:0) (cid:11)\n1 + (cid:13)\n\nThe Hessian matrix is non-negative de\ufb01nite if (cid:12) = 1. The converse direction is deferred to the\nsupplementary material.\n\nUp to a constant factor, the localized PS-divergence with (cid:12) = 1 characterized by Theorem 1 is\ndenotes as S(cid:11);(cid:11)\u2032(p; q) that is de\ufb01ned by\n1\n(cid:11) (cid:0) 1\n\u2032 \u0338= 0. The parameter (cid:11)\n\n\u2032 can be negative if p is positive on X . Clearly, S(cid:11);(cid:11)\u2032 (p; q)\n\nfor (cid:11) > 1 > (cid:11)\nsatis\ufb01es the homogeneity and the weak coincidence axiom as well as S(cid:11);(cid:11)\u2032;(cid:13)(p; q).\n\n1 (cid:0) (cid:11)\u2032 log\u27e8p(cid:11)\n\nS(cid:11);(cid:11)\u2032 (p; q) =\n\np(cid:11)q1(cid:0)(cid:11)\n\nq1(cid:0)(cid:11)\n\nlog\n\n\u2032\u27e9\n\n+\n\n1\n\n\u2032\n\n\u27e9\n\n\u27e8\n\n4 Estimation with the localized pseudo-spherical divergence\n\nGiven the empirical distribution ~p and the unnormalized model q(cid:18), we de\ufb01ne a novel estimator with\nthe localized PS-divergence S(cid:11);(cid:11)\u2032;(cid:13) (or S(cid:11);(cid:11)\u2032). Though the localized PS-divergence plugged-in\nthe empirical distribution is not well-de\ufb01ned when (cid:11)\n< 0, we can formally de\ufb01ne the following\nestimator by restricting the domain X to the observed set of examples Z, even for negative (cid:11)\n\n\u2032:\n\n\u2032\n\n^(cid:18) = argmin\n\n(cid:18)\n\nS(cid:11);(cid:11)\u2032;(cid:13)(~p; q(cid:18))\n\n= argmin\n\n(cid:18)\n\n1\n\n1 + (cid:13)\n(cid:0) log\n\nlog\n\n\u2211\n\nx2Z\n\n(\n)(cid:12)\n\n\u2211\n(\n\nx2Z\nnx\nn\n\n)(cid:11)\n\nnx\nn\nq(cid:18)(x)1(cid:0)(cid:12):\n\nq(cid:18)(x)1(cid:0)(cid:11) +\n\n(cid:13)\n\n1 + (cid:13)\n\nlog\n\n)(cid:11)\n\n\u2211\n\n(\n\nx2Z\n\nnx\nn\n\n(6)\n\n\u2032\n\nq(cid:18)(x)1(cid:0)(cid:11)\n\n\u2032\n\nRemark 3. The summation in (6) is de\ufb01ned on Z and then is computable even when (cid:11); (cid:11)\nAlso the summation includes only Z((cid:20) n) terms and its computational cost is O(n).\nProposition 1. For the unnormalized model (2), the estimator (6) is Fisher consistent.\n\n\u2032\n\n; (cid:12) < 0.\n\nProof. We observe\n\n@\n@(cid:18)\n\nS(cid:11);(cid:11)\u2032;(cid:13)((cid:22)q(cid:18)0 ; q(cid:18))\n\nimplying the Fisher consistency of ^(cid:18).\n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\n=\n\n(cid:18)=(cid:18)0\n\n(\n(cid:12) (cid:0) (cid:11) + (cid:13)(cid:11)\n1 + (cid:13)\n\n\u2032\n\n)\u27e8\n\n\u27e9\n\n\u2032\n(cid:18)0\n\n(cid:22)q(cid:18)0 \n\n= 0\n\n4\n\n\f(cid:12)(cid:12)(cid:12)(cid:12)\n\n(cid:12)(cid:12)(cid:12)(cid:12)\n(cid:12)(cid:12)(cid:12)(cid:12)\n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\nTheorem 2. Let q(cid:18)(x) be the unnormalized model (2), and (cid:18)0 be the true parameter of underlying\ndistribution p(x) = (cid:22)q(cid:18)0(x). Then an asymptotic distribution of the estimator (6) is written as\n\np\nn( ^(cid:18) (cid:0) (cid:18)0) (cid:24) N (0; I((cid:18)0)\n\n(cid:0)1)\n\nwhere I((cid:18)0) = V(cid:22)q(cid:18)0\n\n[ \n\n\u2032\n(cid:18)0\n\n] is the Fisher information matrix.\n\nProof. We shall sketch a proof and the detailed proof is given in supplementary material. Let us\nassume that the empirical distribution is written as\nNote that \u27e8\u03f5\u27e9 = 0 because ~p; (cid:22)q(cid:18)0\nthe estimator (6) around (cid:18) = (cid:18)0 leads to\n\n2 P. The asymptotic expansion of the equilibrium condition for\n\n~p(x) = (cid:22)q(cid:18)0 (x) + \u03f5(x):\n\n0 =\n\n@\n@(cid:18)\n\n=\n\n@\n@(cid:18)\n\nS(cid:11);(cid:11)\u2032;(cid:13)(~p; q(cid:18))\n\n(cid:18)= ^(cid:18)\n\nS(cid:11);(cid:11)\u2032;(cid:13)(~p; q(cid:18))\n\n(cid:18)=(cid:18)0\nBy the delta method [15], we have\n\n@2\n\n+\n\n@(cid:18)@(cid:18)T S(cid:11);(cid:11)\u2032;(cid:13)(~p; q(cid:18))\n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\n(cid:18)=(cid:18)0\n\n( ^(cid:18) (cid:0) (cid:18)0) + O(jj ^(cid:18) (cid:0) (cid:18)0jj2)\n\u27e9\n\n\u27e8\n\n\u2243 (cid:0)\n\n(cid:13)\n\n(1 + (cid:13))2 ((cid:11) (cid:0) (cid:11)\n\n\u2032\n\n)2\n\n\u2032\n(cid:18)0\n\n\u03f5\n\n \n\n@\n@(cid:18)\n\nS(cid:11);(cid:11)\u2032;(cid:13)(~p; q(cid:18))\n\nS(cid:11);(cid:11)\u2032;(cid:13)(p; q(cid:18))\n\nand from the central limit theorem, we observe that\n\n(cid:0) @\n@(cid:18)\n\n\u27e9\n\n\u2032\n(cid:18)0\n\n\u03f5\n\n \n\n=\n\n(cid:18)=(cid:18)0\n\n\u27e8\n\np\nn\n\nn\u2211\n\n(\n\n(cid:18)=(cid:18)0\n\n(xi) (cid:0)\u27e8\n\n\u2032\n(cid:18)0\n\n \n\ni=1\n\np\nn\n\n1\nn\n\n\u27e9)\n\n\u2032\n(cid:18)0\n\n(cid:22)q(cid:18)0 \n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\nasymptotically follows the normal distribution with mean 0, and variance I((cid:18)0) = V(cid:22)q(cid:18)0\n[ \nis known as the Fisher information matrix. Also from the law of large numbers, we have\n\n@2\n\n@(cid:18)@(cid:18)T S(cid:11);(cid:11)\u2032;(cid:13)(~p; q(cid:18))\n\n( ^(cid:18) (cid:0) (cid:18)0) ! (cid:13)\n\n(1 + (cid:13))2 ((cid:11) (cid:0) (cid:11)\n\n\u2032\n\n)2I((cid:18)0);\n\n(cid:18)=(cid:18)0\n\nin the limit of n ! 1. Consequently, we observe that (2).\nRemark 4. The asymptotic distribution of (6) is equal to that of the MLE, and its variance does not\ndepend on (cid:11); (cid:11)\nRemark 5. As shown in Remark 1, the normalized model (1) is a special case of the unnormalized\nmodel (2) and then Theorem 2 holds for the normalized model.\n\n; (cid:13).\n\n\u2032\n\n\u2032\n(cid:18)0\n\n], which\n\n5 Characterization of localized pseudo-spherical divergence S(cid:11);(cid:11)\u2032\n\nThroughout this section, we assume that (cid:12) = 1 holds and investigate properties of the localized PS-\n\u2032 and characterization of the localized\ndivergence S(cid:11);(cid:11)\u2032. We discuss in\ufb02uence of selection of (cid:11); (cid:11)\nPS-divergence S(cid:11);(cid:11)\u2032 in the following subsections.\n\n5.1\n\nIn\ufb02uence of selection of (cid:11); (cid:11)\n\n\u2032\n\nWe investigate in\ufb02uence of selection of (cid:11); (cid:11)\nthe estimating equation. The estimator ^(cid:18) derived from S(cid:11);(cid:11)\u2032 satis\ufb01es\n\u2032\n^(cid:18)\n\n@S(cid:11);(cid:11)\u2032(~p; q(cid:18))\n\n~r(cid:11)\u2032; ^(cid:18) \n\n~r(cid:11); ^(cid:18) \n\n(cid:0)\n\n\u2032\n^(cid:18)\n\n\u27e9\n\n\u27e8\n\n@(cid:18)\n\n\u2032 for the localized PS-divergence S(cid:11);(cid:11)\u2032 with a view of\n\nwhich is a moment matching with respect to two distributions ~r(cid:11);(cid:18) and ~r(cid:11)\u2032;(cid:18) ((cid:11); (cid:11)\nother hand, the estimating equation of the MLE is written as\n\n/\n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\u27e8\n\u27e9 (cid:0) \u27e8(cid:22)q(cid:18)mle (cid:18)mle\n\n(cid:18)= ^(cid:18)\n\n\u27e8\n\n(7)\n\u2032 \u0338= 0; 1). On the\n\u27e9\n\n= 0;\n\n(8)\n\n\u2032\n(cid:18)mle\n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\n/\u27e8\n\n@L((cid:18))\n\n@(cid:18)\n\n(cid:18)=(cid:18)mle\n\n\u2032\n(cid:18)mle\n\n~p \n\nwhich is a moment matching with respect to the empirical distribution ~p = ~r1;(cid:18)mle and the\nnormalized model (cid:22)q(cid:18) = ~r0;(cid:18)mle. While the localized PS-divergence S(cid:11);(cid:11)\u2032 is not de\ufb01ned with\n) = (0; 1), comparison of (7) with (8) implies that behavior the estimator ^(cid:18) becomes simi-\n((cid:11); (cid:11)\nlar to that of the MLE in the limit of (cid:11) ! 1 and (cid:11)\n\n\u2032 ! 0.\n\n\u2032\n\n= 0:\n\n\u27e9\n\u27e9 (cid:0)\u27e8\n\n\u27e9 =\n\n~r1;(cid:18)mle \n\n\u2032\n(cid:18)mle\n\n~r0;(cid:18)mle \n\n5\n\n\f5.2 Relationship with the (cid:11)-divergence\nThe (cid:11)-divergence between two positive measures f; g 2 M is de\ufb01ned as\n(cid:11)f + (1 (cid:0) (cid:11))g (cid:0) f (cid:11)g1(cid:0)(cid:11)\n\nD(cid:11)(f; g) =\n\n\u27e8\n\n1\n\n\u27e9\n\n;\n\n(cid:11)(1 (cid:0) (cid:11))\n\nwhere (cid:11) is a real number. Note that D(cid:11)(f; g) (cid:21) 0 and 0 if and only if f = g, and the (cid:11)-divergence\nreduces to KL(f; g) and KL(g; f ) in the limit of (cid:11) ! 1 and 0, respectively.\nRemark 6. An estimator de\ufb01ned by minimizing (cid:11)-divergence D(cid:11)(~p; (cid:22)q(cid:18)) between the empirical\ndistribution and normalized model, satis\ufb01es\n\nand requires calculation proportional to jXj which is infeasible. Also the same hold for an estimator\nde\ufb01ned by minimizing (cid:11)-divergence D(cid:11)(~p; q(cid:18)) between the empirical distribution and unnormalized\nmodel, satisfying @D(cid:11)( ~p;q(cid:18))\n\u2032\n\n\u0338= 0; 1 and consider a trick to cancel out the term \u27e8g\u27e9 by mixing two\n\n(cid:0) ~p(cid:11)q1(cid:0)(cid:11)\n\n= 0.\n\n\u2032\n(cid:18)\n\n@(cid:18)\n\n(cid:18)\n\nHere, we assume that (cid:11); (cid:11)\n(cid:11)-divergences as follows.\n\n\u27e9\n\u27e9)\n\n= 0\n\n\u2032\n(cid:18)\n\n(cid:0) \u27e8(cid:22)q(cid:18) \n\u27e9\n\n@D(cid:11)(~p; (cid:22)q(cid:18))\n\n~p(cid:11)q1(cid:0)(cid:11)\n\n(cid:18)\n\n\u2032\n(cid:18)\n\n( \n\n/\u27e8\n\n@(cid:18)\n\n/\u27e8\n((cid:0)(cid:11)\n\n(1 (cid:0) (cid:11))q(cid:18) \n)\n\n\u2032\n\n)\n\n\u27e9\n\n\u27e8(\n\nD(cid:11);(cid:11)\u2032 (f; g) =D(cid:11)(f; g) +\n(cid:0)\n\n=\n\n1\n1 (cid:0) (cid:11)\n\nD(cid:11)\u2032(f; g)\nf (cid:0)\n\n\u2032\n\n(cid:11)\n(cid:11)(1 (cid:0) (cid:11)\u2032)\n\n(cid:11)\n\n1\n\n(cid:11)(1 (cid:0) (cid:11))\n\n\u2032\n\n\u2032\n\n:\n\n1\n\nf (cid:11)\n\ng1(cid:0)(cid:11)\n\n(cid:11)(1 (cid:0) (cid:11)\u2032)\n\nf (cid:11)g1(cid:0)(cid:11) +\n< 0 holds, i.e., D(cid:11);(cid:11)\u2032(f; g) (cid:21) 0 and\n\u2032 for\n(\n\n}\n\n)(cid:11)\n\nRemark 7. D(cid:11);(cid:11)\u2032(f; g) (cid:21) 0 is divergence when (cid:11)(cid:11)\nD(cid:11);(cid:11)\u2032 (f; g) = 0 if and only if f = g. Without loss of generality, we assume (cid:11) > 0 > (cid:11)\nD(cid:11);(cid:11)\u2032.\n\n\u2032\n\n{\n\n\u2211\n\n(\n\n)(cid:11)\n\n\u2032\n\n(cid:18)\n\n1\n\nmin\n\nx2Z\n\nFirstly, we consider an estimator de\ufb01ned by the minmizer of\n\u2032 (cid:0) 1\n1 (cid:0) (cid:11)\n\nq(cid:18)(x)1(cid:0)(cid:11)\n\n1 (cid:0) (cid:11)\u2032\n\n:\nNote that the summation in (9) includes only Z((cid:20) n) terms. We remark the following.\nRemark 8. Let (cid:22)q(cid:18)0(x) be the underlying distribution and q(cid:18)(x) be the unnormalized model (2).\nThen an estimator de\ufb01ned by minimizing D(cid:11);(cid:11)\u2032 ((cid:22)q(cid:18)0 ; q(cid:18)) is not in general Fisher consistent, i.e.,\n@D(cid:11);(cid:11)\u2032((cid:22)q(cid:18)0 ; q(cid:18))\n\n)\u27e8\n\nq(cid:18)(x)1(cid:0)(cid:11)\n\nnx\nn\n\nnx\nn\n\n\u27e9\n\n\u27e8\n\n(\n\n(9)\n\n\u2032\n\n\u2032\n\n/\n\n(cid:18)0 q1(cid:0)(cid:11)\n(cid:22)q(cid:11)\n\n(cid:18)0\n\n\u2032\n(cid:18)0\n\n \n\n(cid:0) (cid:22)q(cid:11)\n\n(cid:18)0q1(cid:0)(cid:11)\n\n(cid:18)0\n\n\u2032\n(cid:18)0\n\n \n\n\u27e9 \u0338= 0:\n\n\u27e8q(cid:18)0\n\n\u27e9(cid:0)(cid:11)\n\n\u2032 (cid:0) \u27e8q(cid:18)0\n\n\u27e9(cid:0)(cid:11)\n\n=\n\n\u2032\n(cid:18)0\n\nq(cid:18)0 \n\n(cid:12)(cid:12)(cid:12)(cid:12)\n\n@(cid:18)\n\n(cid:18)=(cid:18)0\n\nThis remark shows that an estimator associated with D(cid:11);(cid:11)\u2032 (~p; q(cid:18)) does not have suitable properties\nsuch as (asymptotic) unbiasedness and consistency while required computational cost is drastically\nreduced. Intuitively, this is because the (mixture of) (cid:11)-divergence satis\ufb01es the coincidence axiom.\nTo overcome this drawback, we consider the following minimization problem for estimation of the\nparameter (cid:18) of model (cid:22)q(cid:18)(x).\n\n( ^(cid:18); ^r) = argmin\n\nD(cid:11);(cid:11)\u2032(~p; rq(cid:18))\n\nwhere r is a constant corresponding to an inverse of the normalization term Z(cid:18) = \u27e8q(cid:18)\u27e9.\nProposition 2. Let q(cid:18)(x) be the unnormalized model (2). For (cid:11) > 1 and 0 > (cid:11)\nof D(cid:11);(cid:11)\u2032 (~p; rq(cid:18)) is equivalent to the minimization of\n\n(cid:18);r\n\n\u2032, the minimization\n\nProof. For a given (cid:18), we observe that\n\n^r(cid:18) = argmin\n\nr\n\nD(cid:11);(cid:11)\u2032(~p; rq(cid:18)) =\n\n6\n\nS(cid:11);(cid:11)\u2032 (~p; q(cid:18)):\n\n( \u27e8\n\u27e8\n\n~p(cid:11)q1(cid:0)(cid:11)\nq1(cid:0)(cid:11)\u2032\n~p(cid:11)\u2032\n\n(cid:18)\n\n(cid:18)\n\n\u27e9) 1\n\u27e9\n\n(cid:11)(cid:0)(cid:11)\u2032\n\n:\n\n(10)\n\n\fNote that computation of (10) requires only sample order O(n) calculation. By plugging (10) into\nD(cid:11);(cid:11)\u2032 (~p; rq(cid:18)), we observe\n\n^(cid:18) = argmin\n\n(cid:18)\n\nD(cid:11);(cid:11)\u2032(~p; ^r(cid:18)q(cid:18)) = argmin\n\n(cid:18)\n\nS(cid:11);(cid:11)\u2032 (~p; q(cid:18)):\n\n(11)\n\n\u2032\n\nIf (cid:11) > 1 and (cid:11)\n< 0 hold, the estimator (11) is equivalent to the estimator associated with the\nlocalized PS-divergence S(cid:11);(cid:11)\u2032, implying that S(cid:11);(cid:11)\u2032 is characterized by the mixture of (cid:11)-divergences.\nRemark 9. From a viewpoint of the information geometry [11], a metric (information geometrical\nstructure) induced by the (cid:11)-divergence is the Fisher metric induced by the KL-divergence. This im-\nplies that the estimation based on the (mixture of) (cid:11)-divergence is Fisher ef\ufb01cient and is an intuitive\nexplanation of the Theorem 2. The localized PS divergence S(cid:11);(cid:11)\u2032;(cid:13) and S(cid:11);(cid:11)\u2032 with (cid:11)(cid:11)\n> 0 can be\ninterpreted as an extension of the (cid:11)-divergence, which preserves Fisher ef\ufb01ciency.\n\n\u2032\n\n6 Experiments\n\nWe especially focus on a setting of (cid:12) = 1, i.e., convexity of the risk function with the unnormalized\nmodel exp((cid:18)T \u03d5(x)) holds (Theorem 1) and examined performance of the proposed estimator.\n\n6.1 Fully visible Boltzmann machine\n\n\u2032\n\n\u2032\n\n(cid:3)\n\nIn the \ufb01rst experiment, we compared the proposed estimator with parameter settings ((cid:11); (cid:11)\n) =\n(1:01; 0:01); (1:01;(cid:0)0:01); (2;(cid:0)1), with the MLE and the ratio matching method [8]. Note that\nthe ratio matching method also does not require calculation of the normalization constant, and the\n) = (1:01;(cid:6)0:01) may behave like the MLE as discussed in section\nproposed method with ((cid:11); (cid:11)\n5.1.\nAll methods were optimized with the optim function in R language [16]. The dimension d of input\nwas set to 10 and the synthetic dataset was randomly generated from the second order Boltzmann\n(cid:3) (cid:24) N (0; I). We repeated comparison 50 times and\nmachine (Example 2) with a parameter (cid:18)\nobserved averaged performance. Figure 1 (a) shows median of the root mean square errors (RMSEs)\nand ^(cid:18) of each method over 50 trials, against the number n of examples. We observe that\nbetween (cid:18)\nthe proposed estimator works well and is superior to the ratio matching method. In this experiment,\nthe MLE outperforms the proposed method contrary to the prediction of Theorem 2. This is because\nobserved patterns were only a small portion of all possible patterns, as shown in Figure 1 (b). Even\nin such a case, the MLE can take all possible patterns (210 = 1024) into account through the\nnormalization term log Z(cid:18) \u2243 Const + 1\njj(cid:18)jj2 that works like a regularizer. On the other hand, the\nproposed method genuinely uses only the observed examples, and the asymptotic analysis would not\nbe relevant in this case. Figure 1 (c) shows median of computational time of each method against\nn. The computational time of the MLE does not vary against n because the computational cost is\ndominated by the calculation of the normalization constant. Both the proposed estimator and the\nratio matching method are signi\ufb01cantly faster than the MLE, and the ratio matching method is faster\nthan the proposed estimator while the RMSE of the proposed estimator is less than that of the ratio\nmatching.\n\n2\n\n6.2 Boltzmann machine with hidden variables\n\n\u2032\n\nIn this subsection, we applied the proposed estimator for the Boltzmann machine with hidden vari-\nables whose associated function is written as (3). The proposed estimator with parameter settings\n) = (1:01; 0:01); (1:01;(cid:0)0:01); (2;(cid:0)1) was compared with the MLE. The dimension d1 of\n((cid:11); (cid:11)\n(cid:3)\nobserved variables was \ufb01xed to 10 and d2 of hidden variables was set to 2, and the parameter (cid:18)\n(cid:3) (cid:24) N (0; I) including parameters corresponding to hidden variables. Note\nwas generated as (cid:18)\nthat the Boltzmann machine with hidden variables is not identi\ufb01able and different values of the pa-\nrameter do not necessarily generate different probability distributions, implying that estimators are\nin\ufb02uenced by local minimums. Then we measured performance of each estimator by the averaged\n\n7\n\n\f\u2211\n\nn\n\nFigure 1:\n(a) Median of RMSEs of each method against n, in log scale. (b) Box-whisker plot of\nnumber jZj of unique patterns in the dataset D against n. (c) Median of computational time of each\nmethod against n, in log scale.\n\ni=1 log (cid:22)q ^(cid:18)(xi) rather than the RMSE. An initial value of the parameter was set\nlog-likelihood 1\nby N (0; I) and commonly used by all methods. We repeated the comparison 50 times and ob-\nn\nserved the averaged performance. Figure 2 (a) shows median of averaged log-likelihoods of each\nmethod over 50 trials, against the number n of example. We observe that the proposed estimator is\ncomparable with the MLE when the number n of examples becomes large. Note that the averaged\nlog-likelihood of MLE once decreases when n is samll, and this is due to over\ufb01tting of the model.\nFigure 2 (b) shows median of averaged log-likelihoods of each method for test dataset consists of\n10000 examples, over 50 trials. Figure 2 (c) shows median of computational time of each method\nagainst n, and we observe that the proposed estimator is signi\ufb01cantly faster than the MLE.\n\nFigure 2: (a) Median of averaged log-likelihoods of each method against n. (b) Median of averaged\nlog-likelihoods of each method calculated for test dataset against n. (c) Median of computational\ntime of each method against n, in log scale.\n\n7 Conclusions\n\nWe proposed a novel estimator for probabilistic model on discrete space, based on the unnormal-\nized model and the localized PS-divergence which has the homogeneous property. The proposed\nestimator can be constructed without calculation of the normalization constant and is asymptotically\nef\ufb01cient, which is the most important virtue of the proposed estimator. Numerical experiments show\nthat the proposed estimator is comparable to the MLE and required computational cost is drastically\nreduced.\n\n8\n\n05000100001500020000250000.10.20.51.0nRMSEMLERatio matchinga1=1.01,a2=0.01a1=1.01,a2=\u22120.01a1=2,a2=\u221211002004008001600320064001280025600050100150200250300nNumber |Z| of unique patterns050001000015000200002500025102050100200500nTime[s]MLERatio matchinga1=1.01,a2=0.01a1=1.01,a2=\u22120.01a1=2,a2=\u221210500010000150002000025000\u221215\u221210\u22125nAveraged Log likelihoodMLEa1=1.01,a2=0.01a1=1.01,a2=\u22120.01a1=2,a2=\u221210500010000150002000025000\u221215\u221210\u22125nAveraged Log likelihoodMLEa1=1.01,a2=0.01a1=1.01,a2=\u22120.01a1=2,a2=\u22121050001000015000200002500051020501002005001000nTime[s]MLEa1=1.01,a2=0.01a1=1.01,a2=\u22120.01a1=2,a2=\u22121\fReferences\n[1] Hinton, G. E. & Sejnowski, T. J. (1986) Learning and relearning in boltzmann machines. MIT\n\nPress, Cambridge, Mass, 1:282\u2013317.\n\n[2] Ackley, D. H., Hinton, G. E. & Sejnowski, T. J. (1985) A learning algorithm for boltzmann\n\nmachines. Cognitive Science, 9(1):147\u2013169.\n\n[3] Amari, S., Kurata, K. & Nagaoka, H. (1992) Information geometry of Boltzmann machines.\n\nIn IEEE Transactions on Neural Networks, 3: 260\u2013271.\n\n[4] Hinton, G. E. & Salakhutdinov, R. R. (2012) A better way to pretrain deep boltzmann ma-\nchines. In Advances in Neural Information Processing Systems, pp. 2447\u20132455 Cambridge,\nMA: MIT Press.\n\n[5] Opper, M. & Saad, D. (2001) Advanced Mean Field Methods: Theory and Practice. MIT\n\nPress, Cambridge, MA.\n\n[6] Hinton, G.E. (2002) Training Products of Experts by Minimizing Contrastive Divergence.\n\nNeural Computation, 14(8):1771\u20131800.\n\n[7] Hyv\u00a8arinen, A. (2005) Estimation of non-normalized statistical models by score matching.\n\nJournal of Machine Learning Research, 6:695\u2013708.\n\n[8] Hyv\u00a8arinen, A. (2007) Some extensions of score matching. Computational statistics & data\n\nanalysis, 51(5):2499\u20132512.\n\n[9] Dawid, A. P., Lauritzen, S. & Parry, M. (2012) Proper local scoring rules on discrete sample\n\nspaces. The Annals of Statistics, 40(1):593\u2013608.\n\n[10] Gutmann, M. & Hirayama, H. (2012) Bregman divergence as general framework to estimate\n\nunnormalized statistical models. arXiv preprint arXiv:1202.3727.\n\n[11] Amari, S & Nagaoka, H. (2000) Methods of Information Geometry, volume 191 of Transla-\n\ntions of Mathematical Monographs. Oxford University Press.\n\n[12] Sejnowski, T. J. (1986) Higher-order boltzmann machines. In American Institute of Physics\n\nConference Series, 151:398\u2013403.\n\n[13] Good, I. J. (1971) Comment on \u201cmeasuring information and uncertainty,\u201d by R. J. Buehler.\nIn Godambe, V. P. & Sprott, D. A. editors, Foundations of Statistical Inference, pp. 337\u2013339,\nToronto: Holt, Rinehart and Winston.\n\n[14] Fujisawa, H. & Eguchi, S. (2008) Robust parameter estimation with a small bias against heavy\n\ncontamination. Journal of Multivariate Analysis, 99(9):2053\u20132081.\n\n[15] Van der Vaart, A. W. (1998) Asymptotic Statistics. Cambridge University Press.\n[16] R Core Team. (2013) R: A Language and Environment for Statistical Computing. R Foundation\n\nfor Statistical Computing, Vienna, Austria.\n\n9\n\n\f", "award": [], "sourceid": 529, "authors": [{"given_name": "Takashi", "family_name": "Takenouchi", "institution": "Future University Hakodate"}, {"given_name": "Takafumi", "family_name": "Kanamori", "institution": "Nagoya University"}]}