{"title": "Products of Gaussians", "book": "Advances in Neural Information Processing Systems", "page_first": 1017, "page_last": 1024, "abstract": null, "full_text": "Products of Gaussians \n\nChristopher K. I. Williams \n\nFelix V. Agakov \n\nDivision of Informatics \nUniversity of Edinburgh \nEdinburgh EH1 2QL, UK \n\nc. k. i. williams@ed.ac.uk \n\nhttp://anc.ed.ac.uk \n\nSystem Engineering Research Group \nChair of Manufacturing Technology \n\nUniversitiit Erlangen-Niirnberg \n\n91058 Erlangen, Germany \n\nF.Agakov@lft\u00b7uni-erlangen.de \n\nStephen N. Felderhof \nDivision of Informatics \nUniversity of Edinburgh \nEdinburgh EH1 2QL, UK \n\nstephenf@dai.ed.ac.uk \n\nAbstract \n\nRecently Hinton (1999) has introduced the Products of Experts \n(PoE) model in which several individual probabilistic models for \ndata are combined to provide an overall model of the data. Be(cid:173)\nlow we consider PoE models in which each expert is a Gaussian. \nAlthough the product of Gaussians is also a Gaussian, if each Gaus(cid:173)\nsian has a simple structure the product can have a richer structure. \nWe examine (1) Products of Gaussian pancakes which give rise to \nprobabilistic Minor Components Analysis, (2) products of I-factor \nPPCA models and (3) a products of experts construction for an \nAR(l) process. \n\nRecently Hinton (1999) has introduced the Products of Experts (PoE) model in \nwhich several individual probabilistic models for data are combined to provide an \noverall model of the data. In this paper we consider PoE models in which each \nexpert is a Gaussian. It is easy to see that in this case the product model will \nalso be Gaussian. However, if each Gaussian has a simple structure, the product \ncan have a richer structure. Using Gaussian experts is attractive as it permits a \nthorough analysis of the product architecture, which can be difficult with other \nmodels, e.g. models defined over discrete random variables. \n\nBelow we examine three cases of the products of Gaussians construction: (1) Prod(cid:173)\nucts of Gaussian pancakes (PoGP) which give rise to probabilistic Minor Compo(cid:173)\nnents Analysis (MCA), providing a complementary result to probabilistic Principal \nComponents Analysis (PPCA) obtained by Tipping and Bishop (1999); (2) Prod(cid:173)\nucts of I-factor PPCA models; (3) A products of experts construction for an AR(l) \nprocess. \n\n\fProducts of Gaussians \n\nIf each expert is a Gaussian pi(xI8i ) '\" N(J1i' ( i), the resulting distribution of the \nproduct of m Gaussians may be expressed as \n\nBy completing the square in the exponent it may be easily shown that p(xI8) \nN(/1;E, (2:), where (E l = 2::1 (i l . To simplify the following derivations we will \nassume that pi(xI8i ) '\" N(O, (i) and thus that p(xI8) '\" N(O, (2:). J12: i \u00b0 can be \n\nobtained by translation of the coordinate system. \n\n1 Products of Gaussian Pancakes \n\nA Gaussian \"pancake\" (GP) is a d-dimensional Gaussian, contracted in one dimen(cid:173)\nsion and elongated in the other d - 1 dimensions. In this section we show that the \nmaximum likelihood solution for a product of Gaussian pancakes (PoGP) yields a \nprobabilistic formulation of Minor Components Analysis (MCA). \n\n1.1 Covariance Structure of a GP Expert \n\nConsider a d-dimensional Gaussian whose probability contours are contracted \nin the direction w and equally elongated in mutually orthogonal directions \nVI , ... , vd-l.We call this a Gaussian pancake or GP. Its inverse covariance may be \nwritten as \n\nd - l \n\n( - 1 = L ViV; /30 + wwT /3,;;, \n\ni = l \n\n(1) \n\nwhere VI, ... ,V d - l, W form a d x d matrix of normalized eigenvectors of the covari(cid:173)\nance C. /30 = 0\"0 2 , /3,;; = 0\";;2 define inverse variances in the directions of elongation \nand contraction respectively, so that 0\"5 2 0\"1. Expression (1) can be re-written in \na more compact form as \n\n(2) \nwhere w = wJ/3,;; - /30 and Id C jRdxd is the identity matrix. Notice that according \nto the constraint considerations /30 < /3,;;, and all elements of ware real-valued. \nNote the similarity of (2) with expression for the covariance of the data of a 1-\nfactor probabilistic principal component analysis model ( = 0\"21d + wwT (Tipping \nand Bishop, 1999) , where 0\"2 is the variance of the factor-independent spherical \nGaussian noise. The only difference is that it is the inverse covariance matrix for \nthe constrained Gaussian model rather than the covariance matrix which has the \nstructure of a rank-1 update to a multiple of Id . \n\n1.2 Covariance of the PoGP Model \n\nWe now consider a product of m GP experts, each of which is contracted in a single \ndimension. We will refer to the model as a (I,m) PoGP, where 1 represents the \nnumber of directions of contraction of each expert. We also assume that all experts \nhave identical means. \n\n\fFrom (1), the inverse covariance of the the resulting (I,m) PoGP model can be \nexpressed as \n\nm \n\nC;;l = L Cil \n\ni=l \n\n(3) \n\nwhere columns of We Rdxm correspond to weight vectors of the m PoGP experts, \nand (3E = 2::1 (3~i) > o. \n\n1.3 Maximum-Likelihood Solution for PoGP \n\nComparing (3) with m-factor PPCA we can make a conjecture that in contrast \nwith the PPCA model where ML weights correspond to principal components of \nthe data covariance (Tipping and Bishop, 1999), weights W of the PoGP model \ndefine projection onto m minor eigenvectors of the sample covariance in the visible \nd-dimensional space, while the distortion term (3E Id explains larger variationsl . This \nis indeed the case. \nIn Williams and Agakov (2001) it is shown that stationarity of the log-likelihood \nwith respect to the weight matrix Wand the noise parameter (3E results in three \nclasses of solutions for the experts' weight matrix, namely \n\nW \n5 \n5W \n\n0; \nCE ; \nCEW, W:j:. 0, 5:j:. CE, \n\n(4) \n\nwhere 5 is the covariance matrix of the data (with an assumed mean of zero). The \nfirst two conditions in (4) are the same as in Tipping and Bishop (1999), but for \nPPCA the third condition is replaced by C-l W = 5- l W (assuming that 5- 1 exists). \nIn Appendix A and Williams and Agakov (2001) it is shown that the maximum \nlikelihood solution for W ML is given by: \n\n(5) \n\nwhere R c Rmxm is an arbitrary rotation matrix, A is a m x m matrix containing \nthe m smallest eigenvalues of 5 and U = [Ul , ... ,u m ] c Rdxm is a matrix of the \ncorresponding eigenvectors of 5. Thus, the maximum likelihood solution for the \nweights of the (1, m) PoG P model corresponds to m scaled and rotated minor \neigenvectors of the sample covariance 5 and leads to a probabilistic model of minor \ncomponent analysis. As in the PPCA model, the number of experts m is assumed \nto be lower than the dimension of the data space d. \n\nThe correctness of this derivation has been confirmed experimentally by using a \nscaled conjugate gradient search to optimize the log likelihood as a function of W \nand (3E. \n\n1.4 Discussion of PoGP model \n\nAn intuitive interpretation of the PoGP model is as follows: Each Gaussian pancake \nimposes an approximate linear constraint in x space. Such a linear constraint is that \nx should lie close to a particular hyperplane. The conjunction of these constraints \nis given by the product of the Gaussian pancakes. If m \u00ab d it will make sense to \nlBecause equation 3 has the form of a factor analysis decomposition, but for the inverse \n\ncovariance matrix, we sometimes refer to PoGP as the rotcaf model. \n\n\fdefine the resulting Gaussian distribution in terms of the constraints. However, if \nthere are many constraints (m > d/2) then it can be more efficient to describe the \ndirections of large variability using a PPCA model, rather than the directions of \nsmall variability using a PoGP model. This issue is discussed by Xu et al. (1991) in \nwhat they call the \"Dual Subspace Pattern Recognition Method\" where both PCA \nand MCA models are used (although their work does not use explicit probabilistic \nmodels such as PPCA and PoGP). \n\nMCA can be used, for example, for signal extraction in digital signal processing \n(Oja, 1992), dimensionality reduction, and data visualization. Extraction of the \nminor component is also used in the Pisarenko Harmonic Decomposition method \nfor detecting sinusoids in white noise (see, e.g. Proakis and Manolakis (1992), p. \n911). Formulating minor component analysis as a probabilistic model simplifies \ncomparison of the technique with other dimensionality reduction procedures, per(cid:173)\nmits extending MCA to a mixture of MCA models (which will be modeled as a \nmixture of products of Gaussian pancakes) , permits using PoGP in classification \ntasks (if each PoGP model defines a class-conditional density) , and leads to a num(cid:173)\nber of other advantages over non-probabilistic MCA models (see the discussion of \nadvantages of PPCA over PCA in Tipping and Bishop (1999)). \n\n2 Products of PPCA \n\nIn this section we analyze a product of m I-factor PPCA models, and compare it \nto am-factor PPCA model. \n\n2.1 \n\nI-factor PPCA model \n\nConsider a I-factor PPCA model, having a latent variable Si and visible variables x. \nThe joint distribution is given by P(Si, x) = P(si) P(xlsi). We set P(Si) '\" N(O, 1) \nand P(XI Si) '\" N(WiSi' (]\"2) . Integrating out Si we find that Pi(x) '\" N(O, Ci ) where \nC = wiwT + (]\"21d and \n\nwhere (3 = (]\"-2 and \"(i = (3/(1 + (3 llwi W). (3 and \"(i are the inverse variances in the \ndirections of contraction and elongation respectively. \nThe joint distribution of Si and x is given by \n\n(6) \n\n- 2x WiSi + X X \n\nT \n\nT \n\n. \n\n(7) \n\n(8) \n\n] \n\n(3 [s; \n-\nexp - -\n2 \n\"(i \n\nTipping and Bishop (1999) showed that the general m-factor PPCA model (m(cid:173)\nPPCA) has covariance C = (]\"21d + WWT , where W is the d x m matrix of factor \nloadings. When fitting this model to data, the maximum likelihood solution is to \nchoose W proportional to the principal components of the data covariance matrix. \n\n\f2.2 Products of I-factor PPCA models \n\nWe now consider the product of m I-factor PPCA models, which we denote a \n(1, m)-PoPPCA model. The joint distribution over 5 = (Sl' ... ,Srn)T and x is \n\nP(x,s) ex: exp -\"2 L ---:- - 2x W iSi + X X \n\n13 m [s; \n\nT \n\nT \n\n] \n\ni=l \n\n,,(, \n\n\u2022 \n\n(9) \n\nLet zT d~f (xT , ST). Thus we see that the distribution of z is Gaussian with inverse \ncovariance matrix 13M, where \n\n-W) \nr - 1 \n\n, \n\n(10) \n\nand r = diag(\"(l , ... ,\"(m)' Using the inversion equations for partitioned matrices \n(Press et al., 1992, p. 77) we can show that \n\nwhere ~xx is the covariance of the x variables under this model. It is easy to confirm \nthat this is also the result obtained from summing (6) over i = 1, ... ,m. \n\n(11) \n\n2.3 Maximum Likelihood solution for PoPPCA \nAm-factor PPCA model has covariance a21d + WWT and thus, by the Woodbury \nj3W(a2 lm + WT W) - lWT . The maximum \nformula, it has inverse covariance j3 ld -\nlikelihood solution for a m-PPCA model is similar to (5), i.e. W = U(A _a2Im)1/2 RT, \nbut now A is a diagonal matrix of the m principal eigenvalues, and U is a matrix \nof the corresponding eigenvectors. If we choose RT = I then the columns of W are \northogonal and the inverse covariance of the maximum likelihood m-PPCA model \nhas the form j3 ld - j3WrwT. Comparing this to (11) (with W = W) we see that the \ndifference is that the first term of the RHS of (11) is j3m1d , while for m-PPCA it is \nj3 ld. \nIn section 3.4 and Appendix C.3 of Agakov (2000) it is shown that (for m :::=: 2) we \nobtain the m-factor PPCA solution when \n-\nA<A' < - -A \n' m -I ' \n\ni = 1, ... ,m, \n\n(12) \n\nm \n\n-\n\n-\n\nwhere A is the mean of the d - m discarded eigenvalues, and Ai is a retained eigen(cid:173)\nvalue; it is the smaller eigenvalues that are discarded. We see that the covariance \nmust be nearly spherical for this condition to hold. For covariance matrices sat(cid:173)\nisfying (12) , this solution was confirmed by numerical experiments as detailed in \n(Agakov, 2000, section 3.5). \nTo see why this is true intuitively, observe that Ci 1 for each I-factor PPCA expert \nwill be large (with value 13) in all directions except one. If the directions of con(cid:173)\ntraction for each Ci 1 are orthogonal, we see that the sum of the inverse covariances \nwill be at least (m - 1)13 in a contracted direction and m j3 in a direction in which \nno contraction occurs. The above shows that for certain types of sample covari(cid:173)\nance matrix the (1 , m) PoPPCA solution is not equivalent to the m-factor PPCA \nsolution. However, it is interesting to note that by relaxing the constraint on the \nisotropy of each expert's noise the product of m one-factor factor analysis models \ncan be shown to be equivalent to an m-factor factor analyser (Marks and Movellan, \n2001). \n\n\f\u2022 \u2022 \u2022 \u2022 \n\n(c) \n\n\u2022 \n\u2022 \u2022 \n\n\u2022 \n\n(d) \n\n(b) \n\nFigure 1: (a) Two experts. The upper one depicts 8 filled circles (visible units) and \n4 latent variables (open circles), with connectivity as shown. The lower expert also \nhas 8 visible and 4 latent variables, but shifted by one unit (with wraparound). (b) \nCovariance matrix for a single expert. (c) Inverse covariance matrix for a single \nexpert. (d) Inverse covariance for product of experts. \n\n3 A Product of Experts Representation for an AR(l) \n\nProcess \n\nFor the PoPPCA case above we have considered models where the latent variables \nhave unrestricted connectivity to the visible variables. We now consider a product \nof experts model with two experts as shown in Figure l(a). The upper figure depicts \n8 filled circles (visible units) and 4 latent variables (open circles), with connectivity \nas shown. The lower expert also has 8 visible and 4 latent variables, but shifted \nby one unit (with wraparound) with respect to the first expert. The 8 units are, \nof course, only for illustration-\nthe construction is valid for any even number of \nvisible units. \n\nConsider one hidden unit and its two visible children. Denote the hidden unit by s \nthe visible units as Xl and Xr (l, r for left and right). Set s '\" N(O, 1) and \n\nXl = as + bWI \n\nXr = \u00b1as + bwr , \n\n(13) \n\nwhere WI and Wr are independent N(O , 1) random variables, and a, b are constants. \n(This is a simple example of a Gaussian tree-structured process, as studied by a \nnumber of groups including that led by Prof. Willsky at MIT; see e.g. Luettgen \net al. (1993).) Then (xf) = (x;) = a2 + b2 and (XIX r ) = \u00b1a2 \u2022 The corresponding \n2 x 2 inverse covariance matrix has diagonal entries of (a2 + b2 )j ~ and off-diagonal \nentries of =t=a2 j~ , where ~ = b2 (b2 + 2a2 ). \nGraphically, the covariance matrix of a single expert has the form shown in Figure \nl(b) (where we have used the + rather than - choice from (13) for all variables). \nFigure l(c) shows the corresponding inverse covariance for the single expert, and \nFigure 1 (d) shows the resulting inverse covariance for the product of the two experts, \nwith diagonal elements 2(a2 + b2 )j ~ and off-diagonal entries of =t=a2 j~. \nAn AR(l) process of the circle with d nodes has the form Xi = aXi - 1 (mod d) + Zi, \n\n\fwhere Zi ~ N(O,v). Thusp(X) <X exp-21v L:i(Xi-aXi- 1 (mod d))2 and the inverse \ncovariance matrix has a circulant tridiagonal structure with diagonal entries of \n(1 + ( 2 )/v and off-diagonal entries of -a/v. The product of experts model defined \nabove can be made equivalent to the circular AR(I) process by setting \n\n(14) \n\nThe \u00b1 is needed in (13) as when a is negative we require Xr = -as + bWr to match \nthe inverse covariances. \n\nWe have shown that there is an exact construction to represent a stationary cir(cid:173)\ncular AR(I) process as a product of two Gaussian experts. The approximation \nof other Gaussian processes by products of tree-structured Gaussian processes is \nfurther studied in (Williams and Felderhof, 2001). Such constructions are interest(cid:173)\ning because they may allow fast approximate inference in the case that d is large \n(and the target process may be 2 or higher dimensional) and exact inference is not \ntractable. Such methods have been developed by Willsky and coauthors, but not \nfor products of Gaussians constructions. \n\nAcknowledgements \n\nThis work is partially supported by EPSRC grant GR/L78161 Probabilistic Models \nfor Sequences. Much of the work on PoGP was carried out as part of the MSc \nproject of FVA at the Division of Informatics, University of Edinburgh. CW thanks \nSam Roweis, Geoff Hinton and Zoubin Ghahramani for helpful conversations on the \nrotcaf model during visits to the Gatsby Computational Neuroscience Unit. FVA \ngratefully acknowledges the support of the Royal Dutch Shell Group of Companies \nfor his MSc studies in Edinburgh through a Centenary Scholarship. SNF gratefully \nacknowledges additional support from BAE Systems. \n\nReferences \nAgakov, F. (2000). \n\nInvestigations of Gaussian Products-of-Experts Models. Master's \nthesis, Division of Informatics, The University of Edinburgh. Available at http://'iI'iI'iI . \ndai.ed.ac.uk/homes/felixa/all.ps.gz. \n\nHinton, G. E . (1999) . Products of experts. In Proceedings of the Ninth International \n\nConference on Artificial Neural Networks (ICANN gg), pages 1- 6. \n\nLuettgen, M. , Karl, W. , and Willsky, A. (1993). Multiscale Representations of Markov \n\nRandom Fields. IEEE Trans. Signal Processing, 41(12):3377- 3395. \n\nMarks , T. and Movellan, J. (2001). Diffusion Networks, Products of Experts, and Factor \nAnalysis. In Proceedings of the 3rd International Conference on Independent Component \nAnalysis and Blind Source Separation. \n\nOJ a, E. (1992). Principal Components, Minor Components, and Linear Neural Networks. \n\nNeural N etworks, 5:927 - 935. \n\nPress, W. H. , Teukolsky, S. A., Vetterling, W. T., and Flannery, B. P. (1992). Num erical \n\nRecipes in C. Cambridge University Press, Second edition. \n\nProakis, J. G. and Manolakis, D. G. (1992). Digital Signal Processing: Principles, Algo(cid:173)\n\nrithms and Applications. Macmillan. \n\nTipping, M. E. and Bishop, C. M. (1999). Probabilistic principal components analysis. J. \n\nRoy. Statistical Society B, 61(3) :611- 622. \n\nWilliams, C. K. I. and Agakov, F. V. (2001). Products of Gaussians and Probabilistic \n\nMinor Components Analysis. Technical Report EDI-INF-RR-0043, Division of Infor(cid:173)\nmatics, University of Edinburgh. Available at http://'iI'iI'iI. informatics. ed. ac. ukl \npublications/report/0043.html. \n\n\fWilliams, C. K. I. and Felderhof, S. N. (2001). Products and Sums of Tree-Structured \nGaussian Processes. In Proceedings of the ICSC Symposium on Soft Computing 2001 \n(SOCO 2001). \n\nXu, L. and Krzyzak, A. and Oja, E. (1991). Neural Nets for Dual Subspace Pattern \n\nRecogntion Method. International Journal of Neural Systems, 2(3):169- 184. \n\nA ML Solutions for PoGP \n\nHere we analyze the three classes of solutions for the model covariance matrix which \nresult from equation (4) of section 1.3. \nThe first case W = 0 corresponds to a minimum of the log-likelihood. \nIn the second case, the model covariance e~ is equal to the sample covariance 5. \nFrom expression (3) for e i;l we find WWT = 5- 1 - ;3~ ld. This has the known \nsolution W = Um(A - 1 - ;3~ lm)1 /2 RT , where Um is the matrix of the m eigenvectors \nof 5 with the smallest eigenvalues and A is the corresponding diagonal matrix of the \neigenvalues. The sample covariance must be such that the largest d - m eigenvalues \nare all equal to ;3~; the other m eigenvalues are matched explicitly. \nFinally, for the case of approximate model covariance (5W = e~w, 5 =f. e~) we, by \nanalogy with Tipping and Bishop (1999), consider the singular value decomposition \nof the weight matrix, and establish dependencies between left singular vectors of \nW = ULRT and eigenvectors of the sample covariance 5. U = [U1 , U2 , ... , um] C \nlRdxm is a matrix of left singular vectors of W with columns constituting an or(cid:173)\nthonormal basis, L = diag(h,l2, ... ,lm) C lRmxm is a diagonal matrix of the sin(cid:173)\ngular values of Wand R C lRmxm defines an arbitrary rigid rotation of W. For this \ncase equation (4) can be written as 5UL = e ~ UL , where e~ is obtained from (3) by \napplying the matrix inversion lemma [see e.g. Press et al. (1992)]. This leads to \n\n5UL = e~UL \n\n(;3i;lld - ;3i;l W(;3~ + WTW) -l WT)UL \nU(;3i;l lm - ;3i;l LRT(;3~ lm + RL2RT) -l RL)L \nU(;3i; l lm - ;3i;l(;3~ L -2 + Im) -l) L. \n\n(15) \nNotice that the term ;3;1 1m - ;3;l(;3~ L -2 + Im)-l in the r.h.s. of equation (15) is \njust a scaling factor of U. Equation (15) defines the matrix form of the eigenvector \nequation, with both sides post-multiplied by the diagonal matrix L. \nIf li =f. 0 then (15) implies that \n\ne~ U i = 5Ui = AiUi, Ai = ;3i;1(1 - (;3~li2 + 1) - 1), \n\n(16) \nwhere Ui is an eigenvector of 5, and Ai is its corresponding eigenvalue. The scaling \nfactor li of the ith retained expert can be expressed as li = (Ail - ;3~)1/2 . \nObviously, if li = 0 then Ui is arbitrary. If li = 0 we say that the direction corre(cid:173)\nsponding to Ui is discarded, i.e. the variance in that direction is explained merely \nby noise. Otherwise we say that Ui is retained. All potential solutions of W may \nthen be expressed as \n\nW = Um(D - ;3~ lm)1 /2 RT , \n\n(17) \nwhere R C lRmxm is a rotation matrix, Um = [U1U2 ... um] C lRdxm is a matrix whose \ncolumns correspond to m eigenvectors of 5, and D = diag( d1 , d2 , ... , dm ) C lRm x m \nsuch that di = Ail if Ui is retained and di = ;3~ if Ui is discarded. \nIt may further be shown (Williams and Agakov (2001)) that the optimal solution \nfor the likelihood is reached when W corresponds to the minor eigenvectors of the \nsample covariance 5. \n\n\f", "award": [], "sourceid": 2102, "authors": [{"given_name": "Christopher", "family_name": "Williams", "institution": null}, {"given_name": "Felix", "family_name": "Agakov", "institution": null}, {"given_name": "Stephen", "family_name": "Felderhof", "institution": null}]}