{"title": "SDP Relaxation with Randomized Rounding for Energy Disaggregation", "book": "Advances in Neural Information Processing Systems", "page_first": 4978, "page_last": 4986, "abstract": "We develop a scalable, computationally efficient method for the task of energy disaggregation for home appliance monitoring. In this problem the goal is to estimate the energy consumption of each appliance based on the total energy-consumption signal of a household. The current state of the art models the problem as inference in factorial HMMs, and finds an approximate solution to the resulting quadratic integer program via quadratic programming. Here we take a more principled approach, better suited to integer programming problems, and find an approximate optimum by combining convex semidefinite relaxations with randomized rounding, as well as with a scalable ADMM method that exploits the special structure of the resulting semidefinite program. Simulation results demonstrate the superiority of our methods both in synthetic and real-world datasets.", "full_text": "SDP Relaxation with Randomized Rounding for\n\nEnergy Disaggregation\n\nKiarash Shaloudegi\n\nImperial College London\n\nk.shaloudegi16@imperial.ac.uk\n\nAndr\u00e1s Gy\u00f6rgy\n\nImperial College London\n\na.gyorgy@imperial.ac.uk\n\nCsaba Szepesv\u00e1ri\nUniversity of Alberta\n\nszepesva@ualberta.ca\n\nWilsun Xu\n\nUniversity of Alberta\nwxu@ualberta.ca\n\nAbstract\n\nWe develop a scalable, computationally ef\ufb01cient method for the task of energy\ndisaggregation for home appliance monitoring. In this problem the goal is to\nestimate the energy consumption of each appliance over time based on the total\nenergy-consumption signal of a household. The current state of the art is to model\nthe problem as inference in factorial HMMs, and use quadratic programming to\n\ufb01nd an approximate solution to the resulting quadratic integer program. Here we\ntake a more principled approach, better suited to integer programming problems,\nand \ufb01nd an approximate optimum by combining convex semide\ufb01nite relaxations\nrandomized rounding, as well as a scalable ADMM method that exploits the special\nstructure of the resulting semide\ufb01nite program. Simulation results both in synthetic\nand real-world datasets demonstrate the superiority of our method.\n\n1\n\nIntroduction\n\nEnergy ef\ufb01ciency is becoming one of the most important issues in our society. Identifying the\nenergy consumption of individual electrical appliances in homes can raise awareness of power\nconsumption and lead to signi\ufb01cant saving in utility bills. Detailed feedback about the power\nconsumption of individual appliances helps energy consumers to identify potential areas for energy\nsavings, and increases their willingness to invest in more ef\ufb01cient products. Notifying home owners\nof accidentally running stoves, ovens, etc., may not only result in savings but also improves safety.\nEnergy disaggregation or non-intrusive load monitoring (NILM) uses data from utility smart meters\nto separate individual load consumptions (i.e., a load signal) from the total measured power (i.e., the\nmixture of the signals) in households.\nThe bulk of the research in NILM has mostly concentrated on applying different data mining and\npattern recognition methods to track the footprint of each appliance in total power measurements.\nSeveral techniques, such as arti\ufb01cial neural networks (ANN) [Prudenzi, 2002, Chang et al., 2012,\nLiang et al., 2010], deep neural networks [Kelly and Knottenbelt, 2015], k-nearest neighbor (k-NN)\n[Figueiredo et al., 2012, Weiss et al., 2012], sparse coding [Kolter et al., 2010], or ad-hoc heuristic\nmethods [Dong et al., 2012] have been employed. Recent works, rather than turning electrical events\ninto features fed into classi\ufb01ers, consider the temporal structure of the data[Zia et al., 2011, Kolter\nand Jaakkola, 2012, Kim et al., 2011, Zhong et al., 2014, Egarter et al., 2015, Guo et al., 2015],\nresulting in state-of-the-art performance [Kolter and Jaakkola, 2012]. These works usually model the\nindividual appliances by independent hidden Markov models (HMMs), which leads to a factorial\nHMM (FHMM) model describing the total consumption.\n\n30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain.\n\n\fFHMMs, introduced by Ghahramani and Jordan [1997], are powerful tools for modeling times series\ngenerated from multiple independent sources, and are great for modeling speech with multiple people\nsimultaneously talking [Rennie et al., 2009], or energy monitoring which we consider here [Kim et al.,\n2011]. Doing exact inference in FHMMs is NP hard; therefore, computationally ef\ufb01cient approximate\nmethods have been the subject of study. Classic approaches include sampling methods, such as\nMCMC or particle \ufb01ltering [Koller and Friedman, 2009] and variational Bayes methods [Wainwright\nand Jordan, 2007, Ghahramani and Jordan, 1997]. In practice, both methods are nontrivial to make\nwork and we are not aware of any works that would have demonstrated good results in our application\ndomain with the type of FHMMs we need to work and at practical scales.\nIn this paper we follow the work of Kolter and Jaakkola [2012] to model the NILM problem by\nFHMMs. The distinguishing features of FHMMs in this setting are that (i) the output is the sum of\nthe output of the underlying HMMs (perhaps with some noise), and (ii) the number of transitions\nare small in comparison to the signal length. FHMMs with the \ufb01rst property are called additive. In\nthis paper we derive an ef\ufb01cient, convex relaxation based method for FHMMs of the above type,\nwhich signi\ufb01cantly outperforms the state-of-the-art algorithms. Our approach is based on revisiting\nrelaxations to the integer programming formulation of Kolter and Jaakkola [2012]. In particular,\nwe replace the quadratic programming relaxation of Kolter and Jaakkola, 2012 with a relaxation\nto an semi-de\ufb01nite program (SDP), which, based on the literature of relaxations is expected to be\ntighter and thus better. While SDPs are convex and could in theory be solved using interior-point\n(IP) methods in polynomial time [Malick et al., 2009], IP scales poorly with the size of the problem\nand is thus unsuitable to our large scale problem which may involve as many a million variables. To\naddress this problem, capitalizing on the structure of our relaxation coming from our FHMM model,\nwe develop a novel variant of ADMM [Boyd et al., 2011] that uses Moreau-Yosida regularization\nand combine it with a version of randomized rounding that is inspired by the the recent work of\nPark and Boyd [2015]. Experiments on synthetic and real data con\ufb01rm that our method signi\ufb01cantly\noutperforms other algorithms from the literature, and we expect that it may \ufb01nd its applications in\nother FHMM inference problems, too.\n\n1.1 Notation\nThroughout the paper, we use the following notation: R denotes the set of real numbers, Sn\n+ denotes\nthe set of n \u21e5 n positive semide\ufb01nite matrices, I{E} denotes the indicator function of an event E\n(that is, it is 1 if the event is true and zero otherwise), 1 denotes a vector of appropriate dimension\nwhose entries are all 1. For an integer K, [K] denotes the set {1, 2, . . . , K}. N (\u00b5, \u2303) denotes the\nGaussian distribution with mean \u00b5 and covariance matrix \u2303. For a matrix A, trace(A) denotes its\ntrace and diag(A) denotes the vector formed by the diagonal entries of A.\n\n2 System Model\n\nFollowing Kolter and Jaakkola [2012], the energy usage of the household is modeled using an additive\nfactorial HMM [Ghahramani and Jordan, 1997]. Suppose there are M appliances in a household.\nEach of them is modeled via an HMM: let Pi 2 RKi\u21e5Ki denote the transition-probability matrix of\nappliance i 2 [M ], and assume that for each state s 2 [Ki], the energy consumption of the appliance\nis constant \u00b5i,s (\u00b5i denotes the corresponding Ki-dimensional column vector (\u00b5i,1, . . . , \u00b5i,Ki)>).\nDenoting by xt,i 2{ 0, 1}Ki the indicator vector of the state st,i of appliance i at time t (i.e.,\nxt,i,s = I{st,i=s}), the total power consumption at time t isPi2[M ] \u00b5>i xt,i, which we assume is\nobserved with some additive zero mean Gaussian noise of variance 2: yt \u21e0N (Pi2[M ] \u00b5>i xt,i, 2).1\n\nGiven this model, the maximum likelihood estimate of the appliance state vector sequence can be\nobtained by minimizing the log-posterior function\n\narg min\nxt,i\n\nsubject to\n\ni=1 x>t,i\u00b5i)2\n22\n\n(yt PM\nTXt=1\nxt,i 2{ 0, 1}Ki, 1>xt,i = 1, i 2 [M ] and t 2 [T ],\n\nT1Xt=1\n\nx>t,i(log Pi)xt+1,i\n\n\n\nMXi=1\n\n(1)\n\n1Alternatively, we can assume that the power consumption yt,iof each appliance is normally distributed with\n\nmean \u00b5>i xt,i and variance 2\n\ni , where 2 =Pi2[M ] 2\n\ni , and yt =Pi2[M ] yt,i.\n\n2\n\n\fwhere log Pi denotes a matrix obtained from Pi by taking the logarithm of each entry.\nIn our particular application, in addition to the signal\u2019s temporal structure, large changes in total power\n(in comparison to signal noise) contain valuable information that can be used to further improve the\ninference results (in fact, solely this information was used for energy disaggregation, e.g., by Dong\net al., 2012, 2013, Figueiredo et al., 2012). This observation was used by Kolter and Jaakkola [2012]\nto amend the posterior with a term that tries to match the large signal changes to the possible changes\nin the power level when only the state of a single appliance changes.\nFormally, let yt = yt+1  yt, \u00b5(i)\nm,k = \u00b5i,k  \u00b5i,m, and de\ufb01ne the matrices Et,i 2 RKi\u21e5Ki\nby (Et,i)m,k = ( yt  \u00b5(i)\ndiff), for some constant diff > 0. Intuitively, (Et,i)m,k is\nm,k)2/(22\nthe negative log-likelihood (up to a constant) of observing a change yt in the power level when\nappliance i transitions from state m to state k under some zero-mean Gaussian noise with variance\ndiff. Making the heuristic approximation that the observation noise and this noise are independent\n2\n(which clearly does not hold under the previous model), Kolter and Jaakkola [2012] added the term\n\ni=1 x>t,iEt,ixt+1,i) to the objective of (1), arriving at\n\n(PT1\n\nt=1 PM\n\narg min\nxt,i\n\n(yt PM\n\nTXt=1\n\nf (x1, . . . , xT ) :=\n\ni=1 x>t,i\u00b5i)2\n22\n\nT1Xt=1\nMXi=1\nsubject to xt,i 2{ 0, 1}Ki, 1>xt,i = 1, i 2 [M ] and t 2 [T ] .\nIn the rest of the paper we derive an ef\ufb01cient approximate solution to (2), and demonstrate that it is\nsuperior to the approximate solution derived by Kolter and Jaakkola [2012] with respect to several\nmeasures quantifying the accuracy of load disaggregation solutions.\n\nx>t,i(Et,i + log Pi)xt+1,i\n\n(2)\n\n\n\n3 SDP Relaxation and Randomized Rounding\n\nThere are two major challenges to solve the optimization problem (2) exactly: (i) the optimization is\nover binary vectors xt,i; and (ii) the objective function f, even when considering its extension to a\nconvex domain, is in general non-convex (due to the second term). As a remedy we will relax (2) to\nmake it an integer quadratic programming problem, then apply an SDP relaxation and randomized\nrounding to solve approximately the relaxed problem. We start with reviewing the latter methods.\n\n3.1 Approximate Solutions for Integer Quadratic Programming\n\nIn this section we consider approximate solutions to the integer quadratic programming problem\n\nminimize\nsubject to\n\nf (x) = x>Dx + 2d>x\nx 2{ 0, 1}n,\n\n(3)\n\n+ is positive semide\ufb01nite, and d 2 Rn. While an exact solution of (3) can be found\nwhere D 2 Sn\nby enumerating all possible combination of binary values within a properly chosen box or ellipsoid,\nthe running time of such exact methods is nearly exponential in the number n of binary variables,\nmaking these methods un\ufb01t for large scale problems.\nOne way to avoid exponential running times is to replace (3) with a convex problem with the hope that\nthe solutions of the convex problems can serve as a good starting point to \ufb01nd high-quality solutions\nto (3). The standard approach to this is to linearize (3) by introducing a new variable X 2 Sn\n+\ntied to x trough X = xx>, so that x>Dx = trace(DX), and then relax the nonconvex constraints\nX = xx>, x 2{ 0, 1}n to X \u232b xx>, diag(X) = x, x 2 [0, 1]n. This leads to the relaxed SDP\nproblem\n\nminimize\n\nsubject to\n\ntrace(D>X) + 2d>x\n\n\uf8ff1 x>\nx X \u232b 0,\n\ndiag(X) = x,\n\nx 2 [0, 1]n\n\n(4)\n\n3\n\n\fBy introducing \u02c6X =\uf8ff1 x>\n\nx X this can be written in the compact SDP form\n\nminimize\nsubject to\n\ntrace( \u02c6D> \u02c6X)\n\u02c6X \u232b 0, A \u02c6X = b .\n\nwhere \u02c6D =\uf8ff0 d>\n\nd D 2 Sn+1\n\n+ , b 2 Rm and A : Sn\n\n+ ! Rm is an appropriate linear operator. This\ngeneral SDP optimization problem can be solved with arbitrary precision in polynomial time using\ninterior-point methods [Malick et al., 2009, Wen et al., 2010]. As discussed before, this approach\nbecomes impractical in terms of both the running time and the required memory if either the number\nof variables or the optimization constraints are large [Wen et al., 2010]. We will return to the issue of\nbuilding scaleable solvers for NILM in Section 5.\nNote that introducing the new variable X, the problem is projected into a higher dimensional space,\nwhich is computationally more challenging than just simply relaxing the integrality constraint in (3),\nbut leads to a tighter approximation of the optimum (c.f., Park and Boyd, 2015; see also Lov\u00e1sz and\nSchrijver, 1991, Burer and Vandenbussche, 2006).\nTo obtain a feasible point of (3) from the solution of (5), we still need to change the solution x to\na binary vector. This can be done via randomized rounding [Park and Boyd, 2015, Goemans and\nWilliamson, 1995]: Instead of letting x 2 [0, 1]n, the integrality constraint x 2{ 0, 1}n in (3) can be\nreplaced by the inequalities xi(xi  1)  0 for all i 2 [n]. Although these constraints are nonconvex,\nthey admit an interesting probabilistic interpretation: the optimization problem\n\nminimize\nsubject to\n\nis equivalent to\n\nEw\u21e0N (\u00b5,\u2303)[w>Dw + 2d>w]\nEw\u21e0N (\u00b5,\u2303)[wi(wi  1)]  0,\n\ni 2 [n], \u00b5 2 Rn, \u2303 \u232b 0\n\n(5)\n\n(6)\n\nminimize\nsubject to\n\ntrace((\u2303 + \u00b5\u00b5>)D) + 2d>\u00b5\n\u2303i,i + \u00b52\n\ni  \u00b5i  0,\n\ni 2 [n],\n\nwhich is in the form of (4) with X =\u2303+ \u00b5\u00b5> and x = \u00b5 (above, Ex\u21e0P [f (x)] stands for\nR f (x)dP (x)). This leads to the rounding procedure: starting from a solution (x\u21e4, X\u21e4) of (4),\nwe randomly draw several samples w(j) from N (x\u21e4, X\u21e4  x\u21e4x\u21e4>), round w(j)\nto 0 or 1 to obtain\nx(j), and keep the x(j) with the smallest objective value. In a series of experiments, Park and Boyd\n[2015] found this procedure to be better than just naively rounding the coordinates of x\u21e4.\n\ni\n\n4 An Ef\ufb01cient Algorithm for Inference in FHMMs\n\nTo arrive at our method we apply the results of the previous subsection to (2). To do so, as mentioned\nat the beginning of the section, we need to change the problem to a convex one, since the elements of\nthe second term in the objective of (2), x>t,i(Et,i + log Pi)xt+1,i are not convex. To address this\nissue, we relax the problem by introducing new variables Zt,i = xt,ix>t+1,i and replace the constraint\nZt,i = xt,ix>t+1,i with two new ones:\n\nZt,i1 = xt,i\n\nand Z>t,i1 = xt+1,i.\n\nTo simplify the presentation, we will assume that Ki = K for all i 2 [M ]. Then problem (2) becomes\n\narg min\nxt,i\n\nsubject to\n\nTXt=1\u21e2 1\n\n22yt  x>t \u00b52\n\n p>t zt\n\nxt 2{ 0, 1}M K,\n\u02c6zt 2{ 0, 1}M KK,\n1>xt,i = 1,\nZt,i1> = xt,i,\n\nt 2 [T ],\nt 2 [T  1],\n\nt 2 [T ] and i 2 [M ],\nZ>t,i1> = xt+1,i ,\n\n4\n\n(7)\n\nt 2 [T  1] and i 2 [M ],\n\n\fAlgorithm 1 ADMM-RR: Randomized rounding algorithm for suboptimal solution to (2)\n\nt\n\nt\n\n:= X\u21e4t for t = 1, . . . , T\n\nGiven: number of iterations: itermax, length of input data: T\nSolve the optimization problem (8): Run Algorithm 2 to get X\u21e4t and z\u21e4t\n:= z\u21e4t and X best\nSet xbest\nfor t = 2, . . . , T  1 do\nt1 >, xbest\nSet x := [xbest\nSet X := block(X best\narguments\nSet f best := 1\nForm the covariance matrix \u2303:= X  xxT and \ufb01nd its Cholesky factorization LL> =\u2303 .\nfor k = 1, 2, . . . , itermax do\n\nt >, xbest\nt1 , X best\n\nt+1 >]>\n\nt\n\n, X best\n\nt+1 ) where block(\u00b7,\u00b7) constructs block diagonal matrix from input\n\nRandom sampling: zk := x + Lw, where w \u21e0N (0, I)\nRound zk to the nearest integer point xk that satis\ufb01es the constraints of (7)\nIf f best > ft(xk) then update xbest\nrespectively\n\nand X best\n\nt\n\nt\n\nfrom the corresponding entries of xk and xkxk>,\n\nend for\n\nend for\n\nwhere x>t = [x>t,1, . . . , x>t,M ], \u00b5> = [\u00b5>1 , . . . , \u00b5>M ], z>t = [vec(Zt,1)>, . . . , vec(Zt,M )>] and\np>t = [vec(Et,1 + log P1), . . . , vec(log PT )], with vec(A) denoting the column vector obtained\nby concatenating the columns of A for a matrix A. Expanding the \ufb01rst term of (7) and following the\nrelaxation method of Section 3.1, we get the following SDP problem:2\n\narg min\nXt,zt\nsubject to\n\nTXt=1\n\ntrace(D>t Xt) + d>t zt\n\n(8)\n\n+\n\n+\n\nyt\u00b5\n\nBXt + Czt + EXt+1 = g,\nXt, zt  0 .\n\nAXt = b,\nXt \u232b 0,\n! Rm0 and C2 RM KK\u21e5m0 are all appropriate linear\n! Rm, B,E : SM K+1\nHere A : SM K+1\noperators, and the integers m and m0 are determined by the number of equality constraints, while\n\u00b5\u00b5>  and dt = pt. Notice that (8) is a simple, though huge-dimensional SDP\n22\uf8ff 0\nyt\u00b5>\nDt = 1\nproblem in the form of (5) where \u02c6D has a special block structure.\nNext we apply the randomized rounding method from Section 3.1 to provide an approximate solution\nto our original problem (2). Starting from an optimal solution (z\u21e4, X\u21e4) of (8) , and utilizing that\nwe have an SDP problem for each time step t, we obtain Algorithm 1 that performs the rounding\nsequentially for t = 1, 2, . . . , T . However we run the randomized method for three consecutive time\nsteps, since Xt appears at both time steps t  1 and t + 1 in addition to time t (cf., equation 9).\nFollowing Park and Boyd [2015], in the experiments we introduce a simple greedy search within\nAlgorithm 1: after \ufb01nding the initial point xk, we greedily try to objective the target value by change\nthe status of a single appliance at a single time instant. The search stops when no such improvement\nis possible, and we use the resulting point as the estimate.\n\n5 ADMM Solver for Large-Scale, Sparse Block-Structured SDP Problems\n\nGiven the relaxation and randomized rounding presented in the previous subsection all that remains\nis to \ufb01nd X\u21e4t , z\u21e4t to initialize Algorithm 1. Although interior point methods can solve SDP problems\nef\ufb01ciently, even for problems with sparse constraints as (4), the running time to obtain an \u270f optimal\nsolution is of the order of n3.5 log(1/\u270f) [Nesterov, 2004, Section 4.3.3], which becomes prohibitive\nin our case since the number of variables scales linearly with the time horizon T .\nAs an alternative solution, \ufb01rst-order methods can be used for large scale problems [Wen et al., 2010].\nSince our problem (8) is an SDP problem where the objective function is separable, ADMM is a\npromising candidate to \ufb01nd a near-optimal solution. To apply ADMM, we use the Moreau-Yosida\nquadratic regularization [Malick et al., 2009], which is well suited for the primal formulation we\n\n2The only modi\ufb01cation is that we need to keep the equality constraints in (7) that are missing from (3).\n\n5\n\n\fAlgorithm 2 ADMM for sparse SDPs of the form (8)\n\nGiven: length of input data: T , number of iterations: itermax.\nSet the initial values to zero. W 0\nSet \u00b5 = 0.001 {Default step-size value}\nfor k = 0, 1, . . . , itermax do\n\nt , S0 = 0, 0\n\nt = 0, \u232b0\n\nt , P 0\n\nfor t = 1, 2, . . . , T do\n\nt = 0, and r0\n\nt , h0\n\nt = 0\n\nUpdate P k\n\nt , W k\n\nt , k, Sk\n\nt , rk\n\nt , hk\n\nt , and \u232bk\n\nt , respectively, according to (11) (Appendix A).\n\nend for\n\nend for\n\nconsider. When implementing ADMM over the variables (Xt, zt)t, the sparse structure of our\nconstraints allows to consider the SDP problems for each time step t sequentially:\n\narg min\nXt,zt\nsubject to\n\ntrace(D>t Xt) + d>t zt\nAXt = b,\nBXt + Czt + EXt+1 = g,\nBXt1 + Czt1 + EXt = g,\nXt \u232b 0,\n\nXt, zt  0 .\n\nThe regularized Lagrangian function for (9) is3\n\n(9)\n\n(10)\n\nL\u00b5 =trace(D>X) + d>z +\n\n1\n2\u00b5kX  Sk2\n+ \u232b>(g B X C z E X+) + \u232b>\n trace(W >X)  trace(P >X)  h>z,\n\nF +\n\n1\n2\u00b5kz  rk2\n\n2 + >(b A X)\n\n (g B X C z E X)\n\nwhere , \u232b, W  0, P \u232b 0, and h  0 are dual variables, and \u00b5 > 0 is a constant. By taking the\nderivatives of L\u00b5 and computing the optimal values of X and z, one can derive the standard ADMM\nupdates, which, due to space constraints, are given in Appendix A. The \ufb01nal algorithm, which updates\nthe variables for each t sequentially, is given by Algorithm 2.\nAlgorithms 1 and 2 together give an ef\ufb01cient algorithm for \ufb01nding an approximate solution to (2) and\nthus also to the inference problem of additive FHMMs.\n\n6 Learning the Model\n\nThe previous section provided an algorithm to solve the inference part of our energy disaggregation\nproblem. However, to be able to run the inference method, we need to set up the model. To learn\nthe HMMs describing each appliance, we use the method of Kontorovich et al. [2013] to learn the\ntransition matrix, and the spectral learning method of Anandkumar et al. [2012] (following Mattfeld,\n2014) to determine the emission parameters.\nHowever, when it comes to the speci\ufb01c application of NILM, the problem of unknown, time-varying\nbias also needs to be addressed, which appears due to the presence of unknown/unmodeled appliances\nin the measured signal. A simple idea, which is also followed by Kolter and Jaakkola [2012], is to\nuse a \u201cgeneric model\u201d whose contribution to the objective function is downweighted. Surprisingly,\nincorporating this idea in the FHMM inference creates some unexpected challenges.4\nTherefore, in this work we come up with a practical, heuristic solution tailored to NILM. First we\nidentify all electric events de\ufb01ned by a large change yt in the power usage (using some ad-hoc\nthreshold). Then we discard all events that are similar to any possible level change \u00b5(i)\nm,k. The\nremaining large jumps are regarded as coming from a generic HMM model describing the unregistered\nappliances: they are clustered into K  1 clusters, and an HMM model is built where each cluster is\nregarded as power usage coming from a single state of the unregistered appliances. We also allow an\n\u201coff state\u201d with power usage 0.\n\n3We drop the subscript t and replace t + 1 and t  1 with + and  signs, respectively.\n4For example, the incorporation of this generic model breaks the derivation of the algorithm of Kolter and\n\nJaakkola [2012]. See Appendix B for a discussion of this.\n\n6\n\n\f7 Experimental Results\n\nWe evaluate the performance of our algorithm in two setups:5 we use a synthetic dataset to test\nthe inference method in a controlled environment, while we used the REDD dataset of Kolter and\nJohnson [2011] to see how the method performs on non-simulated, \u201creal\u201d data. The performance of\nour algorithm is compared to the structured variational inference (SVI) method of Ghahramani and\nJordan [1997], the method of Kolter and Jaakkola [2012] and that of Zhong et al. [2014]; we shall\nrefer to the last two algorithms as KJ and ZGS, respectively.\n\n7.1 Experimental Results: Synthetic Data\n\nThe synthetic dataset was generated randomly (the exact procedure is described in Appendix C).\nTo evaluate the performance, we use normalized disaggregation error as suggested by Kolter and\nJaakkola [2012] and also adopted by Zhong et al. [2014]. This measures the reconstruction error for\neach individual appliance. Given the true output yt,i and the estimated output \u02c6yt,i (i.e. \u02c6yt,i = \u00b5>i \u02c6xt,i),\nthe error measure is de\ufb01ned as\n\nNDE =qPt,i(yt,i  \u02c6yt,i)2/Pt,i (yt,i)2 .\n\nFigures 1 and 2 show the performance of the algorithms as the number HMMs (M) (resp., number of\nstates, K) is varied. Each plot is a report for T = 1000 steps averaged over 100 random models and\nrealizations, showing the mean and standard deviation of NDE. Our method, shown under the label\nADMM-RR, runs ADMM for 2500 iterations, runs the local search at the end of each 250 iterations,\nand chooses the result that has the maximum likelihood. ADMM is the algorithm which applies naive\nrounding. It can be observed that the variational inference method is signi\ufb01cantly outperformed by\nall other methods, while our algorithm consistently obtained better results than its competitors, KJ\ncoming second and ZGS third.\n\nNumber of states: 3; Data length T=1000; Number of samples: 100\n5\n\nNumber of states: 3; Data length T=1000; Number of samples: 100\n1\n\nADMM-RR\nKJ method\nADMM\nZGS method\n\nADMM-RR\nKJ method\nADMM\nVariational Approx.\nZGS method\n\nr\no\nr\nr\ne\nd\ne\nz\n\n \n\ni\nl\n\na\nm\nr\no\nN\n\n0.8\n\n0.6\n\n0.4\n\n0.2\n\n0\n\n-0.2\n\n2\n\n3\n\n4\n\n5\n\n6\n\n7\n\n8\n\n9\n\n2\n\n3\n\n4\n\n5\n\n6\n\n7\n\n8\n\n9\n\nFigure 1: Disaggregation error varying the number of HMMs.\n\nr\no\nr\nr\ne\nd\ne\nz\n\n \n\ni\nl\n\na\nm\nr\no\nN\n\n4\n\n3\n\n2\n\n1\n\n0\n\n-1\n\n3\n\nr\no\nr\nr\ne\n \nd\ne\nz\n\ni\nl\n\na\nm\nr\no\nN\n\n2.5\n\n2\n\n1.5\n\n1\n\n0.5\n\n0\n\n-0.5\n\nNumber of appliances: 5; Data length T=1000; Number of samples: 100\n\nADMM-RR\nKJ method\nADMM\nVariational Approx.\nZGS method\n\nNumber of appliances: 5; Data length T=1000; Number of samples: 100\n0.8\n\nADMM-RR\nKJ method\nADMM\nZGS method\n\nr\no\nr\nr\ne\n\n \n\nd\ne\nz\n\ni\nl\n\na\nm\nr\no\nN\n\n0.6\n\n0.4\n\n0.2\n\n0\n\n-0.2\n\n2\n\n3\n\n4\n\n5\n\n6\n\n2\n\n3\n\n4\n\n5\n\n6\n\nFigure 2: Disaggregation error varying the number of states.\n\n7.2 Experimental Results: Real Data\n\nIn this section, we also compared the 3 best methods on the real dataset REDD [Kolter and Johnson,\n2011]. We use the \ufb01rst half of the data for training and the second half for testing. Each HMM (i.e.,\n\n5Our code is available online at https://github.com/kiarashshaloudegi/FHMM_inference.\n\n7\n\n\fAppliance\n1 Oven-3\n2 Fridge\n3 Microwave\n4 Bath. GFI-12\n5 Kitch. Out.-15\n6 Wash./Dry.-20-A\n7 Unregistered-A\n8 Oven-4\n9 Dishwasher-6\n10 Wash./Dryer-10\n11 Kitch. Out.-16\n12 Wash./Dry.-20-B\n13 Unregistered-B\nAverage\n\nKJ method\n\nZGS method\n5.35/15.04%\n46.89/87.10%\n\nADMM-RR\n61.70/78.30% 27.62/72.32%\n90.22/97.63% 41.20/97.46%\n12.40/74.74%\n50.88/60.25% 12.87/51.46%\n69.23/98.85% 16.66/79.47%\n70.41/98.19%\n98.23/93.80%\n85.35/25.91%\n94.27/87.80%\n25.41/76.37%\n13.60/78.59%\n54.53/90.91%\n25.20/98.72%\n21.92/63.58% 18.63/25.79%\n8.87/100%\n17.88/79.04%\n72.13/77.10%\n98.19/28.31%\n97.78/91.73%\n96.92/73.97%\n60.97/78.56% 38.68/75.02%\n\n13.40/96.32% 4.55/45.07%\n6.16/42.67%\n5.69/26.72%\n15.91/35.51%\n57.43/99.31%\n9.52/12.05%\n29.42/31.01%\n7.79/3.01%\n0.00/0.00%\n27.44/71.25%\n33.63/99.98%\n17.97/36.22%\n\nTable 1: Comparing the disaggregation performance of three different algorithms: precision/recall.\nBold numbers represent statistically better performance on both measures.\n\nappliance) is trained separately using the associated circuit level data, and the HMM corresponding\nto unregistered appliances is trained using the main panel data. In this set of experiments we monitor\nappliances consuming more than 100 watts. ADMM-RR is run for 1000 iterations, and the local\nsearch is run at the end of each 250 iterations, and the result with the largest likelihood is chosen.\nTo be able to use the ZGS method on this data, we need to have some prior information about the\nusage of each appliance; the authors suggestion is to us national energy surveys, but in the lack of\nthis information (also about the number of residents, type of houses, etc.) we used the training data to\nextract this prior knowledge, which is expected to help this method.\nDetailed results about the precision and recall of estimating which appliances are \u2018on\u2019 at any given\ntime are given in Table 1. In Appendix D we also report the error of the total power usage assigned\nto different appliances (Table 2), as well as the amount of assigned power to each appliance as\na percentage of total power (Figure 3). As a summary, we can see that our method consistently\noutperformed the others, achieving an average precision and recall of 60.97% and 78.56%, with about\n50% better precision than KJ with essentially the same recall (38.68/75.02%), while signi\ufb01cantly\nimproving upon ZGS (17.97/36.22%). Considering the error in assigning the power consumption to\ndifferent appliances, our method achieved about 30  35% smaller error (ADMM-RR: 2.87%, KJ:\n4.44%, ZGS: 3.94%) than its competitors.\nIn our real-data experiments, there are about 1 million decision variables: M = 7 or 6 appliances\n(for phase A and B power, respectively) with K = 4 states each and for about T = 30, 000 time\nsteps for one day, 1 sample every 6 seconds. KJ and ZGS solve quadratic programs, increasing their\nmemory usage (14GB vs 6GB in our case). On the other hand, our implementation of their method,\nusing the commercial solver MOSEK inside the Matlab-based YALMIP [L\u00f6fberg, 2004], runs in 5\nminutes, while our algorithm, which is purely Matlab-based takes 5 hours to \ufb01nish. We expect that an\noptimized C++ version of our method could achieve a signi\ufb01cant speed-up compared to our current\nimplementation.\n\n8 Conclusion\n\nFHMMs are widely used in energy disaggregation. However, the resulting model has a huge\n(factored) state space, making standard inference FHMM algorithms infeasible even for only a\nhandful of appliances. In this paper we developed a scalable approximate inference algorithm, based\non a semide\ufb01nite relaxation combined with randomized rounding, which signi\ufb01cantly outperformed\nthe state of the art in our experiments. A crucial component of our solution is a scalable ADMM\nmethod that utilizes the special block-diagonal-like structure of the SDP relaxation and provides a\ngood initialization for randomized rounding. We expect that our method may prove useful in solving\nother FHMM inference problems, as well as in large scale integer quadratic programming.\n\nAcknowledgements\n\nThis work was supported in part by the Alberta Innovates Technology Futures through the Alberta Ingenuity\nCentre for Machine Learning and by NSERC. K. is indebted to Pooria Joulani and Mohammad Ajallooeian,\nwhom provided much useful technical advise, while all authors are grateful for Zico Kolter for sharing his code.\n\n8\n\n\fReferences\nA. Anandkumar, D. Hsu, and S. M. Kakade. A Method of Moments for Mixture Models and Hidden Markov\n\nModels. In COLT, volume 23, pages 33.1\u201333.34, 2012.\n\nS. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein. Distributed Optimization and Statistical Learning via\n\nthe Alternating Direction Method of Multipliers. FTML, 3(1):1\u2013122, 2011.\n\nS. Burer and D. Vandenbussche. Solving Lift-and-Project Relaxations of Binary Integer Programs. SIAM Journal\n\non Optimization, 16(3):726\u2013750, 2006.\n\nH.-H. Chang, K.-L. Chen, Y.-P. Tsai, and W.-J. Lee. A New Measurement Method for Power Signatures of\nNonintrusive Demand Monitoring and Load Identi\ufb01cation. IEEE T. on Industry Applications, 48:764\u2013771,\n2012.\n\nM. Dong, P. C. M. Meira, W. Xu, and W. Freitas. An Event Window Based Load Monitoring Technique for\n\nSmart Meters. IEEE Transactions on Smart Grid, 3(2):787\u2013796, June 2012.\n\nM. Dong, Meira, W. Xu, and C. Y. Chung. Non-Intrusive Signature Extraction for Major Residential Loads.\n\nIEEE Transactions on Smart Grid, 4(3):1421\u20131430, Sept. 2013.\n\nD. Egarter, V. P. Bhuvana, and W. Elmenreich. PALDi: Online Load Disaggregation via Particle Filtering. IEEE\n\nTransactions on Instrumentation and Measurement, 64(2):467\u2013477, 2015.\n\nM. Figueiredo, A. de Almeida, and B. Ribeiro. Home Electrical Signal Disaggregation for Non-intrusive Load\n\nMonitoring (NILM) Systems. Neurocomputing, 96:66\u201373, Nov. 2012.\n\nZ. Ghahramani and M. Jordan. Factorial Hidden Markov Models. Machine learning, 29(2):245\u2013273, 1997.\nM. X. Goemans and D. P. Williamson. Improved Approximation Algorithms for Maximum Cut and Satis\ufb01ability\n\nProblems Using Semide\ufb01nite Programming. J. of the ACM, 42(6):1115\u20131145, 1995.\n\nZ. Guo, Z. J. Wang, and A. Kashani. Home Appliance Load Modeling From Aggregated Smart Meter Data.\n\nIEEE Transactions on Power Systems, 30(1):254\u2013262, Jan. 2015.\n\nJ. Kelly and W. Knottenbelt. Neural NILM: Deep Neural Networks Applied to Energy Disaggregation. In\n\nBuildSys, pages 55\u201364, 2015.\n\nH. Kim, M. Marwah, M. F. Arlitt, G. Lyon, and J. Han. Unsupervised Disaggregation of Low Frequency Power\n\nMeasurements. In ICDM, volume 11, pages 747\u2013758, 2011.\n\nD. Koller and N. Friedman. Probabilistic graphical models: principles and techniques. Adaptive computation\n\nand machine learning. MIT Press, Cambridge, MA, 2009.\n\nJ. Z. Kolter and T. Jaakkola. Approximate Inference in Additive Factorial HMMs with Application to Energy\n\nDisaggregation. In AISTATS, pages 1472\u20131482, 2012.\n\nJ. Z. Kolter and M. J. Johnson. REDD: A Public Data Set for Energy Disaggregation Research. In Workshop on\n\nData Mining Applications in Sustainability (SIGKDD), pages 59\u201362, 2011.\n\nJ. Z. Kolter, S. Batra, and A. Y. Ng. Energy Disaggregation via Discriminative Sparse Coding. In Advances in\n\nNeural Information Processing Systems, pages 1153\u20131161, 2010.\n\nA. Kontorovich, B. Nadler, and R. Weiss. On Learning Parametric-Output HMMs. In ICML, pages 702\u2013710,\n\nJ. Liang, S. K. K. Ng, G. Kendall, and J. W. M. Cheng. Load Signature Study -Part I: Basic Concept, Structure,\n\nand Methodology. IEEE Transactions on Power Delivery, 25(2):551\u2013560, Apr. 2010.\n\nJ. L\u00f6fberg. YALMIP : A Toolbox for Modeling and Optimization in MATLAB. In CACSD, 2004.\nL. Lov\u00e1sz and A. Schrijver. Cones of Matrices and Set-functions and 0-1 Optimization. SIAM Journal on\n\nOptimization, 1(2):166\u2013190, 1991.\n\nJ. Malick, J. Povh, F. Rendl, and A. Wiegele. Regularization Methods for Semide\ufb01nite Programming. SIAM\n\nJournal on Optimization, 20(1):336\u2013356, Jan. 2009. ISSN 1052-6234, 1095-7189.\n\nC. Mattfeld. Implementing spectral methods for hidden Markov models with real-valued emissions. arXiv\n\n2013.\n\nY. Nesterov. Introductory Lectures on Convex Optimization: A Basic Course. Kluwer, 2004.\nJ. Park and S. Boyd. A Semide\ufb01nite Programming Method for Integer Convex Quadratic Minimization. arXiv\n\npreprint arXiv:1404.7472, 2014.\n\npreprint arXiv:1504.07672, 2015.\n\nA. Prudenzi. A neuron nets based procedure for identifying domestic appliances pattern-of-use from energy\n\nrecordings at meter panel. In PESW, volume 2, pages 941\u2013946, 2002.\n\nS. J. Rennie, J. R. Hershey, and P. Olsen. Single-channel speech separation and recognition using loopy belief\n\npropagation. In ICASSP, pages 3845\u20133848, 2009.\n\nM. J. Wainwright and M. I. Jordan. Graphical Models, Exponential Families, and Variational Inference. FTML,\n\nM. Weiss, A. Helfenstein, F. Mattern, and T. Staake. Leveraging smart meter data to recognize home appliances.\n\n1(1\u20132):1\u2013305, 2007.\n\nIn PerCom, pages 190\u2013197, 2012.\n\nZ. Wen, D. Goldfarb, and W. Yin. Alternating direction augmented Lagrangian methods for semide\ufb01nite\n\nprogramming. Mathematical Programming Computation, 2(3-4):203\u2013230, Dec. 2010.\n\nM. Zhong, N. Goddard, and C. Sutton. Signal Aggregate Constraints in Additive Factorial HMMs, with\n\nApplication to Energy Disaggregation. In NIPS, pages 3590\u20133598, 2014.\n\nT. Zia, D. Bruckner, and A. Zaidi. A hidden Markov model based procedure for identifying household electric\n\nloads. In IECON, pages 3218\u20133223, 2011.\n\n9\n\n\f", "award": [], "sourceid": 2533, "authors": [{"given_name": "Kiarash", "family_name": "Shaloudegi", "institution": "Imperial College London"}, {"given_name": "Andr\u00e1s", "family_name": "Gy\u00f6rgy", "institution": "Imperial College London"}, {"given_name": "Csaba", "family_name": "Szepesvari", "institution": "U. Alberta"}, {"given_name": "Wilsun", "family_name": "Xu", "institution": "University of Alberta"}]}