{"title": "Structured Local Minima in Sparse Blind Deconvolution", "book": "Advances in Neural Information Processing Systems", "page_first": 2322, "page_last": 2331, "abstract": "Blind deconvolution is a ubiquitous problem of recovering two unknown signals from their convolution. Unfortunately, this is an ill-posed problem in general. This paper focuses on the {\\em short and sparse} blind deconvolution problem, where the one unknown signal is short and the other one is sparsely and randomly supported. This variant captures the structure of the unknown signals in several important applications. We assume the short signal to have unit $\\ell^2$ norm and cast the blind deconvolution problem as a nonconvex optimization problem over the sphere. We demonstrate that (i) in a certain region of the sphere, every local optimum is close to some shift truncation of the ground truth, and (ii) for a generic short signal of length $k$, when the sparsity of activation signal $\\theta\\lesssim k^{-2/3}$ and number of measurements $m\\gtrsim\\poly\\paren{k}$, a simple initialization method together with a descent algorithm which escapes strict saddle points recovers a near shift truncation of the ground truth kernel.", "full_text": "Structured Local Minima in\nSparse Blind Deconvolution\n\nYuqian Zhang, Han-Wen Kuo, John Wright\n\nDepartment of Electrical Engineer and Data Science Institute\n\nColumbia University, New York, NY 10027\n{yz2409, hk2673, jw2966}@columbia.edu\n\nAbstract\n\nBlind deconvolution is a ubiquitous problem of recovering two unknown\nsignals from their convolution. Unfortunately, this is an ill-posed problem\nin general. This paper focuses on the short and sparse blind deconvolu-\ntion problem, where the one unknown signal is short and the other one\nis sparsely and randomly supported. This variant captures the structure\nof the unknown signals in several important applications. We assume the\nshort signal to have unit (cid:96)2 norm and cast the blind deconvolution problem\nas a nonconvex optimization problem over the sphere. We demonstrate that\n(i) in a certain region of the sphere, every local optimum is close to some\nshift truncation of the ground truth, and (ii) for a generic short signal of\nlength k, when the sparsity of activation signal \u03b8 (cid:46) k\u22122/3 and number of\nmeasurements m (cid:38) poly (k), a simple initialization method together with a\ndescent algorithm which escapes strict saddle points recovers a near shift\ntruncation of the ground truth kernel.\n\nIntroduction\n\n1\nBlind deconvolution is the problem of recovering two unknown signals a0 and x0 from their\nconvolution y = a0 \u2217 x0. This fundamental problem recurs across several \ufb01elds, including\nastronomy, microscopy data processing [1], neural spike sorting [2], computer vision [3], etc.\nHowever, this problem is ill-posed without further priors on the unknown signals, as there\nare in\ufb01nitely many pairs of signals (a, x) whose convolution equals a given observation y.\nFortunately, in practice, the target signals (a, x) are often structured. In particular, a number\nof practical applications exhibit a common short-and-sparse structure:\nIn Neural spike sorting: Neurons in the brain \ufb01re brief voltage spikes when stimulated. The\nsignatures of the spikes encode critical features of the neuron and the occurrence of such\nspikes are usually sparse and random in time [2, 4].\nIn Microscopy data analysis: Some nanoscale materials are contaminated by randomly and\nsparsely distributed \u201cdefects\u201d, which change the electronic structure of the material [1].\nIn Image deblurring: Blurred images due to camera shake can be modeled as a convolution of\nthe latent sharp image and a kernel capturing the motion of the camera. Although natural\nimages are not sparse, they typically have (approximately) sparse gradients [5, 6].\nIn the above applications, the observation signal y \u2208 Rm is generated via the convolution\nof a short kernel a0 \u2208 Rk(k (cid:28) m) and a sparse activation coe\ufb03cient x0 \u2208 Rm ((cid:107)x0(cid:107)0 (cid:28) m).\nWithout loss of generality, we let y denote the circular convolution of a0 and x0\n\ny = a0 (cid:126) x0 =(cid:102)a0 (cid:126) x0,\n\n(1)\n\n32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montr\u00e9al, Canada.\n\n\fFigure 1: Local Minimum.\nTop: observation y = a0 (cid:126) x0,\nand ground truth a0, and x0;\nBottom: recovered a (cid:126) x, a,\nand x at one local minimum of\na natural formulation in [16].\n\nwith(cid:102)a0 \u2208 Rm denoting the zero padded m-length version of a0, which can be expressed as\n(cid:102)a0 = \u03b9ka0. Here, \u03b9k : Rk \u2192 Rm is a zero padding operator. Its adjoint \u03b9\u2217\nk : Rm \u2192 Rk acts as\n\na projection onto the lower dimensional space by keeping the \ufb01rst k components.\nThe short-and-sparse blind deconvolution problem exhibits a scaled-shift ambiguity, which\nderives from the basic properties of a convolution operator. Namely, for any observation\nsignal y, and any nonzero scalar \u03b1 and integer shift \u03c4, the following equality always holds\n(2)\n\ny = (\u00b1\u03b1s\u03c4 [(cid:102)a0]) (cid:126)(cid:0)\u00b1\u03b1\u22121s\u2212\u03c4 [x0](cid:1) .\n\nHere, s\u2212\u03c4 [v] denotes the cyclic shift of the vector v by \u03c4 entries:\n\ns\u03c4 [v](i) = v ([i \u2212 \u03c4 \u2212 1]m + 1) ,\n\n\u2200 i \u2208 {1,\u00b7\u00b7\u00b7 , m} .\n\nks\u03c4 [(cid:102)a0] (cid:126) s\u2212\u03c4 [x0] \u2248 y.\n\nks\u03c4 [(cid:102)a0] can be convolved with\n\n(3)\nClearly, both scaling and cyclic shifts preserve the short-and-sparse structure of (a0, x0). This\nscaled-shift symmetry raises nontrivial challenges for computation, making straightforward\nconvexi\ufb01cation approaches ine\ufb00ective.1\nNonconvex algorithms for sparse blind deconvolution have been well developed and prac-\nticed, especially in computer vision [12, 13, 14, 15]. Despite its empirical success, little was\nknown about its working mechanism. Recently, [16] studies the optimization landscape of\nthe natural nonconvex formulation for sparse blind deconvolution, assuming the kernel\na \u2208 Rk to have unit Frobenius norm (denote as a \u2208 Sk\u22121). [16] argues that under conditions,\nthis problem has well-structured local optima, in the sense that every local optimum is close to\nsome shift truncation of the ground truth (Figure 1).\nThe presence of these local optima can be viewed as a result of the shift symmetry associated\nto the convolution operator: the shifted and truncated kernel \u03b9\u2217\nthe sparse signal s\u2212\u03c4 [x0] (shifted in the other direction) to produce a near approximation to\nthe observation \u03b9\u2217\nIn [16], this geometric insight about local optima is corroborated with a lot of experiments,\nbut rigorous proof is only available in the \u201cdilute limit\u201d in which the sparse coe\ufb03cient signal\nx0 is a single spike. In this paper, we adopt the unit Frobenius norm constraint for the short\nconvolution kernel a as in [16], but consider a di\ufb00erent objective function. We formulate the\nsparse blind deconvolution problem as the following optimization problem over the sphere:\n(4)\nHere, \u02c7y denotes the reversal of y2 and ry (q) is a preconditioner which we will discuss in\ndetail later. Convolution y (cid:126) ry (q) approximates the reversed underlying activation signal\nx0, and \u2212(cid:107)\u00b7(cid:107)4\nWe demonstrate that even when x0 is relatively dense, any local minimum in certain region\nof the sphere is close to a shift truncation \u03b9\u2217\ncontains the sub-level set of small objective value. Algorithmically, if initialized at a point\nwith small enough objective value, then a descent algorithm always decreases the objective\nvalue and hence stays in this region. Speci\ufb01cally, for a generic kernel3 a0 \u2208 Sk\u22121, if the\n1A number of works [7, 8, 9, 10, 11] have developed provable methods for blind deconvolution\nunder the assumption that a0 and x0 belong to random subspaces, or are sparse in random dictionaries.\nThese random models exhibit simpler geometry than the short-and-sparse model. Because our target\nsignal is sparse in the standard basis, the aforementioned results are not applicable in our setting.\n2Denote y = [y1, y2,\u00b7\u00b7\u00b7 , ym\u22121, ym]T , then its reversal \u02c7y = [y1, ym, ym\u22121,\u00b7\u00b7\u00b7 , y2]T with y1 not\n3In this paper, we refer a kernel sampled following a uniform distribution over the sphere as a\n\nks\u03c4 [(cid:102)a0] of the ground truth. This benign region\n\n4 serves as the sparsity penalty.\n\nmin \u2212(cid:107) \u02c7y (cid:126) ry (q)(cid:107)4\n\n4\n\ns. t.\n\n(cid:107)q(cid:107)F = 1.\n\nmoved.\n\ngeneric kernel on the sphere.\n\n2\n\nConvolution YKernel AActivation X00.511.5200.20.40.60.811.21.41.61.82\fsparsity rate4 \u03b8 (cid:46) k\u22122/3 and the number of measurement m (cid:38) poly(k), initializing at some\npreconditioned k consecutive entries of y, and applying any descent method that converges\nto a local minimizer under a strict saddle hypothesis [17, 18], produces a near shift-truncation\nof the ground truth.5\nAssumptions and Notations We assume that x0 \u2208 Rm follows Bernoulli-Gaussian (BG)\nmodel with sparsity level \u03b8: x0 (i) = \u03c9igi with \u03c9i \u223c Ber (\u03b8) and gi \u223c N (0, 1), where all the\ndi\ufb00erent random variables are jointly independent. For simplicity, we write x0 \u223ci.i.d. BG (\u03b8).\nThroughout this paper, a vector v \u2208 Rk is indexed as v = [v1, v2,\u00b7\u00b7\u00b7 , vk], and [\u00b7]m denotes the\nmodulo operator of m. We use (cid:107)\u00b7(cid:107)op, (cid:107)\u00b7(cid:107)F , and (cid:107)\u00b7(cid:107)p to denote operator norm, Frobenius norm,\nand entry wise (cid:96)p norm respectively. PS [\u00b7]\ndenotes projection onto the Frobenius\n\u25e6p is the entry wise p-th order exponent operator. We use C, c to denote positive\nsphere. (\u00b7)\nconstants, and their value change across the paper.\n2 Problem Formulation\nIn the short-and-sparse blind deconvolution problem, any k consecutive entries in y only\ndepend on 2k \u2212 1 consecutive entries in x0:\n\n= \u00b7\n.\n(cid:107)\u00b7(cid:107)F\n\n(cid:3)T\n\n=\n\nk\u22121(cid:88)\n\nks\u03c4 [(cid:102)a0]\n\n\u00b7 \u03b9\u2217\n\nx1+[i+\u03c4\u22121]m\n\nyi =(cid:2)yi,\u00b7\u00b7\u00b7 , y1+[i+k\u22121]m\n\uf8ee\uf8ef\uf8ef\uf8ef\uf8ef\uf8f0\n\n=\n\nak\u22121\nak\n...\n0\n0\n\nak\n0\n...\n0\n0\n\na1\na2\n...\nak\u22121\nak\n\n\u00b7\u00b7\u00b7\n\u00b7\u00b7\u00b7\n. . .\n\u00b7\u00b7\u00b7\n\u00b7\u00b7\u00b7\nA0\u2208Rk\u00d7(2k\u22121)\n\n(cid:123)(cid:122)\n\n(cid:124)\n\n\u03c4 =\u2212(k\u22121)\n\u00b7\u00b7\u00b7\n\u00b7\u00b7\u00b7\n. . .\n\u00b7\u00b7\u00b7\n\u00b7\u00b7\u00b7\n\n0\n0\n0\n0\n...\n...\na1\n0\na2 a1\n\n\uf8f9\uf8fa\uf8fa\uf8fa\uf8fa\uf8fb\n\n(cid:125)\n\n\uf8ee\uf8ef\uf8ef\uf8ef\uf8ef\uf8ef\uf8ef\uf8f0\n\n(cid:124)\n\n\uf8f9\uf8fa\uf8fa\uf8fa\uf8fa\uf8fa\uf8fa\uf8fb\n\n(cid:125)\n\n.\n\nx1+[i\u2212k]m\n\n...\nxi\n...\n\n(cid:123)(cid:122)\n\nx1+[i+k\u22122]m\nxi\u2208R(2k\u22121)\u00d71\n\n(5)\n\n(6)\n\n(8)\n\nWrite Y = [y1, y2, . . . , ym] \u2208 Rk\u00d7m and X0 = [x1, . . . , xm] \u2208 R2k\u22121\u00d7m. Using the above\nexpression, we have that\n(7)\nEach column xi of X0 only contains some 2k \u2212 1 entries of x0. The rows of X0 are cyclic\nshifts of the reversal of x0:\n\n(cid:34) s0[ \u02c7x0]\n\nY = A0X0.\n\n(cid:35)\n\nX0 =\n\n...\n\ns2k\u22122[ \u02c7x0]\n\n.\n\nmin\n\nThe shifts of \u02c7x0 are sparse vectors in the linear subspace row(X0). Note that if we could\nrecover some shift s\u03c4 [x0], we could subsequently determine s\u2212\u03c4 [a0] by solving a linear\nsystem of equations, and hence solve the deconvolution problem, up to the shift ambiguity.6\n2.1 Finding a Shifted Sparse Signal\nIn light of the above observations, a natural computational approach to sparse blind de-\nconvolution is to attempt to \ufb01nd x0 by searching for a sparse vector in the linear subspace\nrow(X0), e.g., by solving an optimization problem\n\nv \u2208 row (X0) , (cid:107)v(cid:107)2 = 1,\n\nof the observation.\n\n(9)\n4This equivalently says there could be as many as O(k1/3) shifts of the kernel in a k-length window\n5[16] proposes to solve the short-and-sparse blind deconvolution problem with a two phase al-\ngorithm which \ufb01rst recovers a shift truncation, and then recovers the ground truth kernel with an\nannealing algorithm. We present additional experimental results on the recovery of the ground truth\nin the supplementary material.\n6[19] considers the multi-channel blind deconvolution problem, where many independent observa-\ntions yp = a0 \u2217 xp are available. [19] shows how to formulate this problem as searching for a sparse\nvector in a linear subspace. Our approach is also inspired by the idea of looking for a sparse/spiky\nvector in a subspace. However, it pertains to a di\ufb00erent problem, in which only a single observation is\navailable. The short and sparse problem exhibits a more complicated optimization landscape, due to\nthe signed shift ambiguity.\n\n(cid:107)v(cid:107)(cid:63)\n\ns. t.\n\n3\n\n\fwhere (cid:107)\u00b7(cid:107)(cid:63) is chosen to encourage sparsity of the target signal [20, 21, 22, 23].\nIn sparse blind deconvolution, we do not have access to the row space of X0. Instead, we only\nobserve the subspace row(Y ) \u2282 row(X0). The subspace row(Y ) does not necessarily contain\nthe desired sparse vector eT\ni X0, but it does contain some approximately sparse vectors. In\nparticular, consider following vector in row(Y ),\n\nv = Y T a0 = \u02c7x0\nsparse\n\n+\n\n(cid:104)a0, si[a0](cid:105) si[ \u02c7x0]\n\n.\n\n(10)\n\n(cid:88)\n(cid:124)\n\ni(cid:54)=0\n\n(cid:123)(cid:122)\n\n\u201cnoise\u201d z\n\n(cid:125)\n\n0\n\n0\n\nmin \u03c8 (q)\n\n0 (A0AT\n\n(12)\n\n(14)\n\nq \u2248 el,\n\n0\n\nq\n\n4\n\n\u03b6 = AT\n0\n\n.\n\nA\n\nq\n\n4\n\ns. t.\n\n= \u2212 1\n.\n4m\n\nq\n\n= \u2212 1\n4m\n\n4\n\n\u223c (cid:107) \u02c7x0 (cid:126) \u03b6(cid:107)4\n\nmin \u2212 1\n\n4 (cid:107)v(cid:107)4\n\n4\n\nl \u2208 {1,\u00b7\u00b7\u00b7 , 2k \u2212 1} .\n\nv \u2208 row (Y ) , (cid:107)v(cid:107)2 = 1.\n\n(11)\nq, with (cid:107)v(cid:107)2 = (cid:107)q(cid:107)2.\n\nThe vector v is a superposition of a sparse signal \u02c7x0 and its scaled shifts (cid:104)a0, si[a0](cid:105) si[ \u02c7x0].\nIf the shift-coherence |(cid:104)a0, s\u03c4 [a0](cid:105)| is small7 and x0 is sparse enough, z can be viewed as\nsmall noise.8 The vector v is not sparse, but it is spiky: a few of its entries are much larger\nthan the rest. We deploy a milder sparsity penalty \u2212(cid:107)\u00b7(cid:107)4\n4 to recover such a spiky vector, as\n(cid:107)\u00b7(cid:107)4\n\n4 is very \ufb02at around 0 and insensitive to small noise in the signal.9 This gives\n\nFor simplicity, we de\ufb01ne the preconditioned convolution matrix\n\nThis leads to the following equivalent optimization problem over the sphere\n(cid:107)q(cid:107)2 = 1.\n\nWe can express a generic unit vector v \u2208 row(Y ) as v = Y T(cid:0)Y Y T(cid:1)\u22121/2\n(cid:13)(cid:13)(cid:13)4\n\n(cid:13)(cid:13)(cid:13)4\n(cid:13)(cid:13)(cid:13)Y T(cid:0)Y Y T(cid:1)\u22121/2\n(cid:13)(cid:13)(cid:13)4\n(cid:13)(cid:13)(cid:13) \u02c7x0 (cid:126) AT\n(cid:13)(cid:13)(cid:13) \u02c7y (cid:126)(cid:0)Y Y T(cid:1)\u22121/2\n(cid:0)Y Y T(cid:1)\u22121/2\n(cid:0)A0AT\n(cid:1)\u22121/2\n=(cid:0)A0AT\n(cid:1)\u22121/2\n\nInterpretation: preconditioned shifts. This objective \u03c8 (q) can be rewritten as\n\u03c8 (q) = \u2212 1\n4 , (13)\n4m\n0 )\u22121/2q. This approximation becomes accurate as m grows.10 This\nwhere \u03b6 = AT\nobjective encourages the convolution of \u02c7x0 and \u03b6 to be as spiky as possible. Reasoning\nanalogous to (10) suggests that \u02c7x0 (cid:126) \u03b6 will be spiky if\n\n(15)\n= maxi(cid:54)=j |(cid:104)ai, aj(cid:105)|. Then \u03b6 can\n.\nwith column coherence (preconditioned shift coherence) \u00b5\nalso be interpreted as measuring the inner products of q with columns of A. Making this\nintuition rigorous, we will show that minimizing this objective over a certain region of the\nsphere yields a preconditioned shift truncate al, from which we can recover a shift truncate\nof the original signal a0.\n2.2 Structured Local Minima\nWe will show that in a certain region RC(cid:63) \u2282 Sk\u22121, the precondi-\ntioned shift truncations al are the only local minimizers. Moreover,\nthe other critical points in RC(cid:63) can be interpreted as resulting from\ncompetition between several of these local minima (Figure 2). At\nany saddle point, there exists strict negative curvature in the direc-\ntion of a nearby local minimizer which breaks the balance in favor\nof some particular al. The region RC(cid:63) is de\ufb01ned as follows:\nDe\ufb01nition 2.1. For \ufb01xed C(cid:63) > 0, letting \u03ba denote the condition number\n= maxi(cid:54)=j |(cid:104)ai, aj(cid:105)| the column coherence of A, we de\ufb01ne\n.\nof A0, and \u00b5\n(cid:110)\nq \u2208 Sk\u22121|(cid:13)(cid:13)AT q(cid:13)(cid:13)6\ntwo regions RC(cid:63), \u02c6RC(cid:63) \u2282 Sk\u22121, as\n(cid:110)\nq \u2208 Sk\u22121|(cid:13)(cid:13)AT q(cid:13)(cid:13)6\n\n(17)\n\u221a\n7For a generic kernel a0, the shift-coherence is bounded by |(cid:104)a0, s\u03c4 [a0](cid:105)| \u2248 1/\nk for any shift \u03c4.\n8In particular, under a Bernoulli-Gaussian model, for each j, E[z2\n9In comparison, the classical choice (cid:107)\u00b7(cid:107)(cid:63) = (cid:107)\u00b7(cid:107)1 is a strict sparsity penalty that essentially encourages\n10As Ex0\u223ci.i.d.BG(\u03b8)[Y Y T ] = Ex0\u223ci.i.d.BG(\u03b8)[A0X0X T\n\n(cid:111)\n4 \u2265 C(cid:63)\u00b5\u03ba2(cid:13)(cid:13)AT q(cid:13)(cid:13)3\n4 \u2265 C(cid:63)\u00b5\u03ba2(cid:111) \u2286 RC(cid:63) .\nj ] = \u03b8(cid:80)\n\nFigure 2: Saddles points\nare approximately super-\npositions of local minima.\n\nall small entries to be 0.\n\ni(cid:54)=0 (cid:104)a0, si[a0](cid:105)2.\n\nRC(cid:63)\n\u02c6RC(cid:63)\n\n.\n=\n\n.\n=\n\n\u00b7\u00b7\u00b7 a2k\u22121] ,\n\n0 AT\n\n0 ] = \u03b8mA0AT\n0 .\n\nA0 = [a1 a2\n\n.\n\n3\n\n(16)\n\ns. t.\n\n4\n\n\u03b9\u2217sj[a0]\u03d5(a)\u03b9\u2217si[a0]\u2212\u03c080\u03c08\fcan be viewed as a sub-level set for \u2212(cid:13)(cid:13)AT q(cid:13)(cid:13)4\n\nA simpler and smaller region \u02c6RC(cid:63) is also introduced in De\ufb01nition (2.1). This region \u02c6RC(cid:63)\n4, which is proportional to the objective value\n\u03c8 (q) assuming m is su\ufb03ciently large11. Therefore, once initialized within \u02c6RC(cid:63), the iterates\nproduced by a descent algorithm will stay in \u02c6RC(cid:63).\nIn particular, at any stationary point q \u2208 R10, the local optimization landscape can be\ncharacterized in terms of the number of spikes (entries with nontrivial magnitude12) in \u03b6. If\nthere is only one spike in \u03b6, then such stationary point q is a local minimum that is close to\none local minimizer; if there are more than two spikes in \u03b6, then such stationary point q is\nsaddle point. Based on the above characterizations of stationary points in RC(cid:63) with C(cid:63) \u2265 10,\nwe can deduce that any local minimum is close to al for some integer l, a preconditioned\nshift truncation of the ground truth a0.\nTheorem 2.2 (Main Result). Assuming observation y \u2208 Rm is the circulant convolution of\na0 \u2208 Rk and x0 \u223ci.i.d. BG (\u03b8) \u2208 Rm, where the convolutional matrix A0 has minimum singular\nvalue \u03c3min > 0 and condition number \u03ba \u2265 1, and A has column incoherence 0 \u2264 \u00b5 < 1. There\nexists a positive constant C such that whenever the number of measurements\n\nmin(cid:8)\u00b5\u22124/3, \u03ba2k2(cid:9)\n\n(1 \u2212 \u03b8)2 \u03c32\n\nmin\n\nm \u2265 C\n\n(cid:18)\n\n(cid:19)\n\n\u03ba8k4 log3\n\n\u03bak\n\n(1 \u2212 \u03b8) \u03c3min\n\n(18)\n\nand \u03b8 \u2265 log k/k, then with high probability, any local optima \u00afq \u2208 \u02c6R2C(cid:63) satis\ufb01es\n\n(19)\n\n|(cid:104)\u00afq,PS [al](cid:105)| \u2265 1 \u2212 c(cid:63)\u03ba\u22122\nfor some integer 1 \u2264 l \u2264 2k \u2212 1. Here, C(cid:63) \u2265 10 and c(cid:63) = 1/C(cid:63).\nThis theorem says that any local minimum in \u02c6R2C(cid:63) is close to some normalized column of\nA given polynomially many observation. The parameters \u03c3min, \u03ba and \u00b5 e\ufb00ectively measure\nthe spectrum \ufb02atness of the ground truth kernel a0 and characterize how broad the results\nhold. A random like kernel usually has big \u03c3min, small \u03ba and \u00b5, which equivalently implies\nthe result holds in a large sub-level set \u02c6R2C(cid:63) even with fewer observations.\nHence, once assuring the algorithm \ufb01nds a local minimum in \u02c6R2C(cid:63), then some shifted\ntruncation of the ground truth kernel a0 can be recovered. In other words, if we can \ufb01nd\nan initialization point with small objective value then a descent algorithm minimizing the\nobjective function guarantees that q always stays in \u02c6R2C(cid:63) in proceeding iterations. Therefore,\nany descent algorithm that escapes a strict saddle point can be applied to \ufb01nd some al, or\nsome shift truncation of a0.\n2.3\nRecall that yi = A0xi, which is a sparse superposition of about 2\u03b8k columns of A0. Intuitively\nspeaking, such qinit already encodes certain preferences towards a few preconditioned shift\ntruncations of the ground truth. Therefore, we randomly choose an index i and set the\ninitialization point as\nqinit = PS\n\n\u03b6init = AT qinit \u2248 PS(cid:2)AT Axi\n\nxi can be approximately preserved, that PS(cid:2)AT Axi\n\n(20)\nFor a generic kernel a0 \u2208 Sk\u22121, AT A is close to a diagonal matrix, as the magnitudes of\no\ufb00-diagonal entries are bounded by column incoherence \u00b5. Hence, the sparse property of\n4. By\nleveraging the sparsity level \u03b8, one can make sure such initialization point qinit falls in \u02c6R2C(cid:63).\nTherefore, we propose Algorithm 1 for solving sparse blind deconvolution with its working\nconditions stated in Corollary 2.3. For the choice of descent algorithms which escape strict\nsaddle points, there are several such algorithms specially tailored for sphere constrained\noptimization problems [24, 25].\n\n(cid:3) is spiky vector with small \u2212(cid:107)\u00b7(cid:107)4\n\n(cid:104)(cid:0)Y Y T(cid:1)\u22121/2\n\nInitialization with a Random Sample\n\n(cid:3) .13\n\n(cid:105)\n\nyi\n\n,\n\n11Please refer to Section 3 for more arguments.\n12We call any \u03b6l with magnitude no smaller than 2\u00b5(cid:107)\u03b6(cid:107)3\n13As Ex0\u223ci.i.d.BG(\u03b8)[Y Y T ] = \u03b8mA0AT\n0 .\n\nreasonings to later sections.\n\n3 /(cid:107)\u03b6(cid:107)4\n\n4 to be nontrivial and defer technical\n\n5\n\n\fAlgorithm 1 Short and Sparse Blind Deconvolution\nInput: Observations y \u2208 Rm and kernel size k.\nOutput: Recovered Kernel \u00afa.\n1: Generate random index i \u2208 [1, m] and set qinit = PS\n2: Solve following nonconvex optimization problem with a descent algorithm that escapes\n3: Set \u00afa = PS\n\nsaddle point and \ufb01nd a local minimizer \u00afq = arg minq\u2208Sk\u22121 \u03d5 (q)\n\n(cid:104)(cid:0)Y Y T(cid:1)\u22121/2\n\n(cid:104)(cid:0)Y Y T(cid:1)1/2\n\n(cid:105)\n\n(cid:105)\n\nyi\n\n\u00afq\n\n.\n\n.\n\n1\n\nCorollary 2.3. Suppose the ground truth a0 kernel has preconditioned shift coherence 0 \u2264 \u00b5 \u2264\n\u22123/2 (k) and sparse coe\ufb03cient x0 \u223ci.i.d. BG (\u03b8) \u2208 Rm. There exist positive constants\n8\u00d748 log\nC \u2265 25604 and C(cid:48) such that whenever the sparsity level\n4 \u2212 640\n64k\u22121 log k \u2264 \u03b8 \u2264 min\n(cid:19)\n\n(cid:1)(cid:0)3C(cid:63)\u00b5\u03ba2(cid:1)\u22122/3\nmin(cid:8)\u00b5\u22124/3, \u03ba2k2(cid:9)\n\n(cid:110) 1\nk3(cid:0)1 + 36\u00b52k log k(cid:1)4\n\n\u22122 k,(cid:0) 1\n(cid:18) \u03bak\n\nk\u22121(cid:0)1 + 36\u00b52k log k(cid:1)\u22122(cid:111)\n\nand signal length\nm \u2265C(cid:48) max\n\n482 \u00b5\u22122k\u22121 log\n\n\u03ba8k4 log3\n\n(cid:40)\n\n(cid:18)\n\nlog\n\nC1/4\n\n\u03bak\n\n,\n\n(cid:19)(cid:41)\n\n,\n\n\u03c3min\n\n(1 \u2212 \u03b8)2 \u03c32\n\nmin\n\n(1 \u2212 \u03b8) \u03c3min\n\n\u03b82\u03ba6\n\u03c32\n\nmin\n\nthen with high probability, Algorithm 1 recovers \u00afa such that\n\u221a\n\n(cid:107)\u00afa \u00b1 PS [\u03b9ks\u03c4 [(cid:102)a0]](cid:107)2 \u2264 2\n\n2c(cid:63)\n\nfor some integer shift \u2212 (k \u2212 1) \u2264 \u03c4 \u2264 k \u2212 1.\nFor a generic a0 \u2208 Sk\u22121, plugging in the numerical estimation of the parameters \u03c3min, \u03ba and\n\u00b5 (Figure 3), accurate recovery can be obtained with m (cid:38) \u03b82k6 poly log (k) measurements\nand sparsity level \u03b8 (cid:46) k\u22122/3 poly log (k). For bandpass kernels a0, \u03c3min is smaller and \u03ba, \u00b5\nare larger, and so our results require x0 to be longer and sparser.\n3 Optimization Function Landscape\nWe next brie\ufb02y present the key elements in deriving the main results of this paper. We \ufb01rst\ninvestigate the stationary points of the \u201cpopulation\u201d objective Ex0 [\u03c8(q)]. We demonstrate\nthat any local minimizer in RC(cid:63) is close to a signed column of A, a preconditioned shift\ntruncation of a0. We then demonstrate that when m is su\ufb03ciently large, the \u201c\ufb01nite sample\u201d\nobjective \u03c8(q) has similar properties.\nUsing E[Y Y T ] = \u03b8mA0AT\napproximated as follows:\n\u2212 1\nm\n\n0 again, the expectation of the objective function \u03c8 (q) can be\n\n(cid:13)(cid:13)(cid:13)Y T(cid:0)\u03b8mA0AT\n\n(cid:13)(cid:13)AT q(cid:13)(cid:13)4\n\n= \u2212 3 (1 \u2212 \u03b8)\n\nE [\u03c8(q)] \u2248 E\n\n4 \u2212 3\nm2 .\n\n(22)\n\n(cid:21)\n\n(cid:20)\n\n\u03b8m2\n\n4\n\n0\n\nThis approximation can be made rigorous (see Lemma 2.1 of the supplementary material),\nallowing us to study the critical points of E[\u03c8] by studying the simpler problem\n\nq\n\n(cid:13)(cid:13)(cid:13)4\n(cid:1)\u22121/2\n(cid:13)(cid:13)AT q(cid:13)(cid:13)4\n\n(21)\n\n(23)\n\n(24)\n\nmin\nq\u2208Rk\u22121\n\n\u03d5 (q)\n\n= \u2212 1\n.\n4\n\n4 = \u2212 1\n\n4\n\n(cid:107)\u03b6(cid:107)4\n4 .\n\nThe Euclidean gradient and Riemannian gradient [26] of \u03d5 are\n\n\u2207\u03d5(q) = \u2212A\u03b6\u25e63,\n\ngrad [\u03d5] (q) = \u2212A\u03b6\u25e63 + q (cid:107)\u03b6(cid:107)4\n4 .\n\n3.1 Critical Points of the Population Objective\nWe wish to argue that every local minimizer of \u03d5 is close to a preconditioned shift-truncation\nai. We do this by showing that at any other critical point, there is a direction of strict negative\ncurvature. We will show that at any critical point q \u2208 R4, the correlation \u03b6 exhibits a very\nspecial structure:\n\n6\n\n\f(P) The entries \u03b6i = (cid:104)ai, q(cid:105) are either close to zero, or have magnitude |\u03b6i|\nclose to (cid:107)\u03b6(cid:107)4\n\n4 /(cid:107)ai(cid:107)2.\n\n+\n\n\u03b1i\n\n2\n\n\u03b2i\n\nj(cid:54)=i\n\n4 = 0.\n\nj\n\n= 0.\n\n(26)\n\n2 \u03b6 3\n\ni +\n\n(cid:107)\u03b6(cid:107)4\n(cid:107)ai(cid:107)2\n\n4\n\n2\n\ni \u2212 \u03b6i\n\n(cid:88)\n\n(cid:104)ai, aj(cid:105) \u03b6 3\n\nj \u2212 \u03b6i (cid:107)\u03b6(cid:107)4\n\nA\u03b6\u25e63 \u2212 q (cid:107)\u03b6(cid:107)4\n\nWe can demonstrate this property directly from the stationarity condition grad [\u03d5] (q) = 0.\n(25)\n\n4 = 0 \u21d2 AT A\u03b6\u25e63 \u2212 AT q (cid:107)\u03b6(cid:107)4\n(cid:80)\nj(cid:54)=i (cid:104)ai, aj(cid:105) \u03b6 3\n(cid:124)\n(cid:125)\n\nThe i-th entry \u03b6i of the correlation \u03b6 therefore satis\ufb01es the following cubic equation\n(cid:107)ai(cid:107)2\n\n(cid:123)(cid:122)\n(cid:107)ai(cid:107)2\n\u03b1i (cid:29) \u03b2i obtains whenever(cid:13)(cid:13)AT q(cid:13)(cid:13)6\nIf \u03b1i (cid:29) \u03b2i, the roots of (26) are either very close to 0, or very close to \u00b1\u221a\n\n\u03b1i. The condition\n3, and hence on R4, every critical point\n\n4 \u2265 4\u00b5(cid:13)(cid:13)AT q(cid:13)(cid:13)3\n\n4 = 0 \u21d2 \u03b6 3\n\n(cid:124) (cid:123)(cid:122) (cid:125)\n\n3 /(cid:107)\u03b6(cid:107)4\n\n3 /(cid:107)\u03b6(cid:107)4\n4.\n\nsatis\ufb01es property (P).\n3.2 Asymptotic Function Landscape on RC(cid:63)\nThe local optimization landscape around any stationary point q is characterized by the\nRiemannian Hessian. In particular, at a stationary point q, if Hess [\u03d5] (q) is positive semidef-\ninite, then the function is convex and q is a local minimum; if Hess [\u03d5] (q) has a negative\neigenvalue, then there exists a direction along which the objective value decreases and q is a\nsaddle point. Technically, on RC(cid:63) with C(cid:63) \u2265 10, the minimum eigenvalue of the Riemannian\nHessian can be controlled based on the spikiness of \u03b6.\nFirst, we demonstrate that once constrained in RC(cid:63) with C(cid:63) \u2265 10, then any stationary point\nmust have cross correlation \u03b6 with entries of nontrivial magnitude, or entries of \u03b6 cannot\nbe simultaneously close to 0. Geometrically, this implies that any stationary point q \u2208 RC(cid:63)\nshould be \"close\" to certain preconditioned shift truncations.\nLemma 3.1. For any stationary point q \u2208 RC(cid:63) with C(cid:63) \u2265 10, magnitude of vector \u03b6 = AT q\ncannot be uniformly bounded by 2\u00b5(cid:107)\u03b6(cid:107)3\nLocal Minima\ngle entry \u03b6l with magnitude larger than 2\u00b5(cid:107)\u03b6(cid:107)3\nHess[\u03d5] (q) is always positive de\ufb01nite, and the function is locally convex.\n\nIf q is a stationary point in RC(cid:63) with C(cid:63) \u2265 10, and \u03b6 only has one sin-\n4, then the Riemannian Hessian\nIn addition,\n\n(cid:63) \u03ba\u22122(cid:1)(cid:107)al(cid:107)2, hence such q is one local minimum near al.\n\n|(cid:104)q, al(cid:105)| >(cid:0)1 \u2212 2C\u22121\n\nLemma 3.2. Suppose q is a stationary point in RC(cid:63) with C(cid:63) \u2265 10, and \u03b6 = AT q has only one\nentry \u03b6l of magnitude no smaller than 2\u00b5(cid:107)\u03b6(cid:107)3\n4, then q is a local minimum near al such that\n|(cid:104)q,PS [al](cid:105)| > 1 \u2212 2c(cid:63)\u03ba\u22122 with c(cid:63) = 1/C(cid:63).\nIf q is a stationary point in RC(cid:63) with C(cid:63) \u2265 10, and \u03b6 has more than one\nSaddle Points\nnontrivial entry, then the Riemannian Hessian Hess \u03d5 (q) has negative eigenvalue(s) and\nhence q is a saddle point. Especially, denoting any two nontrivial entries of \u03b6 with \u03b6l and \u03b6l(cid:48),\nthen there exists a negative curvature in the span of al and al(cid:48).\nLemma 3.3. Suppose q is a stationary point in RC(cid:63) with C(cid:63) \u2265 10, and \u03b6 = AT q has at least two\nentries \u03b6l and \u03b6l(cid:48) with magnitude larger than 2\u00b5(cid:107)\u03b6(cid:107)3\n4, then the Riemannian Hessian at q has\nnegative eigenvalue(s) and q is a saddle point.\n3.3 Finite Sample Concentration\nWe argue that the critical points of the \ufb01nite sample objective function \u03c8(q) are similar to\nthose of the asymptotic objective function \u03d5(q):\nCritical points are close. The Riemannian gradient concentrates, such that there is a bijection\nbetween critical points qpop of \u03d5 and critical points qfs of \u03c8, with (cid:107)qpop \u2212 qfs(cid:107)2 small.\nCurvature is preserved. The Riemannian Hessian concentrates, such that Hess[\u03c8](qfs) has a\nnegative eigenvalue if and only if Hess[\u03d5](qpop) has a negative eigenvalue, and Hess[\u03c8](qfs)\nis positive de\ufb01nite if and only if Hess[\u03d5](qpop) is positive de\ufb01nite.\nThis implies that every local minimizer of the \ufb01nite sample objective function is close to a\npreconditioned shift-truncation. While conceptually straightforward, the proofs of these\n\n3 /(cid:107)\u03b6(cid:107)4\n\n3 /(cid:107)\u03b6(cid:107)4\n\n7\n\n\fproperties are somewhat involved, due to the presence of the preconditioner (Y Y T )\u22121/2.\nWe give rigorous versions of all of the above statements, and a complete proof, in the\nsupplementary appendix.\n\n4 Experiments\nIn our main result, the sparsity rate \u03b8 depends on the\nProperties of a Random Kernel.\ncondition number \u03ba and induced column coherence \u00b5. Figure 3 plots the average values\n(over 100 independent simulations) of \u03ba and \u00b5 for generic unit kernels of varying dimension\nk = 10, 20,\u00b7\u00b7\u00b7 , 1000.\n\nFigure 3: Coherence of random kernels. Average of \u03c3min (left), \u03ba (middle), and \u00b5 (right) over 100\nindependent trials, for varying kernel length k.\n\nThese simulations suggest the following estimates:\n\n\u03c3min \u223c log\n\n\u22121 (k) ,\n\n\u03ba \u223c log4/3 k, \u00b5 \u223c(cid:112)log (k) /k.\n\n(27)\n\nHence, reliable recovery of the shift truncation of a generic kernel can be guaranteed even\nwhen the sparse signal is relatively dense (\u03b8 \u223c k\u22122/3). On the other hand, if the convolution\nkernel a0 is lowpass, then \u03c3min decreases, and \u03ba, \u00b5 increase, then more observations m and\nsmaller sparsity level \u03b8 is required for the proposed algorithm to perform as desired.\nRecovery Error of the Proposed Algorithm We present the performance of Algorithm 1\nunder varying settings. We de\ufb01ne the recover error as err = 1 \u2212 max\u03c4 |(cid:104)\u00afa,PS [\u03b9\u2217\nand calculate the average error from 50 independent experiments. The left \ufb01gure plots the\naverage error when we \ufb01x the kernel size k = 50, and vary the dimension m and the sparsity\n\u03b8 of x0.14 The right \ufb01gure plots the average error when we vary the dimensions k, m of both\nconvolution signals, and set the sparsity as \u03b8 = k\u22122/3.\n\nks\u03c4 [(cid:102)a0]](cid:105)|,\n\nFigure 4: Recovery Error of the Shift Truncated Kernel by Algorithm 1.\n\nAcknowledgement The authors gratefully acknowledge support from NSF 1343282, NSF\nCCF 1527809, NSF CCF 1740833, and NSF IIS 1546411.\n\n14Note that the x-axis is indexed with overlapping ratio k \u00b7 \u03b8, which indicates how many copies of\n\na0 present in a k-length window of y on average.\n\n8\n\nAverage error, k = 50 Overlapping ratio k\"32.664.135.607.07Signal length m2500200015001000 5000.20.10.0Average error, sparsity = k-2/3Kernel size k1020304050607080Signal length m22001900160013001000 700 400 1000.20.10.0\fReferences\n[1] Sky Cheung, Yenson Lau, Zhengyu Chen, Ju Sun, Yuqian Zhang, John Wright, and Abhay\nPasupathy. Beyond the fourier transform: A nonconvex optimization approach to microscopy\nanalysis. Submitted, 2017.\n\n[2] M. S. Lewicki. A review of methods for spike sorting: the detection and classi\ufb01cation of neural\n\naction potentials. Network: Computation in Neural Systems, 9(4):53\u201378, 1998.\n\n[3] D. Kundur and D. Hatzinakos. Blind image deconvolution. Signal Processing Magazine, IEEE,\n\n13(3):43\u201364, May 1996.\n\n[4] Chaitanya Ekanadham, Daniel Tranchina, and Eero P. Simoncelli. A blind sparse deconvolution\nmethod for neural spike identi\ufb01cation. In Advances in Neural Information Processing Systems 24,\npages 1440\u20131448. 2011.\n\n[5] T. Chan and C. Wong. Total variation blind deconvolution. IEEE Transactions on Image Processing,\n\n7(3):370\u2013375, Mar 1998.\n\n[6] A. Levin, Y. Weiss, F. Durand, and W. Freeman. Understanding blind deconvolution algorithms.\n\nIEEE Transactions on Pattern Analysis and Machine Intelligence, 33(12):2354\u20132367, Dec 2011.\n\n[7] A. Ahmed, B. Recht, and J. Romberg. Blind deconvolution using convex programing. arXiv\n\npreprint:1211.5608, 2012.\n\n[8] Xiaodong Li, Shuyang Ling, Thomas Strohmer, and Ke Wei. Rapid, robust, and reliable blind\n\ndeconvolution via nonconvex optimization. preprint, 2016.\n\n[9] Shuyang Ling and Thomas Strohmer. Self-calibration and biconvex compressive sensing. Inverse\n\nProblems, 31(11):115002, 2015.\n\n[10] Shuyang Ling and Thomas Strohmer. Blind deconvolution meets blind demixing: Algorithms\n\nand performance bounds. IEEE Transactions on Information Theory, 63(7):4497\u20134520, 2017.\n\n[11] Yuejie Chi. Guaranteed blind sparse spikes deconvolution via lifting and convex optimization.\n\nIEEE Journal of Selected Topics in Signal Processing, 10(4):782\u2013794, June 2016.\n\n[12] A. Benichoux, E. Vincent, and R. Gribonval. A fundamental pitfall in blind deconvolution with\nsparse and shift-invariant priors. 38th International Conference on Acoustics, Speech, and Signal\nProcessing, May 2013.\n\n[13] Daniele Perrone and Paolo Favaro. Total variation blind deconvolution: The devil is in the details.\n\nIn IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014.\n\n[14] David Wipf and Haichao Zhang. Revisiting bayesian blind deconvolution. arXiv preprint:1305.2362,\n\n2013.\n\n[15] Haichao Zhang, David Wipf, and Yanning Zhang. Multi-image blind deblurring using a coupled\nadaptive sparse prior. IEEE Conference on Computer Vision and Pattern Recognition (CVPR),\nJanuary 2013.\n\n[16] Yuqian Zhang, Yenson Lau, Han-wen Kuo, Sky Cheung, Abhay Pasupathy, and John Wright. On\nthe global geometry of sphere-constrained sparse blind deconvolution. In The IEEE Conference on\nComputer Vision and Pattern Recognition (CVPR), July 2017.\n\n[17] Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan. How to escape\n\nsaddle points e\ufb03ciently. arXiv preprint arXiv:1703.00887, 2017.\n\n[18] Peng Xu, Farbod Roosta-Khorasani, and Michael W. Mahoney. Second-order optimization for\n\nnon-convex machine learning: An empirical study. arXiv preprint arXiv:1708.07827, 2017.\n\n[19] L. Wang and Y. Chi. Blind deconvolution from multiple sparse inputs. IEEE Signal Processing\n\nLetters, 23(10):1384\u20131388, Oct 2016.\n\n[20] Daniel Spielman, Huan Wang, and John Wright. Exact recovery of sparsely-used dictionaries.\n\npreprint, 2012.\n\n[21] Ju Sun, Qing Qu, and John Wright. Complete dictionary recovery over the sphere. preprint, 2015.\n\n9\n\n\f[22] Qing Qu, Ju Sun, and John Wright. Finding a sparse vector in a subspace: linear sparsity using\n\nalternating directions. IEEE Transactions on Information Theory, 2016.\n\n[23] Samuel B. Hopkins, Tselil Schrammand, Jonathan Shi, and David Steurer. Fast spectral algorithms\nfrom sum-of-squares proofs: Tensor decomposition and planted sparse vectors. In Proceedings of\nthe Forty-eighth Annual ACM Symposium on Theory of Computing, STOC \u201916, pages 178\u2013191, 2016.\n[24] P.-A. Absil, C.G. Baker, and K.A. Gallivan. Trust-region methods on riemannian manifolds.\n\nFoundations of Computational Mathematics, 7(3):303\u2013330, Jul 2007.\n\n[25] Donald Goldfarb, Zaiwen Wen, and Wotao Yin. A curvilinear search method for p-harmonic\n\n\ufb02ows on spheres. SIAM J. Imaging Sciences, 2(1):84\u2013109, 2009.\n\n[26] P.-A. Absil, R. Mahony, and R. Sepulchre. Optimization Algorithms on Matrix Manifolds. Princeton\n\nUniversity Press, Princeton, NJ, USA, 2007.\n\n10\n\n\f", "award": [], "sourceid": 1169, "authors": [{"given_name": "Yuqian", "family_name": "Zhang", "institution": "Cornell University"}, {"given_name": "Han-wen", "family_name": "Kuo", "institution": "Columbia University"}, {"given_name": "John", "family_name": "Wright", "institution": "Columbia University"}]}