Tensor decompositions for learning latent variable models

Anima Anandkumar, Rong Ge, Daniel Hsu, Sham M. Kakade, Matus Telgarsky

Introduction

The method of moments is a classical parameter estimation technique (Pearson, 1894) from statistics which has proved invaluable in a number of application domains. The basic paradigm is simple and intuitive: (i) compute certain statistics of the data—often empirical moments such as means and correlations—and (ii) find model parameters that give rise to (nearly) the same corresponding population quantities. In a number of cases, the method of moments leads to consistent estimators which can be efficiently computed; this is especially relevant in the context of latent variable models, where standard maximum likelihood approaches are typically computationally prohibitive, and heuristic methods can be unreliable and difficult to validate with high-dimensional data. Furthermore, the method of moments can be viewed as complementary to the maximum likelihood approach; simply taking a single step of Newton-Raphson on the likelihood function starting from the moment based estimator (Le Cam, 1986) often leads to the best of both worlds: a computationally efficient estimator that is (asymptotically) statistically optimal.

The primary difficulty in learning latent variable models is that the latent (hidden) state of the data is not directly observed; rather only observed variables correlated with the hidden state are observed. As such, it is not evident the method of moments should fare any better than maximum likelihood in terms of computational performance: matching the model parameters to the observed moments may involve solving computationally intractable systems of multivariate polynomial equations. Fortunately, for many classes of latent variable models, there is rich structure in low-order moments (typically second- and third-order) which allow for this inverse moment problem to be solved efficiently (Cattell, 1944; Cardoso, 1991; Chang, 1996; Mossel and Roch, 2006; Hsu et al., 2012b; Anandkumar et al., 2012c, a; Hsu and Kakade, 2013). What is more is that these decomposition problems are often amenable to simple and efficient iterative methods, such as gradient descent and the power iteration method.

In this work, we observe that a number of important and well-studied latent variable models—including Gaussian mixture models, hidden Markov models, and Latent Dirichlet allocation—share a certain structure in their low-order moments, and this permits certain tensor decomposition approaches to parameter estimation. In particular, this decomposition can be viewed as a natural generalization of the singular value decomposition for matrices.

While much of this (or similar) structure was implicit in several previous works (Chang, 1996; Mossel and Roch, 2006; Hsu et al., 2012b; Anandkumar et al., 2012c, a; Hsu and Kakade, 2013), here we make the decomposition explicit under a unified framework. Specifically, we express the observable moments as sums of rank-one terms, and reduce the parameter estimation task to the problem of extracting a symmetric orthogonal decomposition of a symmetric tensor derived from these observable moments. The problem can then be solved by a variety of approaches, including fixed-point and variational methods.

One approach for obtaining the orthogonal decomposition is the tensor power method of Lathauwer et al. (2000, Remark 3). We provide a convergence analysis of this method for orthogonally decomposable symmetric tensors, as well as a detailed perturbation analysis for a robust (and a computationally tractable) variant (Theorem 5.1). This perturbation analysis can be viewed as an analogue of Wedin’s perturbation theorem for singular vectors of matrices (Wedin, 1972), providing a bound on the error of the recovered decomposition in terms of the operator norm of the tensor perturbation. This analysis is subtle in at least two ways. First, unlike for matrices (where every matrix has a singular value decomposition), an orthogonal decomposition need not exist for the perturbed tensor. Our robust variant uses random restarts and deflation to extract an approximate decomposition in a computationally tractable manner. Second, the analysis of the deflation steps is non-trivial; a naïve argument would entail error accumulation in each deflation step, which we show can in fact be avoided. When this method is applied for parameter estimation in latent variable models previously discussed, improved sample complexity bounds (over previous work) can be obtained using this perturbation analysis.

Finally, we also address computational issues that arise when applying the tensor decomposition approaches to estimating latent variable models. Specifically, we show that the basic operations of simple iterative approaches (such as the tensor power method) can be efficiently executed in time linear in the dimension of the observations and the size of the training data. For instance, in a topic modeling application, the proposed methods require time linear in the number of words in the vocabulary and in the number of non-zero entries of the term-document matrix. The combination of this computational efficiency and the robustness of the tensor decomposition techniques makes the overall framework a promising approach to parameter estimation for latent variable models.

2 Related Work

The connection between tensor decompositions and latent variable models has a long history across many scientific and mathematical disciplines. We review some of the key works that are most closely related to ours.

The role of tensor decompositions in the context of latent variable models dates back to early uses in psychometrics (Cattell, 1944). These ideas later gained popularity in chemometrics, and more recently in numerous science and engineering disciplines, including neuroscience, phylogenetics, signal processing, data mining, and computer vision. A thorough survey of these techniques and applications is given by Kolda and Bader (2009). Below, we discuss a few specific connections to two applications in machine learning and statistics, independent component analysis and latent variable models (between which there is also significant overlap).

Tensor decompositions have been used in signal processing and computational neuroscience for blind source separation and independent component analysis (ICA) (Comon and Jutten, 2010). Here, statistically independent non-Gaussian sources are linearly mixed in the observed signal, and the goal is to recover the mixing matrix (and ultimately, the original source signals). A typical solution is to locate projections of the observed signals that correspond to local extrema of the so-called “contrast functions” which distinguish Gaussian variables from non-Gaussian variables. This method can be effectively implemented using fast descent algorithms (Hyvarinen, 1999). When using the excess kurtosis (i.e., fourth-order cumulant) as the contrast function, this method reduces to a generalization of the power method for symmetric tensors (Lathauwer et al., 2000; Zhang and Golub, 2001; Kofidis and Regalia, 2002). This case is particularly important, since all local extrema of the kurtosis objective correspond to the true sources (under the assumed statistical model) (Delfosse and Loubaton, 1995); the descent methods can therefore be rigorously analyzed, and their computational and statistical complexity can be bounded (Frieze et al., 1996; Nguyen and Regev, 2009; Arora et al., 2012b).

Higher-order tensor decompositions have also been used to develop estimators for commonly used mixture models, hidden Markov models, and other related latent variable models, often using the the algebraic procedure of R. Jennrich (as reported in the article of Harshman, 1970), which is based on a simultaneous diagonalization of different ways of flattening a tensor to matrices. Jennrich’s procedure was employed for parameter estimation of discrete Markov models by Chang (1996) via pair-wise and triple-wise probability tables; and it was later used for other latent variable models such as hidden Markov models (HMMs), latent trees, Gaussian mixture models, and topic models such as latent Dirichlet allocation (LDA) by many others (Mossel and Roch, 2006; Hsu et al., 2012b; Anandkumar et al., 2012c, a; Hsu and Kakade, 2013). In these contexts, it is often also possible to establish strong identifiability results, without giving an explicit estimators, by invoking the non-constructive identifiability argument of Kruskal (1977)—see the article by Allman et al. (2009) for several examples.

Related simultaneous diagonalization approaches have also been used for blind source separation and ICA (as discussed above), and a number of efficient algorithms have been developed for this problem (Bunse-Gerstner et al., 1993; Cardoso and Souloumiac, 1993; Cardoso, 1994; Cardoso and Comon, 1996; Corless et al., 1997; Ziehe et al., 2004). A rather different technique that uses tensor flattening and matrix eigenvalue decomposition has been developed by Cardoso (1991) and later by De Lathauwer et al. (2007). A significant advantage of this technique is that it can be used to estimate overcomplete mixtures, where the number of sources is larger than the observed dimension.

The relevance of tensor analysis to latent variable modeling has been long recognized in the field of algebraic statistics (Pachter and Sturmfels, 2005), and many works characterize the algebraic varieties corresponding to the moments of various classes of latent variable models (Drton et al., 2007; Sturmfels and Zwiernik, 2013). These works typically do not address computational or finite sample issues, but rather are concerned with basic questions of identifiability.

The specific tensor structure considered in the present work is the symmetric orthogonal decomposition. This decomposition expresses a tensor as a linear combination of simple tensor forms; each form is the tensor product of a vector (i.e., a rank- $1$ tensor), and the collection of vectors form an orthonormal basis. An important property of tensors with such decompositions is that they have eigenvectors corresponding to these basis vectors. Although the concepts of eigenvalues and eigenvectors of tensors is generally significantly more complicated than their matrix counterpart—both algebraically (Qi, 2005; Cartwright and Sturmfels, 2013; Lim, 2005) and computationally (Hillar and Lim, 2013; Kofidis and Regalia, 2002)—the special symmetric orthogonal structure we consider permits simple algorithms to efficiently and stably recover the desired decomposition. In particular, a generalization of the matrix power method to symmetric tensors, introduced by Lathauwer et al. (2000, Remark 3) and analyzed by Kofidis and Regalia (2002), provides such a decomposition. This is in fact implied by the characterization of Zhang and Golub (2001), which shows that iteratively obtaining the best rank- $1$ approximation of such orthogonally decomposable tensors also yields the exact decomposition. We note that in general, obtaining such approximations for general (symmetric) tensors is NP-hard (Hillar and Lim, 2013).

2.2 Latent Variable Models

This work focuses on the particular application of tensor decomposition methods to estimating latent variable models, a significant departure from many previous approaches in the machine learning and statistics literature. By far the most popular heuristic for parameter estimation for such models is the Expectation-Maximization (EM) algorithm (Dempster et al., 1977; Redner and Walker, 1984). Although EM has a number of merits, it may suffer from slow convergence and poor quality local optima (Redner and Walker, 1984), requiring practitioners to employ many additional heuristics to obtain good solutions. For some models such as latent trees (Roch, 2006) and topic models (Arora et al., 2012a), maximum likelihood estimation is NP-hard, which suggests that other estimation approaches may be more attractive. More recently, algorithms from theoretical computer science and machine learning have addressed computational and sample complexity issues related to estimating certain latent variable models such as Gaussian mixture models and HMMs (Dasgupta, 1999; Arora and Kannan, 2005; Dasgupta and Schulman, 2007; Vempala and Wang, 2004; Kannan et al., 2008; Achlioptas and McSherry, 2005; Chaudhuri and Rao, 2008; Brubaker and Vempala, 2008; Kalai et al., 2010; Belkin and Sinha, 2010; Moitra and Valiant, 2010; Hsu and Kakade, 2013; Chang, 1996; Mossel and Roch, 2006; Hsu et al., 2012b; Anandkumar et al., 2012c; Arora et al., 2012a; Anandkumar et al., 2012a). See the works by Anandkumar et al. (2012c) and Hsu and Kakade (2013) for a discussion of these methods, together with the computational and statistical hardness barriers that they face. The present work reviews a broad range of latent variables where a mild non-degeneracy condition implies the symmetric orthogonal decomposition structure in the tensors of low-order observable moments.

Notably, another class of methods, based on subspace identification (Overschee and Moor, 1996) and observable operator models/multiplicity automata (Schützenberger, 1961; Jaeger, 2000; Littman et al., 2001), have been proposed for a number of latent variable models. These methods were successfully developed for HMMs by Hsu et al. (2012b), and subsequently generalized and extended for a number of related sequential and tree Markov models models (Siddiqi et al., 2010; Bailly, 2011; Boots et al., 2010; Parikh et al., 2011; Rodu et al., 2013; Balle et al., 2012; Balle and Mohri, 2012), as well as certain classes of parse tree models (Luque et al., 2012; Cohen et al., 2012; Dhillon et al., 2012). These methods use low-order moments to learn an “operator” representation of the distribution, which can be used for density estimation and belief state updates. While finite sample bounds can be given to establish the learnability of these models (Hsu et al., 2012b), the algorithms do not actually give parameter estimates (e.g., of the emission or transition matrices in the case of HMMs).

3 Organization

The rest of the paper is organized as follows. Section 2 reviews some basic definitions of tensors. Section 3 provides examples of a number of latent variable models which, after appropriate manipulations of their low order moments, share a certain natural tensor structure. Section 4 reduces the problem of parameter estimation to that of extracting a certain (symmetric orthogonal) decomposition of a tensor. We then provide a detailed analysis of a robust tensor power method and establish an analogue of Wedin’s perturbation theorem for the singular vectors of matrices. The discussion in Section 6 addresses a number of practical concerns that arise when dealing with moment matrices and tensors.

Preliminaries

Note that if $A$ is a matrix ( $p=2$ ), then

where $I$ is the $n\times n$ identity matrix. As a final example of this notation, observe

The notion of tensor (symmetric) rank is considerably more delicate than matrix (symmetric) rank. For instance, it is not clear a priori that the symmetric rank of a tensor should even be finite (Comon et al., 2008). In addition, removal of the best rank- $1$ approximation of a (general) tensor may increase the tensor rank of the residual (Stegeman and Comon, 2010).

Throughout, we use $\|v\|=(\sum_{i}v_{i}^{2})^{1/2}$ to denote the Euclidean norm of a vector $v$ , and $\|M\|$ to denote the spectral (operator) norm of a matrix. We also use $\|T\|$ to denote the operator norm of a tensor, which we define later.

Tensor Structure in Latent Variable Models

In this section, we give several examples of latent variable models whose low-order moments can be written as symmetric tensors of low symmetric rank; some of these examples can be deduced using the techniques developed in the text by McCullagh (1987). The basic form is demonstrated in Theorem 3.1 for the first example, and the general pattern will emerge from subsequent examples.

One advantage of this encoding of words is that the (cross) moments of these random vectors correspond to joint probabilities over words. For instance, observe that

The second advantage of the vector encoding of words is that the conditional expectation of $x_{t}$ given $h=j$ is simply $\mu_{j}$ , the vector of word probabilities for topic $j$ :

(where $[\mu_{j}]_{i}$ is the $i$ -th entry in the vector $\mu_{j}$ ). Because the words are conditionally independent given the topic, we can use this same property with conditional cross moments, say, of $x_{1}$ and $x_{2}$ :

This and similar calculations lead one to the following theorem.

As we will see in Section 4.3, the structure of $M_{2}$ and $M_{3}$ revealed in Theorem 3.1 implies that the topic vectors $\mu_{1},\mu_{2},\dotsc,\mu_{k}$ can be estimated by computing a certain symmetric tensor decomposition. Moreover, due to exchangeability, all triples (resp., pairs) of words in a document—and not just the first three (resp., two) words—can be used in forming $M_{3}$ (resp., $M_{2}$ ); see Section 6.1.

2 Beyond Raw Moments

In the single topic model above, the raw (cross) moments of the observed words directly yield the desired symmetric tensor structure. In some other models, the raw moments do not explicitly have this form. Here, we show that the desired tensor structure can be found through various manipulations of different moments.

We now consider a mixture of $k$ Gaussian distributions with spherical covariances. We start with the simpler case where all of the covariances are identical; this probabilistic model is closely related to the (non-probabilistic) $k$ -means clustering problem (MacQueen, 1967).

2.2 Spherical Gaussian Mixtures: Differing Covariances

2.3 Independent Component Analysis (ICA)

Let $\mu_{i}$ denote the $i$ -th column of the mixing matrix $A$ .

where $T$ is the fourth-order tensor with

Note that $\kappa_{i}$ corresponds to the excess kurtosis, a measure of non-Gaussianity as $\kappa_{i}=0$ if $h_{i}$ is a standard normal random variable. Furthermore, note that $A$ is not identifiable if $h$ is a multivariate Gaussian.

We may derive forms similar to that of $M_{2}$ and $M_{3}$ from Theorem 3.1 using $M_{4}$ by observing that

2.4 Latent Dirichlet Allocation (LDA)

The parameter $\alpha_{0}$ (the sum of the “pseudo-counts”) characterizes the concentration of the distribution. As $\alpha_{0}\rightarrow 0$ , the distribution degenerates to a single topic model (i.e., the limiting density has, with probability $1$ , exactly one entry of $h$ being $1$ and the rest are ). At the other extreme, if $\alpha=(c,c,\dotsc,c)$ for some scalar $c>0$ , then as $\alpha_{0}=ck\to\infty$ , the distribution of $h$ becomes peaked around the uniform vector $(1/k,1/k,\dotsc,1/k)$ (furthermore, the distribution behaves like a product distribution). We are typically interested in the case where $\alpha_{0}$ is small (e.g., a constant independent of $k$ ), whereupon $h$ typically has only a few large entries. This corresponds to the setting where the documents are mainly comprised of just a few topics.

Note that $\alpha_{0}$ needs to be known to form $M_{2}$ and $M_{3}$ from the raw moments. This, however, is a much weaker than assuming that the entire distribution of $h$ is known (i.e., knowledge of the whole parameter vector $\alpha$ ).

3 Multi-View Models

We first note the form for the raw (cross) moments.

The cross moments do not possess a symmetric tensor form when the conditional distributions are different. Nevertheless, the moments can be “symmetrized” via a simple linear transformation of $x_{1}$ and $x_{2}$ (roughly speaking, this relates $x_{1}$ and $x_{2}$ to $x_{3}$ ); this leads to an expression from which the conditional means of $x_{3}$ (i.e., $\mu_{3,1},\mu_{3,2},\dotsc,\mu_{3,k}$ ) can be recovered. For simplicity, we assume $d_{1}=d_{2}=d_{3}=k$ ; the general case (with $d_{t}\geq k$ ) is easily handled using low-rank singular value decompositions.

Assume that $\{\mu_{v,1},\mu_{v,2},\dotsc,\mu_{v,k}\}$ are linearly independent for each $v\in\{1,2,3\}$ . Define

We now discuss three examples (taken mostly from Anandkumar et al., 2012c) where the above observations can be applied. The first two concern mixtures of product distributions, and the last one is the time-homogeneous hidden Markov model.

For a mixture of product distributions, any partitioning of the dimensions $[n]$ into three groups creates three (possibly asymmetric) “views” which are conditionally independent once the mixture component is selected. However, recall that Theorem 3.6 requires that for each view, the $k$ conditional means be linearly independent. In general, this may not be achievable; consider, for instance, the case $\mu_{i}=e_{i}$ for each $i\in[k]$ . Such cases, where the component means are very aligned with the coordinate basis, are precluded by the incoherence condition.

Define $\operatorname{coherence}(A):=\max_{i\in[n]}\{e_{i}^{\scriptscriptstyle\top}\Pi_{A}e_{i}\}$ to be the largest diagonal entry of the orthogonal projector to the range of $A$ , and assume $A$ has rank $k$ . The coherence lies between $k/n$ and $1$ ; it is largest when the range of $A$ is spanned by the coordinate axes, and it is $k/n$ when the range is spanned by a subset of the Hadamard basis of cardinality $k$ . The incoherence condition requires, for some $\varepsilon,\delta\in(0,1)$ , $\operatorname{coherence}(A)\leq(\varepsilon^{2}/6)/\ln(3k/\delta)$ . Essentially, this condition ensures that the non-degeneracy of the component means is not isolated in just a few of the $n$ dimensions. Operationally, it implies the following.

for some $\varepsilon,\delta\in(0,1)$ . With probability at least $1-\delta$ , a random partitioning of the dimensions $[n]$ into three groups (for each $i\in[n]$ , independently pick $t\in\{1,2,3\}$ uniformly at random and put $i$ in group $t$ ) has the following property. For each $t\in\{1,2,3\}$ and $j\in[k]$ , let $\mu_{t,j}$ be the entries of $\mu_{j}$ put into group $t$ , and let $A_{t}:=[\mu_{t,1}|\mu_{t,2}|\dotsb|\mu_{t,k}]$ . Then for each $t\in\{1,2,3\}$ , $A_{t}$ has full column rank, and the $k$ -th largest singular value of $A_{t}$ is at least $\sqrt{(1-\varepsilon)/3}$ times that of $A$ .

Therefore, three asymmetric views can be created by randomly partitioning the observed random vector $x$ into $x_{1}$ , $x_{2}$ , and $x_{3}$ , such that the resulting component means for each view satisfy the conditions of Theorem 3.6.

3.2 Spherical Gaussian Mixtures, Revisited

Consider again the case of spherical Gaussian mixtures (cf. Section 3.2). As we shall see in Section 4.3, the previous techniques (based on Theorem 3.2 and Theorem 3.3) lead to estimation procedures when the dimension of $x$ is $k$ or greater (and when the $k$ component means are linearly independent). We now show that when the dimension is slightly larger, say greater than $3k$ , a different (and simpler) technique based on the multi-view structure can be used to extract the relevant structure.

3.3 Hidden Markov Models

Define $h:=y_{2}$ , where $y_{2}$ is the second hidden state in the Markov chain. Then

$x_{1},x_{2},x_{3}$ are conditionally independent given $h$ ;

the distribution of $h$ is given by the vector $w:=T\pi\in\Delta^{k-1}$ ;

Note the matrix of conditional means of $x_{t}$ has full column rank, for each $t\in\{1,2,3\}$ , provided that: (i) $O$ has full column rank, (ii) $T$ is invertible, and (iii) $\pi$ and $T\pi$ have positive entries.

Orthogonal Tensor Decompositions

We now show how recovering the $\mu_{i}$ ’s in our aforementioned problems reduces to the problem of finding a certain orthogonal tensor decomposition of a symmetric tensor. We start by reviewing the spectral decomposition of symmetric matrices, and then discuss a generalization to the higher-order tensor case. Finally, we show how orthogonal tensor decompositions can be used for estimating the latent variable models from the previous section.

Such a decomposition is guaranteed to exist for every symmetric matrix.

Recovery of the $v_{i}$ ’s and $\lambda_{i}$ ’s can be viewed at least two ways. First, each $v_{i}$ is fixed under the mapping $u\mapsto Mu$ , up to a scaling factor $\lambda_{i}$ :

as $v_{j}^{\scriptscriptstyle\top}v_{i}=0$ for all $j\neq i$ by orthogonality. The $v_{i}$ ’s are not necessarily the only such fixed points. For instance, with the multiplicity $\lambda_{1}=\lambda_{2}=\lambda$ , then any linear combination of $v_{1}$ and $v_{2}$ is similarly fixed under $M$ . However, in this case, the decomposition in (1) is not unique, as $\lambda_{1}v_{1}v_{1}^{\scriptscriptstyle\top}+\lambda_{2}v_{2}v_{2}^{\scriptscriptstyle\top}$ is equal to $\lambda(u_{1}u_{1}^{\scriptscriptstyle\top}+u_{2}u_{2}^{\scriptscriptstyle\top})$ for any pair of orthonormal vectors, $u_{1}$ and $u_{2}$ spanning the same subspace as $v_{1}$ and $v_{2}$ . Nevertheless, the decomposition is unique when $\lambda_{1},\lambda_{2},\dotsc,\lambda_{k}$ are distinct, whereupon the $v_{j}$ ’s are the only directions fixed under $u\mapsto Mu$ up to non-trivial scaling.

The second view of recovery is via the variational characterization of the eigenvalues. Assume $\lambda_{1}>\lambda_{2}>\dotsb>\lambda_{k}$ ; the case of repeated eigenvalues again leads to similar non-uniqueness as discussed above. Then the Rayleigh quotient

is maximized over non-zero vectors by $v_{1}$ . Furthermore, for any $s\in[k]$ , the maximizer of the Rayleigh quotient, subject to being orthogonal to $v_{1},v_{2},\dotsc,v_{s-1}$ , is $v_{s}$ . Another way of obtaining this second statement is to consider the deflated Rayleigh quotient

and observe that $v_{s}$ is the maximizer.

Efficient algorithms for finding these matrix decompositions are well studied (Golub and van Loan, 1996, Section 8.2.3), and iterative power methods are one effective class of algorithms.

We remark that in our multilinear tensor notation, we may write the maps $u\mapsto Mu$ and $u\mapsto u^{\scriptscriptstyle\top}Mu/\|u\|_{2}^{2}$ as

2 The Tensor Case

Decomposing general tensors is a delicate issue; tensors may not even have unique decompositions. Fortunately, the orthogonal tensors that arise in the aforementioned models have a structure which permits a unique decomposition under a mild non-degeneracy condition. We focus our attention to the case $p=3$ , i.e., a third order tensor; the ideas extend to general $p$ with minor modifications.

Note that since we are focusing on odd-order tensors ( $p=3$ ), we have added the requirement that the $\lambda_{i}$ be positive. This convention can be followed without loss of generality since $-\lambda_{i}v_{i}^{\otimes p}=\lambda_{i}(-v_{i})^{\otimes p}$ whenever $p$ is odd. Also, it should be noted that orthogonal decompositions do not necessarily exist for every symmetric tensor.

In analogy to the matrix setting, we consider two ways to view this decomposition: a fixed-point characterization and a variational characterization. Related characterizations based on optimal rank- $1$ approximations are given by Zhang and Golub (2001).

For a tensor $T$ , consider the vector-valued map

which is the third-order generalization of (2). This can be explicitly written as

Observe that (5) is not a linear map, which is a key difference compared to the matrix case.

(To simplify the discussion, we assume throughout that eigenvectors have unit norm; otherwise, for scaling reasons, we replace the above equation with $T(I,u,u)=\lambda\|u\|u$ .) This concept was originally introduced by Lim (2005) and Qi (2005). For orthogonally decomposable tensors $T=\sum_{i=1}^{k}\lambda_{i}v_{i}^{\otimes 3}$ ,

By the orthogonality of the $v_{i}$ , it is clear that $T(I,v_{i},v_{i})=\lambda_{i}v_{i}$ for all $i\in[k]$ . Therefore each $(v_{i},\lambda_{i})$ is an eigenvector/eigenvalue pair.

There are a number of subtle differences compared to the matrix case that arise as a result of the non-linearity of (5). First, even with the multiplicity $\lambda_{1}=\lambda_{2}=\lambda$ , a linear combination $u:=c_{1}v_{1}+c_{2}v_{2}$ may not be an eigenvector. In particular,

may not be a multiple of $c_{1}v_{1}+c_{2}v_{2}$ . This indicates that the issue of repeated eigenvalues does not have the same status as in the matrix case. Second, even if all the eigenvalues are distinct, it turns out that the $v_{i}$ ’s are not the only eigenvectors. For example, set $u:=(1/\lambda_{1})v_{1}+(1/\lambda_{2})v_{2}$ . Then,

so $u/\|u\|$ is an eigenvector. More generally, for any subset $S\subseteq[k]$ , the vector

starting from $\theta$ converges to $u$ . Note that the map (6) rescales the output to have unit Euclidean norm. Robust eigenvectors are also called attracting fixed points of (6) (see, e.g., Kolda and Mayo, 2011).

The following theorem implies that if $T$ has an orthogonal decomposition as given in (4), then the set of robust eigenvectors of $T$ are precisely the set $\{v_{1},v_{2},\ldots v_{k}\}$ , implying that the orthogonal decomposition is unique. (For even order tensors, the uniqueness is true up to sign-flips of the $v_{i}$ .)

Let $T$ have an orthogonal decomposition as given in (4).

The set of robust eigenvectors of $T$ is equal to $\{v_{1},v_{2},\dotsc,v_{k}\}$ .

The proof of Theorem 4.1 is given in Appendix A.1, and follows readily from simple orthogonality considerations. Note that every $v_{i}$ in the orthogonal tensor decomposition is robust, whereas for a symmetric matrix $M$ , for almost all initial points, the map $\bar{\theta}\mapsto\frac{M\bar{\theta}}{\|M\bar{\theta}\|}$ converges only to an eigenvector corresponding to the largest magnitude eigenvalue. Also, since the tensor order is odd, the signs of the robust eigenvectors are fixed, as each $-v_{i}$ is mapped to $v_{i}$ under (6).

2.2 Variational Characterization

We now discuss a variational characterization of the orthogonal decomposition. The generalized Rayleigh quotient (Zhang and Golub, 2001) for a third-order tensor is

Let $T$ have an orthogonal decomposition as given in (4), and consider the optimization problem

The stationary points are eigenvectors of $T$ .

A stationary point $u$ is an isolated local maximizer if and only if $u=v_{i}$ for some $i\in[k]$ .

The proof of Theorem 4.2 is given in Appendix A.2. It is similar to local optimality analysis for ICA methods using fourth-order cumulants (e.g., Delfosse and Loubaton, 1995; Frieze et al., 1996).

Again, we see similar distinctions to the matrix case. In the matrix case, the only local maximizers of the Rayleigh quotient are the eigenvectors with the largest eigenvalue (and these maximizers take on the globally optimal value). For the case of orthogonal tensor forms, the robust eigenvectors are precisely the isolated local maximizers.

An important implication of the two characterizations is that, for orthogonally decomposable tensors $T$ , (i) the local maximizers of the objective function $u\mapsto T(u,u,u)/(u^{\scriptscriptstyle\top}u)^{3/2}$ correspond precisely to the vectors $v_{i}$ in the decomposition, and (ii) these local maximizers can be reliably identified using a simple fixed-point iteration (i.e., the tensor analogue of the matrix power method). Moreover, a second-derivative test based on $T(I,I,u)$ can be employed to test for local optimality and rule out other stationary points.

3 Estimation via Orthogonal Tensor Decompositions

We now demonstrate how the moment tensors obtained for various latent variable models in Section 3 can be reduced to an orthogonal form. For concreteness, we take the specific form from the exchangeable single topic model (Theorem 3.1):

(The more general case allows the weights $w_{i}$ in $M_{2}$ to differ in $M_{3}$ , but for simplicity we keep them the same in the following discussion.) We now show how to reduce these forms to an orthogonally decomposable tensor from which the $w_{i}$ and $\mu_{i}$ can be recovered. See Appendix D for a discussion as to how previous approaches (Mossel and Roch, 2006; Anandkumar et al., 2012c, a; Hsu and Kakade, 2013) achieved this decomposition through a certain simultaneous diagonalization method.

Throughout, we assume the following non-degeneracy condition.

Observe that Condition 4.1 implies that $M_{2}\succeq 0$ is positive semidefinite and has rank $k$ . This is often a mild condition in applications. When this condition is not met, learning is conjectured to be generally hard for both computational (Mossel and Roch, 2006) and information-theoretic reasons (Moitra and Valiant, 2010). As discussed by Hsu et al. (2012b) and Hsu and Kakade (2013), when the non-degeneracy condition does not hold, it is often possible to combine multiple observations using tensor products to increase the rank of the relevant matrices. Indeed, this observation has been rigorously formulated in very recent works of Bhaskara et al. (2014) and Anderson et al. (2014) using the framework of smoothed analysis (Spielman and Teng, 2009).

As the following theorem shows, the orthogonal decomposition of $\widetilde{M}_{3}$ can be obtained by identifying its robust eigenvectors, upon which the original parameters $w_{i}$ and $\mu_{i}$ can be recovered. For simplicity, we only state the result in terms of robust eigenvector/eigenvalue pairs; one may also easily state everything in variational form using Theorem 4.2.

Assume Condition 4.1 and take $\widetilde{M}_{3}$ as defined above.

The theorem follows by combining the above discussion with the robust eigenvector characterization of Theorem 4.1. Recall that we have taken as convention that eigenvectors have unit norm, so the $\mu_{i}$ are exactly determined from the robust eigenvector/eigenvalue pairs of $\widetilde{M}_{3}$ (together with the pseudoinverse of $W^{\scriptscriptstyle\top}$ ); in particular, the scale of each $\mu_{i}$ is correctly identified (along with the corresponding $w_{i}$ ). Relative to previous works on moment-based estimators for latent variable models (e.g., Anandkumar et al., 2012c, a; Hsu and Kakade, 2013), Theorem 4.3 emphasizes the role of the special tensor structure, which in turn makes transparent the applicability of methods for orthogonal tensor decomposition.

3.2 Local Maximizers of (Cross Moment) Skewness

The variational characterization provides an interesting perspective on the robust eigenvectors for these latent variable models. Consider the exchangeable single topic models (Theorem 3.1), and the objective function

In this case, every local maximizer $u^{*}$ satisfies $M_{2}(I,u^{*})=\sqrt{w_{i}}\mu_{i}$ for some $i\in[k]$ . The objective function can be interpreted as the (cross moment) skewness of the random vectors $x_{1},x_{2},x_{3}$ along direction $u$ .

Tensor Power Method

In this section, we consider the tensor power method of Lathauwer et al. (2000, Remark 3) for orthogonal tensor decomposition. We first state a simple convergence analysis for an orthogonally decomposable tensor $T$ .

When only an approximation $\hat{T}$ to an orthogonally decomposable tensor $T$ is available (e.g., when empirical moments are used to estimate population moments), an orthogonal decomposition need not exist for this perturbed tensor (unlike for the case of matrices), and a more robust approach is required to extract the approximate decomposition. Here, we propose such a variant in Algorithm 1 and provide a detailed perturbation analysis. We note that alternative approaches such as simultaneous diagonalization can also be employed (see Appendix D).

The following lemma establishes the quadratic convergence of the tensor power method—i.e., repeated iteration of (6)—for extracting a single component of the orthogonal decomposition. Note that the initial vector $\theta_{0}$ determines which robust eigenvector will be the convergent point. Computation of subsequent eigenvectors can be computed with deflation, i.e., by subtracting appropriate terms from $T$ .

That is, repeated iteration of (6) starting from $\theta_{0}$ converges to $v_{1}$ at a quadratic rate.

To obtain all eigenvectors, we may simply proceed iteratively using deflation, executing the power method on $T-\sum_{j}\lambda_{j}v_{j}^{\otimes 3}$ after having obtained robust eigenvector / eigenvalue pairs $\{(v_{j},\lambda_{j})\}$ .

Proof Let $\overline{\theta}_{0},\overline{\theta}_{1},\overline{\theta}_{2},\dotsc$ be the sequence given by $\overline{\theta}_{0}:=\theta_{0}$ and $\overline{\theta}_{t}:=T(I,\theta_{t-1},\theta_{t-1})$ for $t\geq 1$ . Let $c_{i}:=v_{i}^{\scriptscriptstyle\top}\theta_{0}$ for all $i\in[k]$ . It is easy to check that (i) $\theta_{t}=\overline{\theta}_{t}/\|\overline{\theta}_{t}\|$ , and (ii) $\overline{\theta}_{t}=\sum_{i=1}^{k}\lambda_{i}^{2^{t}-1}c_{i}^{2^{t}}v_{i}$ . (Indeed, $\overline{\theta}_{t+1}=\sum_{i=1}^{k}\lambda_{i}(v_{i}^{\scriptscriptstyle\top}\overline{\theta}_{t})^{2}v_{i}=\sum_{i=1}^{k}\lambda_{i}(\lambda_{i}^{2^{t}-1}c_{i}^{2^{t}})^{2}v_{i}=\sum_{i=1}^{k}\lambda_{i}^{2^{t+1}-1}c_{i}^{2^{t+1}}v_{i}$ .) Then

Since $\lambda_{1}>0$ , we have $v_{1}^{\scriptscriptstyle\top}\theta_{t}>0$ and hence $\|v_{1}-\theta_{t}\|^{2}=2(1-v_{1}^{\scriptscriptstyle\top}\theta_{t})\leq 2(1-(v_{1}^{\scriptscriptstyle\top}\theta_{t})^{2})$ as required.

2 Perturbation Analysis of a Robust Tensor Power Method

Now we consider the case where we have an approximation $\hat{T}$ to an orthogonally decomposable tensor $T$ . Here, a more robust approach is required to extract an approximate decomposition. We propose such an algorithm in Algorithm 1, and provide a detailed perturbation analysis. For simplicity, we assume the tensor $\hat{T}$ is of size $k\times k\times k$ as per the reduction from Section 4.3. In some applications, it may be preferable to work directly with a $n\times n\times n$ tensor of rank $k\leq n$ (as in Lemma 5.1); our results apply in that setting with little modification.

In our latent variable model applications, $\hat{T}$ is the tensor formed by using empirical moments, while $T$ is the orthogonally decomposable tensor derived from the population moments for the given model. In the context of parameter estimation (as in Section 4.3), $E$ must account for any error amplification throughout the reduction, such as in the whitening step (see, e.g., Hsu and Kakade, 2013, for such an analysis).

The following theorem is similar to Wedin’s perturbation theorem for singular vectors of matrices (Wedin, 1972) in that it bounds the error of the (approximate) decomposition returned by Algorithm 1 on input $\hat{T}$ in terms of the size of the perturbation, provided that the perturbation is small enough.

(Note that the condition on $L$ holds with $L=\operatorname{poly}(k)\log(1/\eta)$ .) Suppose that Algorithm 1 is iteratively called $k$ times, where the input tensor is $\hat{T}$ in the first call, and in each subsequent call, the input tensor is the deflated tensor returned by the previous call. Let $(\hat{v}_{1},\hat{\lambda}_{1}),(\hat{v}_{2},\hat{\lambda}_{2}),\dotsc,(\hat{v}_{k},\hat{\lambda}_{k})$ be the sequence of estimated eigenvector/eigenvalue pairs returned in these $k$ calls. With probability at least $1-\eta$ , there exists a permutation $\pi$ on $[k]$ such that

The proof of Theorem 5.1 is given in Appendix B.

One important difference from Wedin’s theorem is that this is an algorithm dependent perturbation analysis, specific to Algorithm 1 (since the perturbed tensor need not have an orthogonal decomposition). Furthermore, note that Algorithm 1 uses multiple restarts to ensure (approximate) convergence—the intuition is that by restarting at multiple points, we eventually start at a point in which the initial contraction towards some eigenvector dominates the error $E$ in our tensor. The proof shows that we find such a point with high probability within $L=\operatorname{poly}(k)$ trials. It should be noted that for large $k$ , the required bound on $L$ is very close to linear in $k$ .

In general, it is possible, when run on a general symmetric tensor (e.g., $\hat{T}$ ), for the tensor power method to exhibit oscillatory behavior (Kofidis and Regalia, 2002, Example 1). This is not in conflict with Theorem 5.1, which effectively bounds the amplitude of these oscillations; in particular, if $\hat{T}=T+E$ is a tensor built from empirical moments, the error term $E$ (and thus the amplitude of the oscillations) can be driven down by drawing more samples. The practical value of addressing these oscillations and perhaps stabilizing the algorithm is an interesting direction for future research (Kolda and Mayo, 2011).

A final consideration is that for specific applications, it may be possible to use domain knowledge to choose better initialization points. For instance, in the topic modeling applications (cf. Section 3.1), the eigenvectors are related to the topic word distributions, and many documents may be primarily composed of words from just single topic. Therefore, good initialization points can be derived from these single-topic documents themselves, as these points would already be close to one of the eigenvectors.

Discussion

In this section, we discuss some practical and application-oriented issues related to the tensor decomposition approach to learning latent variable models.

A number of practical concerns arise when dealing with moment matrices and tensors. Below, we address two issues that are especially pertinent to topic modeling applications (Anandkumar et al., 2012c, a) or other settings where the observations are sparse.

It can be checked that this quantity is equal to

where the sum is over all ordered word triples in the document. A similar expression is easily derived for the contribution of the document to the empirical second-order moment matrix:

Note that the word count vector $c$ is generally a sparse vector, so this representation allows for efficient multiplication by the moment matrices and tensors in time linear in the size of the document corpus (i.e., the number of non-zero entries in the term-document matrix).

1.2 Dimensionality Reduction

Another serious concern regarding the use of tensor forms of moments is the need to operate on multidimensional arrays with $\Omega(d^{3})$ values (it is typically not exactly $d^{3}$ due to symmetry). When $d$ is large (e.g., when it is the size of the vocabulary in natural language applications), even storing a third-order tensor in memory can be prohibitive. Sparsity is one factor that alleviates this problem. Another approach is to use efficient linear dimensionality reduction. When this is combined with efficient techniques for matrix and tensor multiplication that avoid explicitly constructing the moment matrices and tensors (such as the procedure described above), it is possible to avoid any computational scaling more than linear in the dimension $d$ and the training sample size.

2 Computational Complexity

It is worth noting that the running times differ by roughly a factor of $\Theta(k^{1+\delta})$ , which can be accounted for by the random restarts. This gap can potentially be alleviated or removed by using a more clever method for initialization. Moreover, using special structure in the problem (as discussed above) can also improve the running time of the tensor power method.

3 Sample Complexity Bounds

Previous work on using linear algebraic methods for estimating latent variable models crucially rely on matrix perturbation analysis for deriving sample complexity bounds (Mossel and Roch, 2006; Hsu et al., 2012b; Anandkumar et al., 2012c, a; Hsu and Kakade, 2013). The learning algorithms in these works are plug-in estimators that use empirical moments in place of the population moments, and then follow algebraic manipulations that result in the desired parameter estimates. As long as these manipulations can tolerate small perturbations of the population moments, a sample complexity bound can be obtained by exploiting the convergence of the empirical moments to the population moments via the law of large numbers. As discussed in Appendix D, these approaches do not directly lead to practical algorithms due to a certain amplification of the error (a polynomial factor of $k$ , which is observed in practice).

Using the perturbation analysis for the tensor power method, improved sample complexity bounds can be obtained for all of the examples discussed in Section 3. The underlying analysis remains the same as in previous works (e.g., Anandkumar et al., 2012a; Hsu and Kakade, 2013), the main difference being the accuracy of the orthogonal tensor decomposition obtained via the tensor power method. Relative to the previously cited works, the sample complexity bound will be considerably improved in its dependence on the rank parameter $k$ , as Theorem 5.1 implies that the tensor estimation error (e.g., error in estimating $\widetilde{M}_{3}$ from Section 4.3) is not amplified by any factor explicitly depending on $k$ (there is a requirement that the error be smaller than some factor depending on $k$ , but this only contributes to a lower-order term in the sample complexity bound). See Appendix D for further discussion regarding the stability of the techniques from these previous works.

4 Other Perspectives

The tensor power method is simply one approach for extracting the orthogonal decomposition needed in parameter estimation. The characterizations from Section 4.2 suggest that a number of fixed point and variational techniques may be possible (and Appendix D provides yet another perspective based on simultaneous diagonalization). One important consideration is that the model is often misspecified, and therefore approaches with more robust guarantees (e.g., for convergence) are desirable. Our own experience with the tensor power method (as applied to exchangeable topic modeling) is that while model misspecification does indeed affect convergence, the results can be very reasonable even after just a dozen or so iterations (Anandkumar et al., 2012a). Nevertheless, robustness is likely more important in other applications, and thus the stabilization approaches (Kofidis and Regalia, 2002; Regalia and Kofidis, 2003; Erdogan, 2009; Kolda and Mayo, 2011) may be advantageous.

We thank Boaz Barak, Dean Foster, Jon Kelner, and Greg Valiant for helpful discussions. We are also grateful to Hanzhang Hu, Drew Bagnell, and Martial Hebert for alerting us of an issue with Theorem 4.2 and suggesting a simple fix. This work was completed while DH was a postdoctoral researcher at Microsoft Research New England, and partly while AA, RG, and MT were visiting the same lab. AA is supported in part by the NSF Award CCF-1219234, AFOSR Award FA9550-10-1-0310 and the ARO Award W911NF-12-1-0404.

A Fixed-Point and Variational Characterizations of Orthogonal Tensor Decompositions

We give detailed proofs of Theorems 4.1 and 4.2 in this section for completeness.

Let $T$ have an orthogonal decomposition as given in (4).

The set of robust eigenvectors of $T$ is $\{v_{1},v_{2},\dotsc,v_{k}\}$ .

We now prove the second claim. First, we show that every $v_{i}$ is a robust eigenvector. Pick any $i\in[k]$ , and note that for a sufficiently small ball around $v_{i}$ , we have that for all $\theta$ in this ball, $\lambda_{i}v_{i}^{\scriptscriptstyle\top}\theta$ is strictly greater than $\lambda_{j}v_{j}^{\scriptscriptstyle\top}\theta$ for $j\in[k]\setminus\{i\}$ . Thus by Lemma 5.1, $v_{i}$ is a robust eigenvector. Now we show that the $v_{i}$ are the only robust eigenvectors. Suppose there exists some robust eigenvector $u$ not equal to $v_{i}$ for any $i\in[k]$ . Then there exists a positive measure set around $u$ such that all points in this set converge to $u$ under repeated iteration of (6). This contradicts the first claim.

A.2 Proof of Theorem 4.2

Let $T$ have an orthogonal decomposition as given in (4), and consider the optimization problem

The stationary points are eigenvectors of $T$ .

A stationary point $u$ is an isolated local maximizer if and only if $u=v_{i}$ for some $i\in[k]$ .

Now we characterize the isolated local maximizers. Observe that if $u\neq 0$ and $T(I,u,u)=\lambda u$ for $\lambda<0$ , then $T(u,u,u)<0$ . Therefore $u^{\prime}=(1-\delta)u$ for any $\delta\in(0,1)$ satisfies $T(u^{\prime},u^{\prime},u^{\prime})=(1-\delta)^{3}T(u,u,u)>T(u,u,u)$ . So such a $u$ cannot be a local maximizer. Moreover, if $\|u\|<1$ and $T(I,u,u)=\lambda u$ for $\lambda>0$ , then $u^{\prime}=(1+\delta)u$ for a small enough $\delta\in(0,1)$ satisfies $\|u^{\prime}\|\leq 1$ and $T(u^{\prime},u^{\prime},u^{\prime})=(1+\delta)^{3}T(u,u,u)>T(u,u,u)$ . Therefore a local maximizer must have $T(I,u,u)=\lambda u$ for some $\lambda\geq 0$ , and $\|u\|=1$ whenever $\lambda>0$ .

The point $u$ is an isolated local maximum if the above quantity is strictly negative for all unit vectors $w$ orthogonal to $u$ . We now consider three cases depending on the cardinality of $\Omega$ and the sign of $\lambda$ .

Case 1: $|\Omega|=1$ and $\lambda>0$ . This means $u=v_{i}$ for some $i\in[k]$ (as $u=-v_{i}$ implies $\lambda=-\lambda_{i}<0$ ). In this case,

Case 2: $|\Omega|\geq 2$ and $\lambda>0$ . Since $|\Omega|\geq 2$ , we may pick a strict non-empty subset $S\subsetneq\Omega$ and set

Since $\epsilon\leq\delta$ , for small enough $\delta$ , the RHS is strictly greater than $\lambda$ . This implies that $u$ is not an isolated local maximizer.

for sufficiently small $\delta$ . Thus $u$ is not an isolated local maximizer.

From these exhaustive cases, we conclude that a stationary point $u$ is an isolated local maximizer if and only if $u=v_{i}$ for some $i\in[k]$ .

We are grateful to Hanzhang Hu, Drew Bagnell, and Martial Hebert for alerting us of an issue with our original statement of Theorem 4.2 and its proof, and for suggesting a simple fix. The original statement used the optimization constraint $\|u\|=1$ (rather than $\|u\|\leq 1$ ), but the characterization of the decomposition with this constraint is then only given by isolated local maximizers $u$ with the additional constraint that $T(u,u,u)>0$ —that is, there can be isolated local maximizers with $T(u,u,u)\leq 0$ that are not vectors in the decomposition. The suggested fix of Hu, Bagnell, and Herbert is to relax to $\|u\|\leq 1$ , which eliminates isolated local maximizers with $T(u,u,u)\leq 0$ ; this way, the characterization of the decomposition is simply the isolated local maximizers under the relaxed constraint.

B Analysis of Robust Power Method

In this section, we prove Theorem 5.1. The proof is structured as follows. In Appendix B.1, we show that with high probability, at least one out of $L$ random vectors will be a good initializer for the tensor power iterations. An initializer is good if its projection onto an eigenvector is noticeably larger than its projection onto other eigenvectors. We then analyze in Appendix B.2 the convergence behavior of the tensor power iterations. Relative to the proof of Lemma 5.1, this analysis is complicated by the tensor perturbation. We show that there is an initial slow convergence phase (linear rate rather than quadratic), but as soon as the projection of the iterate onto an eigenvector is large enough, it enters the quadratic convergence regime until the perturbation dominates. Finally, we show how errors accrue due to deflation in Appendix B.3, which is rather subtle and different from deflation with matrix eigendecompositions. This is because when some initial set of eigenvectors and eigenvalues are accurately recovered, the additional errors due to deflation are effectively only lower-order terms. These three pieces are assembled in Appendix B.4 to complete the proof of Theorem 5.1.

There exists an absolute constant $c>0$ such that if positive integer $L\geq 2$ satisfies

It suffices to show that with probability at least $1/2$ , there is a column $j^{*}\in[L]$ such that

Since $\max_{j\in[L]}|Z_{1,j}|$ is a $1$ -Lipschitz function of $L$ independent $\mathcal{N}(0,1)$ random variables, it follows that

Observe that the cumulative distribution function of $\max_{j\in[L]}Z_{1,j}$ is given by $F(z)=\Phi(z)^{L}$ , where $\Phi$ is the standard Gaussian CDF. Since $F(m)=1/2$ , it follows that $m=\Phi^{-1}(2^{-1/{L}})$ . It can be checked that

for some absolute constant $c>0$ . Also, let $j^{*}:=\arg\max_{j\in[L]}|Z_{1,j}|$ .

Now for each $j\in[L]$ , let $|Z_{2:k,j}|:=\max\{|Z_{2,j}|,|Z_{3,j}|,\dotsc,|Z_{k,j}|\}$ . Again, since $|Z_{2:k,j}|$ is a $1$ -Lipschitz function of $k-1$ independent $\mathcal{N}(0,1)$ random variables, it follows that

Since $|Z_{2:k,j}|$ is independent of $|Z_{1,j}|$ for all $j\in[L]$ , it follows that the previous two displayed inequalities also hold with $j$ replaced by $j^{*}$ .

Therefore we conclude with a union bound that with probability at least $1/2$ ,

Since $L$ satisfies (10) by assumption, in this event, the $j^{*}$ -th random vector is $\gamma$ -separated.

B.2 Tensor Power Iterations

where $\{v_{1},v_{2},\dotsc,v_{k}\}$ is an orthonormal basis, and, without loss of generality,

for all $i\in[k]$ . Combining (17) and (18) gives

Moreover, by the triangle inequality and Hölder’s inequality,

In terms of $R_{t+1}$ , $R_{t}$ , $\gamma_{t}$ , and $\delta_{t}$ , this reads

where the last inequality follows from Proposition B.1.

and $\gamma_{t}>2(1+2\kappa\rho^{2})\delta_{t}$ .

If $r_{i,t}^{2}\leq 2\rho^{2}$ , then $r_{i,t+1}\geq|r_{i,t}|\bigl{(}1+\frac{\gamma_{t}}{2}\bigr{)}$ .

If $\rho^{2}<r_{i,t}^{2}$ , then $r_{i,t+1}\geq\min\{r_{i,t}^{2}/\rho,\ \frac{1-\delta_{t}-1/\rho}{\kappa\delta_{t}}\}$ .

$\gamma_{t+1}\geq\min\{\gamma_{t},1-1/\rho\}$ .

If $R_{t}\leq 1+2\kappa\rho^{2}$ , then $R_{t+1}\geq R_{t}\bigl{(}1+\frac{\gamma_{t}}{3}\bigr{)}$ , $\theta_{1,t+1}^{2}\geq\theta_{1,t}^{2}$ , and $\delta_{t+1}\leq\delta_{t}$ .

Proof Consider two (overlapping) cases depending on $r_{i,t}^{2}$ .

Case 1: $r_{i,t}^{2}\leq 2\rho^{2}$ . By (15) from Proposition B.2,

where the last inequality uses the assumption $\gamma_{t}>2(1+2\kappa\rho^{2})\delta_{t}$ . This proves the first claim.

Case 2: $\rho^{2}<r_{i,t}^{2}$ . We split into two sub-cases. Suppose $r_{i,t}^{2}\leq(\rho(1-\delta_{t})-1)/(\kappa\delta_{t})$ . Then, by (15),

Now suppose instead $r_{i,t}^{2}>(\rho(1-\delta_{t})-1)/(\kappa\delta_{t})$ . Then

Observe that if $\min_{i\neq 1}r_{i,t}^{2}\leq(\rho(1-\delta_{t})-1)/(\kappa\delta_{t})$ , then $r_{i,t+1}\geq|r_{i,t}|$ for all $i\in[k]$ , and hence $\gamma_{t+1}\geq\gamma_{t}$ . Otherwise we have $\gamma_{t+1}>1-\frac{\kappa\delta_{t}}{1-\delta_{t}-1/\rho}>1-1/\rho$ . This proves the third claim.

If $\min_{i\neq 1}r_{i,t}^{2}>(\rho(1-\delta_{t})-1)/(\kappa\delta_{t})$ , then we may apply the inequality (20) from the second sub-case of Case 2 above to get

Finally, for the last claim, if $R_{t}\leq 1+2\kappa\rho^{2}$ , then by (16) from Proposition B.2 and the assumption $\gamma_{t}>2(1+2\kappa\rho^{2})\delta_{t}$ ,

This in turn implies that $\theta_{1,t+1}^{2}\geq\theta_{1,t}^{2}$ via Proposition B.1, and thus $\delta_{t+1}\leq\delta_{t}$ .

Assume $0\leq\delta_{t}<1/2$ and $\gamma_{t}>0$ . Pick any $\beta>\alpha>0$ such that

If $R_{t}\geq 1/\alpha$ , then $R_{t+1}\geq 1/\alpha$ .

If $1/\alpha>R_{t}\geq 1/\beta$ , then $R_{t+1}\geq\min\{R_{t}^{2}/(2\kappa),\ 1/\alpha\}$ .

Now consider the following cases depending on $R_{t}$ .

Case 1: $R_{t}\geq 1/\alpha$ . In this case, we have

by (21) (with $c=\alpha$ ) and the condition on $\alpha$ . Combining this with (16) from Proposition B.2 gives

Case 2: $1/\beta\leq R_{t}<1/\alpha$ . In this case, we have

by (21) (with $c=\beta$ ) and the conditions on $\alpha$ and $\beta$ . If $\delta_{t}\geq 1/(2+R_{t}^{2}/\kappa)$ , then (16) implies

If instead $\delta_{t}<1/(2+R_{t}^{2}/\kappa)$ , then (16) implies

Proof Assume without loss of generality that $i^{*}=1$ . We consider three phases: (i) iterations before the first time $t$ such that $R_{t}>1+2\kappa\rho^{2}=1+8\kappa$ (using $\rho:=2$ ), (ii) the subsequent iterations before the first time $t$ such that $R_{t}\geq 1/\alpha$ (where $\alpha$ will be defined below), and finally (iii) the remaining iterations.

and hence the preconditions on $\delta_{t}$ and $\gamma_{t}$ of Lemma B.2 hold for $t=0$ . For all $t\in T_{1}$ satisfying the preconditions, Lemma B.2 implies that $\delta_{t+1}\leq\delta_{t}$ and $\gamma_{t+1}\geq\min\{\gamma_{t},1-1/\rho\}$ , so the next iteration also satisfies the preconditions. Hence by induction, the preconditions hold for all iterations in $T_{1}$ . Moreover, for all $i\in[k]$ , we have

iterations in $T_{1}$ . As soon as $\min_{i\neq 1}r_{i,t}^{2}>\frac{1-\delta_{t}-1/\rho}{\kappa\delta_{t}}$ , we have that in the next iteration,

and all the while $R_{t}$ is growing at a linear rate (given in Lemma B.2, claim 5). Therefore, there are at most an additional

iterations in $T_{1}$ over that counted in (22). Therefore, by combining the counts in (22) and (23), we have that the number of iterations in the first phase satisfies

We now analyze the second phase, i.e., the iterates in $T_{2}:=\{t\geq 0:t\notin T_{1},\ R_{t}<1/\alpha\}$ . Define

Therefore the total number of iterations before $R_{t}\geq 1/\alpha$ is

After $R_{t^{\prime\prime}}\geq 1/\alpha$ (for $t^{\prime\prime}:=\max(T_{1}\cup T_{2})+1$ ), we have

we also replace $\alpha$ and $\beta$ with $\overline{\alpha}$ and $\overline{\beta}$ , which we set to

Therefore, the preconditions of Lemma B.3 are satisfied for the initial iteration $t^{\prime\prime}$ in this final phase, and by the same arguments as before, the preconditions hold for all subsequent iterations $t\geq t^{\prime\prime}$ . Initially, we have $R_{t^{\prime\prime}}\geq 1/\alpha\geq 1/\overline{\beta}$ , and by Lemma B.3, we have that $R_{t}$ increases at a quadratic rate in this final phase until $R_{t}\geq 1/\overline{\alpha}$ . So the number of iterations before $R_{t}\geq 1/\overline{\alpha}$ can be bounded as

Once $R_{t}\geq 1/\overline{\alpha}$ , we have

Since $\operatorname{sign}(\theta_{1,t})=r_{1,t}\geq r_{1,t-1}^{2}\cdot(1-\overline{\delta}_{t-1})/(1+\kappa\overline{\delta}_{t-1}r_{1,t-1}^{2})=(1-\overline{\delta}_{t-1})/(1+\kappa\overline{\delta}_{t-1})>0$ by Proposition B.2, we have $\theta_{1,t}>0$ . Therefore we can conclude that

B.3 Deflation

for all $i\in[t]$ , then for any unit vector $u\in S^{k-1}$ ,

Proof For any unit vector $u$ and $i\in[t]$ , the error term

lives in $\operatorname{span}\{v_{i},\hat{v}_{i}\}$ ; this space is the same as $\operatorname{span}\{v_{i},\hat{v}_{i}^{\perp}\}$ , where

is the projection of $\hat{v}_{i}$ onto the subspace orthogonal to $v_{i}$ . Since $\|\hat{v}_{i}-v_{i}\|^{2}=2(1-v_{i}^{\scriptscriptstyle\top}\hat{v}_{i})$ , it follows that

(the inequality follows from the assumption $\|\hat{v}_{i}-v_{i}\|\leq\sqrt{2}$ , which in turn implies $0\leq c_{i}\leq 1$ ). By the Pythagorean theorem and the above inequality for $c_{i}$ ,

Later, we will also need the following bound, which is easily derived from the above inequalities and the triangle inequality:

We now express $\mathcal{E}_{i}(I,u,u)$ in terms of the coordinate system defined by $v_{i}$ and $\hat{v}_{i}^{\perp}$ , depicted below. Define

(Note that the part of $u$ living in $\operatorname{span}\{v_{i},\hat{v}_{i}^{\perp}\}^{\perp}$ is irrelevant for analyzing $\mathcal{E}_{i}(I,u,u)$ .) We have

The overall error can also be expressed in terms of the $A_{i}$ and $B_{i}$ :

where the first inequality uses the fact $(x+y)^{2}\leq 2(x^{2}+y^{2})$ and the triangle inequality, and the second inequality uses the orthonormality of the $v_{i}$ and the triangle inequality.

and therefore (via $(x+y)^{2}\leq 2(x^{2}+y^{2})$ )

The second term, $|B_{i}|$ , is bounded similarly:

Therefore, using the inequality from (25) and again $(x+y)^{2}\leq 2(x^{2}+y^{2})$ ,

B.4 Proof of the Main Theorem

Proof We prove by induction that for each $i\in[k]$ (corresponding to the $i$ -th call to Algorithm 1), with probability at least $1-i\eta/k$ , there exists a permutation $\pi$ on $[k]$ such that the following assertions hold.

For all $j\leq i$ , $\|v_{\pi(j)}-\hat{v}_{j}\|\leq 8\epsilon/\lambda_{\pi(j)}$ and $|\lambda_{\pi(j)}-\hat{\lambda}_{j}|\leq 12\epsilon$ .

We actually take $i=0$ as the base case, so we can ignore the first assertion, and just observe that for $i=0$ ,

Now fix some $i\in[k]$ , and assume as the inductive hypothesis that, with probability at least $1-(i-1)\eta/k$ , there exists a permutation $\pi$ such that two assertions above hold for $i-1$ (call this $\mathsf{Event}_{i-1}$ ). The $i$ -th call to Algorithm 1 takes as input

which is intended to be an approximation to

By Lemma B.1, with conditional probability at least $1-\eta/k$ given $\mathsf{Event}_{i-1}$ , at least one of $\theta_{0}^{(\tau)}$ for $\tau\in[L]$ is $\gamma$ -separated relative to $\pi(j_{\max})$ , where $j_{\max}:=\arg\max_{j\geq i}\lambda_{\pi(j)}$ , (for $\gamma=0.01$ ; call this $\mathsf{Event}_{i}^{\prime}$ ; note that the application of Lemma B.1 determines $C_{3}$ ). Therefore $\Pr[\mathsf{Event}_{i-1}\cap\mathsf{Event}_{i}^{\prime}]=\Pr[\mathsf{Event}_{i}^{\prime}|\mathsf{Event}_{i-1}]\Pr[\mathsf{Event}_{i-1}]\geq(1-\eta/k)(1-(i-1)\eta/k)\geq 1-i\eta/k$ . It remains to show that $\mathsf{Event}_{i-1}\cap\mathsf{Event}_{i}^{\prime}\subseteq\mathsf{Event}_{i}$ ; so henceforth we condition on $\mathsf{Event}_{i-1}\cap\mathsf{Event}_{i}^{\prime}$ .

On the other hand, by the triangle inequality,

where $j^{*}:=\arg\max_{j\geq i}\lambda_{\pi(j)}|\theta_{\pi(j),N}|$ . Therefore

Squaring both sides and using the fact that $\theta_{\pi(j^{*}),N}^{2}+\theta_{\pi(j),N}^{2}\leq 1$ for any $j\neq j^{*}$ ,

This means that $\theta_{N}$ is $(1/4)$ -separated relative to $\pi(j^{*})$ . Also, observe that

Since $\hat{v}_{i}=\hat{\theta}$ and $\hat{\lambda}_{i}=\hat{\lambda}$ , the first assertion of the inductive hypothesis is satisfied, as we can modify the permutation $\pi$ by swapping $\pi(i)$ and $\pi(j^{*})$ without affecting the values of $\{\pi(j):j\leq i-1\}$ (recall $j^{*}\geq i$ ).

To prove that (27) holds, pick any unit vector $u\in S^{k-1}$ such that there exists $j^{\prime}\geq i+1$ with $(u^{\scriptscriptstyle\top}v_{\pi(j^{\prime})})^{2}\geq 1-(168\epsilon/\lambda_{\pi(j^{\prime})})^{2}$ . We have, via the second bound on $C_{1}$ in (28) and the corresponding assumed bound $\epsilon\leq C_{1}\cdot\frac{\lambda_{\min}}{k}$ ,

From the last induction step ( $i=k$ ), it is also clear from (29) that $\|T-\sum_{j=1}^{k}\hat{\lambda}_{j}\hat{v}_{j}^{\otimes 3}\|\leq 55\epsilon$ (in $\mathsf{Event}_{k-1}\cap\mathsf{Event}_{k}^{\prime}$ ). This completes the proof of the theorem.

C Variant of Robust Power Method that uses a Stopping Condition

In this section we analyze a variant of Algorithm 1 that uses a stopping condition. The variant is described in Algorithm 2. The key difference is that the inner for-loop is repeated until a stopping condition is satisfied (rather than explicitly $L$ times). The stopping condition ensures that the power iteration is converging to an eigenvector, and it will be satisfied within $\operatorname{poly}(k)$ random restarts with high probability. The condition depends on one new quantity, $r$ , which should be set to $r:=k-\text{\# deflation steps so far}$ (i.e., the first call to Algorithm 2 uses $r=k$ , the second call uses $r=k-1$ , and so on).

For a matrix $A$ , we use $\|A\|_{F}:=(\sum_{i,j}A_{i,j}^{2})^{1/2}$ to denote its Frobenius norm. For a third-order tensor $A$ , we use $\|A\|_{F}:=(\sum_{i}\|A(I,I,e_{i})\|_{F}^{2})^{1/2}=(\sum_{i}\|A(I,I,v_{i})\|_{F}^{2})^{1/2}$ .

Combining these bounds on $|\phi_{1}|$ gives

which in turn implies (for $\alpha\in(0,1/20)$ )

Now we prove the final claim. This is done by (i) showing that $\theta$ has a large projection onto $u_{1}$ , (ii) using an SVD perturbation argument to show that $\pm u_{1}$ is close to $v_{1}$ , and (iii) concluding that $\theta$ has a large projection onto $v_{1}$ .

so we may apply Wedin’s theorem to obtain

It remains to show that $\theta_{1}=v_{1}^{\scriptscriptstyle\top}\theta$ is large. Indeed, by the triangle inequality, Cauchy-Schwarz, and the above inequalities on $(u_{1}^{\scriptscriptstyle\top}v_{1})^{2}$ and $(u_{1}^{\scriptscriptstyle\top}\theta)^{2}$ ,

so $\operatorname{sign}(\theta_{1})>-1$ , meaning $\theta_{1}>0$ . Therefore $\theta_{1}=|\theta_{1}|\geq 1-2\alpha$ . This proves the final claim.

To the conclusion of Lemma B.4, it can be added that the stopping condition (31) is satisfied by $\theta=\theta_{t}$ .

Proof Without loss of generality, assume $i^{*}=1$ . By the triangle inequality and Cauchy-Schwarz,

Using the definition of the tensor Frobenius norm, we have

Combining this with the above inequality implies

Therefore the stopping condition (31) is satisfied.

C.2 Sketch of Analysis of Algorithm 2

The analysis of Algorithm 2 is very similar to the proof of Theorem 5.1 for Algorithm 1, so here we just sketch the essential differences.

First, the guarantee afforded to Algorithm 2 is somewhat different than Theorem 5.1. Specifically, it is of the following form: (i) under appropriate conditions, upon termination, the algorithm returns an accurate decomposition, and (ii) the algorithm terminates after $\operatorname{poly}(k)$ random restarts with high probability.

The conditions on $\epsilon$ and $N$ are the same (but for possibly different universal constants $C_{1},C_{2}$ ). In Lemma C.1 and Lemma C.2, there is reference to a condition on the Frobenius norm of $E$ , but we may use the inequality $\|E\|_{F}\leq k\|E\|\leq k\epsilon$ so that the condition is subsumed by the $\epsilon$ condition.

Now we outline the differences relative to the proof of Theorem 5.1. The basic structure of the induction argument is the same. In the induction step, we argue that (i) if the stopping condition is satisfied, then by Lemma C.1 (with $\alpha=0.05$ and $\beta=1/2$ ), we have a vector $\theta_{N}$ such that, for some $j^{*}\geq i$ ,

$\lambda_{\pi(j^{*})}\geq\lambda_{\pi(j_{\max})}/(4\sqrt{k})$ ;

$\theta_{N}$ is $(1/4)$ -separated relative to $\pi(j^{*})$ ;

and (ii) the stopping condition is satisfied within $\operatorname{poly}(k)$ random restarts (via Lemma B.1 and Lemma C.2) with high probability. We now invoke Lemma B.4 to argue that executing another $N$ power iterations starting from $\theta_{N}$ gives a vector $\hat{\theta}$ that satisfies

The main difference here, relative to the proof of Theorem 5.1, is that we use $\kappa:=4\sqrt{k}$ (rather than $\kappa=O(1)$ ), but this ultimately leads to the same guarantee after taking into consideration the condition $\epsilon\leq C_{1}\lambda_{\min}/k$ . The remainder of the analysis is essentially the same as the proof of Theorem 5.1.

D Simultaneous Diagonalization for Tensor Decomposition

As discussed in the introduction, another standard approach to certain tensor decomposition problems is to simultaneously diagonalize a collection of similar matrices obtained from the given tensor. We now examine this approach in the context of our latent variable models, where

Let $V:=[\mu_{1}|\mu_{2}|\dotsb|\mu_{k}]$ and $D(\eta):=\operatorname{diag}(\mu_{1}^{\scriptscriptstyle\top}\eta,\mu_{2}^{\scriptscriptstyle\top}\eta,\dotsc,\mu_{k}^{\scriptscriptstyle\top}\eta)$ , so

Thus, the problem of determining the $\mu_{i}$ can be cast as a simultaneous diagonalization problem: find a matrix $X$ such that $X^{\scriptscriptstyle\top}M_{2}X$ and $X^{\scriptscriptstyle\top}M_{3}(I,I,\eta)X$ (for all $\eta$ ) are diagonal. It is easy to see that if the $\mu_{i}$ are linearly independent, then the solution $X^{\scriptscriptstyle\top}=V^{{\dagger}}$ is unique up to permutation and rescaling of the columns.

With exact moments, a simple approach is as follows. Assume for simplicity that $d=k$ , and define

It should be noted, however, that the use of a single random choice of $\eta$ is quite restrictive, and it is easy to see that a simultaneous diagonalization of $M(\eta)$ for several choices of $\eta$ can be beneficial. While the uniqueness of the eigendecomposition of $M(\eta)$ is only guaranteed when the diagonal entries of $D(\eta)$ are distinct, the simultaneous diagonalization of $M(\eta^{(1)}),M(\eta^{(2)}),\dotsc,M(\eta^{(m)})$ for vectors $\eta^{(1)},\eta^{(2)},\dotsc,\eta^{(m)}$ is unique as long as the columns of

are distinct (i.e., for each pair of column indices $i,j$ , there exists a row index $r$ such that the $(r,i)$ -th and $(r,j)$ -th entries are distinct). This is a much weaker requirement for uniqueness, and therefore may translate to an improved perturbation analysis. In fact, using the techniques discussed in Section 4.3, we may even reduce the problem to an orthogonal simultaneous diagonalization, which may be easier to obtain. Furthermore, a number of robust numerical methods for (approximately) simultaneously diagonalizing collections of matrices have been proposed and used successfully in the literature (e.g., Bunse-Gerstner et al., 1993; Cardoso and Souloumiac, 1993; Cardoso, 1994; Cardoso and Comon, 1996; Ziehe et al., 2004). Another alternative and a more stable approach compared to full diagonalization is a Schur-like method which finds a unitary matrix $U$ which simultaneously triangularizes the respective matrices (Corless et al., 1997). It is an interesting open question whether these techniques can yield similar improved learnability results and also enjoy the attractive computational properties of the tensor power method.