The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices

Florent Benaych-Georges, Raj Rao Nadakuditi

Introduction

Let $X_{n}$ be an $n\times n$ symmetric (or Hermitian) matrix with eigenvalues $\lambda_{1}(X_{n}),\ldots,\lambda_{n}(X_{n})$ and $P_{n}$ be an $n\times n$ symmetric (or Hermitian) matrix with rank $r\leq n$ and non-zero eigenvalues $\theta_{1},\ldots,\theta_{r}$ . A fundamental question in matrix analysis is the following :

How are the eigenvalues and eigenvectors of $X_{n}+P_{n}$ related to the eigenvalues and eigenvectors of $X_{n}$ and $P_{n}$ ?

When $X_{n}$ and $P_{n}$ are diagonalized by the same eigenvectors, we have $\lambda_{i}(X_{n}+P_{n})=\lambda_{j}(X_{n})+\lambda_{k}(P_{n})$ for appropriate choice of indices $i,j,k\in\{1,\ldots,n\}$ . In the general setting, however, the answer is complicated by the fact that the eigenvalues and eigenvectors of their sum depend on the relationship between the eigenspaces of the individual matrices.

In this scenario, one can use Weyl’s interlacing inequalities and Horn inequalities to obtain coarse bounds for the eigenvalues of the sum in terms of the eigenvalues of $X_{n}$ . When the norm of $P_{n}$ is small relative to the norm of $X_{n}$ , tools from perturbation theory (see [22, Chapter 6] or ) can be employed to improve the characterization of the bounded set in which the eigenvalues of the sum must lie. Exploiting any special structure in the matrices allows us to refine these bounds but this is pretty much as far as the theory goes. Instead of exact answers we must resort to a system of coupled inequalities. Describing the behavior of the eigenvectors of the sum is even more complicated.

Surprisingly, adding some randomness to the eigenspaces permits further analytical progress. Specifically, if the eigenspaces are assumed to be “in generic position with respect to each other”, then in place of eigenvalue bounds we have simple, exact answers that are to be interpreted probabilistically. These results bring into focus a phase transition phenomenon of the kind illustrated in Figure 1 for the eigenvalues and eigenvectors of $X_{n}+P_{n}$ and $X_{n}\times(I_{n}+P_{n})$ . A precise statement of the results may be found in Section 2.

The development of the eigenvector aspect is another contribution that we would like to highlight. Generally speaking, the eigenvector question has received less attention in random matrix theory and in free probability theory. A notable exception is the recent body of work on the eigenvectors of spiked Wishart matrices which corresponds to $\mu_{X}$ being the Marčenko-Pastur measure. In this paper, we extend their results for multiplicative models of the kind $(I+P_{n})^{1/2}X_{n}(I+P_{n})^{1/2}$ to the setting where $\mu_{X}$ is an arbitrary probability measure and obtain new results for the eigenvectors for additive models of the form $X_{n}+P_{n}$ .

Our proofs rely on the derivation of master equation representations of the eigenvalues and eigenvectors of the perturbed matrix and the subsequent application of concentration inequalities for random vectors uniformly distributed on high dimensional unit spheres (such as the ones appearing in ) to these implicit master equation representations. Consequently, our technique is simpler, more general and brings into focus the source of the phase transition phenomenon. The underlying methods can and have been adapted to study the extreme singular values and singular vectors of deformations of rectangular random matrices, as well as the fluctuations and the large deviations of our model.

The paper is organized as follows. In Section 2, we state the main results and present the integral transforms alluded to above. Section 3 presents some examples. An outline of the proofs is presented in Section 4. Exact master equation representations of the eigenvalues and the eigenvectors of the perturbed matrices are derived in Section 5 and utilized in Section 6 to prove the main results. Technical results needed in these proofs have been relegated to the Appendix.

Main results

Let $X_{n}$ be an $n\times n$ symmetric (or Hermitian) random matrix whose ordered eigenvalues we denote by $\lambda_{1}(X_{n})\geq\cdots\geq\lambda_{n}(X_{n})$ . Let $\mu_{X_{n}}$ be the empirical eigenvalue distribution, i.e., the probability measure defined as

Assume that the probability measure $\mu_{X_{n}}$ converges almost surely weakly, as $n\longrightarrow\infty$ , to a non-random compactly supported probability measure $\mu_{X}$ . Let $a$ and $b$ be, respectively, the infimum and supremum of the support of $\mu_{X}$ . We suppose the smallest and largest eigenvalue of $X_{n}$ converge almost surely to $a$ and $b$ .

For a given $r\geq 1$ , let $\theta_{1}\geq\cdots\geq\theta_{r}$ be deterministic non-zero real numbers, chosen independently of $n$ . For every $n$ , let $P_{n}$ be an $n\times n$ symmetric (or Hermitian) random matrix having rank $r$ with its $r$ non-zero eigenvalues equal to $\theta_{1},\ldots,\theta_{r}$ . Let the index $s\in\{0,\ldots,r\}$ be defined such that $\theta_{1}\geq\cdots\geq\theta_{s}>0>\theta_{s+1}\geq\cdots\geq\theta_{r}$ .

Recall that a symmetric (or Hermitian) random matrix is said to be orthogonally invariant (or unitarily invariant) if its distribution is invariant under the action of the orthogonal (or unitary) group under conjugation.

We suppose that $X_{n}$ and $P_{n}$ are independent and that either $X_{n}$ or $P_{n}$ is orthogonally (or unitarily) invariant.

2. Notation

we also let $\overset{\textrm{a.s.}}{\longrightarrow}$ denote almost sure convergence. The ordered eigenvalues of an $n\times n$ Hermitian matrix $M$ will be denoted by $\lambda_{1}(M)\geq\cdots\geq\lambda_{n}(M)$ . Lastly, for a subspace $F$ of a Euclidian space $E$ and a vector $x\in E$ , we denote the norm of the orthogonal projection of $x$ onto $F$ by $\langle x,F\rangle$ .

3. Extreme eigenvalues and eigenvectors under additive perturbations

Consider the rank $r$ additive perturbation of the random matrix $X_{n}$ given by

The extreme eigenvalues of $\widetilde{X}_{n}$ exhibit the following behavior as $n\longrightarrow\infty$ . We have that for each $1\leq i\leq s$ ,

while for each fixed $i>s$ , $\lambda_{i}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}b$ .

Similarly, for the smallest eigenvalues, we have that for each $0\leq j<r-s$ ,

while for each fixed $j\geq r-s$ , $\lambda_{n-j}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}a$ .

is the Cauchy transform of $\mu_{X}$ , $G_{\mu_{X}}^{-1}(\cdot)$ is its functional inverse so that $1/\textpm\pm\infty$ stands for .

Consider ${i_{0}}\in\{1,\ldots,r\}$ such that $1/\theta_{i_{0}}\in(G_{\mu_{X}}(a^{-}),G_{\mu_{X}}(b^{+}))$ . For each $n$ , define

and let $\widetilde{u}$ be a unit-norm eigenvector of $\widetilde{X}_{n}$ associated with the eigenvalue $\widetilde{\lambda}_{i_{0}}$ . Then we have, as $n\longrightarrow\infty$ ,

When $r=1$ , let the sole non-zero eigenvalue of $P_{n}$ be denoted by $\theta$ . Suppose that

For each $n$ , let $\widetilde{u}$ be a unit-norm eigenvector of $\widetilde{X}_{n}$ associated with either the largest or smallest eigenvalue depending on whether $\theta>0$ or $\theta<0$ , respectively. Then we have

The following proposition allows to assert that in many classical matrix models, such as Wigner or Wishart matrices, the above phase transitions actually occur with a finite threshold. The proposition is phrased in terms of $b$ , the supremum of the support of $\mu_{X}$ , but also applies for $a$ , the infimum of the support of $\mu_{X}$ . The proof relies on a straightforward computation which we omit.

Assume that the limiting eigenvalue distribution $\mu_{X}$ has a density $f_{\mu_{X}}$ with a power decay at $b$ , i.e., that, as $t\to b$ with $t<b$ , $f_{\mu_{X}}(t)\sim c(b-t)^{\alpha}$ for some exponent $\alpha>-1$ and some constant $c$ . Then:

so that the phase transitions in Theorems 2.1 and 2.3 manifest for $\alpha=1/2$ .

Under additional hypotheses on the manner in which the empirical eigenvalue distribution of $X_{n}\overset{\textrm{a.s.}}{\longrightarrow}\mu_{X}$ as $n\longrightarrow\infty$ , Theorem 2.2 can be generalized to any eigenvalue with limit $\rho$ equal either to $a$ or $b$ such that $G_{\mu_{X}}^{\prime}(\rho)$ is finite. In the same way, Theorem 2.3 can be generalized for any value of $r$ . The specific hypothesis has to do with requiring the spacings between the $\lambda_{i}(X_{n})$ ’s to be more “random matrix like” and exhibit repulsion instead of being “independent sample like” with possible clumping. We plan to develop this line of inquiry in a separate paper.

4. Extreme eigenvalues and eigenvectors under multiplicative perturbations

We maintain the same hypotheses as before so that the limiting probability measure $\mu_{X}$ , the index $s$ and the rank $r$ matrix $P_{n}$ are defined as in Section 2.1. In addition, we assume that for every $n$ , $X_{n}$ is a non-negative definite matrix and that the limiting probability measure $\mu_{X}$ is not the Dirac mass at zero.

Consider the rank $r$ multiplicative perturbation of the random matrix $X_{n}$ given by

The extreme eigenvalues of $\widetilde{X}_{n}$ exhibit the following behavior as $n\longrightarrow\infty$ . We have that for $1\leq i\leq s$ ,

while for each fixed $i>s$ , $\lambda_{i}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}b$ .

In the same way, for the smallest eigenvalues, for each $0\leq j<r-s$ ,

while for each fixed $j\geq r-s$ , $\lambda_{n-j}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}a$ .

is the T-transform of $\mu_{X}$ , $T_{\mu_{X}}^{-1}(\cdot)$ is its functional inverse and $1/\pm\infty$ stands for .

Consider ${{i_{0}}}\in\{1,\ldots,r\}$ such that $1/\theta_{i_{0}}\in(T_{\mu_{X}}(a^{-}),T_{\mu_{X}}(b^{+}))$ . For each $n$ , define

and let $\widetilde{u}$ be a unit-norm eigenvector of $\widetilde{X}_{n}$ associated with the eigenvalue $\widetilde{\lambda}_{i_{0}}$ . Then we have, as $n\longrightarrow\infty$ ,

When $r=1$ , let the sole non-zero eigenvalue of $P_{n}$ be denoted by $\theta$ . Suppose that

For each $n$ , let $\widetilde{u}$ be the unit-norm eigenvector of $\widetilde{X}_{n}$ associated with either the largest or smallest eigenvalue depending on whether $\theta>0$ or $\theta<0$ , respectively. Then, we have

Assume that the limiting eigenvalue distribution $\mu_{X}$ has a density $f_{\mu_{X}}$ with a power decay at $b$ (or $a$ or both), i.e., that, as $t\to b$ with $t<b$ , $f_{\mu_{X}}(t)\sim c(b-t)^{\alpha}$ for some exponent $\alpha>-1$ and some constant $c$ . Then:

so that the phase transitions in Theorems 2.6 and 2.8 manifest for $\alpha=1/2$ .

The analogue of Remark 2.5 also applies here.

Consider the matrix $S_{n}=(I_{n}+P_{n})^{1/2}X_{n}(I_{n}+P_{n})^{1/2}$ . The matrices $S_{n}$ and $\widetilde{X}_{n}=X_{n}(I_{n}+P_{n})$ are related by a similarity transformation and hence share the same eigenvalues and consequently the same limiting eigenvalue behavior in Theorem 2.6. Additionally, if $\widetilde{u}_{i}$ is a unit-norm eigenvector of $\widetilde{X}_{n}$ then $\widetilde{w}_{i}=(I_{n}+P_{n})^{1/2}\widetilde{u}_{i}$ is an eigenvector of $S_{n}$ and the unit-norm eigenvector $\widetilde{v}_{i}=\widetilde{w}_{i}/\|\widetilde{w}_{i}\|$ satisfies

It follows that we obtain the same phase transition behavior and that when $1/\theta_{i}\in(T_{\mu_{X}}(a^{-}),T_{\mu_{X}}(b^{+}))$ ,

so that the analogue of Theorems 2.7 and 2.8 for the eigenvectors of $S_{n}$ holds.

5. The Cauchy and T transforms in free probability theory

The Cauchy transform of a compactly supported probability measure $\mu$ on the real line is defined as:

If $[a,b]$ denotes the convex hull of the support of $\mu$ , then

exist in $[-\infty,0)$ and $(0,+\infty]$ , respectively and $G_{\mu}(\cdot)$ realizes decreasing homeomorphisms from $(-\infty,a)$ onto $(G_{\mu}(a^{-}),0)$ and from $(b,+\infty)$ onto $(0,G_{\mu}(b^{+}))$ . Throughout this paper, we shall denote by $G_{\mu}^{-1}(\cdot)$ the inverses of these homeomorphisms, even though $G_{\mu}$ can also define other homeomorphisms on the holes of the support of $\mu$ .

is the analogue of the logarithm of the Fourier transform for free additive convolution. The free additive convolution of probability measures on the real line is denoted by the symbol $\boxplus$ and can be characterized as follows.

Let $A_{n}$ and $B_{n}$ be independent $n\times n$ symmetric (or Hermitian) random matrices that are invariant, in law, by conjugation by any orthogonal (or unitary) matrix. Suppose that, as $n\longrightarrow\infty$ , $\mu_{A_{n}}\longrightarrow\mu_{A}$ and $\mu_{B_{n}}\longrightarrow\mu_{B}$ . Then, free probability theory states that $\mu_{A_{n}+B_{n}}\longrightarrow\mu_{A}\boxplus\mu_{B}$ , a probability measure which can be characterized in terms of the $R$ -transform as

The connection between free additive convolution and $G_{\mu}^{-1}$ (via the $R$ -transform) and the appearance of $G_{\mu}^{-1}$ in Theorem 2.1 could be of independent interest to free probabilists.

5.2. The T𝑇T-transform and its relation to multiplicative free convolution

In the case where $\mu\neq\delta_{0}$ and the support of $\mu$ is contained in $[0,+\infty)$ , one also defines its $T$ -transform

which realizes decreasing homeomorphisms from $(-\infty,a)$ onto $(T_{\mu}(a^{-}),0)$ and from $(b,+\infty)$ onto $(0,T_{\mu}(b^{+}))$ . Throughout this paper, we shall denote by $T_{\mu}^{-1}$ the inverses of these homeomorphisms, even though $T_{\mu}$ can also define other homeomorphisms on the holes of the support of $\mu$ .

is the analogue of the Fourier transform for free multiplicative convolution $\boxtimes$ . The free multiplicative convolution of two probability measures $\mu_{A}$ and $\mu_{B}$ is denoted by the symbols $\boxtimes$ and can be characterized as follows.

Let $A_{n}$ and $B_{n}$ be independent $n\times n$ symmetric (or Hermitian) positive-definite random matrices that are invariant, in law, by conjugation by any orthogonal (or unitary) matrix. Suppose that, as $n\longrightarrow\infty$ , $\mu_{A_{n}}\longrightarrow\mu_{A}$ and $\mu_{B_{n}}\longrightarrow\mu_{B}$ . Then, free probability theory states that $\mu_{A_{n}\cdot B_{n}}\longrightarrow\mu_{A}\boxtimes\mu_{B}$ , a probability measure which can be characterized in terms of the $S$ -transform as

The connection between free multiplicative convolution and $T_{\mu}^{-1}$ (via the $S$ -transform) and the appearance of $T_{\mu}^{-1}$ in Theorem 2.6 could be of independent interest to free probabilists.

6. Extensions

Theorem 2.1 can easily be adapted to describe the phase transition in the eigenvalues of $X_{n}+P_{n}$ which fall in the “holes” of the support of $\mu_{X}$ . Consider $c<d$ such that almost surely, for $n$ large enough, $X_{n}$ has no eigenvalue in the interval $(c,d)$ . It implies that $G_{\mu_{X}}$ induces a decreasing homeomorphism, that we shall denote by $G_{\mu_{X},(c,d)}$ , from the interval $(c,d)$ onto the interval $(G_{\mu_{X}}(d^{-}),G_{\mu_{X}}(c^{+}))$ . Then it can be proved that almost surely, for $n$ large enough, $X_{n}+P_{n}$ has no eigenvalue in the interval $(c,d)$ , except if some of the $1/\theta_{i}$ ’s are in the interval $(G_{\mu_{X}}(d^{-}),G_{\mu_{X}}(c^{+}))$ , in which case for each such index $i$ , one eigenvalue of $X_{n}+P_{n}$ has limit $G_{\mu_{X},(c,d)}^{-1}(1/\theta_{i})$ as $n\longrightarrow\infty$ .

Theorem 2.1 can also easily be adapted to the case where $X_{n}$ itself has isolated eigenvalues in the sense that some of its eigenvalues have limits out of the support of $\mu_{X}$ . More formally, let us replace the assumption that the smallest and largest eigenvalues of $X_{n}$ tend to the infimum $a$ and the supremum $b$ of the support of $\mu_{X}$ by the following one.

Moreover, $\lambda_{1+p^{+}}(X_{n})\overset{\textrm{a.s.}}{\longrightarrow}b$ and $\lambda_{n-(1+p^{-})}(X_{n})\overset{\textrm{a.s.}}{\longrightarrow}a$ .

The previous remark forms the basis for an iterative application of our theorems to other perturbational models, such as $\widetilde{X}=\sqrt{X}(I+P)\sqrt{X}+Q$ for example. Another way to deal with such perturbations is to first derive the corresponding master equations representations that describe how the eigenvalues and eigenvectors of $\widetilde{X}$ are related to the eigenvalues and eigenvectors of $X$ and the perturbing matrices, along the lines of Proposition 5.1 for additive or multiplicative perturbations of Hermitian matrices.

Let $G$ be an $n\times m$ Gaussian random matrix with independent real (or complex) entries that are normally distributed with mean and variance $1$ . Then the matrix $X=GG^{*}/m$ is orthogonally (or unitarily) invariant. Hence one can choose an orthonormal basis $(U_{1},\ldots,U_{n})$ of eigenvectors of $X$ such that the matrix $U$ with columns $U_{1},\ldots,U_{n}$ is Haar-distributed. When $G$ is a Gaussian-like matrix, in the sense that its entries are i.i.d. with mean zero and variance one, then upon placing adequate restrictions on the higher order moments, for non-random unit norm vector $x_{n}$ , the vector $U^{*}x_{n}$ will be close to uniformly distributed on the unit real (or complex) sphere . Since our proofs rely heavily on the properties of unit norm vectors uniformly distributed on the $n$ -sphere, they could possibly be adapted to the setting where the unit norm vectors are close to uniformly distributed.

Suppose that $P_{n}$ is a random matrix independent of $X_{n}$ , with exactly $r$ non-zero eigenvalues given by $\theta_{1}^{(n)},\ldots,\theta_{r}^{(n)}$ . Let $\theta_{i}^{(n)}\overset{\textrm{a.s.}}{\longrightarrow}\theta_{i}$ as $n\longrightarrow\infty$ . Using [22, Cor. 6.3.8] as in Section 6.2.3, one can easily see that our results will also apply in this case.

The analogues of Remarks 2.11, 2.12, 2.14 and 2.15 for the multiplicative setting also hold here. In particular, Wishart matrices with $c>1$ (cf Section 3.2) gives an illustration of the case where there is a hole in the support of $\mu_{X}$ .

Examples

We now illustrate our results with some concrete computations. The key to applying our results lies in being able to compute the Cauchy or $T$ transforms of the probability measure $\mu_{X}$ and their associated functional inverses. In what follows, we focus on settings where the transforms and their inverses can be expressed in closed form. In settings where the transforms are algebraic so that they can be represented as solutions of polynomial equations, the techniques and software developed in can be utilized. In more complicated settings, one will have to resort to numerical techniques.

Let $X_{n}$ be an $n\times n$ symmetric (or Hermitian) matrix with independent, zero mean, normally distributed entries with variance $\sigma^{2}/n$ on the diagonal and $\sigma^{2}/(2n)$ on the off diagonal. It is known that the spectral measure of $X_{n}$ converges almost surely to the famous semi-circle distribution with density

It is known that the extreme eigenvalues converge almost surely to the endpoints of the support . Associated with the spectral measure, we have

$G_{\mu_{X}}(\pm 2\sigma)=\pm\sigma$ and $G_{\mu_{X}}^{-1}(1/\theta)=\theta+\frac{\sigma^{2}}{\theta}$ .

Thus for a $P_{n}$ with $r$ non-zero eigenvalues $\theta_{1}\geq\cdots\geq\theta_{s}>0>\theta_{s+1}\geq\cdots\geq\theta_{r}$ , by Theorem 2.1, we have for $1\leq i\leq s$ ,

as $n\longrightarrow\infty$ . This result has already been established in for the symmetric case and in for the Hermitian case. Remark 2.14 explains why our results should hold for Wigner matrices of the sort considered in .

In the setting where $r=1$ and $P=\theta\,uu^{*}$ , let $\widetilde{u}$ be a unit-norm eigenvector of $X_{n}+P_{n}$ associated with its largest eigenvalue. By Theorems 2.2 and 2.3, we have

2. Multiplicative perturbation of a random Wishart matrix

Let $G_{n}$ be an $n\times m$ real (or complex) matrix with independent, zero mean, normally distributed entries with variance $1$ . Let $X_{n}=G_{n}G_{n}^{*}/m$ . It is known that, as $n,m\longrightarrow\infty$ with $n/m\to c>0$ , the spectral measure of $X_{n}$ converges almost surely to the famous Marčenko-Pastur distribution with density

where $a=(1-\sqrt{c})^{2}$ and $b=(1+\sqrt{c})^{2}$ . It is known that the extreme eigenvalues converge almost surely to the endpoints of this support.

Associated with this spectral measure, we have

$T_{\mu_{X}}(b^{+})=1/\sqrt{c}$ , $T_{\mu_{X}}(a^{-})\;=\;-1/\sqrt{c}$ and

When $c>1$ , there is an atom at zero so that the smallest eigenvalue of $X_{n}$ is identically zero. For simplicity, let us consider the setting when $c<1$ so that the extreme eigenvalues of $X_{n}$ converge almost surely to $a$ and $b$ . Thus for $P_{n}$ with $r$ non-zero eigenvalues $\theta_{1}\geq\cdots\geq\theta_{s}>0>\theta_{s+1}\geq\cdots\geq\theta_{r}$ , with $l_{i}:=\theta_{i}+1$ , for $c<1$ , by Theorem 2.6, we have for $1\leq i\leq s$ ,

as $n\longrightarrow\infty$ . An analogous result for the smallest eigenvalue may be similarly derived by making the appropriate substitution for $a$ in Theorem 2.6. Consider the matrix $S_{n}=(I_{n}+P_{n})^{1/2}X_{n}(I_{n}+P_{n})^{1/2}$ . The matrix $S_{n}$ may be interpreted as a Wishart distributed sample covariance matrix with “spiked” covariance $I_{n}+P_{n}$ . By Remark 2.10, the above result applies for the eigenvalues of $S_{n}$ as well. This result for the largest eigenvalue of spiked sample covariance matrices was established in and for the extreme eigenvalues in .

In the setting where $r=1$ and $P=\theta\,uu^{*}$ , let $l=\theta+1$ and let $\widetilde{u}$ be a unit-norm eigenvector of $X_{n}(I+P_{n})$ associated with its largest (or smallest, depending on whether $l>1$ or $l<1$ ) eigenvalue. By Theorem 2.8, we have

Let $\widetilde{v}$ be a unit eigenvector of $S_{n}=(I_{n}+P_{n})^{1/2}X_{n}(I_{n}+P_{n})^{1/2}$ associated with its largest (or smallest, depending on whether $l>1$ or $l<1$ ) eigenvalue. Then, by Theorem 2.8 and Remark 2.10, we have

The result has been established in for the eigenvector associated with the largest eigenvalue. We generalize it to the eigenvector associated with the smallest one.

We note that symmetry considerations imply that when $X$ is a Wigner matrix then $-X$ is a Wigner matrix as well. Thus an analytical characterization of the largest eigenvalue of a Wigner matrix directly yields a characterization of the smallest eigenvalue as well. This trick cannot be applied for Wishart matrices since Wishart matrices do not exhibit the symmetries of Wigner matrices. Consequently, the smallest and largest eigenvalues and their associated eigenvectors of Wishart matrices have to be treated separately. Our results facilitate such a characterization.

Outline of the proofs

We now provide an outline of the proofs. We focus on Theorems 2.1, 2.2 and 2.3, which describe the phase transition in the extreme eigenvalues and associated eigenvectors of $X+P$ (the index $n$ in $X_{n}$ and $P_{n}$ has been suppressed for brevity). An analogous argument applies for the multiplicative perturbation setting.

Consider the setting where $r=1$ , so that $P=\theta\,uu^{*}$ , with $u$ being a unit norm column vector. Since either $X$ or $P$ is assumed to be invariant, in law, under orthogonal (or unitary) conjugation, one can, without loss of generality, suppose that $X=\operatorname{diag}(\lambda_{1},\ldots,\lambda_{n})$ and that $u$ is uniformly distributed on the unit $n$ -sphere.

The eigenvalues of $X+P$ are the solutions of the equation

Equivalently, for $z$ so that $zI-X$ is invertible, we have

Consequently, a simple argument reveals that the $z$ is an eigenvalue of $X+P$ and not an eigenvalue of $X$ if and only if $1$ is an eigenvalue of the matrix $(zI-X)^{-1}P$ . But $(zI-X)^{-1}P=(zI-X)^{-1}\theta\,uu^{*}$ has rank one, so its only non-zero eigenvalue will equal its trace, which in turn is equal to $\theta G_{\mu_{n}}(z)$ , where ${\mu_{n}}$ is a “weighted” spectral measure of $X$ , defined by

Thus any $z$ outside the spectrum of $X$ is an eigenvalue of $X+P$ if and only if

Equation (1) describes the relationship between the eigenvalues of $X+P$ and the eigenvalues of $X$ and the dependence on the coordinates of the vector $u$ (via the measure $\mu_{n}$ ).

This is where randomization simplifies analysis. Since $u$ is a random vector with uniform distribution on the unit $n$ -sphere, we have that for large $n$ , $|u_{k}|^{2}\approx\frac{1}{n}$ with high probability. Consequently, we have $\mu_{n}\approx\mu_{X}$ so that $G_{\mu_{n}}(z)\approx G_{\mu_{X}}(z)$ . Inverting equation (1) after substituting these approximations yields the location of the largest eigenvalue to be $G_{\mu_{X}}^{-1}(1/\theta)$ as in Theorem 2.1.

The phase transition for the extreme eigenvalues emerges because under our assumption that the limiting probability measure $\mu_{X}$ is compactly supported on $[a,b]$ , the Cauchy transform $G_{\mu_{X}}$ is defined outside $[a,b]$ and unlike what happens for $G_{\mu_{n}}$ , we do not always have $G_{\mu_{X}}(b^{+})=+\infty$ . Consequently, when $1/\theta<G_{\mu_{X}}(b^{+})$ , we have that $\lambda_{1}(\widetilde{X})\approx G_{\mu_{X}}^{-1}(1/\theta)$ as before. However, when $1/\theta\geq G_{\mu_{X}}(b^{+})$ , the phase transition manifests and $\lambda_{1}(\widetilde{X})\approx\lambda_{1}(X)=b$ .

An extension of these arguments for fixed $r>1$ yields the general result and constitutes the most transparent justification, as sought by the authors in , for the emergence of this phase transition phenomenon in such perturbed random matrix models. We rely on concentration inequalities to make the arguments rigorous.

2. Eigenvectors phase transition

Let $\widetilde{u}$ be a unit eigenvector of $X+P$ associated with the eigenvalue $z$ that satisfies (1). From the relationship $(X+P)\widetilde{u}=z\widetilde{u}$ , we deduce that, for $P=\theta\,uu^{*}$ ,

implying that $\widetilde{u}$ is proportional to $(zI-X)^{-1}u$ .

Equation (2) describes the relationship between the eigenvectors of $X+P$ and the eigenvalues of $X$ and the dependence on the coordinates of the vector $u$ (via the measure $\mu_{n}$ ).

Here too, randomization simplifies analysis since for large $n$ , we have $\mu_{n}\approx\mu_{X}$ and $z\approx\rho$ . Consequently,

so that when $1/\theta<G_{\mu_{X}}(b^{+})$ , which implies that $\rho>b$ , we have

whereas when $1/\theta\geq G_{\mu_{X}}(b^{+})$ and $G_{\mu_{X}}$ has infinite derivative at $\rho=b$ , we have

An extension of these arguments for fixed $r>1$ yields the general result and brings into focus the connection between the eigenvalue phase transition and the associated eigenvector phase transition. As before, concentration inequalities allow us to make these arguments rigorous.

The exact master equations for the perturbed eigenvalues and eigenvectors

In this section, we provide the $r$ -dimensional analogues of the master equations (1) and (2) employed in our outline of the proof.

Let us fix some positive integers $1\leq r\leq n$ . Let $X_{n}=\operatorname{diag}(\lambda_{1},\ldots,\lambda_{n})$ be a diagonal $n\times n$ matrix and $P_{n}=U_{n,r}\Theta U_{n,r}^{*}$ , with $\Theta=\operatorname{diag}(\theta_{1},\textdagger\ldots,\theta_{r})$ an $r\times r$ diagonal matrix and $U_{n,r}$ an $n\times r$ matrix with orthonormal columns, i.e., $U_{n,r}^{*}U_{n,r}=I_{r}$ ).

a) Then any $z\notin\{\lambda_{1},\textdagger\ldots,\lambda_{n}\}$ is an eigenvalue of ${\widetilde{X}}_{n}:=X_{n}+P_{n}$ if and only if the $r\times r$ matrix

and for all $x\in\ker(zI_{n}-{\widetilde{X}}_{n})$ , we have $U_{n,r}^{*}x\in\ker M_{n}(z)$ and

b) Let $u^{(n)}_{k,l}$ denote the $(k,l)-$ th element of the $n\times r$ matrix $U_{n,r}$ for $k=1,\ldots,n$ and $l=1,\ldots,r$ . Then for all $i,j=1,\ldots,r$ , the $(i,j)$ -th entry of the matrix $I_{r}-U_{n,r}^{*}(zI_{n}-X_{n})^{-1}U_{n,r}\Theta$ can be expressed as

where $\mu_{i,j}^{(n)}$ is the complex measure defined by

and $G_{\mu_{i,j}^{(n)}}$ is the Cauchy transform of $\mu_{i,j}^{(n)}$ .

c) In the setting where $\widetilde{X}=X_{n}\times(I_{n}+P_{n})$ and $P_{n}=U_{n,r}\Theta U_{n,r}^{*}$ as before, we obtain the analog of a) by replacing every occurrence, in (4) and (6), of $(zI_{n}-X_{n})^{-1}$ with $(zI_{n}-X_{n})^{-1}X_{n}$ . We obtain the analog of b) by replacing the Cauchy transform in (7) with the $T$ -transform.

Proof. Part a) is proved, for example, in [2, Th. 2.3]. Part b) follows from a straightforward computation of the $(i,j)$ -th entry of $U_{n,r}^{*}(zI_{n}-X_{n})^{-1}U_{n,r}\Theta$ . Part c) can be proved in the same way. $\square$

Proof of Theorem 2.1

The sequence of steps described below yields the desired proof:

Then, we utilize the “master equations” of Section 5 to express the extreme eigenvalues of $\widetilde{X}_{n}$ as the $z$ ’s such that a certain random $r\times r$ matrix $M_{n}(z)$ is singular.

We then exploit convergence properties of certain analytical functions (derived in the appendix) to prove that almost surely, $M_{n}(z)$ converges to a certain diagonal matrix $M_{G_{\mu_{X}}}(z)$ , uniformly in $z$ .

We then invoke a continuity lemma (see Lemma 6.1 - derived next) to claim that almost surely, the $z$ ’s such that $M_{n}(z)$ is singular (i.e. the extreme eigenvalues of $\widetilde{X}_{n}$ ) converge to the $z$ ’s such that $M_{G_{\mu_{X}}}(z)$ is singular.

We conclude the proof by noting that, for our setting, the $z$ ’s such that $M_{G_{\mu_{X}}}(z)$ is singular are precisely the $z$ ’s such that for some $i\in\{1,\ldots,r\}$ , $G_{\mu_{X}}(z)=\frac{1}{\theta_{i}}$ . Part (ii) of Lemma 6.1, about the rank of $M_{n}(z)$ , will be useful to assert that when the $\theta_{i}$ ’s are pairwise distinct, the multiplicities of the isolated eigenvalues are all equal to one.

We now prove a continuity lemma that will be used in the proof of Theorem 2.1. We note that nothing in its hypotheses is random. As hinted earlier, we will invoke it to localize the extreme eigenvalues of $\widetilde{X}_{n}$ .

$G(z)\longrightarrow 0$ as $|z|\longrightarrow\infty$ .

and denote by $z_{1}>\cdots>z_{p}$ the $z$ ’s such that $M_{G}(z)$ is singular, where $p\in\{0,\ldots,r\}$ is identically equal to the number of $i$ ’s such that $G(a^{-})<1/\theta_{i}<G(b^{+})$ .

for $n$ large enough, for each $i$ , $M_{n}(z_{n,i})$ has rank $r-1$ .

Proof. Note firstly that the $z$ ’s such that $M_{G}(z)$ is singular are the $z$ ’s such that for a certain $j\in\{1,\textdagger\ldots,r\}$ ,

Since the $\theta_{j}$ ’s are pairwise distinct, for any $z$ , there cannot exist more than one $j\in\{1,\ldots,r\}$ such that (9) holds. As a consequence, for all $z$ , the rank of $M_{G}(z)$ is either $r$ or $r-1$ . Since the set of matrices with rank at least $r-1$ is open in the set of $r\times r$ matrices, once (i) will be proved, (ii) will follow.

Let us now prove (i). Note firstly that by c), there exists $R>\max\{|a|,|b|\}$ such that for $z$ such that $|z|\geq R$ , $|G(z)|\leq\min_{i}\frac{1}{2|\theta_{i}|}$ . For any such $z$ , $|\det M_{G}(z)|>2^{-r}$ . By e), it follows that for $n$ large enough, the $z$ ’s such that $M_{n}(z)$ is singular satisfy $|z|>R$ . By d), it even follows that the $z$ ’s such that $M_{n}(z)$ is singular satisfy $z\in[-R,R]$ .

the number of $z$ ’s in $(c,d)$ such that $\det M_{n}(z)=0$ , denoted by $\operatorname{Card}_{c,d}(n)$ tends to $\operatorname{Card}_{c,d}$ , the cardinality of the $i$ ’s in $\{1,\ldots,p\}$ such that $c<z_{i}<d$ .

To prove (H), by additivity, one can suppose that $c$ and $d$ are close enough to have $\operatorname{Card}_{c,d}=0$ or $1$ . Let us define $\gamma$ to be the circle with diameter $[c,d]$ . By a) and since $c,d\notin\{z_{1},\ldots,z_{p}\}$ , $\det M_{G}(\cdot)$ does not vanish on $\gamma$ , thus

the last equality following from e). It follows that for $n$ large enough, $\operatorname{Card}_{c,d}(n)=\operatorname{Card}_{c,d}$ (note that since $\operatorname{Card}_{c,d}=0$ or $1$ , no ambiguity due to the orders of the zeros has to be taken into account here). $\square$

2. Proof of Theorem 2.1

Note that Weyl’s interlacing inequalities imply that for all $1\leq i\leq n$ ,

where we employ the convention that $\lambda_{k}(X_{n})=-\infty$ is $k>n$ and $+\infty$ if $k\leq 0$ . It follows that the empirical spectral measure of $\widetilde{X}_{n}\overset{\textrm{a.s.}}{\longrightarrow}\mu_{X}$ because the empirical spectral measure of $X_{n}$ does as well.

Since $a$ and $b$ belong to the support of $\mu_{X}$ , we have, for all $i\geq 1$ fixed,

it follows that for all $i\geq 1$ fixed, $\lambda_{i}(X_{n})\overset{\textrm{a.s.}}{\longrightarrow}b$ and $\lambda_{n+1-i}(X_{n})\overset{\textrm{a.s.}}{\longrightarrow}a$ .

By (10), we deduce both following relation (11) and (12): for all $i\geq 1$ fixed, we have

and for all $i>s$ (resp. $i\geq r-s$ ) fixed, we have

In this section, we assume that the eigenvalues $\theta_{1},\ldots,\theta_{r}$ of the perturbing matrix $P_{n}$ to be pairwise distinct. In the next section, we shall remove this hypothesis by an approximation process.

For a momentarily fixed $n$ , let the eigenvalues of $X_{n}$ be denoted by $\lambda_{1}\geq\ldots\geq\lambda_{n}$ . Consider orthogonal (or unitary) $n\times n$ matrices $U_{X}$ , $U_{P}$ that diagonalize $X_{n}$ and $P_{n}$ , respectively, such that

The spectrum of $X_{n}+P_{n}$ is identical to the spectrum of the matrix

Since we have assumed that $X_{n}$ or $P_{n}$ is orthogonally (or unitarily) invariant and that they are independent, this implies that $U_{n}$ is a Haar-distributed orthogonal (or unitary) matrix that is also independent of $(\lambda_{1},\ldots,\lambda_{n})$ (see the first paragraph of the proof of [21, Th. 4.3.5] for additional details).

Recall that the largest eigenvalue $\lambda_{1}(X_{n})\overset{\textrm{a.s.}}{\longrightarrow}b$ , while the smallest eigenvalue $\lambda_{n}(X_{n})\overset{\textrm{a.s.}}{\longrightarrow}a$ . Let us now consider the eigenvalues of $\widetilde{X}_{n}$ which are out of $[\lambda_{n}(X_{n}),\lambda_{1}(X_{n})]$ . By Proposition 5.1-a) and an application of the identity in Proposition 5.1-b) these eigenvalues are precisely the numbers $z\notin[\lambda_{n}(X_{n}),\lambda_{1}(X_{n})]$ such that the $r\times r$ matrix

is singular. Recall that in (14), $G_{\mu_{i,j}^{(n)}}(z)$ , for $i,j=1,\ldots,r$ is the Cauchy transform of the random complex measure defined by

where $u_{k,i}$ and $u_{k,j}$ are the $(k,i)-$ th and $(k,j)$ -th entries of the orthogonal (or unitary) matrix $U_{n}$ in (13) and $\lambda_{k}$ is the $k$ -th largest eigenvalue of $X_{n}$ as in the first term in (13).

We now note that Hypotheses a), b) and c) of Lemma 6.1 are satisfied and follow from the definition of the Cauchy transform $G_{\mu_{X}}$ . Hypothesis d) of Lemma 6.1 follows from the fact that $\widetilde{X}_{n}$ is Hermitian while hypothesis e) has been established in (17).

Let us recall that the eigenvalues of $\widetilde{X}_{n}$ which are out of $[\lambda_{n}(X_{n}),\lambda_{1}(X_{n})]$ are precisely those values $z_{n}$ where the matrix $M_{n}(z_{n})$ is singular. As a consequence, we are now in a position where Theorem 2.1 follows by invoking Lemma 6.1. Indeed, by Lemma 6.1, if

then their exists some sequences $(z_{n,1})$ ,…, $(z_{n,p})$ converging respectively to $z_{1},\ldots,z_{p}$ such that for any $\varepsilon>0$ small enough, for $n$ large enough, the eigenvalues of $\widetilde{X}_{n}$ that are out of $[a-\varepsilon,b+\varepsilon]$ are exactly $z_{n,1},\ldots,z_{n,p}$ . Moreover, (5) and Lemma 6.1-(ii) ensure that for $n$ large enough, these eigenvalues have multiplicity one.

We now treat the case where the $\theta_{i}$ ’s are not supposed to be pairwise distinct.

We want to prove that for all $1\leq i\leq s$ , $\lambda_{i}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}\rho_{\theta_{i}}$ and that for all $0\leq j<r-s$ , $\lambda_{n-j}(\widetilde{X}_{n})\overset{\textrm{a.s.}}{\longrightarrow}\rho_{\theta_{r-j}}$ .

We shall treat only the case of largest eigenvalues (the case of smallest ones can be treated in the same way). So let us fix $1\leq i\leq s$ and $\varepsilon>0$ .

There is $\eta>0$ such that $|\rho_{\theta}-\rho_{\theta_{i}}|\leq\varepsilon$ whenever $|\theta-\theta_{i}|\leq\eta$ . Consider pairwise distinct non zero real numbers $\theta^{\prime}_{1}>\cdots>\theta^{\prime}_{r}$ such that for all $j=1,\ldots,r$ , $\theta_{j}$ and $\theta^{\prime}_{j}$ have the same sign and

It implies that $|\rho_{\theta^{\prime}_{i}}-\rho_{\theta_{i}}|\leq\varepsilon$ . With the notation in Section 6.2.2, for each $n$ , we define

Note that by [22, Cor. 6.3.8], we have, for all $n$ ,

Theorem 2.1 can applied to $X_{n}+P^{\prime}_{n}$ (because the $\theta_{1}^{\prime},\ldots,\theta_{r}^{\prime}$ are pairwise distinct). It follows that almost surely, for $n$ large enough,

By the triangular inequality, almost surely, for $n$ large enough,

so that $\lambda_{i}(X_{n}+P_{n})\overset{\textrm{a.s.}}{\longrightarrow}\rho_{\theta_{i}}$ . $\square$

Proof of Theorem 2.2

The eigenvectors of $X_{n}+P_{n}$ , are precisely $U_{X}$ times the eigenvectors of

Consequently, we have proved Theorem 2.2 by proving the result in the setting where $X_{n}=\operatorname{diag}(\lambda_{1},\ldots,\lambda_{n})$ and $P_{n}=U_{n}\operatorname{diag}(\theta_{1},\ldots,\theta_{r},0,\ldots,0)U_{n}^{*}$ , where $U_{n}$ is a Haar-distributed orthogonal (or unitary) matrix.

Let $r_{0}$ be the number of $i$ ’s such that $\theta_{i}=\theta_{i_{0}}$ . Up to a reindexing of the $\theta_{i}$ ’s (which are then no longer decreasing - this fact does not affect our proof), one can suppose that ${{i_{0}}}=1$ , $\theta_{1}=\cdots=\theta_{r_{0}}$ . This choice implies that, for each $n$ , $\ker(\theta_{1}I_{n}-P_{n})$ is the linear span of the ${r_{0}}$ first columns $u_{1}$ , …, $u_{r_{0}}$ of $U_{n}$ . By construction, these columns are orthonormal. Hence, we will have proved Theorem 2.2 if we can prove that as $n\longrightarrow\infty$ ,

As before, for every $n$ and for all $z$ outside the spectrum of $X_{n}$ , define the $r\times r$ random matrix:

where, for all $i,j=1,\ldots,r$ , $\mu_{i,j}^{(n)}$ is the random complex measure defined by (15).

We have established in Theorem 2.1 that because $\theta_{i_{0}}>1/G_{\mu_{X}}(b^{+})$ , $\widetilde{\lambda}_{i_{0}}\overset{\textrm{a.s.}}{\longrightarrow}\rho=G_{\mu_{X}}^{-1}(1/\theta_{i_{0}})\notin[a,b]$ as $n\longrightarrow\infty$ . It follows that:

Proposition 5.1 a) states that for $n$ large enough so that such that $\widetilde{\lambda}_{i_{0}}$ is not an eigenvalue of $X_{n}$ , the $r\times 1$ vector

is in the kernel of the $r\times r$ matrix $M_{n}(z_{n})$ with $||U_{n,r}^{*}\widetilde{u}||_{2}\leq 1$ .

Thus by (20), any limit point of $U_{n,r}^{*}\widetilde{u}$ is in the kernel of the matrix on the right hand side of (21), i.e. has its $r-r_{0}$ last coordinates equal to zero.

Thus (19) holds and we have proved Theorem 2.2-b). We now establish (18).

By (6), one has that for all $n$ , the eigenvector $\widetilde{u}$ of $\widetilde{X}_{n}$ associated with the eigenvalue $\widetilde{\lambda}_{i_{0}}$ can be expressed as:

As $\widetilde{\lambda}_{i_{0}}\overset{\textrm{a.s.}}{\longrightarrow}\rho\notin[a,b]$ , the sequence $(\widetilde{\lambda}_{i_{0}}I_{n}-X_{n})^{-1}$ is bounded in operator norm so that by (19), $\|\widetilde{u}^{\prime\prime}\|\overset{\textrm{a.s.}}{\longrightarrow}0$ . Since $\|\widetilde{u}\|=1$ , this implies that $\|\widetilde{u}^{\prime}\|\overset{\textrm{a.s.}}{\longrightarrow}1$ .

Since we assumed that $\theta_{i_{0}}=\theta_{1}=\cdots=\theta_{r_{0}}$ , we must have that:

By Proposition 9.3, we have that for all $i\neq j$ , $\mu_{i,j}^{(n)}\overset{\textrm{a.s.}}{\longrightarrow}\delta_{0}$ while for all $i$ , $\mu_{i,i}^{(n)}\overset{\textrm{a.s.}}{\longrightarrow}\mu_{X}$ . Thus, since we have that $z_{n}\overset{\textrm{a.s.}}{\longrightarrow}\rho\notin[a,b]$ , we have that for all $i,j=1,\ldots,r_{0}$ ,

Combining the relationship in (22) with the fact that $\|\widetilde{u}^{\prime}\|\overset{\textrm{a.s.}}{\longrightarrow}1$ , yields (18) and we have proved Theorem 2.2-a). $\square$

Proof of Theorem 2.3

Let us assume that $\theta>0$ . The proof supplied below can be easily ported to the setting where $\theta<0$ .

We denote the coordinates of $u$ by $u^{(n)}_{1},\ldots,u^{(n)}_{n}$ and define, for each $n$ , the random probability measure

The $r=1$ setting of Proposition 5.1-b) states that the eigenvalues of $X_{n}+P_{n}$ which are not eigenvalue of $X_{n}$ are the solutions of

Since $G_{\mu_{X}^{(n)}}(z)$ decreases from $+\infty$ to for increasing values of $z\in(\lambda_{1},+\infty)$ , we have that $\lambda_{1}(X_{n}+P_{n})=:\widetilde{\lambda}_{1}>\lambda_{1}$ . Reproducing the arguments leading to (3) in Section (4.2), yields the relationship:

By Theorem 2.1, we have that $\widetilde{\lambda}_{1}\overset{\textrm{a.s.}}{\longrightarrow}b$ so that

so that by (23), $\langle\widetilde{u},\ker(\theta I_{n}-P_{n})\rangle|^{2}\overset{\textrm{a.s.}}{\longrightarrow}0$ thereby proving Theorem 2.3. $\square$

We omit the details of the proofs of Theorems 2.6–2.8, since these are straightforward adaptations of the proofs of Theorems 2.1-Theorem 2.3 that can obtained by following the prescription in Proposition 5.1-c).

Appendix: convergence of weighted spectral measures

We now establish a lemma on the weak convergence of complex measures that will be useful in proving Proposition 9.3. We note that the counterpart of this lemma for probability measures is well known. We did not find any reference in standard literature to the “complex measures version” stated next, so we provide a short proof.

which can be made arbitrarily small by appropriately choosing $g$ . The tightness hypothesis ensures that such a $g$ can always be found. This proves that $(\mu_{n})$ converges weakly to $\mu$ . The uniform convergence follows from a straightforward application of Ascoli’s Theorem. $\square$

2. Convergence of weighted spectral measures

b) Suppose that $\frac{1}{n}(x_{1}+x_{2}+\cdots+x_{n})$ converges almost surely to a deterministic limit $l$ . Then

Proof. We use Lemma 9.1. Note first that almost surely, since $\sup_{n,k}|\lambda_{k}|<\infty$ , both sequences are tight. Moreover, we have