Fluctuations of the extreme eigenvalues of finite rank deformations of random matrices

Florent Benaych-Georges, Alice Guionnet, Mylène Maïda

Introduction

Most of the spectrum of a large matrix is not much altered if one adds a finite rank perturbation to the matrix, simply because of Weyl’s interlacement properties of the eigenvalues. But the extreme eigenvalues, depending on the strength of the perturbation, can either stick to the extreme eigenvalues of the non-perturbed matrix or deviate to some larger values. This phenomenon was made precise in , where a sharp phase transition, known as the BBP transition , was exhibited for finite rank perturbations of a complex Gaussian Wishart matrix. In this case, it was shown that if the strength of the perturbation is above a threshold, the largest eigenvalue of the perturbed matrix deviates away from the bulk and has then Gaussian fluctuations, otherwise it sticks to the bulk and fluctuates according to the Tracy-Widom law. The fluctuations of the extreme eigenvalues which deviate from the bulk were studied as well when the non-perturbed matrix is a Wishart (or Wigner) matrix with non-Gaussian entries; they were shown to be Gaussian if the perturbation is chosen randomly with i.i.d. entries in , or with completely delocalised eigenvectors , whereas in , a non-Gaussian behaviour was exhibited when the perturbation has localised eigenvectors. The influence of the localisation of the eigenvectors of the perturbation was studied more precisely in .

In this paper, we also focus on the behaviour of the extreme eigenvalues of a finite rank perturbation of a large matrix, this time in the framework where the large matrix is deterministic whereas the perturbation has delocalised random eigenvectors. We show that the eigenvalues which deviate away from the bulk have Gaussian fluctuations, whereas those which stick to the bulk are extremely close to the extreme eigenvalues of the non-perturbed matrix. In a one-dimensional perturbation situation, we can as well study the fluctuations of the next eigenvalues, for instance showing that if the first eigenvalue deviates from the bulk, the second eigenvalue will stick to the first eigenvalue of the non-perturbed matrix, whereas if the first eigenvalue sticks to the bulk, the second eigenvalue will be very close to the second eigenvalue of the non-perturbed matrix. Hence, for a one dimensional perturbation, the eigenvalues which stick to the bulk will fluctuate as the eigenvalues of the non-perturbed matrix. We can also extend these results beyond the case when the non-perturbed matrix is deterministic. In particular, if the non-perturbed matrix is a Wishart (or Wigner) matrix with rather general entries, or a matrix model, we can use the universality of the fluctuations of the extreme eigenvalues of these random matrices to show that the $p$ th extreme eigenvalue which sticks to the bulk fluctuates according to the $p$ th dimensional Tracy-Widom law. This proves the universality of the BBP transition at the fluctuation level, provided the perturbation is delocalised and random. The reader should notice however that we do not deal with the asymptotics of eigenvalues corresponding to critical deformations. This probably requires a case-by-case analysis and may depend on the model under consideration.

Let us now describe more precisely the models we will be dealing with. We consider a deterministic self-adjoint matrix $X_{n}$ with eigenvalues $\lambda_{1}^{n}\leq\cdots\leq\lambda_{n}^{n}$ satisfying the following hypothesis.

The spectral measure $\mu_{n}:=n^{-1}\sum_{l=1}^{n}\delta_{\lambda_{l}^{n}}$ of $X_{n}$ converges towards a deterministic probability measure $\mu_{X}$ with compact support. Moreover, the smallest and largest eigenvalues of $X_{n}$ converge respectively to $a$ and $b$ , the lower and upper bounds of the support of $\mu_{X}$ .

We consider now a random vector $v^{n}=\frac{1}{\sqrt{n}}(x_{1},\ldots,x_{n})^{T}$ with $(x_{i})_{1\leq i\leq n}$ i.i.d. real or complex random variables with law $\nu$ . Then

Either the $u_{i}^{n}$ ’s ( $i=1,\ldots,r$ ) are independent copies of $v^{n}$

Or $(u_{i}^{n})_{1\leq i\leq r}$ are obtained by the Gram-Schmidt orthonormalisation of $r$ independent copies of a vector $v^{n}.$

We shall refer to the model (1) as the i.i.d. model and to the model (2) as the orthonormalised model.

Before giving a rough statement of our results, let us make a few remarks. We first recall that a probability measure $\nu$ is said to satisfy a logarithmic Sobolev inequality with constant $c$ if, for any differentiable funtion $f$ in $L^{2}(\nu),$

Consequently, in the sequel, we shall restrict ourselves to the event when the model (2) is well defined without mentioning it explicitly.

In this work, we study the asymptotics of the eigenvalues of $\widetilde{X_{n}}$ outside the spectrum of $X_{n}$ .

It has already been observed in similar situations, see , that these eigenvalues converge to the boundary of the support of $X_{n}$ if the $\theta_{i}$ ’s are small enough, whereas for sufficiently large values of the $\theta_{i}$ ’s, they stay away from the bulk of $X_{n}$ . More precisely, if we let $G_{\mu_{X}}$ be the Cauchy-Stieltjes transform of $\mu_{X}$ , defined, for $z<a$ or $z>b,$ by the formula

then the eigenvalues of $\widetilde{X_{n}}$ outside the bulk converge to the solutions of $G_{\mu_{X}}(z)=\theta_{i}^{-1}$ if they exist.

then we have the following theorem. Let $r_{0}\in\{0,\ldots,r\}$ be such that

Assume that Hypothesis 1.1 and Assumption 1.2 are satisfied. For all $i\in\{1,\ldots,r_{0}\}$ , we have

and for all $i\in\{{r_{0}+1},\ldots,r\}$ ,

Moreover, for all $i>r_{0}$ (resp. for all $i\geq r-r_{0}$ ) independent of $n$ ,

The uniform case was proved in [10, Theorem 2.1] and we will follow a similar strategy to prove Theorem 1.3 under our assumptions in Section 2.

The main object of this paper is to study the fluctuations of the extreme eigenvalues of $\widetilde{X_{n}}.$ Precise statements will be given in Theorems 3.2, 3.4, 4.3, 4.4 and 4.5. For any $x$ such that $x\leq a$ or $x\geq b,$ we denote by $I_{x}$ the set of indices $i$ such that $\rho_{\theta_{i}}=x.$ The results roughly state as follows.

Let $\alpha_{1}<\cdots<\alpha_{q}$ be the different values of the $\theta_{i}$ ’s such that $\rho_{\theta_{i}}\notin\{a,b\}$ and denote, for each $j$ , $k_{j}=|I_{\rho_{\alpha_{j}}}|$ and $q_{0}$ the largest index so that $\alpha_{q_{0}}<0$ . Then, the law of the random vector

converges to the law of the eigenvalues of $(c_{\alpha_{j}}M_{j})_{1\leq j\leq q}$ with the $M_{j}$ ’s being independent matrices following the law of a $k_{j}\times k_{j}$ matrix from the GUE or the GOE, depending whether $\nu$ is supported on the complex plane or the real line. The constant $c_{\alpha_{j}}$ is explicitly defined in Equation (6).

If none of the $\theta_{i}$ ’s are critical (i.e. equal to $\underline{\theta}$ or $\overline{\theta}$ ), with overwhelming probability, the extreme eigenvalues converging to $a$ or $b$ are at distance at most $n^{-1+{\epsilon}}$ of the extreme eigenvalues of $X_{n}$ for some ${\epsilon}>0$ .

If $r=1$ and $\theta_{1}=\theta>0$ , we have the following more precise picture about the extreme eigenvalues:

If $\rho_{\theta}>b$ , $\sqrt{n}(\widetilde{\lambda}^{n}_{n}-\rho_{\theta})$ converges towards a Gaussian variable, whereas $n^{1-{\epsilon}}(\widetilde{\lambda}^{n}_{n-i}-\lambda_{n-i+1})$ vanishes in probability as $n$ goes to infinity for any fixed $i\geq 1$ and some ${\epsilon}>0$ .

If $\rho_{\theta}=b$ and $\theta\neq\overline{\theta},$ $n^{1-{\epsilon}}(\widetilde{\lambda}^{n}_{n-i}-\lambda_{n-i})$ vanishes in probability as $n$ goes to infinity for any fixed $i\geq 1$ and some ${\epsilon}>0$ .

For any fixed $j\geq 1$ , $n^{1-{\epsilon}}(\widetilde{\lambda}^{n}_{j}-\lambda_{j})$ vanishes in probability as $n$ goes to infinity for some ${\epsilon}>0$ .

These different behaviours are illustrated in Figure 1 below.

𝑋diag𝜃0…0\widetilde{X_{n}}=X+\operatorname{diag}(\theta,0,\ldots,0) (above the dotted line). In the left picture, $\theta=0.5<\overline{\theta}=1$ and as predicted, $\widetilde{\lambda}_{1}\approx b=2$ , whereas in the right one, $\theta=1.5>\overline{\theta}$ , which indeed implies that $\widetilde{\lambda}_{1}\approx\rho_{\theta}=\theta+\frac{1}{\theta}=2.17$ and $\widetilde{\lambda}_{2}\approx b$ . Moreover, in the left picture, we have, for all $i$ , $\widetilde{\lambda}_{i}\approx\lambda_{i}$ , with some deviations In the same way, in the right picture, $i$ , $\widetilde{\lambda}_{i+1}\approx\lambda_{i}$ , with some deviations At last, here, in the right picture, we have $\widetilde{\lambda}_{1}\approx 2.167$ , which gives $\frac{\sqrt{n}(\widetilde{\lambda}_{1}-\rho_{\theta})}{c_{\theta}}\approx 0.040$ , reasonable value for a standard Gaussian variable. The first part of this theorem will be proved in Section 3, whereas Section 4 will be devoted to the study of the eigenvalues sticking to the bulk, i.e. to the proof of the second and third parts of the theorem. Moreover, our results can be easily generalised to non-deterministic self-adjoint matrices $X_{n}$ that satisfy our hypotheses with probability tending to one. This will allow us to study in Section 5 the deformations of various classical models. This will include the study of the Gaussian fluctuations away from the bulk for rather general Wigner and Wishart matrices, hence providing a new proof of the first part of [18, Theorem 1.1] and of [5, Theorem 3.1] but also a new generalisation to non-white ensembles. The study of the eigenvalues that stick to the bulk requires a finer control on the eigenvalues of $X_{n}$ in the vicinity of the edges of the bulk, which we prove for random matrices such as Wigner and Wishart matrices with entries having a sub-exponential tail. This result complements [18, Theorem 1.1], where the fluctuations of the largest eigenvalue of a non-Gaussian Wishart matrix perturbed by a delocalised but deterministic rank one perturbation was studied. One should remark that our result depends very little on the law $\nu$ (only through its fourth moment in fact).

Our approach is based upon a determinant computation (see Lemma 6.1), which shows that the eigenvalues of $\widetilde{X_{n}}$ we are interested in are the solutions of the equation

and hence it is clear that one should expect the eigenvalues of $\widetilde{X_{n}}$ outside of the bulk to converge to the solutions of $G_{\mu_{X}}(z)=\theta_{i}^{-1}$ if they exist. Studying the fluctuations of these eigenvalues amounts to analyse the behavior of the solutions of (3) around their limit. Such an approach was already developed in several papers (see e.g or ). However, to our knowledge, the model we consider, with a fixed deterministic matrix $X_{n}$ , was not yet studied and the fluctuations of the eigenvalues which stick to the bulk of $X_{n}$ was never achieved in such a generality.

For the sake of clarity, throughout the paper, we will call “hypothesis” any hypothesis we need to make on the deterministic part of the model $X_{n}$ and “assumption” any hypothesis we need to make on the deformation $R_{n}.$ Moreover, because of concentration considerations that are developed in the Appendix of the paper, the proofs will be quite similar in the i.i.d. and orthonormalised models. Therefore, we will detail each proof in the i.i.d. model, which is simpler and then check that the argument is the same in the orthonormalised model or detail the slight changes to make in the proofs.

$\bullet$ $p_{+}$ is the number of $i$ ’s such that $\rho_{\theta_{i}}>b$ , $p_{-}$ is the number of $i$ ’s such that $\rho_{\theta_{i}}<a$ and $\alpha_{1}<\cdots<\alpha_{q}$ are the different values of the $\theta_{i}$ ’s such that $\rho_{\theta_{i}}\notin\{a,b\}$ (so that $q\leq p_{-}+p_{+}$ , with equality in the particular case where the $\theta_{i}$ ’s are pairwise distinct), $\bullet$ $\gamma_{1}^{n},\ldots\ldots\gamma_{p_{-}+p_{+}}^{n}$ are the rescaled differences between the eigenvalues with limit out of $[a,b]$ and their limits:

$\bullet$ for any $x$ such that $x\leq a$ or $x\geq b,$ $I_{x}$ is the set of indices $i$ such that $\rho_{\theta_{i}}=x$ , $\bullet$ for any $j=1,\ldots,q$ , $k_{j}$ is the number of indices $i$ such that $\theta_{i}=\alpha_{j}$ , i.e. $k_{j}=|I_{\rho_{\alpha_{j}}}|$ .

Almost sure convergence of the extreme eigenvalues

For the sake of completeness, in this section, we prove Theorem 1.3. In fact, we shall even prove the more general following result.

Assume that Hypothesis 1.1 and Assumption 1.2 are satisfied.

Let us fix, independently of $n$ , an integer $i\geq 1$ and $V$ , a neighborhood of $\rho_{\theta_{i}}$ if $i\leq r_{0}$ and of $a$ if $i>r_{0}$ . Then $\widetilde{\lambda}_{i}^{n}\in V$ with overwhelming probability.

The analogue result exists for largest eigenvalues: for any fixed integer $i\geq 0$ and $V$ , a neighborhood of $\rho_{\theta_{r-i}}$ if $i<r-r_{0}$ and of $b$ if $i\geq r-r_{0}$ , $\widetilde{\lambda}_{n-i}^{n}\in V$ with overwhelming probability.

By Lemma 6.1, the eigenvalues of $\widetilde{X_{n}}$ which are not in the spectrum of $X_{n}$ are the solutions of the equation

the functions $G_{s,t}^{n}(\cdot)$ being defined in (4). For $z$ out of the support of $\mu_{X}$ , let us introduce the $r\times r$ matrix

The key point, to prove Theorem 2.1, is the following lemma. For $A=[A_{i,j}]_{i,j=1}^{r}$ and $r\times r$ matrix, we set $|A|_{\infty}:=\sup_{i,j}|A_{i,j}|$ .

Assume that Hypothesis 1.1 and Assumption 1.2 are satisfied. For any $\delta,\varepsilon>0,$ with overwhelming probability,

In the case where the $\theta_{i}$ ’s are pairwise distinct, Theorem 2.1 follows directly from this lemma, because the $z$ ’s such that $\det(M(z))=0$ are precisely the $z$ ’s such that for some $i$ , $G_{\mu_{X}}(z)=\frac{1}{\theta_{i}}$ and because close continuous functions on an interval have close zeros. The case where the $\theta_{i}$ ’s are not pairwise distinct can then be deduced by an approximation procedure similar to the one of Section 6.2.3 of .

Then since the support of $\mu_{X}$ is contained in $[a,b]$ and for $n$ large enough, the eigenvalues of $X_{n}$ are all in $[a-\delta/2,b+\delta/2]$ , it suffices to prove that with overwhelming probability,

Now, fix some $z$ such that $|z|\leq R$ , $d(z,[a,b])>\delta,$ and $n$ large enough. By Proposition 6.2 with $A=(z-X_{n})^{-1}$ , whose operator norm is bounded by $2\delta^{-1},$ we find that for any ${\epsilon}>0$ , there exists $c>0$ such that

It follows that there are $c,\eta>0$ such that for all $z$ such that $|z|\leq R$ , $d(z,[a,b])>\delta,$

As a consequence, since the number of $z$ ’s such that $|z|\leq R$ and $nz$ have integer real and imaginary parts has order $n^{2}$ , there is a constant $C$ such that

This concludes the proof for the i.i.d. model.

The orthonormalised model can be treated similarly, by writing $U_{n}=W^{n}G_{n}$ with $\sqrt{n}W^{n}$ a matrix converging almost surely to the identity by Proposition 6.3. $\square$

Fluctuations of the eigenvalues away from the bulk

Let $p_{+}$ be the number of $i$ ’s such that $\rho_{\theta_{i}}>b$ and $p_{-}$ be the number of $i$ ’s such that $\rho_{\theta_{i}}<a$ . In this section, we study the fluctuations of the eigenvalues of $\widetilde{X_{n}}$ with limit out of the bulk, that is $(\widetilde{\lambda}_{1}^{n},\ldots,\widetilde{\lambda}_{p_{-}}^{n},\widetilde{\lambda}_{n-p_{+}+1}^{n},\ldots,\widetilde{\lambda}_{n}^{n})$ . We shall assume throughout this section that the spectral measure of $X_{n}$ converges to $\mu_{X}$ faster than $1/\sqrt{n}.$ More precisely,

For all $z\in\{\rho_{\alpha_{1}},\ldots,\rho_{\alpha_{q}}\}$ , $\sqrt{n}(G_{\mu_{n}}(z)-G_{\mu_{X}}(z))$ converges to .

Our theorem deals with the limiting joint distribution of the variables $\gamma^{n}_{1},\ldots,\gamma^{n}_{p_{-}+p_{+}}$ , the rescaled differences between the eigenvalues with limit out of $[a,b]$ and their limits:

Let us recall that for $k\geq 1$ , $\operatorname{GOE}(k)$ (resp. $\operatorname{GUE}(k)$ ) is the distribution of a $k\times k$ symmetric (resp. Hermitian) random matrix $[g_{i,j}]_{i,j=1}^{k}$ such that the random variables $\{\frac{1}{\sqrt{2}}g_{i,i}\,;\,1\leq i\leq k\}\cup\{g_{i,j}\,;\,1\leq i<j\leq k\}$ (resp. $\{g_{i,i}\,;\,1\leq i\leq k\}\cup\{\sqrt{2}\Re(g_{i,j})\,;\,1\leq i<j\leq k\}\cup\{\sqrt{2}\Im(g_{i,j})\,;\,1\leq i<j\leq k\}$ ) are independent standard Gaussian variables.

The limiting behaviour of the eigenvalues with limit outside the bulk will depend on the law $\nu$ through the following quantity, called the fourth cumulant of $\nu$

Note that if $\nu$ is Gaussian standard, then $\kappa_{4}(\nu)=0$ .

The definitions of the $\alpha_{j}$ ’s and of the $k_{j}$ ’s have been given in Theorem 1.4 and recalled in the Notations gathered at the end of the introduction above.

Suppose that Assumption 1.2 holds with $\kappa_{4}(\nu)=0,$ as well as Hypotheses 1.1 and 3.1. Then the law of

converges to the law of $(\lambda_{i,j},1\leq i\leq k_{j})_{1\leq j\leq q},$ with $\lambda_{i,j}$ the $i$ th largest eigenvalue of $c_{\alpha_{j}}M_{j}$ with $(M_{1},\ldots,M_{q})$ being independent matrices, $M_{j}$ following the GUE $(k_{j})$ (resp. GOE $(k_{j})$ ) distribution if $\nu$ is supported on the complex plane (resp. the real line). The constant $c_{\alpha}$ is given by

When $\kappa_{4}(\nu)\neq 0,$ we need a bit more than Hypothesis 3.1, namely

In the case when Assumption 1.2 holds with $\kappa_{4}(\nu)\neq 0,$ under Hypotheses 1.1, 3.1 and 3.3, Theorem 3.2 stays true, replacing the matrices $c_{\alpha_{j}}M_{j}$ by matrices $c_{\alpha_{j}}M_{j}+D_{j}$ where the $D_{j}$ ’s are independent diagonal random matrices, independent of the $M_{j}$ ’s, and such that for all $j$ , the diagonal entries of $D_{j}$ are independent centred real Gaussian variables, with variance $-l({\rho_{\alpha_{j}}})\kappa_{4}(\nu)/G_{\mu_{X}}^{\prime}(\rho_{\alpha_{j}})$ .

2. Proof of Theorems 3.2 and 3.4

We prove hereafter Theorem 3.2 and we will indicate briefly at the end of this section the minor changes to make to get Theorem 3.4. The main ingredient will be a central limit theorem for quadratic forms, stated in Theorem 6.4 in the appendix.

We set $\rho_{n}^{i}(x):=\rho_{\alpha_{i}}+\frac{x}{\sqrt{n}}.$

The first step of the proof will be to get the asymptotic behavior of $M^{n}(i,x).$

with $(n_{s,t})_{s,t=1,\ldots,r}$ a family of independent Gaussian variables with $n_{s,s}\sim\mathcal{N}(0,2)$ and $n_{s,t}\sim\mathcal{N}(0,1)$ when $s\neq t$ in the real case (resp. $n_{s,s}\sim\mathcal{N}(0,1)$ and $\Re(n_{s,t}),\Im(n_{s,t})\sim\mathcal{N}(0,1/2)$ and independent in the complex case).

From (5), we know that for $s\notin I_{\rho_{\alpha_{i}}}$ ,

Let $s\in I_{\rho_{\alpha_{i}}}.$ We write the decomposition

The asymptotics of the first term is given by Theorem 6.4 with a variance given by

As $\rho_{\alpha_{i}}$ is at distance of order one from the support of $X_{n}$ , we can expand $x/\sqrt{n}$ in $M_{s,t}^{n,2}(i,x)$ to deduce that

Equations (8), (9), (10) and (11) prove the lemma (using the fact that the distribution of the Gaussian variables $n_{s,s}$ and $n_{s,t}$ are symmetric). ∎

The last point to check is a result of asymptotic independence, from which the independence of the matrices $M_{1},\ldots,M_{q}$ will be inherited. In fact, the matrices $(M^{n}(1,x_{1}),\ldots,M^{n}(q,x_{q}))$ won’t be asymptotically independent but their determinants will.

Then, as the set of indices $I_{\rho_{\alpha_{1}}},\ldots,I_{\rho_{\alpha_{q}}}$ are disjoint, the submatrices involved in the main terms are independent in the i.i.d case and asymptotically independent in the orthonormalised case.

Let us now show (12). Firstly, note that by the convergence of $M^{n}_{s,t}(i,x)$ obtained in the proof of the Lemma 3.5, we have for all $s,t\in\{1,\ldots,r\}$ such that $s\neq t$ or $s\in I_{\rho_{\alpha_{i}}}$ , for all $\kappa<1/2$ ,

it suffices to prove that for any $\sigma\in S_{r}$ such that for some $i_{0}\in\{1,\ldots,r\}\backslash I_{\rho_{\alpha_{i}}}$ , $\sigma(i_{0})\neq i_{0}$ ,

It follows immediately from (13) since for any $\kappa<1/2$ , in the above product, all the terms with index in $I_{\rho_{\alpha_{i}}}$ are of order at most $n^{-\kappa}$ , giving a contribution $n^{-k_{i}\kappa}$ , and $i_{0}$ is not in $I_{\rho_{\alpha_{i}}}$ and satisfies $\sigma(i_{0})\neq i_{0}$ , yielding another term of order at most $n^{-\kappa}$ . Hence, the other terms being bounded because $\rho_{n}^{i}(x)$ stays bounded away from $[a,b]$ , the above product is at most of order $n^{-\kappa(k_{i}+1)}$ and so taking $\kappa\in(\frac{k_{i}}{2(k_{i}+1)},\frac{1}{2})$ proves (14). ∎

we can deduce from the lemmata above the following

Under the hypothesis of Theorem 3.2, the random process

converges weakly, as $n$ goes to infinity to the random process

in the sense of finite dimensional marginals, with the constants $c_{\alpha_{i}}$ and the joint distribution of $(M_{1},\ldots,M_{q})$ as in the statement of Theorem 3.2.

From there, the proof of Theorem 3.2 is straightforward.

To prove Theorem 3.4, the only substantial change to make is in the definition (7), in the case when $s\in I_{\rho_{\alpha_{i}}},$ we have to put

The convergence of $\left[M^{n}(i,x)\right]_{s,t}$ to $\left[\mathcal{M}(i,x)\right]_{s,t}$ is again obtained by applying Theorem 6.4.

The sticking eigenvalues

To study the fluctuations of the eigenvalues which stick to the bulk, we need a more precise information on the eigenvalues of $X_{n}$ in the vicinity of their extremes. More explicitly, we shall need the following additional hypothesis, which depends on a positive integer $p$ and a real number $\alpha\in(0,1)$ . Note that this hypothesis has two versions: Hypothesis 4.1 $[p,\alpha,a]$ is adapted to the study of the smallest eigenvalues (it is the version detailed below) and Hypothesis 4.1 $[p,\alpha,b]$ is adapted to the study of the largest eigenvalues (this version is only outlined below).

$[p,\alpha,a]$ There exists a sequence $m_{n}$ of positive integers tending to infinity such that $m_{n}=O(n^{\alpha})$ ,

and there exist $\eta_{2}>0$ and $\eta_{4}>0$ , so that for $n$ large enough

Hypothesis 4.1. $[p,\alpha,b]$ is the same hypothesis where we replace $\lambda_{p}^{n}-\lambda_{i}^{n}$ by $\lambda_{n-p+1}^{n}-\lambda_{n-i+1}^{n}$ , and (15) becomes

For many matrix models, the behaviors of largest and smallest eigenvalues are similar, and Hypothesis 4.1 $[p,\alpha,a]$ is satisfied if and only if Hypothesis 4.1 $[p,\alpha,b]$ is satisfied. In such cases, we shall simply say that Hypothesis 4.1 $[p,\alpha]$ is satisfied.

For rank one perturbations and in the i.i.d. model, we will only require the two first conditions (15) and (16) whereas for higher rank perturbations, we will need in addition (17) to control the off-diagonal terms of the determinant.

Moreover, we shall not study the critical case where for some $i$ , $\theta_{i}\in\{\underline{\theta},\overline{\theta}\}$ .

For all $i$ , $\theta_{i}\neq\underline{\theta}$ and $\theta_{i}\neq\overline{\theta}$ .

In fact, Assumption 4.2 can be weakened into: for all $i$ , $\theta_{i}\neq\underline{\theta}$ (resp. $\theta_{i}\neq\overline{\theta}$ ) if we only study the smallest (resp. largest) eigenvalues.

The fact that the eigenvalues of the non-perturbed matrix are sufficiently spread at the edges to insure the above hypothesis allow the eigenvalues of the perturbed matrix to be very close to them, as stated in the following theorem.

Let $I_{a}=\{i\in[1,r]:\rho_{\theta_{i}}=a\}=[p_{-}+1,r_{0}]$ (resp. $I_{b}=\{i\in[1,r]:\rho_{\theta_{i}}=b\}=[r_{0}+1,r-p_{+}]$ ) be the set of indices corresponding to the eigenvalues $\widetilde{\lambda}_{i}^{n}$ (resp. $\widetilde{\lambda}_{n-r+i}^{n}$ ) converging to the lower (resp. upper) bound of the support of $\mu_{X}$ . Let us suppose Hypothesis 1.1, Hypothesis 4.1 $[r,\alpha,a]$ (resp. Hypothesis 4.1 $[r,\alpha,b]$ ) and Assumptions 1.2 and 4.2 to hold. Then for any $\alpha^{\prime}>\alpha,$ we have, for all $i\in I_{a}$ (resp. $i\in I_{b}$ ),

Moreover, in the case where the perturbation has rank one, we can locate exactly in the neighborhood of which eigenvalues of the non-perturbed matrix the eigenvalues of the perturbed matrix lie.

We state hereafter the result for the smallest eigenvalues, but of course a similar statement holds for the largest ones.

Let $(\widetilde{\lambda}_{i}^{n})_{i\geq 1}$ be the eigenvalues of $X_{n}+\theta u_{1}u_{1}^{*}$ , with $\theta<0$ . Then, under Assumption 1.2 and Hypothesis 1.1, if (15) and (16) in Hypothesis 4.1 $[p,\alpha,a]$ hold for some $\alpha\in(0,1)$ and a positive integer $p$ , then for any $\alpha^{\prime}>\alpha$ , we have

if $\theta<\underline{\theta}$ , $\widetilde{\lambda}_{1}^{n}$ converges to $\rho_{\theta}<a$ whereas $n^{1-\alpha^{\prime}}(\widetilde{\lambda}_{i+1}^{n}-\lambda_{i}^{n})_{1\leq i\leq p-1}$ vanishes in probability as $n$ goes to infinity,

if $\theta\in(\underline{\theta},0)$ , $n^{1-\alpha^{\prime}}(\widetilde{\lambda}_{i}^{n}-\lambda_{i}^{n})_{1\leq i\leq p}$ vanishes in probability as $n$ goes to infinity,

if, instead of (15) and (16) in Hypothesis 4.1 $[p,\alpha,a]$ , one supposes (15) and (16) in Hypothesis 4.1 $[p,\alpha,b]$ to hold, then $n^{1-\alpha^{\prime}}(\widetilde{\lambda}_{n-i}^{n}-\lambda_{n-i}^{n})_{0\leq i<p}$ vanishes in probability as $n$ goes to infinity.

Consider the i.i.d. model and let $(\widetilde{\lambda}_{i}^{n})_{i\geq 1}$ be the eigenvalues of $X_{n}+\sum_{i=1}^{r}\theta_{i}u_{i}u_{i}^{*}$ . Let $p_{-}$ (resp. $p_{+}$ ) be the number of indices $i$ so that $\rho_{\theta_{i}}<a$ (resp. $\rho_{\theta_{i}}>b$ ). We assume that Assumptions 1.2 and 4.2, Hypothesis 1.1, and (15) and (16) in Hypotheses 4.1 $[p,\alpha,a]$ and $[q,\alpha,b]$ hold for some $\alpha\in(0,1)$ and integers $p,q$ . Then, for all $\alpha^{\prime}>\alpha$ , for all fixed $1\leq i\leq p-(p_{-}+r)$ and $0\leq j<p-(p_{+}+r)$ ,

both vanish in probability as $n$ goes to infinity.

Note that if $p-(p_{-}+r)\leq 0$ (resp. if $p-(p_{+}+r)<0$ ), then the statement of the theorem is empty as far as $i$ ’s (resp. $j$ ’s) are concerned. The same convention is made throughout the proof.

2. Proofs

Let us first prove Theorem 4.3. Let us choose $i_{0}\in I_{a}$ and study the behaviour of $\widetilde{\lambda}_{i_{0}}^{n}$ (the case of the largest eigenvalues can be treated similarly). We assume throughout the section that Hypotheses 1.1, 4.1 $[r,\alpha,a]$ and Assumptions 1.2 and 4.2 are satisfied. We also fix $\alpha^{\prime}>\alpha.$

We know, by Lemma 6.1, that the eigenvalues of $\widetilde{X_{n}}$ which are not eigenvalues of $X_{n}$ are the $z$ ’s such that

Recall that by Weyl’s interlacing inequalities (see [1, Th. A.7])

Let $\zeta$ be a fixed constant such that $\max_{1\leq i\leq p_{-}}\rho_{\theta_{i}}<\zeta<a$ . By Theorem 2.1, we know that

With overwhelming probability, $\widetilde{\lambda}^{n}_{i_{0}}>\zeta$ .

We want to show that (18) is not possible on

The following lemma deals with the asymptotic behaviour of the off-diagonal terms of the matrix $M_{n}(z)$ of (19).

For $s\neq t$ and $\kappa>0$ small enough,

The following lemma deals with the asymptotic behaviour of the diagonal terms of the matrix $M_{n}(z)$ of (19).

For all $s=1,\ldots,r$ , for all $\delta>0$ , any $\delta>0$ ,

Let us assume these lemmas proven for a while and complete the proof of Theorem 4.3. By these two lemmas, for $z\in\Omega_{n}$ , we find by expanding the determinant that with overwhelming probability,

where the $O(n^{-\kappa})$ is uniform on $z\in\Omega_{n}$ . Indeed, in the second term of the right hand side of

each diagonal term is bounded and each non diagonal term is $O(n^{-\kappa})$ .

Since for all $i$ , $\theta_{i}\neq\underline{\theta}$ , (22) and Lemma 4.8 allow to assert that with overwhelming probability, for all $z\in\Omega_{n}$ , $\det(M_{n}(z))\neq 0$ . It completes the proof of the theorem. ∎

Proof of Lemma 4.7. Let us consider $z\in\Omega_{n}$ ( $z$ might depend on $n$ , but for notational brevity, we omit to denote it by $z_{n}$ ). We treat simultaneously the orthonormalised model and the i.i.d. model (in the i.i.d. model, one just takes $W^{n}=I$ and replaces $\|(G^{n}(W^{n})^{T})_{s}\|_{2}$ by $\sqrt{n}$ in the proof below). Observe that if we write $X_{n}=O^{*}D_{n}O$ with $D_{n}=(\lambda_{1}^{n},\ldots,\lambda_{n}^{n})$ and $O$ a unitary or orthogonal matrix,

The first step is to show that for any ${\epsilon}>0$ , with overwhelming probability,

Indeed, with $O_{l}$ the $l$ th row vector of $O$ and using the notations of Section 6.2,

But $g\mapsto\langle O_{l},g^{n}_{s}\rangle$ is Lipschitz for the Euclidean norm with constant one. Hence, by concentration inequality due to the log-Sobolev hypothesis (see e.g. [1, section 4.4]), there exists $c>0$ such that for all $\delta>0$ ,

From Proposition 6.3, we know that with overwhelming probability, $\|(G^{n}(W^{n})^{T})_{s}\|_{2}$ is bounded below by $\sqrt{n}n^{-{\epsilon}}$ and the entries of $W^{n}$ are of order one. This gives therefore (23).

Note that as $|(Ou_{s}^{n})_{l}|,1\leq l\leq m_{n}$ , are smaller than $n^{-\frac{1}{2}+{\epsilon}^{\prime}}$ by (23), for any ${\epsilon}^{\prime}>0$ , with overwhelming probability, we have, uniformly on $z\in\Omega_{n}$ ,

We choose $0<{\epsilon}^{\prime}\leq(\alpha^{\prime}-\alpha)/4$ and now study $B_{n}(z)$ which can be written

with $P$ the orthogonal projection onto the linear span of the eigenvectors of $X_{n}$ corresponding to the eigenvalues $\lambda_{m_{n}+1}^{n},\ldots,\lambda_{n}^{n}$ . By the second point in Proposition 6.2, with $z\in\Omega_{n}$ , for all $s\neq t$ ,

Moreover, by Hypothesis 4.1, for $n$ large enough, for all $z\in\Omega_{n}$ ,

We deduce that there is $C,\eta>0$ such that for all $z\in\Omega_{n}$ ,

A similar control is verified for $s=t$ since we have, by Proposition 6.2,

whereas Hypothesis 4.1 insures that the term $\frac{1}{n}{\rm Tr}(P(z-X_{n})^{-1})$ is bounded uniformly on $\Omega_{n}$ . Thus, up to a change of the constants $C$ and $\eta$ , there is a constant $M$ such that for all $z\in\Omega_{n}$ ,

Therefore, with Proposition 6.3 and developing the vectors $u_{s}^{n}$ ’s as the normalised column vectors of $G^{n}(W^{n})^{T}$ , we conclude that, up to a change of the constants $C$ and $\eta$ , for all $z\in\Omega_{n}$ ,

Hence, we have proved that there exists $\kappa>0,C$ and $\eta>0$ so that for all $z\in\Omega_{n}$ ,

We finally obtain this control uniformly on $z\in\Omega_{n}$ by noticing that $z{\rightarrow}G^{n}_{s,t}(z)$ is Lipschitz on $\Omega_{n}$ , with constant bounded by $(\min|z-\lambda_{i}|)^{-2}\leq n^{2-2\alpha^{\prime}}$ . Thus, if we take a grid $(z_{k}^{n})_{0\leq k\leq cn^{2}}$ of $\Omega_{n}$ with mesh $\leq n^{-2+2\alpha^{\prime}-\kappa}$ (there are about $n^{2}$ such $z_{k}^{n}$ ’s) we have

Since there are at most $cn^{2}$ such $k$ and $n^{2}$ possible $i,j$ , we conclude that

Proof of Lemma 4.8. We shall use the decomposition

with $P$ as above the orthogonal projection onto the linear span of the eigenvectors of $X_{n}$ corresponding to the eigenvalues $\lambda_{m_{n}+1}^{n},\ldots,\lambda_{n}^{n}$ , and then prove that for $z\in\Omega_{n}$ ,

Let us now give a formal proof. Again, we first prove the estimate for a fixed $z\in\Omega_{n}$ , the uniform estimate on $z$ being obtained by a grid argument as in the previous proof (a key point being that the constants $C$ and $\eta$ of the definition of overwhelming probability are independent of the choice of $z\in\Omega_{n}$ ).

First, observe that (15) implies that for any sequence $\varepsilon_{n}$ tending to zero,

Indeed, for all ${\epsilon}>0$ , for $n$ such that $\lambda_{n}^{p}$ and $a-\varepsilon_{n}$ are both $\geq a-{\epsilon}$ , we have, for all $z\in[a-\varepsilon_{n},\lambda_{p}^{n}]$ ,

So let us consider $z\in\Omega_{n}$ ( $z$ might depend on $n$ , but for notational brevity, we omit to denote it by $z_{n}$ ). By the inequality $|z-\lambda_{k}^{n}|>n^{-1+\alpha^{\prime}}$ for all $1\leq k\leq m_{n}$ and (27), we have

with, by (24), the off diagonal terms $t\neq v$ of order $n^{-\eta_{2}\wedge\eta_{4}/8}$ with overwhelming probability, whereas the diagonal terms are close to $\frac{1}{n}{\rm Tr}(P(z-X_{n})^{-1})$ with overwhelming probability by (25). Hence, we deduce with Proposition 6.2 that for any $\delta>0$ ,

with overwhelming probability. Hence, by (28), for any $\delta>0$

with overwhelming probability. On the other hand

By Proposition 6.3, the denominator is of order $n$ with overwhelming probability, whereas by Proposition 6.2, the numerator is of order $m_{n}+n^{\epsilon}\sqrt{m_{n}}$ (since ${\rm Tr}(1-P)=m_{n}$ ) with overwhelming probability. As $W^{n}$ is bounded by Proposition 6.3 we conclude that

with overwhelming probability. Putting together Equations (29), (30) and (31), we have proved that for any $z\in\Omega_{n}$ , any $\delta>0$ ,

with overwhelming probability, the constants $C$ and $\eta$ of the definition of overwhelming probability being independent of the choice of $z\in\Omega_{n}$ We do not detail the grid argument used to get a control uniform on $z$ because this argument is similar to what we did in the proof of the previous lemma. ∎

Proof of Theorem 4.4. In the one dimensional case, the eigenvalues of $\widetilde{X}_{n}$ which do not belong to the spectrum of $X_{n}$ are the zeroes of

with $\varepsilon_{n}(g)=1$ or $\|g\|_{2}^{2}/n$ according to the model we are considering. A straightforward study of the function $f_{n}$ tells us that the eigenvalues of $\widetilde{X}_{n}$ are distinct from those of $X_{n}$ as soon as $X_{n}$ has no multiple eigenvalue and

has no null entry, which we can always assume up to modify $X_{n}$ and $g$ so slightly that the fluctuations of the eigenvalues are not affected. We do not detail these arguments but the reader can refer to Lemmas 9.3, 9.4 and 11.2 of for a full proof in the finite rank case. Therefore, (32) characterises all the eigenvalues of $\widetilde{X}_{n}$ . Moreover, by Weyl’s interlacing properties, for $\theta<0$ ,

Theorems 2.1 and 4.3 thus already settle the study of $\widetilde{\lambda}_{1}^{n}$ which either goes to $\rho_{\theta}$ or is at distance $O(n^{-1+\alpha^{\prime}})$ of $\lambda_{1}^{n}$ depending on the strength of $\theta$ . We consider $\alpha^{\prime}>\alpha$ and $i\in\{2,\ldots,p\}$ and define

Note first that if $\Lambda_{n}$ is empty, then the eigenvalue of $\widetilde{X_{n}}$ which lies between $\lambda_{i-1}^{n}$ and $\lambda_{i}^{n}$ is within $n^{-1+\alpha^{\prime}}$ to both $\lambda_{i-1}^{n}$ and $\lambda_{i}^{n}$ , so we have nothing to prove. Now, we want to prove that $f_{n}$ does not vanish on $\Lambda_{n}$ and that according to the sign of $\frac{1}{\theta}-\frac{1}{\underline{\theta}}$ , it vanishes on one side or the other of $\Lambda_{n}$ in $]\lambda_{i-1}^{n},\lambda_{i}^{n}[$ . This will prove (i) and (ii) of the theorem. Part (iii) can be proved in the same way, proving that with overwhelming probability, $f_{n}$ does not vanish in $\left]\lambda_{n-i-1}^{n}+{n^{-1+\alpha^{\prime}}},\lambda_{n-i}^{n}-{n^{-1+\alpha^{\prime}}}\right[$ .

The proof of this fact will follow the same lines as the proof of Lemma 4.8 and we recall that $P$ was defined above as the orthogonal projection onto the linear span of the eigenvectors of $X_{n}$ corresponding to the eigenvalues $\lambda_{m_{n}+1}^{n},\ldots,\lambda_{n}^{n}$ . Then, exactly as for (30), we can show that for all $\delta>0$ ,

with overwhelming probability. Moreover, for any $z\in\Lambda_{n}$ , for any $j=1,\ldots,m_{n}$ , we have

By Proposition 6.2, we deduce that for any ${\epsilon}>0$ ,

with overwhelming probability. We choose ${\epsilon}$ in such a way that the latter right hand side goes to zero. Therefore, we know that uniformly on $\Lambda_{n}$ ,

with overwhelming probability. Since for all $n$ , $f_{n}$ is decreasing, going to $+\infty$ (resp. $-\infty$ ) as $z$ goes to any $\lambda_{i-1}^{n}$ on the right (resp. $\lambda_{i}^{n}$ on the left), it follows that according to the sign of $\frac{1}{\underline{\theta}}-\frac{1}{\theta}$ , the zero of $f_{n}$ in $]\lambda_{i-1}^{n},\lambda_{i}^{n}[$ is either in $]\lambda_{i-1}^{n},\lambda_{i-1}^{n}+{n^{-1+\alpha^{\prime}}}[$ or in $]\lambda_{i}^{n}-{n^{-1+\alpha^{\prime}}},\lambda_{i}^{n}[$ .∎

Let us also choose $\zeta_{a}<a$ and $\zeta_{b}>b$ such that

Application to classical models of matrices

Let $(X_{n})$ be a sequence of random matrices independent of the $u_{i}^{n}$ ’s. Under Assumption 1.2,

If Hypothesis 1.1 holds in probability, Theorem 2.1 holds.

If $\kappa_{4}(\nu)=0$ and Hypotheses 1.1 and 3.1 hold in probability, Theorem 3.2 holds. If $\kappa_{4}(\nu)\neq 0$ and Hypotheses 1.1 and 3.3 hold in probability, Theorem 3.4 holds.

Under Assumption 4.2, if Hypotheses 1.1 and 4.1 hold in probability, Theorem 4.3 holds “with probability converging to one” instead of “with overwhelming probability”; Theorems 4.4 and Corollary 4.5 hold.

The remaining of this section is devoted to showing that such results hold if $X_{n}$ , independent of $(u_{i}^{n})_{1\leq i\leq r}$ , is a Wigner or a Wishart matrix or a random matrix which law has density proportional to $e^{-\operatorname{Tr}V}$ for a certain potential $V$ . In each case, we have to check that the hypotheses hold in probability.

Let $(x_{i,j})_{i,j\geq 1}$ be an infinite Hermitian random matrix which entries are independent up to the condition $x_{j,i}=\overline{x_{i,j}}$ such that the $x_{i,i}$ ’s are distributed according to $\mu_{2}$ and the $x_{i,j}$ ’s ( $i\neq j$ ) are distributed according to $\mu_{1}$ . We take $X_{n}=\frac{1}{\sqrt{n}}\left[{x_{i,j}}\right]_{i,j=1}^{n},$ which is said to be a Wigner matrix. For certain results, we will also need an additional hypothesis, which we present here:

The probability measures $\mu_{1}$ and $\mu_{2}$ have a sub-exponential decay, that is there exists positive constants $C,C^{\prime}$ such that if $X$ is distributed according to $\mu_{1}$ or $\mu_{2},$ for all $t\geq C^{\prime}$ ,

Moreover, $\mu_{1}$ and $\mu_{2}$ are symmetric.

The following Proposition generalises some results of which study the effect of a finite rank perturbation on a non-Gaussian Wigner matrix. In particular, it includes the study of the eigenvalues which stick to the bulk.

Let $X_{n}$ be a Wigner matrix. Assume that Assumption 1.2 holds. The limits of the extreme eigenvalues of $\widetilde{X_{n}}$ are given by Theorem 2.1 and the fluctuations of the ones which limits are out of $[-2\sigma,2\sigma]$ are given by Theorem 3.2, where the parameters $a,b,\rho_{\theta},c_{\alpha}$ are given by the following formulas : $b=-a=2\sigma$ ,

Assume moreover that, for all $i,$ $\theta_{i}\not\in\{-\sigma,\sigma\}$ and Hypothesis 5.2 holds. If the perturbation has rank one, we have the following precise description of the fluctuations of the sticking eigenvalues :

If $\theta>\sigma$ (resp. $\theta<-\sigma$ ), for all $p\geq 2$ , $n^{2/3}(\widetilde{\lambda}_{n-p+1}^{n}-2\sigma)$ (resp. $n^{2/3}(\widetilde{\lambda}_{p}^{n}+2\sigma)$ ) converges in law to the $p-1$ th Tracy Widom law.

If $0\leq\theta<\sigma$ (resp. $-\sigma<\theta\leq 0$ ), for all $p\geq 1,$ $n^{2/3}(\widetilde{\lambda}_{n-p+1}^{n}-2\sigma)$ (resp. $n^{2/3}(\widetilde{\lambda}_{p}^{n}+2\sigma)$ ) converges in law to the $p$ th Tracy Widom law.

If the perturbation is rank more than one and Assumption 4.2 holds, the extreme eigenvalues of $\widetilde{X_{n}}$ are at distance less than $n^{-1+{\epsilon}}$ for any ${\epsilon}>0$ to the extreme eigenvalues of $X_{n},$ which have Tracy-Widom fluctuations. We can localize exactly near which eigenvalue of $X_{n}$ they lie by using Theorem 4.5 in the i.i.d model.

According to Theorem 5.1, it suffices to verify that the hypotheses hold in probability for $(X_{n})_{n\geq 1}$ . We study separately the eigenvalues which stick to the bulk and those which deviate from the bulk.

If $X_{n}$ is a Wigner matrix (that is, with our terminology, with entries having a finite fourth moment), the fact that $X_{n}$ satisfies Hypothesis 1.1 in probability is a well known result (see for example [4, Th. 5.2]) for $\mu_{X}$ the semicircle law with support $[-2\sigma,2\sigma]$ . The formulas for $\rho_{\theta}$ and $c_{\alpha}$ can be checked with the well known formula [1, Sect. 2.4]:

Moreover, [5, Th. 1.1] shows that ${\rm Tr}(f(X_{n}))-n\int f(x)d\sigma(x)$ converges in law to a Gaussian distribution for any function $f$ which is analytic in a neighborhood of $[-2\sigma,2\sigma]$ . For any fixed $z\notin[-2\sigma,2\sigma]$ , applied for $f(t)=\frac{1}{z-t}$ , we get that $n(G_{\mu_{n}}(z)-G_{\mu_{X}}(z))$ converges in law to a Gaussian distribution, hence $\sqrt{n}(G_{\mu_{n}}(z)-G_{\mu_{X}}(z))$ converges in probability to zero, so that Hypothesis 3.1 holds in probability.

We now assume moreover that the laws of the entries satisfy Hypothesis 5.2. In order to lighten the notation, we shall now suppose that $\sigma=1$ . Let us first recall that by , the extreme eigenvalues of the non-perturbed matrix $X_{n}$ , once re-centred and renormalised by $n^{2/3}$ , converge to the Tracy-Widom law (which depends on whether the entries are complex or real). We need to verify that Hypothesis 4.1[p, $\alpha$ ] for any finite $p$ and an $\alpha<1/3$ is fulfilled in probability. By , the spacing between the two smallest eigenvalues of $X_{n}$ is of order greater than $n^{-\gamma}$ for $\gamma>2/3$ with probability going to one and therefore, by the inequality

it is sufficient to prove the first point of Hypothesis 4.1[p, $\alpha$ ]. We shall prove it by replacing first the smallest eigenvalue by the edge $-2$ thanks to a lemma that Benjamin Schlein kindly communicated to us. We will then prove that the sum of the inverse of the distance of the eigenvalues to the edge indeed converges to the announced limit, thanks to both Soshnikov paper (for sub-Gaussian tails) or (for finite moments), and Tao and Vu article .

Suppose the entries of $X_{n}$ have a uniform sub-exponential tail. Then for all $\delta>0$ , for all integer number $p$ ,

Now, for any $K_{2}>K_{1}$ , on the event $\{|\lambda_{p}^{n}+2|<K_{1}n^{-2/3}\}$ , for any $\kappa>0$ , we have

Let us fix $\kappa\in(\frac{2}{3},1)$ . It follows that the first term of the r.h.s. of (39) can be estimated by

Let us now estimate the second term of the r.h.s. of (39). For any positive integer $K_{3}$ , we have

From (38), (39), (41) and (42), we conclude that

for arbitrary $0<K_{1}<K_{3}$ and $K_{3}\geq 1$ . Taking the limit $n\to\infty$ , the last two terms disappear, because by [42, Th. 1.16], the distribution of the smallest $K_{3}$ eigenvalues lives on scales of order $n^{-2/3}\gg n^{-5/6}$ . Therefore,

still for any $0<K_{1}<K_{3}$ and $K_{3}\geq 1$ . Now, note that for $K_{1}$ large enough, the first term can be made as small as we want. Then, keeping $K_{1}$ fixed, $K_{2}$ can be chosen in such a way to make the second term as small as we want too. At last, keeping $K_{2}$ fixed, one can choose $K_{3}$ large enough to make the third term as small as we want (as can be computed since the limit is given by the $K_{3}$ correlation function of the Airy kernel). $\square$

To complete the proof of Hypothesis 4.1, we therefore need to show that

Assume that the entries of $X_{n}$ satisfy Hypothesis 5.2. Then, for any $\delta>0$ , any finite integer number $p$ ,

Proof. Notice that by we know that the $p$ smallest eigenvalues of $X_{n}$ converge in law towards the Tracy-Widom law, so that

Thus, for any finite $p$ , with large probability,

and therefore it is enough to prove the lemma for any particular $p$ . As in the previous proof, we choose $p$ large enough so that $\lambda_{p}^{n}\geq-2+n^{-\frac{2}{3}}$ with probability greater than $1-\delta(p)$ with $\delta(p)$ going to zero as $p$ goes to infinity. We shall prove that with high probability

This is enough to prove the statement as for any $\gamma>0$ , $2+\lambda^{n}_{[n\gamma]}$ converges to $\delta(\gamma)>0$ so that $\mu_{X}([{\delta(\gamma)},2])=1-\gamma$ , see [43, Theorem 1.3],

which converges as $\gamma$ goes to zero to $\int(2+x)^{-1}d\mu_{X}(x)=1$ (by e.g. (37)). To prove (43), we choose $\rho\in(2/3,\sqrt{2/3})$ and write, on the event $\lambda_{j}^{n}+2\geq\lambda_{p}^{n}+2\geq n^{-\frac{2}{3}}\geq n^{-\rho}$ for $j\geq p$ ,

For the first term, we use Sinai-Soshnikov bound, which under the weakest hypothesis are given in [39, Theorem 2.1]. It implies that with probability going to one with $M$ going to infinity, for $s_{n}=o(n^{2/3})$ going to infinity,

This implies, by Tchebychev’s inequality and taking $s_{n}=n^{+\rho^{k+1}}$ that

which goes to zero as $\rho>2/3$ . For the second term $B_{n}$ , note that by [42, Theorem 1.10], for any ${\epsilon}>0$ small enough,

which goes to zero as $n$ goes to infinity and then $\gamma$ goes to zero. ∎

2. Coulomb Gases

We can also consider random matrices $X_{n}$ which law is invariant under the action of the unitary or the orthogonal group and with eigenvalues with law given by

with a polynomial function $V$ of even degree and positive leading coefficient and $\beta=1,2$ or $4$ . We assume moreover that $V$ is such that the limiting spectral measure $\mu_{V}$ of $(X_{n})$ is connected and compact and that its smallest and largest eigenvalues converge to the boundaries of the support. This set of hypotheses is often referred to as the “one-cut assumption”. It holds in particular if $V$ is strictly convex and this includes the classical Gaussian ensembles GOE and GUE (with $V(x)=x^{2}/4$ and $\beta=1,2$ ).

Under the above hypothesis on $V,$ the extreme eigenvalues of $X_{n}$ converge to the boundary of the support. The convergence of the extreme eigenvalues of $\widetilde{X_{n}}$ is given by Theorem 2.1. These eigenvalues have Gaussian fluctuations as stated in Theorem 3.2 if they deviate away from the bulk. Suppose moreover that Assumption 4.2 holds. If the perturbation is of rank one and is strong enough so that the largest eigenvalues deviates from the bulk, for all $k\geq 2,$ the rescaled $k$ th largest eigenvalue $n^{\frac{2}{3}}(\widetilde{\lambda}_{n-k+1}^{n}-b_{V})$ converges weakly towards the $k-1$ -th Tracy Widom law. If the perturbation is of rank one and is weak enough, for all $k\geq 1,$ the rescaled $k$ th largest eigenvalue $n^{\frac{2}{3}}(\widetilde{\lambda}_{n-k+1}^{n}-b_{V})$ converges weakly towards the $k$ -th Tracy Widom law. If the perturbation is of rank more than one, the extreme eigenvalues of $\widetilde{X_{n}}$ sticking to the bulk are at distance less than $n^{-1+{\epsilon}}$ for any ${\epsilon}>0$ from the eigenvalues of $X_{n}.$ In the i.i.d model, Theorem 4.5 prescribes exactly in the neighborhood of which eigenvalues of $X_{n}$ each of them lie.

Proof. As explained above, it suffices to verify that the hypotheses hold in probability for $(X_{n})_{n\geq 1}$ .

Note that the convergence of the spectral measure, of the edges and the fluctuations of the extreme eigenvalues were obtained in . The fact that $\sqrt{n}(G_{\mu_{n}}(z)-G_{\operatorname{sc}}(z))$ converges in probability to zero is a consequence of so that Hypothesis 3.1 holds.

We next check Hypothesis 4.1[p, $\alpha$ ] for the matrix model $P_{n}.$ We shall prove it for any $\alpha>1/3$ and any integer $p$ . We first show that

Indeed, the joint distribution of $(\lambda_{1}^{n},\ldots,\lambda_{n}^{n})$ is

with $\beta=1,2$ or $4$ , $Z_{n}^{\beta}$ is the normalising constant and $\Delta_{n}=\{\lambda_{1}<\cdots<\lambda_{n}\}$ . Therefore,

by integration by parts. Equation (45) follows, since $\lambda^{n}_{p}$ converges almost surely to $a_{V}$ (and concentration inequalities insures $V^{\prime}(\lambda_{p}^{n})$ is uniformly integrable). But, for any ${\epsilon}>0$ ,

with, by convergence of the spectral measure and of $\lambda^{n}_{p}$ , the right hand side converging to $-G_{\mu_{X}}(-a_{V}-{\epsilon})$ which converges as ${\epsilon}$ decreases to zero to $-G_{\mu_{X}}(-a_{V})=-V^{\prime}(a_{V})$ . Hence, $\frac{1}{n}\sum_{i\neq p}\frac{1}{\lambda^{n}_{i}-\lambda^{n}_{p}}$ is bounded below by $-V^{\prime}(a_{V})$ with large probability for large $n$ , and converges in expectation to $-V^{\prime}(a_{V})$ , and therefore converges in probability to $-V^{\prime}(a_{V})$ .

Moreover, by (see in the Gaussian case), the joint law of

converges weakly towards a probability measure which is absolutely continuous with respect to Lebesgue measure. As a consequence, we also deduce from the first point that $n^{-1}\sum_{i<m_{n}}(\lambda_{p}^{n}-\lambda_{i}^{n})^{-1}$ vanishes as $n$ goes to infinity in probability for $m_{n}\ll n^{1/3}$ and therefore (45) proves the lacking point of Hypothesis 4.1.

so that by (45) and Markov’s inequality, Hypothesis 4.1 holds in probability for any $\eta<1/3$ , $\eta_{4}<1$ and $\alpha>1/3$ . $\square$

3. Wishart matrices

Let $n,m$ tend to infinity in such a way that $n/m\to c\in(0,1)$ . The limits of the extreme eigenvalues of $\widetilde{X_{n}}$ are given by Theorem 2.1 and the fluctuations of those which limits are out of $[a,b]$ are given by Theorem 3.2, where the parameters $a,b,\rho_{\theta},c_{\alpha}$ are given by the following formulas: $a=(1-\sqrt{c})^{2}$ , $b=(1+\sqrt{c})^{2}$

Assume now that the law of the entries satisfy Hypothesis 5.2. If the perturbation has rank one, we have the following precise description of the fluctuations of the extreme eigenvalues of $\widetilde{X_{n}}$ :

If $\theta>c+\sqrt{c}$ (resp. $\theta<c-\sqrt{c}$ ), for all $p\geq 2$ , $n^{2/3}(\widetilde{\lambda}_{n-p+1}^{n}-2\sigma)$ (resp. $n^{2/3}(\widetilde{\lambda}_{p}^{n}-2\sigma)$ ) converges in law to the $p-1$ th Tracy Widom law.

If $0\leq\theta<c+\sqrt{c}$ (resp. $c-\sqrt{c}<\theta\leq 0$ ), for all $p\geq 1,$ $n^{2/3}(\widetilde{\lambda}_{n-p+1}^{n}-2\sigma)$ (resp. $n^{2/3}(\widetilde{\lambda}_{p}^{n}-2\sigma)$ ) converges in law to the $p$ th Tracy Widom law.

If the perturbation has rank more than one and for all $i,$ $\theta_{i}\notin\{c+\sqrt{c},c-\sqrt{c}\}$ , the extreme eigenvalues of $\widetilde{X_{n}}$ are at distance less than $n^{-1+{\epsilon}}$ for any ${\epsilon}>0$ to the extreme eigenvalues of $X_{n},$ which have Tracy-Widom fluctuations.

Before getting into the proof, let us make a remark. The Proposition above generalizes some results first appeared in . In these papers, the authors consider models with multiplicative perturbations (in the sense that the population covariance $\Sigma$ matrix is assumed to be a perturbation of the identity). Here, we consider additive perturbations but the two models are in fact similar, since a Wishart matrix can be written as a sum of rank one matrices $\sum_{i=1}^{m}\sigma_{i}Y_{i}Y_{i}^{*},$ with $\sigma_{i}$ the eigenvalues of $\Sigma$ and $Y_{i}$ $n$ -dimensional vectors with i.i.d. entries. So, adding our perturbation $\sum_{i=1}^{r}\theta_{i}U_{i}U_{i}^{*}$ boils down to change $m$ into $m+r$ (the limit of $m/n$ is not changed) and to extend $\Sigma$ with some new eigenvalues $\theta_{1},\ldots,\theta_{r}$ .

Proof. Again, it suffices to verify that the hypotheses hold in probability for $(X_{n})_{n\geq 1}$ .

It is known, , that the spectral measure of $X_{n}$ converges to the so-called Marčenko-Pastur distribution

where $a=(1-\sqrt{c})^{2}$ and $b=(1+\sqrt{c})^{2}$ . It is known, [4, Th. 5.11], that the extreme eigenvalues converge to the bounds of this support. The formula

allows to compute $\rho_{\theta}$ and $c_{\alpha}$ . Moreover, by [3, Th. 1.1] or [4, Th. 9.10], we also know that a central limit theorem holds for the linear statistics of Wishart matrices, giving Hypothesis 3.1 as in the Wigner case.

For Hypothesis 4.1, the proof is similar to the Wigner case. The convergence to the Tracy-Widom law of the non-perturbed matrix is due to S. Péché (see and for the Gaussian case). The approximation of the eigenvalues by the quantiles of the limiting law can be found in [17, Theorem 9.1] whereas the absolute continuity property needed to prove Lemma 5.5 is derived in [17, Lemma 8.1]. This allows to prove Hypothesis 4.1 in this setting as in the Wigner case, we omit the details. $\square$

4. Non-white ensembles

In the case of non-white matrices, we can only study the fluctuations away from the bulk (since we do not have the appropriate information about the top eigenvalues to prove Hypothesis 4.1). We illustrate this generalisation in a few cases, but it is rather clear that Theorem 3.2 applies in a much wider generality.

4.2. Non-white Wigner matrices

There are less results in the literature about the central limit theorem for band matrices (with centring with respect to the limit) and the convergence of the spectrum. We therefore concentrate on a special case, namely a Hermitian matrix $X_{n}$ with independent Gaussian centred entries so that $E[|X_{ij}|^{2}]=n^{-1}\sigma(i/n,j/n)$ with a stepwise constant function

which entails the convergence of the spectrum of $X_{n}$ towards the support of the limiting measure [31, Proposition 11] with exponential speed by [31, Proof of Lemma 14]. Thus $X_{n}$ satisfies Hypothesis 1.1. Hypothesis 3.1 can be checked by modifying slightly the proof of (46) which is based on an integration by parts to be able to take $z$ on the real line but away from the limiting support. Indeed, as in [23, Section 3.3], we can add a smooth cut-off function in the expectation which vanishes outside of the event $A_{n}$ that $X_{n}$ has all its eigenvalues within a small neighborhood of the limiting support. This additional cut-off will only give a small error in the integration by parts due to the previous point. Then, (46), but with an expectation restricted to this event, is proved exactly in the same way, except that $\Im z$ can be replaced by the distance of $z$ to the neighborhood of the limiting support where the eigenvalues of $X_{n}$ lives. Finally, concentration inequalities, in the local version [22, Lemma 5.9 and Part II], insure that on $A_{n}$ ,

is at most of order $n^{-1+{\epsilon}}$ with overwhelming probability. This completes the proof of Hypothesis 3.1.

5. Some models for which our hypothesis are not satisfied

We gather hereafter a few remarks about some models for which the hypothesis we made on $X_{n}$ are not satisfied. For sake of simplicity, we present hereafter only the case of i.i.d. perturbations (1).

We assume that $X_{n}$ is diagonal with i.i.d. entries which law $\mu$ is compactly supported. As in the core of the paper, we denote by a (resp. b) the left (resp. right) edge of the support of $\mu.$ We also denote by $F_{\mu}$ its cumulative distribution function and assume that there is $\kappa>0$ such that for all $c>0,$

In this situation, it is easy to check that Hypothesis 1.1 holds in probability with $\mu_{X}=\mu$ . But Hypothesis 3.1 is not satisfied. Indeed, by classical CLT, we have, for $\rho_{\alpha}\notin[a,b],$

converges in law, as $n$ goes to infinity to a Gaussian variable $W_{\alpha}$ with variance $-G^{\prime}_{\mu}(\rho_{\alpha})-G_{\mu}(\rho_{\alpha})^{2}.$ Moreover,

Nevertheless, Theorem 3.2 holds for this model. Indeed, the whole proof of this theorem goes through in this context, except the proof of Lemma 3.5, where we have to make the following decomposition $M_{s,t}^{n}(i,x)=M_{s,t}^{n,1}(i,x)+M_{s,t}^{n,2}(i,x)+M_{s,t}^{n,3}(i,x)$ with the difference that this time $M_{s,t}^{n,3}$ does not go to zero but converges towards $W_{\alpha_{i}}$ . Hence, the eigenvalues fluctuate according to the distribution of the eigenvalues of $(c_{j}M_{j}+W_{\alpha_{j}}I_{k_{j}})_{1\leq j\leq q}$ , with $c_{j}$ and $M_{j}$ as in the statement of Theorem 3.2 and $I_{k_{j}}$ denotes the $k_{j}\times k_{j}$ identity matrix.

5.2. Coulomb gases with non-convex potentials

In , Pastur showed that for a Coulomb gas law (44) with a potential $V$ so that the equilibrium measure has a disconnected support, the central limit theorem does not hold in the sense that the variance may have different limits according to subsequences (see [35, (3.4)]. Moreover the asymptotics of $\sqrt{n}({\rm Tr}(X_{n})-\mu(x))$ can be computed sometimes and do not lead to a Gaussian limit. We might expect then that also $\sqrt{n}(G_{\mu_{n}}(x)-G_{\mu}(x))$ converges to a non-Gaussian limit, which would then result with non-Gaussian fluctuations for the eigenvalues outside of the bulk.

Appendix

We here state formula (3), which can be deduced from the well known formula $\det\left(\begin{array}[]{cc}A&B\cr C&D\cr\end{array}\right)=\det(D)\det(A-BD^{-1}C)$ .

2. Concentration estimates

Under Assumption 1.2, there exists a constant $c>0$ so that for any matrix $A:=(a_{jk})_{1\leq j,k\leq n}$ with complex entries, for any $\delta>0$ , for any $g=(g_{1},\ldots,g_{n})^{T}$ with i.i.d. entries $(g_{i})_{1\leq i\leq n}$ with law $\nu$ ,

On the other hand, the previous estimate shows that

As a consequence, we deduce the second point of the proposition. $\square$

Let ${G}^{n}=\begin{bmatrix}g_{1}^{n}\cdots g_{r}^{n}\end{bmatrix}$ be an $n\times r$ matrix which columns $g^{n}_{1},\ldots,g^{n}_{r}$ , are independent copies of an $n\times 1$ matrix with i.i.d. entries with law $\nu$ and define

and, for $j\leq i-1$ , if $\det[V^{n}_{k,l}]_{k,l=1}^{i-1}\neq 0$ ,

On $\det[V^{n}_{k,l}]_{k,l=1}^{i-1}=0$ , we give to $W_{i,j}^{n}$ an arbitrary value, say one. Putting $W^{n}_{ii}=1$ and $W^{n}_{ij}=0$ for $j\geq i+1$ , it is a standard linear algebra exercise to check that the column vectors

For any $\gamma>0$ , there exists finite positive constants $c,C$ (depending on $r$ ) so that for $Z^{n}=V^{n}$ or $W^{n}$ ,

Moreover, with $\|v||_{2}^{2}=\sum_{i=1}^{n}|v_{i}|^{2}$ , for any $\gamma\in(0,\sqrt{n}(2^{-r}-\epsilon)$ for some ${\epsilon}>0,$

Proof. We first consider the case $Z^{n}=V^{n}$ . The maximum of $|V_{ij}^{n}-\delta_{ij}|$ is controlled by the previous proposition with $A=n^{-1}I$ , and the result follows from ${\rm Tr}AA^{*}=n^{-1}$ and ${\rm Tr}((AA^{*})^{2})=n^{-3}$ , and choosing $\delta=\gamma/\sqrt{2}$ , $\kappa=\sqrt{n}$ . The result for $W^{n}$ follows as on $\|V^{n}-I\|_{\infty}\leq\gamma n^{-\frac{1}{2}}\leq 1$

For the last point, we just notice that since $\frac{1}{n}\|\sum_{j=1}^{r}Z_{i,j}^{n}g^{n}_{j}\|_{2}^{2}=(ZVZ^{*})_{i,i}$ , we have

for a finite constant $C(r)$ which only depends on $r$ . Thus the result follows from the previous point. $\square$

3. Central Limit Theorem for quadratic forms

Let us fix $r\geq 1$ and let, for each $n$ , $A^{n}(s,t)$ ( $1\leq s,t\leq r$ ) be a family of $n\times n$ real (resp. complex) matrices such that for all $s,t$ , $A^{n}(t,s)=A^{n}(s,t)^{*}$ and such that for all $s,t=1,\ldots,r$ ,

for some finite numbers $\sigma_{s,t},\omega_{s}$ (in the case where $\kappa_{4}(\nu)=0$ , the part of the hypothesis related to $\omega_{s}$ can be removed). For each $n$ , let us define the $r\times r$ random matrix

Then the distribution of $G_{n}$ converges weakly to the distribution of a real symmetric (resp. Hermitian) random matrix $G=[g_{s,t}]_{s,t=1}^{r}$ such that the random variables

are independent and for all $s$ , $g_{s,s}\sim\mathcal{N}(0,2\sigma_{s,s}^{2}+\kappa_{4}(\nu)\omega_{s})$ (resp. $g_{s,s}\sim\mathcal{N}(0,\sigma_{s,s}^{2}+\kappa_{4}(\nu)\omega_{s})$ ) and for all $s\neq t$ , $g_{s,t}\sim\mathcal{N}(0,\sigma_{s,t}^{2})$ (resp. $\Re(g_{s,t}),\Im(g_{s,t})\sim\mathcal{N}(0,\sigma_{s,t}^{2}/2)$ ).

then it follows directly from Theorem 6.4 and from a second moment computation that each finite dimensional marginal of the process

We have to prove that for any real symmetric (resp. Hermitian) matrix $B:=[b_{s,t}]_{s,t=1}^{r}$ , the distribution of $\operatorname{Tr}(BG_{n})$ converges weakly to the distribution of $\operatorname{Tr}(BG)$ . Note that

where $C^{n}$ is the $rn\times rn$ matrix and $U_{n}$ is the $rn\times 1$ random vector defined by

In the real (resp. complex) case, let us now apply Theorem 7.1 of in the case $K=1$ . It follows that the distribution of

converges weakly to a centred real Gaussian law with variance

It completes the proof in the i.i.d. model.

$\bullet$ In the orthonormalised model, we can write $u^{n}_{s}=\frac{1}{\|\sum_{i=1}^{s}W^{n}_{si}g_{i}\|_{2}}\sum_{j=1}^{s}W^{n}_{sj}g_{j}$ , where the matrix $W^{n}$ is the one introduced in this section. It follows that, with

by orthonormalization of the $u_{s}^{n}$ ’s

But, by the previous result, if $i\neq j$ ,

converges in distribution to a Gaussian law, whereas if $i=j$ ,

where both terms converge to a Gaussian. Thus this term is also bounded as $n$ goes to infinity.

Hence, by Proposition 6.3, we may and shall replace $W^{n}$ by the identity (since the error term would be of order at most $n^{-\frac{1}{2}+{\epsilon}}$ ), which yields

so that we are back to the previous setting with $B$ instead of $A$ . $\square$

Acknowledgments: We are very grateful to B. Schlein for communicating us Lemma 5.5. We also thank G. Ben Arous and J. Baik for fruitful discussions. We also thank the referee, who pointed some vagueness in the first version of the paper.