The LASSO risk for gaussian matrices

Mohsen Bayati, Andrea Montanari

Introduction

with $\lambda>0$ . The original signal is estimated by

A large and rapidly growing literature is devoted to developing fast algorithms for solving the optimization problem (1.3) and characterizing the performances and optimality of the estimator $\widehat{x}$ . We refer to Section 1.3 for an unavoidably incomplete overview.

Despite such substantial effort, and many remarkable achievements, our understanding of (1.3) is not even comparable to the one we have of more classical topics in statistics and estimation theory. For instance, the best bound on the mean squared error ( ${\rm MSE}$ ) of the estimator (1.3), i.e. on the quantity $N^{-1}\|\widehat{x}-x_{0}\|^{2}$ , was proved by Candes, Romberg and Tao [CRT06] (who in fact did not consider the LASSO but a related optimization problem). Their result estimates the mean squared error only up to an unknown numerical multiplicative factor. Work by Candes and Tao [CT07] on the analogous Dantzig selector, upper bounds the mean squared error up to a factor $C\log N$ , under somewhat different assumptions.

The objective of this paper is to complement this type of ‘rough but robust’ bounds by proving asymptotically exact expressions for the mean square error. Our asymptotic result holds almost surely for sequences of random matrices $A$ with fixed aspect ratio and independent gaussian entries. While this setting is admittedly specific, the careful study of such matrix ensembles has a long tradition both in statistics and communications theory and has spurred many insights [Joh06, Tel99]. Further, we carried out simulations on real data matrices with continuous entries (gene expression data) and binary feature matrices (hospital medical records). The results appear to be quite encouraging.

Although our rigorous results are asymptotic in the problem dimensions, numerical simulations have shown that they are accurate already on problems with a few hundreds of variables. Further, they seem to enjoy a remarkable universality property and to hold for a fairly broad family of matrices [DMM10]. Both these phenomena are analogous to ones in random matrix theory, where delicate asymptotic properties of gaussian ensembles were subsequently proved to hold for much broader classes of random matrices. Also, asymptotic statements in random matrix theory have been replaced over time by concrete probability bounds in finite dimensions. Of course the optimization problem (1.2) is not immediately related to spectral properties of the random matrix $A$ . As a consequence, universality and non-asymptotic results in random matrix theory cannot be directly exported to the present problem. Nevertheless, we expect such developments to be foreseeable.

Our proofs are based on the analysis of an efficient iterative algorithm first proposed by [DMM09], and called AMP, for approximate message passing. The algorithm is inspired by belief-propagation on graphical models; although the resulting iteration is significantly simpler (and scales linearly in the number of nodes). Extensive simulations [DMM10] showed that, in a number of settings, AMP performances are statistically indistinguishable to the ones of LASSO, while its complexity is essentially as low as the one of the simplest greedy algorithms.

The proof technique just described is new. Earlier literature analyzes the convex optimization problem (1.3) –or similar problems– by a clever construction of an approximate optimum, or of a dual witness. Such constructions are largely explicit. Here instead we prove an asymptotically exact characterization of a rather non-trivial iterative algorithm. The algorithm is then proved to converge to the exact optimum.

As already mentioned, we will consider sequences of instances of increasing sizes, along which the LASSO behavior has a non-trivial limit.

Let us stress that our proof only applies to a subclass of converging sequences, namely for gaussian measurement matrices $A(N)$ . The notion of converging sequences is however important since it defines a class of problem instances to which the ideas developed below might be generalizable. Also, while the measurement matrices $A(N)$ will be random, the signal $x_{0}(N)$ , and noise vectors $w(N)$ will be deterministic.

For a converging sequence of instances, and an arbitrary sequence of thresholds $\{\theta_{t}\}_{t\geq 0}$ (independent of $N$ ), the asymptotic behavior of the recursion (1.8) can be characterized as follows.

where $Z\sim{\sf N}(0,1)$ is independent of $X_{0}$ . Notice that the function ${\sf F}$ depends implicitly on the law $p_{X_{0}}$ . We will see later that the quantity $A^{*}z^{t}+x^{t}$ has the same distribution as $X_{0}+\tau_{t}Z$ . In other words, $\tau_{t}^{2}$ is the MSE of the estimator $A^{*}z^{t}+x^{t}$ for $x_{0}$ .

The next proposition that was conjectured in [DMM09] and proved in [BM11] shows that the behavior of AMP can be tracked by the above one dimensional recursion. We often refer to this prediction by state evolution.

where $Z\sim{\sf N}(0,1)$ is independent of $X_{0}\sim p_{X_{0}}$ .

In order to establish the connection with the LASSO, a specific policy has to be chosen for the thresholds $\{\theta_{t}\}_{t\geq 0}$ . Throughout this paper we will take $\theta_{t}=\alpha\tau_{t}$ with $\alpha$ is fixed. In other words, the sequence $\{\tau_{t}\}_{t\geq 0}$ is given by the recursion

Let us finally discuss why there should be any relation at all between the AMP algorithm (1.8) and the solution of the LASSO. Assume that $\theta_{t}\to\theta$ , and that $(x,z)$ is a fixed point of the corresponding AMP iteration. Let $\omega=\delta^{-1}\langle\eta^{\prime}(x+A^{*}z;\theta)\rangle$ . Then the fixed point condition reads

Notice that $x=\eta(r;\theta)$ if and only if there exists $v(x)\in\partial\|x\|_{1}$ such that $x+\theta v(x)=r$ (here $\partial f$ denotes the subgradient of the function $f$ ). It follows that the fixed point condition can be rewritten as

Comparing with the stationarity condition for the LASSO cost function (1.2) we obtain the following.

Any fixed point $x^{t}=x$ of the AMP iteration with $\theta_{t}=\theta$ is a minimizer of the LASSO cost function with

2 Main result

Before stating our results, we have to describe a calibration mapping between $\alpha$ and $\lambda$ that was introduced in [DMM10]. This mapping is necessary since in the analysis of AMP $\alpha$ plays the role of $\lambda$ . In other words, it can be viewed as regularization parameter and controls sparsity of AMP estimates. In particular, we will show that there exist a one-to-one (monotone) function between values of $\alpha$ and $\lambda$ .

Let us start by stating some convenient properties of the state evolution recursion.

Let $\alpha_{\rm min}=\alpha_{\rm min}(\delta)$ be the unique non-negative solution of the equation

with $\phi(z)\equiv e^{-z^{2}/2}/\sqrt{2\pi}$ the standard gaussian density and $\Phi(z)\equiv\int_{-\infty}^{z}\phi(x)\,{\rm d}x$ .

For any $\sigma^{2}>0$ , $\alpha>\alpha_{\rm min}(\delta)$ , the fixed point equation $\tau^{2}={\sf F}(\tau^{2},\alpha\tau)$ admits a unique solution. Denoting by $\tau_{*}=\tau_{*}(\alpha)$ this solution, we have $\lim_{t\to\infty}\tau_{t}=\tau_{*}(\alpha)$ . Further the convergence takes place for any initial condition and is monotone. Finally $\left|\frac{{\rm d}{\sf F}}{{\rm d}\tau^{2}}(\tau^{2},\alpha\tau)\right|<1$ at $\tau=\tau_{*}$ .

For greater convenience of the reader, a proof of this statement is provided in Appendix A.1.

We then define the function $\alpha\mapsto\lambda(\alpha)$ on $(\alpha_{\rm min}(\delta),\infty)$ , by

This function defines a correspondence (calibration) between the threshold $\alpha\tau_{*}$ and the regularization parameter $\lambda$ . It should be intuitively clear that larger $\lambda$ corresponds to larger thresholds and hence larger $\alpha$ since both cases yield smaller estimates of $x_{0}$ . The specific choice in Eq. (1.18) is motivated by Lemma 1.2.

In the following we will need to invert this function. We thus define $\alpha:(0,\infty)\to(\alpha_{\rm min},\infty)$ in such a way that

The next result implies that the set on the right-hand side is non-empty and therefore the function $\lambda\mapsto\alpha(\lambda)$ is well defined.

The function $\alpha\mapsto\lambda(\alpha)$ is continuous on the interval $(\alpha_{\rm min},\infty)$ with $\lambda(\alpha_{\rm min}+)=-\infty$ and $\lim_{\alpha\to\infty}\lambda(\alpha)=\infty$ .

Therefore the function $\lambda\mapsto\alpha(\lambda)$ satisfying Eq. (1.19) exists.

A proof of this statement is provided in Section A.2. We will denote by ${\cal A}=\alpha((0,\infty))$ the image of the function $\alpha$ . Notice that the definition of $\alpha$ is a priori not unique. We will see that uniqueness follows from our main theorem.

Examples of the mappings $\tau^{2}\mapsto{\sf F}(\tau^{2},\alpha\tau)$ , $\alpha\mapsto\tau_{*}(\alpha)$ and $\alpha\mapsto\lambda(\alpha)$ are presented in Figures 1, 2, and 3 respectively.

2.2 Main results

where $Z\sim{\sf N}(0,1)$ is independent of $X_{0}\sim p_{X_{0}}$ , $\tau_{*}=\tau_{*}(\alpha(\lambda))$ and $\theta_{*}=\alpha(\lambda)\tau_{*}(\alpha(\lambda))$ .

Let us emphasize oonce more that the vectors $x_{0}(N)$ , $w(N)$ are deterministic in this statement, and ‘almost surely’ is understood with respect to the choice of $A(N)$ .

As a corollary, using function $\psi(a,b)\equiv(a-b)^{2}$ we obtain:

Assume the hypothesis of Theorem 1.5. Let $\widehat{x}(\lambda;N)$ be the LASSO estimator for instance $(x_{0}(N),w(N),A(N))$ . Then, almost surely

where $Z\sim{\sf N}(0,1)$ is independent of $X_{0}\sim p_{X_{0}}$ , $\tau_{*}=\tau_{*}(\alpha(\lambda))$ and $\theta_{*}=\alpha(\lambda)\tau_{*}(\alpha(\lambda))$ .

As a second corollary of Theorem 1.5, the function $\lambda\mapsto\alpha(\lambda)$ is indeed uniquely defined.

For any $\lambda,\sigma^{2}>0$ there exists a unique $\alpha>\alpha_{\rm min}$ such that $\lambda(\alpha)=\lambda$ (with the function $\alpha\to\lambda(\alpha)$ defined as in Eq. (1.18).

Hence the function $\lambda\mapsto\alpha(\lambda)$ is continuous non-decreasing with $\alpha((0,\infty))\equiv{\cal A}=(\alpha_{0},\infty)$ .

The proof of this corollary (which uses Theorem 1.5) is provided in Appendix A.3.

We prove Theorem 1.5 by proving the following result in Section 3.

Assume the hypotheses of Theorem 1.5. Let $\widehat{x}(\lambda;N)$ be the LASSO estimator for instance $(x_{0}(N),w(N),A(N))$ , and denote by $\{x^{t}(N)\}_{t\geq 0}$ the sequence of estimates produced by AMP. Then

Let us emphasize that the statement of Theorem 1.8 requires taking the limit of infinite dimensions $N\to\infty$ before the limit of an infinite number of iterations $t\to\infty$ . In this sense it is (informally speaking) a statement about the high-dimensional limit behavior, for a large-but-finite number of iterations. Although this is not a common setting within mathematical optimization, we think that it is particularly compelling from a compressed sensing point of view. It implies that, for any finite tolerance $\varepsilon>0$ , there exists a finite number of iterations $t_{*}(\varepsilon)$ such that for any fixed $t\geq t_{*}(\varepsilon)$ , AMP has mean squared error at most $\varepsilon$ larger than the LASSO, with high probability as $N\to\infty$ . Further, closer analysis of the state evolution recursion [DMM09, DMM10] implies that $t_{*}(\varepsilon)\leq C\,\log(1/\varepsilon)$ for some constant $C$ independent of the dimension, and the signal $x_{0}$ , provided the under-sampling ratio $\delta$ is larger than a phase transition value $\delta_{c}$ . Notice that taking the high dimensional point of view yields us a considerably faster convergence than the optimum rate at fixed dimension, namely $t_{*}(\varepsilon)\leq C/\sqrt{\varepsilon}$ [BT09].

3 Related work

The LASSO was introduced in [Tib96, CD95]. Several papers provide performance guarantees for the LASSO or similar convex optimization methods [CRT06, CT07], by proving upper bounds on the resulting mean squared error. These works assume an appropriate ‘isometry’ condition to hold for $A$ . While such condition hold with high probability for some random matrices, it is often difficult to verify them explicitly. Further, it is only applicable to very sparse vectors $x_{0}$ . These restrictions are intrinsic to the worst-case point of view developed in [CRT06, CT07].

Guarantees have been proved for correct support recovery in [ZY06], under an appropriate ‘incoherence’ assumption on $A$ . While support recovery is an interesting conceptualization for some applications (e.g. model selection), the metric considered in the present paper (mean squared error) provides complementary information and is quite standard in many different fields.

Closer to the spirit of this paper [RFG09] derived expressions for the mean squared error under the same model considered here. Similar results were presented recently in [KWT09, GBS09]. These papers argue that a sharp asymptotic characterization of the LASSO risk can provide valuable guidance in practical applications. For instance, it can be used to evaluate competing optimization methods on large scale applications, or to tune the regularization parameter $\lambda$ .

Unfortunately, these results were non-rigorous and were obtained through the famously powerful ‘replica method’ from statistical physics [MM09].

Let us emphasize that the present paper offers two advantages over these recent developments: $(i)$ It is completely rigorous, thus putting on a firmer basis this line of research; $(ii)$ It is algorithmic in that the LASSO mean squared error is shown to be equivalent to the one achieved by a low-complexity message passing algorithm.

Numerical illustrations

Further, our result is asymptotic, while and one might wonder how accurate it is for instances of moderate dimensions.

We obtained the optimum estimator $\widehat{x}$ using CVX, a package for specifying and solving convex programs [GB10] and OWLQN, a package for solving large-scale versions of LASSO [AG07]. We used several values of $\lambda$ between and $2$ and $N$ equal to $200$ , $500$ , $1000$ , and $2000$ . The aspect ratio of matrices was fixed in all cases to $\delta=0.64$ . For each case, the point $(\lambda,{\rm MSE})$ was plotted and the results are shown in the figures. Continuous lines corresponds to the asymptotic prediction by Corollary 1.6, namely $\delta(\tau_{*}^{2}-\sigma^{2})$ .

The agreement is remarkably good already for $N,n$ of the order of a few hundreds, and deviations are consistent with statistical fluctuations.

The four figures correspond to measurement matrices $A$ :

Figure 4: Data consist of $2253$ measurements of expression level of $7077$ genes.From this matrix we took sub-matrices $A$ of aspect ratio $\delta$ for each $N$ . The entries were continuous variables. We standardized all columns of $A$ to have mean 0 and variance 1.

Figure 5: From a data set of $1932$ patient records we extracted $4833$ binary features describing demographic information, medical history, lab results, medications etc. The - $1$ matrix was sparse (with only $3.1\%$ non-zero entries). Similar to $(i)$ , for each $N$ , the sub-matrices $A$ with aspect ratio $\delta$ were selected and standardized.

Figure 6: Random gaussian matrices with aspect ratio $\delta$ and iid ${\sf N}(0,1/n)$ entries (as in Theorem 1.5);

Figure 7: Random $\pm 1$ matrices with aspect ratio $\delta$ . Each entry is independently equal to $+1/\sqrt{n}$ or $-1/\sqrt{n}$ with equal probability.

Notice the behavior appears to be essentially indistinguishable. Also the asymptotic prediction has a minimum as a function of $\lambda$ . The location of this minimum can be used to select the regularization parameter. Further empirical analysis is presented in [BBM11].

A structural property and proof of the main results

The rest of the paper is devoted to the proof of Theorem 1.8. Section 3.2 proves a structural property that is the key tool in this proof. Section 3.3 uses this property together with a few lemmas to prove Theorem 1.8

The proof of Theorem 1.5 follows immediately from Theorem 1.8.

For any $t\geq 0$ , we have, by the pseudo-Lipschitz property of $\psi$ ,

where the second inequality follows by Cauchy-Schwarz. Next we take the limit $N\to\infty$ followed by $t\to\infty$ . The first term vanishes by Theorem 1.8. For the second term, note that $\|x_{0}\|_{2}^{2}/N$ remains bounded since $(x_{0},w,A)$ is a converging sequence. The two terms $\|x^{t+1}\|_{2}^{2}/N$ and $\|\widehat{x}\|_{2}^{2}/N$ also remain bounded in this limit because of state evolution (as proved in Lemma 3.2 below).

where we used Theorem 1.1 and Proposition 1.3. ∎

For a matrix $M$ we denote its minimum and maximum singular values by $\sigma_{\rm min}(M)$ , $\sigma_{\rm max}(M)$ respectively. We also denote the minimum non-zero singular value of $M$ by $\hat{\sigma}_{\rm min}(M)$ .

2 A structural property of the LASSO cost function

One main challenge in the proof of Theorem 1.5 lies in the fact that the function $x\mapsto{\cal C}_{A,y}(x)$ is not –in general– strictly convex. Hence there can be, in principle, vectors $x$ of cost very close to the optimum and nevertheless far from the optimum.

The following Lemma provides conditions under which this does not happen.

There exists a function $\xi(\varepsilon,c_{1},\dots,c_{5})$ such that the following happens.

There exists $\mbox{\rm sg}({\cal C},x)\in\partial{\cal C}(x)$ with $\|\mbox{\rm sg}({\cal C},x)\|_{2}\leq\sqrt{N}\,\varepsilon$ ;

Let $v\equiv(1/\lambda)[A^{*}(y-Ax)+\mbox{\rm sg}({\cal C},x)]\in\partial\|x\|_{1}$ , and $S(c_{2})\equiv\{i\in[N]:\;|v_{i}|\geq 1-c_{2}\}$ . Then, for any $S^{\prime}\subseteq[N]$ , $|S^{\prime}|\leq c_{3}N$ , we have $\sigma_{\rm min}(A_{S(c_{2})\cup S^{\prime}})\geq c_{4}$ ;

The maximum singular value of $A$ is bounded: $\sigma_{\rm max}(A)^{2}\leq c_{5}$ .

Then $\|r\|_{2}\leq\sqrt{N}\,\xi(\varepsilon,c_{1},\dots,c_{5})$ . Further for any $c_{1},\dots,c_{5}>0$ , $\xi(\varepsilon,c_{1},\dots,c_{5})\to 0$ as $\varepsilon\to 0$ .

Further, if $\ker(A)=\{0\}$ , the same conclusion holds under assumptions 1, 2, 3, 5.

Throughout the proof we denote $\xi_{1},\xi_{2},\dots$ functions of the constants $c_{1},\dots,c_{5}>0$ and of $\varepsilon$ such that $\xi_{i}(\varepsilon)\to 0$ as $\varepsilon\to 0$ (we shall omit the dependence of $\xi_{i}$ on $\varepsilon$ ).

Let $S={\rm supp}(x)\subseteq[N]$ . We have

where $(a)$ follows from hypothesis (2), $(c)$ from the fact that $v_{S}=\mbox{\rm sign}(x_{S})$ since $v\in\partial\|x\|_{1}$ which gives

and $(d)$ follows from the definition of $(v)$ .

Using hypothesis (1) and (3), we get by Cauchy-Schwarz

Each of the three terms on the left-hand side is non-negative. The third one is trivial. The first one is non-negative since

and each $(x_{i}+r_{i})\left[\mbox{\rm sign}(x_{i}+r_{i})-\mbox{\rm sign}(x_{i})\right]$ is either equal to (when $\mbox{\rm sign}(x_{i})=\mbox{\rm sign}(x_{i}+r_{i})$ ) or equal to $2|x_{i}+r_{i}|$ otherwise. The second term in (3.3) is also non-negative since $|r_{i}|-v_{i}r_{i}=|r_{i}|[1-v_{i}\mbox{\rm sign}(r_{i})]$ and $1\geq v_{i}\,\mbox{\rm sign}(r_{i})$ since $|v_{i}|\leq 1$ by definition of subgradient. Therefore,

Since $\|A_{\perp}r^{\perp}\|_{2}^{2}\geq(c_{4}^{2}/4)\|r^{\perp}\|^{2}_{2}$ , we have

In the case $V_{\parallel}=\{0\}$ , the proof is concluded. In the case $V_{\parallel}\neq\{0\}$ , we need to prove an analogous bound for $r^{\parallel}$ . From Eq. (3.4) together with $\|r^{\perp}_{\overline{S}}\|_{1}\leq\sqrt{N}\|r^{\perp}_{\overline{S}}\|_{2}\leq\sqrt{N}\|r^{\perp}\|_{2}\leq(2N/c_{4})\sqrt{\xi_{1}(\varepsilon)}$ , we get

Where (3.9) follows immediately from definition of $A_{\perp}$ and $r^{\parallel}$ . Now, notice that $\overline{S}(c_{2})\subseteq\overline{S}$ . From Eq. (3.8) and definition of $S(c_{2})$ it follows that

To conclude the proof, it is sufficient to prove an analogous bound for $\|r^{\parallel}_{S_{+}}\|_{2}^{2}$ with $S_{+}=[N]\setminus\overline{S}_{+}=S(c_{2})\cup S_{1}$ . Since $|S_{1}|\leq Nc_{3}$ , we have by hypothesis (4) that $\sigma_{\rm min}(A_{S_{+}})\geq c_{4}$ . By Eq. (3.9) we have $A_{\parallel}r^{\parallel}=Ar^{\parallel}=A_{S_{+}}r^{\parallel}_{S_{+}}+A_{\overline{S}_{+}}r^{\parallel}_{\overline{S}_{+}}$ . Therefore

In the last step we used triangular inequality together with the fact that $\sigma_{\rm max}(A_{\overline{S}_{+}})^{2}\leq c_{5}$ (by assumption (5)) and $\sigma_{\rm max}(A_{\parallel})\leq c_{4}/2$ (by construction). Using $\|r^{\parallel}\|^{2}_{2}=\|r^{\parallel}_{S_{+}}\|_{2}^{2}+\|r^{\parallel}_{\overline{S}_{+}}\|_{2}^{2}$ , we get

This finishes the proof when $|\overline{S}(c_{2})|\geq Nc_{3}/2$ . Note that if this assumption does not hold then we can take $\overline{S}_{+}=\emptyset$ and $S_{+}=[N]$ . Hence, the result follows as a special case of above. ∎

3 Proof of Theorem 1.8

The proof is based on a series of Lemmas that are used to check the assumptions of Lemma 3.1

Under the conditions of Theorem 1.5, assume $\lambda>0$ and $\alpha=\alpha(\lambda)$ . Denote by $\widehat{x}(\lambda;N)$ the ${\rm LASSO}$ estimator and by $\{x^{t}(N)\}$ the sequence of AMP estimates. Then there is a constant ${\sf B}$ such that for all $t\geq 0$ , almost surely

The second Lemma implies that the estimates of AMP are approximate minima, in the sense that the cost function ${\cal C}$ admits a small subgradient at $x^{t}$ , when $t$ is large. The proof is deferred to Section 5.2.

Under the conditions of Theorem 1.5, for all $t$ there exists a subgradient $\mbox{\rm sg}({\cal C},x^{t})$ of ${\cal C}$ at point $x^{t}$ such that almost surely,

The next lemma implies that sub-matrices of $A$ constructed using the first $t$ iterations of the AMP algorithm are non-singular (more precisely, have singular values bounded away from ). The proof can be found in Section 5.3.

Let $S\subseteq[N]$ be measurable on the $\sigma$ -algebra $\mathfrak{S}_{t}$ generated by $\{z^{0},\dots,z^{t-1}\}$ and $\{x^{0}+A^{*}z^{0},\dots,x^{t-1}+A^{*}z^{t-1}\}$ and assume $|S|\leq N(\delta-c)$ for some $c>0$ . Then there exists $a_{1}=a_{1}(c)>0$ (independent of $t$ ) and $a_{2}=a_{2}(c,t)>0$ (depending on $t$ and $c$ ) such that

eventually almost surely as $N\to\infty$ .

We will apply this lemma to a specific choice of the set $S$ . Namely, defining

for $\gamma\in(0,1)$ . Our last lemma shows that this sequence of sets $S_{t}(\gamma)$ ‘converges’ in the following sense. The proof can be found in Section 5.4.

Fix $\gamma\in(0,1)$ and let the sequence $\{S_{t}(\gamma)\}_{t\geq 0}$ be defined as in Eq. (3.17) above. For any $\xi>0$ there exists $t_{*}=t_{*}(\xi,\gamma)<\infty$ such that, for all $t_{2}\geq t_{1}\geq t_{*}$ fixed, we have

eventually almost surely as $N\to\infty$ .

The above two lemmas imply the following.

There exist constants $\gamma_{1}\in(0,1)$ , $\gamma_{2}$ , $\gamma_{3}>0$ and $t_{\rm min}<\infty$ such that, for any $t\geq t_{\rm min}$ ,

eventually almost surely as $N\to\infty$ .

First notice that, for any fixed $\gamma$ , the set $S_{t}(\gamma)$ is measurable on $\mathfrak{S}_{t}$ . Indeed by Eq. (1.8) $\mathfrak{S}_{t}$ contains $\{x^{0},\dots,x^{t}\}$ as well, and hence it contains $v^{t}$ which is a linear combination of $x^{t-1}+A^{*}z^{t-1}$ , $x^{t}$ . Finally $S_{t}(\gamma)$ is obviously a measurable function of $v^{t}$ .

Using Lemma F.3(b) the empirical distribution of $(x_{0}-A^{*}z^{t-1}-x^{t-1},x_{0})$ converges weakly to $(\tau_{t-1}Z,X_{0})$ for $Z\sim{\sf N}(0,1)$ independent of $X_{0}\sim p_{X_{0}}$ . (Following the notation of [BM11], we let $h^{t}=x_{0}-A^{*}z^{t-1}-x^{t-1}$ .) Therefore, for any constant $\gamma$ we have almost surely

The last equality follows from the weak convergence of the empirical distribution of $\{(h_{i},x_{0,i})\}_{i\in[N]}$ (from Lemma F.3(b), which takes the same form as Theorem 1.8), together with the absolute continuity of the distribution of $|X_{0}+\tau_{t-1}Z-\eta(X_{0}+\tau_{t-1}Z,\theta_{t-1})|$ .

eventually almost surely as $N\to\infty$ , for all fixed $t$ larger than some $t_{{\rm min},1}(c)$ .

For any $t\geq t_{{\rm min},1}(c)$ we can apply Lemma 3.4 for some $a_{1}(c)$ , $a_{2}(c,t)>0$ . Fix $c>0$ and let $a_{1}=a_{1}(c)$ be fixed as well. Let $t_{\rm min}=\max(t_{{\rm min},1},t_{*}(a_{1}/2,\gamma_{1}))$ (with $t_{*}(\,\cdot\,)$ defined as per Lemma 3.5). Take $a_{2}=a_{2}(c,t_{\rm min})$ . Obviously $t\mapsto a_{2}(c,t)$ is non-increasing. Then we have, by Lemma 3.4

where both events hold eventually almost surely as $N\to\infty$ . The claim follows with $\gamma_{2}=a_{1}(c)/2$ and $\gamma_{3}=a_{2}(c,t_{\rm min})$ . ∎

We are now in position to prove Theorem 1.8.

We apply Lemma 3.1 to $x=x^{t}$ , the AMP estimate and $r=\widehat{x}-x^{t}$ the distance from the ${\rm LASSO}$ optimum. The thesis follows by checking conditions 1–5. Namely we need to show that there exists constants $c_{1},\dots,c_{5}>0$ and, for each $\varepsilon>0$ some $t=t(\varepsilon)$ exists such that 1–5 hold eventually almost surely as $N\to\infty$ .

Condition 2 is immediate since $x+r=\widehat{x}$ minimizes ${\cal C}(\,\cdot\,)$ .

Condition 3 follows from Lemma 3.3 with $\varepsilon$ arbitrarily small for $t$ large enough.

Condition 4. Notice that this condition only needs to be verified for $\delta<1$ .

Take $v=v^{t}$ as defined in Eq. (3.16). Using the definition (1.8), it is easy to check that $|v_{i}^{t}|\leq 1$ if $x_{i}^{t}=0$ and $v_{i}^{t}=\mbox{\rm sign}(x^{t}_{i})$ otherwise. In other words $v^{t}\in\partial\|x\|_{1}$ as required. Further by inspection of the proof of Lemma 3.3, it follows that $v^{t}=(1/\lambda)[A^{*}(y-Ax^{t})+\mbox{\rm sg}({\cal C},x^{t})]$ , with $\mbox{\rm sg}({\cal C},x^{t})$ the subgradient bounded in that lemma (cf. Eq. (5.3)). The condition then holds by Proposition 3.6.

Condition 5 follows from standard limit theorems on the singular values of Wishart matrices (cf. Theorem F.2). ∎

State evolution estimates

This section contains a reminder of the state-evolution method developed in [BM11]. For greater convenience of the reader, we also restate two lemmas from [BM11] (namely, Lemmas F.3 and F.3) in appendix F.3. We will use these two Lemmas throughout our analysis.

We also state some extensions of those results that will be proved in the appendices.

AMP, cf. Eq. (1.8) is a special case of the general iterative procedure given by Eq. (3.1) of [BM11]. This takes the general form

where $\xi_{t}=\langle g^{\prime}(b^{t},w)\rangle$ , $\lambda_{t}=\frac{1}{\delta}\langle f^{\prime}_{t}(h^{t},x_{0})\rangle$ (both derivatives are with respect to the first argument).

and the initial condition is $q^{0}=-x_{0}$ .

Regarding $h^{t},b^{t}$ as column vectors, the equations for $b^{0},\ldots,b^{t-1}$ and $h^{1},\ldots,h^{t}$ can be written in matrix form as:

or in short $Y_{t}=AQ_{t}$ and $X_{t}=A^{*}M_{t}$ .

Following [BM11], we define $\mathfrak{S}_{t}$ as the $\sigma$ -algebra generated by $b^{0},\ldots,b^{t-1}$ , $m^{0},\ldots,m^{t-1}$ , $h^{1},\ldots,h^{t}$ , and $q^{0},\ldots,q^{t}$ . The conditional distribution of the random matrix $A$ given the $\sigma$ -algebra ${\mathfrak{S}_{t}}$ , is given by

Here $P_{M_{t}}^{\perp}=I-P_{M_{t}}$ , $P_{Q_{t}}^{\perp}=I-P_{Q_{t}}$ , and $P_{Q_{t}}$ , $P_{M_{t}}$ are orthogonal projector onto column spaces of $Q_{t}$ and $M_{t}$ respectively.

Before proceeding, it is convenient to introduce the notation

to denote the coefficient of $z^{t-1}$ in Eq. (1.8). Using $h^{t}=x_{0}-A^{*}z^{t-1}-x^{t-1}$ and Lemma F.3(b) (proved in [BM11]) we get, almost surely,

Notice that the function $\eta^{\prime}(\,\cdot\,;\theta_{t-1})$ is discontinuous and therefore Lemma F.3(b) does not apply immediately. On the other hand, this implies that the empirical distribution of $\{(A^{*}z^{t-1}_{i}+x^{t-1}_{i},x_{0,i})\}_{1\leq i\leq N}$ converges weakly to the distribution of $(X_{0}+\tau_{t-1}Z,X_{0})$ . The claim follows from the fact that $X_{0}+\tau_{t-1}Z$ has a density, together with the standard properties of weak convergence.

2 Some consequences and generalizations

We begin with a simple calculation, that will be useful.

If $\{z^{t}\}_{t\geq 0}$ are the AMP residuals, then, almost surely,

Using representation (4.5) and Lemma F.3(b)(c), we get

Next, we need to generalize state evolution to compute large system limits for functions of $x^{t}$ , $x^{s}$ , with $t\neq s$ . To this purpose, we define the covariances $\{{\sf R}_{s,t}\}_{s,t\geq 0}$ recursively by

with $Z_{t}\sim{\sf N}(0,{\sf R}_{t,t})$ independent of $X_{0}$ . This determines by the above recursion ${\sf R}_{t,s}$ for all $t\geq 0$ and for all $s\geq 0$ .

With these definition, we have the following generalization of Theorem 1.1.

Clearly this result reduces to Theorem 1.1 in the case $s=t$ by noting that ${\sf R}_{t,t}=\tau^{2}_{t}$ . The general proof can be found in Appendix B.

The following lemma implies that, asymptotically for large $N$ , the AMP estimates converge.

Under the condition of Theorem 1.5, the estimates $\{x^{t}\}_{t\geq 0}$ and residuals $\{z^{t}\}_{t\geq 0}$ of AMP almost surely satisfy

Proofs of auxiliary lemmas

In order to bound the norm of $x^{t}$ , we use state evolution, Theorem 1.1, for the function $\psi(a,b)=a^{2}$ ,

for $Z\sim{\sf N}(0,1)$ and independent of $X_{0}\sim p_{X_{0}}$ . The expectation on the right hand side is bounded and hence $\lim_{t\to\infty}\lim_{N\to\infty}\langle x^{t},x^{t}\rangle$ is bounded.

The last bound holds almost surely as $N\to\infty$ , using standard asymptotic estimate on the singular values of random matrices (cf. Theorem F.2) implying that $\sigma_{\max}(A)$ has a bounded limit almost surely, together with the fact that $(x_{0},w,A)$ is a converging sequence.

Now, decompose $\widehat{x}$ as $\widehat{x}=\widehat{x}_{\parallel}+\widehat{x}_{\perp}$ where $\widehat{x}_{\parallel}\in\ker(A)$ and $\widehat{x}_{\perp}\in\ker(A)^{\perp}$ (the orthogonal complement of $\ker(A)$ ). Since, $\widehat{x}_{\parallel}$ belongs to the random subspace $\ker(A)$ with dimension $N-n=N(1-\delta)$ , Kashin theorem (cf. Theorem F.1) implies that there exists a positive constant $c_{1}=c_{1}(\delta)$ such that

Hence, by using triangle inequality and Cauchy-Schwarz, we get

By definition of cost function we have $\|\widehat{x}\|_{1}\leq\lambda^{-1}{\cal C}(\widehat{x})$ . Further, limit theorems for the eigenvalues of Wishart matrices (cf. Theorem F.2) imply that there exists a constant $c=c(\delta)$ such that asymptotically almost surely $\|\widehat{x}_{\perp}\|^{2}\leq c\,\|A\widehat{x}_{\perp}\|^{2}$ . Therefore (denoting by $c_{i}:~{}i=2,3,4$ bounded constants), we have

The claim follows by using the Eq. (5.1) to bound ${\cal C}(\widehat{x})/N$ and using $\|Ax_{0}+w\|^{2}\leq\sigma_{\max}(A)^{2}\|x_{0}\|^{2}+\|w\|^{2}\leq 2N{\sf B}_{1}$ to bound the last term. $\Box$

2 Proof of Lemma 3.3

First note that equation $x^{t}=\eta(A^{*}z^{t-1}+x^{t-1};\theta_{t-1})$ of AMP implies

Therefore, the vector $\mbox{\rm sg}({\cal C},x^{t})\equiv\lambda\,s^{t}-A^{*}(y-Ax^{t})$ where

is a valid subgradient of ${\cal C}$ at $x^{t}$ . On the other hand, $y-Ax^{t}=z^{t}-\omega_{t}z^{t-1}$ . We finally get

It is straightforward to see from Eqs. (LABEL:eq:x(t)=eta_interpreted) and (5.3) that $(I)=\lambda(x^{t-1}-x^{t})$ . Hence,

By Lemma 4.3, and the fact that $\sigma_{\max}(A)$ is almost surely bounded as $N\to\infty$ (cf. Theorem F.2), we deduce that the two terms $\lambda\|x^{t}-x^{t-1}\|/(\theta_{t-1}\sqrt{N})$ and $\sigma_{\max}(A)\|z^{t}-z^{t-1}\|^{2}/\sqrt{N}$ converge to when $N\to\infty$ and then $t\to\infty$ . For the third term, using state evolution (see Lemma 4.1), we obtain $\lim_{N\to\infty}\|z^{t-1}\|^{2}/N<\infty$ . Finally, using the calibration relation Eq. (1.18), we get

3 Proof of Lemma 3.4

The proof uses the representation (4.9), together with the expression (4.10) for the conditional expectation. Apart from the matrices $Y_{t}$ , $Q_{t}$ , $X_{t}$ , $M_{t}$ introduced there, we will also use

We state below a somewhat more convenient description.

It is clearly sufficient to prove that, for $v=v_{\parallel}+v_{\perp}$ , $P_{Q}v_{\parallel}=v_{\parallel}$ , $P_{Q}^{\perp}v_{\perp}=v_{\perp}$ , we have

The first identity is an easy consequence of the fact that $X^{*}Q=M^{*}AQ=M^{*}Y$ , while the second one follows immediately from $Q^{*}v_{\perp}=0$ . ∎

The following fact (see Appendix D for a proof) will be used several times.

For any $t$ there exists $c>0$ such that, for $R\in\{Q^{*}Q;\,M^{*}M;\,X^{*}X;\,Y^{*}Y\}$ , eventually almost surely as $N\to\infty$ ,

Given the above remarks, we will immediately see that Lemma 3.4 is implied by the following statement.

Let $S\subseteq[N]$ be given such that $|S|\leq N(\delta-\gamma)$ , for some $\gamma>0$ . Then there exists $\alpha_{1}=\alpha_{1}(\gamma)>0$ (independent of $t$ ) and $\alpha_{2}=\alpha_{2}(\gamma,t)>0$ (depending on $t$ and $\gamma$ ) such that

eventually almost surely as $N\to\infty$ . (With $Ev=Y(Q^{*}Q)^{-1}Q^{*}P_{Q}v+M(M^{*}M)^{-1}X^{*}P_{Q}^{\perp}v$ .)

In the next section we will show that this lemma implies Lemma 3.4. We will then prove the lemma just stated.

By Borel-Cantelli, it is sufficient to show that, for $S$ measurable on $\mathfrak{S}_{t}$ and $|S|\leq N(\delta-c)$ there exist $a_{1}=a_{1}(c)>0$ and $a_{2}=a_{2}(c,t)>0$ , such that

for all $N$ large enough. Conditioning on $\mathfrak{S}_{t}$ and using the union bound, this probability can be estimated as

where $h(p)=-p\log p-(1-p)\log(1-p)$ is the binary entropy function. The union bound calculation indeed proceeds as follows

where ${\sf X}_{S^{\prime}}=\min_{\|v\|=1,{\rm supp}(v)\subseteq S\cup S^{\prime}}\|Av\|$ . Now, fix $a_{1}<c/2$ in such a way that $h(a_{1})\leq\alpha_{1}(c/2)/2$ (with $\alpha_{1}$ defined as per Lemma 5.3). Further choose $a_{2}=\alpha_{2}(c/2,t)/2$ . The above probability is then upper bounded by

Finally, applying Lemma 5.3 and using Lemma 5.1 to estimate $Av$ , we get, for all $N$ large enough,

3.2 Proof of Lemma 5.3

We begin with the following Pythagorean inequality.

Let $S\subseteq[N]$ be given such that $|S|\leq N(\delta-\gamma)$ , for some $\gamma>0$ . Recall that $Ev=Y(Q^{*}Q)^{-1}Q^{*}P_{Q}v+M(M^{*}M)^{-1}X^{*}P_{Q}^{\perp}v$ and consider the event

Here the notation $(u,v)$ refers to the usual scalar product $u^{*}v$ of vectors $u$ and $v$ of the same dimension. Assuming that the claim holds, we have indeed

(Notice that in Proposition E.1 is stated for the equivalent case of a random sub-space of fixed dimension $d$ , and a subspace of dimension scaling linearly with the ambient one.) ∎

Let $S\subseteq[N]$ be given such that $|S|\leq N(\delta-\gamma)$ , for some $\gamma>0$ . Then there exists constant $c_{1}=c_{1}(\gamma)$ , $c_{2}=c_{2}(\gamma)$ such that the event

Let $V$ be the linear space $V={\rm im}(P_{Q}^{\perp}P_{S})$ . Of course the dimension of $V$ is at most $N(\delta-\gamma)$ . Then we have (for all vectors with ${\rm supp}(v)\subseteq S$ )

Finally a simple bound to control the norm of $Ev$ .

There exists a constant $c=c(t)>0$ such that, defining the event,

we have that ${\cal E}_{3}$ holds eventually almost surely as $N\to\infty$ .

The bound $\|EP_{Q}^{\perp}v\|\leq c(t)^{-1}\|P^{\perp}_{Q}v\|$ is proved analogously. ∎

By Lemma 5.6 we can assume that event ${\cal E}_{3}$ holds, for some function $c=c(t)$ (without loss of generality $c<1/2$ ). We will let ${\cal E}$ be the event

Let us assume first that $\|P_{Q}^{\perp}v\|\leq c^{2}/10$ , whence

where the last inequality uses $\|P_{Q}v\|=\sqrt{1-\|P_{Q}^{\perp}v\|^{2}}\geq 1/2$ . Therefore, using Lemma 5.4, we get

Next we assume $\|P_{Q}^{\perp}v\|\geq c^{2}/10$ . Due to Lemma 5.4 and 5.5 we can assume that events ${\cal E}_{1}$ and ${\cal E}_{2}$ hold. Therefore

4 Proof of Lemma 3.5

The key step consists in establishing the following result, which will be instrumental in the proof of Lemma 4.3 as well (and whose proof is deferred to Appendix C.1).

Assume $\alpha>\alpha_{\rm min}(\delta)$ and let $\{{\sf R}_{s,t}\}$ be defined by the recursion (4.13) with initial condition (4.14). Then there exists constants ${\sf B}_{1}$ , ${\sf r}_{1}>0$ such that for all $t\geq 0$

It is also useful to prove the following fact.

For any $\alpha>0$ and $T\geq 0$ , the $T\times T$ matrix $R_{T+1}\equiv\{{\sf R}_{s,t}\}_{0\leq s,t<T}$ is strictly positive definite.

almost surely. Hence, $R_{T+1}\stackrel{{\scriptstyle{\rm a.s.}}}{{=}}\delta\lim_{N\to\infty}(M_{T+1}^{*}M_{T+1}/N)$ . Thus the result follows from Lemma 5.2. ∎

It is then relatively easy to deduce the following.

By triangular inequality and Eq. (5.11), we have

which, together with Eq. (5.14) proves our claim. ∎

We are now in position to prove Lemma 3.5.

We will show that, under the assumptions of the Lemma, $\lim_{N\to\infty}|S_{t_{2}}(\gamma)\setminus S_{t_{1}}(\gamma)|/N\leq\xi$ almost surely, which implies our claim. Indeed, by Theorem 4.2 we have

Let $a\equiv(1-\gamma)\alpha\tau_{*}$ . By Proposition 1.3, for any $\varepsilon>0$ and all $t_{*}$ large enough we have $|(1-\gamma)\theta_{t_{i}-1}-a|\leq\varepsilon$ for $i\in\{1,2\}$ . Then

where the last inequality follows by Lemma 5.9. By taking $\varepsilon=e^{-{\sf r}_{2}\,t_{*}/3}$ we finally get (for some constant $C$ ) $P_{t_{1},t_{2}}\leq C\,e^{-{\sf r}_{2}t_{*}}$ , which implies our claim. ∎

Acknowledgement

It is a pleasure to thank David Donoho and Arian Maleki for many stimulating exchanges. We are also indebted with José Bento who collaborated in preparing Figures 4 to 7.

An earlier version of this paper stated some auxiliary lemmas in terms of convergence in probability. We rectified this to convergence almost sure as for the main theorems (with virtually no change in the proofs). We are grateful to Edgar Dobriban and Weijie Su for pointing out this inconsistency.

This work was partially supported by a Terman fellowship, the NSF CAREER award CCF-0743978 and the NSF grant DMS-0806211.

Appendix A Properties of the state evolution recursion

It is a straightforward calculus exercise to compute the partial derivatives

From these formulae we obtain the total derivative

with the inequality being strict whenever $\alpha>0$ , $u\neq 0$ . It follows that $\tau^{2}\mapsto{\sf F}(\tau^{2},\alpha\tau)$ is concave, and strictly concave provided $\alpha>0$ and $X_{0}$ is not identically .

which is strictly positive for all $\alpha\geq 0$ . To see this, let $f(\alpha)\equiv(1+\alpha^{2})\Phi(-\alpha)-\alpha\,\phi(\alpha)$ , and notice that $f^{\prime}(\alpha)=2\alpha\Phi(-\alpha)-2\phi(\alpha)<0$ , and $f(\infty)=0$ .

Since $\tau^{2}\mapsto{\sf F}(\tau^{2},\alpha\tau)$ is concave, and strictly increasing for $\tau^{2}$ large enough, it also follows that it is increasing everywhere.

Notice that $\alpha\mapsto f(\alpha)$ is strictly decreasing with $f(0)=1/2$ . Hence, for $\alpha>\alpha_{\rm min}(\delta)$ , we have ${\sf F}(\tau^{2},\alpha\tau)>\tau^{2}$ for $\tau^{2}$ small enough and ${\sf F}(\tau^{2},\alpha\tau)<\tau^{2}$ for $\tau^{2}$ large enough. Therefore the fixed point equation admits at least one solution. It follows from the concavity of $\tau^{2}\mapsto{\sf F}(\tau^{2},\alpha\tau)$ that the solution is unique and that the sequence of iterates $\tau_{t}^{2}$ converge to $\tau_{*}$ . $\Box$

A.2 Proof of Proposition 1.4

As a first step, we claim that $\alpha\mapsto\tau_{*}^{2}(\alpha)$ is continuously differentiable on $(0,\infty)$ . Indeed this is defined as the unique solution of

Since $(\tau^{2},\alpha)\mapsto{\sf F}(\tau^{2}_{*},\alpha\tau_{*})$ is continuously differentiable and $0\leq\frac{{\rm d}{\sf F}}{{\rm d}\tau^{2}}(\tau^{2}_{*},\alpha\tau_{*})<1$ (the second inequality being a consequence of concavity plus $\lim_{\tau^{2}\to\infty}\frac{{\rm d}{\sf F}}{{\rm d}\tau^{2}}(\tau^{2},\alpha\tau)<1$ , both shown in the proof of Proposition 1.3), the claim follows from the implicit function theorem applied to the mapping $(\tau^{2},\alpha)\mapsto[\tau^{2}-F(\tau^{2},\alpha)]$ .

Next notice that $\tau_{*}^{2}(\alpha)\to+\infty$ as $\alpha\downarrow\alpha_{\rm min}(\delta)$ . Indeed, introducing the notation ${\sf F}^{\prime}_{\infty}\equiv\lim_{\tau^{2}\to\infty}\frac{{\rm d}{\sf F}}{{\rm d}\tau^{2}}(\tau^{2},\alpha\tau)$ , we have, again by concavity,

i.e. $\tau_{*}^{2}\geq{\sf F}(0,0)/(1-{\sf F}^{\prime}_{\infty})$ . Now ${\sf F}(0,0)\geq\sigma^{2}$ , while ${\sf F}^{\prime}_{\infty}\uparrow 1$ as $\alpha\downarrow\alpha_{\rm min}(\delta)$ (shown in the proof of Proposition 1.3), whence the claim follows.

Next consider the function $(\alpha,\tau^{2})\mapsto g(\alpha,\tau^{2})$ defined by

Notice that $\lambda(\alpha)=g(\alpha,{\tau_{*}}^{2}(\alpha))$ . Since $g$ is continuously differentiable, it follows that $\alpha\mapsto\lambda(\alpha)$ is continuously differentiable as well.

Using the characterization of $\alpha_{\rm min}$ in Eq. (1.17) (and the well known inequality $\alpha\Phi(-\alpha)\leq\phi(\alpha)$ valid for all $\alpha>0$ ), it is immediate to show that $l_{*}<0$ . Therefore

A.3 Proof of Corollary 1.7

By Proposition 1.4, it is sufficient to prove that, for any $\lambda>0$ there exists a unique $\alpha>\alpha_{\rm min}$ such that $\lambda(\alpha)=\lambda$ . Assume by contradiction that there are two distinct such values $\alpha_{1}$ , $\alpha_{2}$ .

Notice that in this case, the function $\alpha(\lambda)$ is not defined uniquely and we can apply Theorem 1.5 to both choices $\alpha(\lambda)=\alpha_{1}$ and $\alpha(\lambda)=\alpha_{2}$ . Using the test function $\psi(x,y)=(x-y)^{2}$ we deduce that

Since the left hand side does not depend on the choice of $\alpha$ , it follows that $\tau_{*}(\alpha_{1})=\tau_{*}(\alpha_{2})$ .

Next apply Theorem 1.5 to the function $\psi(x,y)=|x|$ . We get

Appendix B Proof of Theorem 4.2

First note that using representation (4.2) we have $x^{t}+A^{*}z^{t}=x_{0}-h^{t+1}$ . Furthermore, using Lemma F.3(b) we have almost surely

For $s=t=0$ we have using Lemma F.3(b) almost surely

Induction hypothesis: Assume that for all $s\leq k$ and $t\leq k$ ,

Then we prove Eq. (B.2) for $t=k+1$ (case $s=k+1$ is similar). First assume $s=0$ and $t=k+1$ in which using Lemma F.3(c) we have almost surely

Similarly, for the case $t=k+1$ and $s>0$ , using Lemma F.3(b)(c) we have almost surely

using the induction hypothesis. Hence the result follows.

Appendix C Proof of Lemma 4.3

The proof of Lemma 4.3 relies on Lemma 5.7 which we will prove in the first subsection.

Before proving Lemma 5.7, we state and prove the following property of gaussian random variables.

for $f$ the indicator function of $I$ . Since the Ornstein-Uhlenbeck process is reversible with respect to the standard gaussian measure $\mu_{\rm G}$ , we have

It is convenient to change coordinates and define

Next we will show that by induction on $t$ that the stronger inequality $y_{t,3}<(y_{t,1}+y_{t,2})$ holds for all $t$ . We have indeed

The initial condition implied by Eq. (4.14) is

We will consider the above iteration for arbitrary initialization $y_{0}$ (satisfying $y_{0,3}<y_{0,1}+y_{0,2}$ ) and will show the following three facts:

Fact $(i)$ . As $t\to\infty$ , $y_{t,1},y_{t,2}\to\tau_{*}^{2}$ . Further the convergence is monotone.

Fact $(ii)$ . If $y_{0,1}=y_{0,2}=\tau_{*}^{2}$ and $y_{0,3}\leq 2\tau_{*}^{2}$ , then $y_{t,1}=y_{t,2}=\tau_{*}^{2}$ for all $t$ and $y_{t,3}\to 0$ .

Fact $(iii)$ . The jacobian $J=J_{{\sf G}}(y_{*})$ of ${\sf G}$ at $y_{*}=(\tau_{*}^{2},\tau_{*}^{2},0)$ has spectral radius $\sigma(J)<1$ .

By simple compactness arguments, Facts $(i)$ and $(ii)$ imply $y_{t}\to y_{*}$ as $t\to\infty$ . (Notice that $y_{t,3}$ remains bounded since $y_{t,3}\leq(y_{t,1}+y_{t,2})$ and by the convergence of $y_{t,1},y_{t,2}$ .) Fact $(iii)$ implies that convergence is exponentially fast.

Proof of Fact $(i)$ . Notice that $y_{t,2}$ evolves independently by $y_{t+1,2}={\sf G}_{2}(y_{t})={\sf F}(y_{2,t},\alpha\sqrt{y_{2,t}})$ , with ${\sf F}(\,\cdot\,,\,\cdot\,)$ the state evolution mapping introduced in Eq. (1.9). It follows from Proposition 1.3 that $y_{t,2}\to\tau_{*}^{2}$ monotonically for any initial condition. Since $y_{t+1,1}=y_{t,2}$ , the same happens for $y_{t,1}$ .

where $Z_{t-1}=\sqrt{\tau_{*}^{2}-x^{2}/4}Z+(x/2)W$ , $Z_{t}=\sqrt{\tau_{*}^{2}-x^{2}/4}Z-(x/2)W$ . In particular, by Lemma C.1, $x\mapsto{\sf G}_{*}(x)$ is strictly increasing (notice that the covariance of $Z_{t-1}$ and $Z_{t}$ is $\tau_{*}^{2}-(x/2)$ which is decreasing in $x$ ). Further

Hence, since $\lambda>0$ using Eq. (1.18) we have ${\sf G}^{\prime}(0)<1$ . Finally, by Lemma C.1, $x\mapsto{\sf G}^{\prime}(x)$ is decreasing in $[0,2\tau_{*})$ . It follows that $y_{t,3}\leq{\sf G}^{\prime}(0)^{t}y_{0,3}\to 0$ as claimed.

Proof of Fact $(iii)$ . From the definition of ${\sf G}$ , we have the following expression for the Jacobian

where with an abuse of notation we let ${\sf F}^{\prime}(\tau_{*}^{2})\equiv\left.\frac{{\rm d}\phantom{\tau^{2}}}{{\rm d}\tau^{2}}{\sf F}(\tau^{2},\alpha\tau)\right|_{\tau^{2}=\tau^{2}_{*}}$ . Computing the eigenvalues of the above matrix, we get

Since ${\sf G}_{*}^{\prime}(0)<1$ as proved above, and ${\sf F}(\tau_{*}^{2})<1$ as per Proposition 1.3, the claim follows. ∎

C.2 Lemma 5.7 implies Lemma 4.3

Using representations (4.4) and (4.3) (i.e., $b^{t}=w-z^{t}$ and $q^{t}=x_{0}-x^{t}$ ) and Lemma F.3(c) we obtain,

where the last equality uses $q^{t}=x^{t}-x_{0}$ . Therefore, it is sufficient to prove the thesis for $\|x^{t+1}-x^{t}\|_{2}$ . By state evolution, Theorem 4.2, we have

The first term vanishes as $t\to\infty$ because $\theta_{t}=\alpha\tau_{t}\to\alpha\tau_{*}$ by Proposition 1.3. The second term instead vanishes since ${\sf R}_{t,t}\to\tau_{*}$ , ${\sf R}_{t,t-1}\to\tau_{*}$ by Lemma 5.7.

Appendix D Proof of Lemma 5.2

First note that the upper bound on $\lambda_{\max}(R/N)$ is trivial since using representations (4.7), (4.8), $q^{t}=f_{t}(h^{t},x_{0})$ , $m^{t}=g_{t}(b^{t},w)$ and Lemma F.3(c)(d) all entries of the matrix $R/N$ are bounded as $N\to\infty$ and the matrix has fixed dimensions. Hence, we only focus on the lower-bound for $\lambda_{\min}(R/N)$ .

The result for $R=M^{*}M$ and $R=Q^{*}Q$ follows directly from Lemma F.3(g) and Lemma 8 of [BM11].

For $R=Y^{*}Y$ and $R=X^{*}X$ the proof is by induction on $t$ .

For $t=1$ we have $Y_{t}=b^{0}$ and $X_{t}=h^{1}+\xi_{0}q^{0}=h^{1}-x_{0}$ . Using Lemma F.3(b)(c) we obtain almost surely

Induction hypothesis: Assume that for all $t\leq k$ there exist positive constants $c_{X}(t)$ and $c_{Y}(t)$ such that as $N\to\infty$

First write $\vec{a}_{t}=(a_{1},\ldots,a_{t})$ and denote its first $t-1$ coordinates with $\vec{a}_{t-1}$ . Next, we consider the conditional distribution $A|_{\mathfrak{S}_{t-1}}$ . Using Eqs. (4.9) and (4.10) we obtain (since $Y_{t}=AQ_{t}$ )

Hence, conditional on $\mathfrak{S}_{t-1}$ we have, almost surely

From Lemma F.3(g) we know that $\lim_{N\to\infty}\langle q^{t-1}_{\perp},q^{t-1}_{\perp}\rangle$ is larger than a positive constant $\varsigma_{t}$ . Hence, from representation (D.3) and induction hypothesis (D.1)

To simplify the notation let $c^{\prime}_{t}\equiv\lim_{N\to\infty}N^{-1/2}\|b^{t-1}+\lambda_{t-1}m^{t-2}\|$ . Now if $c^{\prime}_{t}|a_{t}|\leq\sqrt{c_{Y}(t-1)}\|\vec{a}_{t-1}\|/2$ then

which proves the result. Otherwise, we obtain the inequality

Appendix E A concentration estimate

The following proposition follows from standard concentration-of-measure arguments.

Let $N_{d}(\varepsilon/2)$ be a $(\varepsilon/2)$ -net in $S_{d}$ , i.e. a subset of vectors $\{u^{1},\dots,u^{M}\}\in S^{d}$ such that, for any $u\in S^{d}$ , there exists $i\in\{1,\dots,M\}$ such that $\|u-u^{i}\|\leq\varepsilon/2$ . It follows from a standard counting argument [Led01] that there exists an $(\varepsilon/2)$ -net of size $|N_{d}(\varepsilon/2)|\equiv M\leq(100/\varepsilon)^{d}$ . Define

Since $u\mapsto P_{\lambda}Qu$ is Lipschitz with modulus $1$ , we have

which is smaller than $e^{-mc(\varepsilon)/2}$ for all $m$ large enough. ∎

Appendix F Useful reference material

In this appendix we collect a few known results that are used several times in our proof. We also provide some pointers to the literature.

In our proof we make use of the following well-known result of Kashin in the theory of diameters of smooth functions [Kas77].

For any positive number $\upsilon$ there exist a universal constant $c_{\upsilon}$ such that for any $n\geq 1$ , with probability at least $1-2^{-n}$ , for a uniformly random subspace $V_{n,\upsilon}$ of dimension $\lfloor n(1-\upsilon)\rfloor$ ,

F.2 Singular values of random matrices

We will repeatedly make use of limit behavior of extreme singular values of random matrices. A very general result was proved in [BY93] (see also [BS05]).

We will also use the following fact that follows from the standard singular value decomposition

F.3 Two Lemmas from [BM11]

Our proof uses the results of [BM11]. We state copy here the crucial technical lemma in that paper. Notations refer to the general algorithm in Eq. (4.1). General state evolution defines quantities $\{\tau_{t}^{2}\}_{t\geq 0}$ and $\{\sigma_{t}^{2}\}_{t\geq 0}$ via

where $W\sim p_{W}$ and $X_{0}\sim p_{X_{0}}$ are independent of $Z\sim{\sf N}(0,1)$

where $(Z_{0},\ldots,Z_{t})$ and $(\hat{Z}_{0},\ldots,\hat{Z}_{t})$ are two zero-mean gaussian vectors independent of $X_{0}$ , $W$ , with $Z_{i},\hat{Z}_{i}\sim{\sf N}(0,1)$ .

For all $0\leq r,s\leq t$ the following equations hold and all limits exist, are bounded and have degenerate distribution (i.e. they are constant random variables):

Here $\varphi^{\prime}$ denotes derivative with respect to the first coordinate of $\varphi$ .

For all $0\leq r\leq t$ and $0\leq s\leq t-1$ the following limits exist, and there exist strictly positive constants $\rho_{r}$ and $\varsigma_{s}$ (independent of $N$ , $n$ ) such that almost surely

It is also useful to recall some simple properties of gaussian random matrices.

Introduction

2 Main result

2.2 Main results

3 Related work

Numerical illustrations

A structural property and proof of the main results

2 A structural property of the LASSO cost function

3 Proof of Theorem 1.8

State evolution estimates

2 Some consequences and generalizations

Proofs of auxiliary lemmas

2 Proof of Lemma 3.3

3 Proof of Lemma 3.4

3.2 Proof of Lemma 5.3

4 Proof of Lemma 3.5

Acknowledgement

Appendix A Properties of the state evolution recursion

A.2 Proof of Proposition 1.4

A.3 Proof of Corollary 1.7

Appendix B Proof of Theorem 4.2

Appendix C Proof of Lemma 4.3

C.2 Lemma 5.7 implies Lemma 4.3

Appendix D Proof of Lemma 5.2

Appendix E A concentration estimate

Appendix F Useful reference material

F.2 Singular values of random matrices

F.3 Two Lemmas from [BM11]

References