Sparse recovery under weak moment assumptions

Guillaume Lecué, Shahar Mendelson

Introduction and main results

Data acquisition is an important task in diverse fields such as mobile communications, medical imaging, radar detection and others, making the design of efficient data acquisition processes a problem of obvious significance.

The core issue in data acquisition is retaining all the valuable information at one’s disposal, while keeping the ‘acquisition cost’ as low as possible. And while there are several ways of defining that cost, depending on the problem (storage, time, financial cost, etc.), the common denominator of being ‘cost effective’ is ensuring the quality of the data while keeping the number of measurements as small as possible.

The rapidly growing area of Compressed Sensing studies ‘economical’ data acquisition processes. We refer the reader to and to the book for more information on the origins of Compressed Sensing and a survey of the progress that has been made in the area in recent years.

At the heart of Compressed Sensing is a simple idea that has been a recurring theme in Mathematics and Statistics: while complex objects (in this case, data), live in high-dimensional spaces, they can be described effectively using low-dimensional, approximating structures; moreover, randomness may be used to expose these low-dimensional structures. Of course, unlike more theoretical applications of this idea, identifying the low-dimensional structures in the context of Compressed Sensing must be robust and efficient, otherwise, such procedures will be of little practical use.

Because the resulting system of equations is under-determined, there is no hope, in general, of identifying $x_{0}$ . However, if $x_{0}$ is believed to be well approximated by a low-dimensional structure, for example, if $x_{0}$ is supported on at most $s$ coordinates for some $s\leq N$ , the problem becomes more feasible.

Given the measurement matrix $\Gamma$ and the measurements $\Gamma x_{0}=(\bigl{<}X_{i},x_{0}\bigr{>})_{i=1}^{N}$ , Basis Pursuit returns a vector $\hat{x}$ that satisfies

Since one may solve this minimization problem effectively, the focus may be shifted to the quality of the solution: whether one can identify measurement vectors $X_{1},....,X_{N}$ for which (1.1) has a unique solution, which is $x_{0}$ itself, for any $x_{0}$ that is $s$ -sparse (i.e. supported on at most $s$ coordinates).

It follows from Proposition 2.2.18 in that if $\Gamma$ satisfies ER( $s$ ) then necessarily the number of measurements (rows) is at least $N\geq c_{0}s\log\big{(}en/s\big{)}$ , where $c_{0}$ is a suitable absolute constant. On the other hand, there are constructions of (random) matrices $\Gamma$ that satisfy ER( $s$ ) with $N$ proportional to $s\log\big{(}en/s\big{)}$ . From here on and with a minor abuse of notation, we will refer to $s\log(en/s)$ as the optimal number of measurements and ignore the exact dependence on the constant $c_{0}$ .

Unfortunately, the only matrices that are known to satisfy the reconstruction property with an optimal number of measurements are random – which is not surprising, as randomness is one of the most effective tools in exposing low-dimensional, approximating structures. A typical example of an ‘optimal matrix’ is the Gaussian matrix, which has independent standard normal random variables as entries. Other examples of optimal measurement matrices are $\Gamma=N^{-1/2}\sum_{i=1}^{N}\bigl{<}X_{i},\cdot\bigr{>}f_{i}$ where $X_{1},...,X_{N}$ are independent, isotropic and $L$ -subgaussian random vectors:

The optimal behaviour of isotropic, $L$ -subgaussian matrix ensembles and other ensembles like it, occurs because a typical matrix acts on $\Sigma_{s}$ in an isomorphic way when $N\geq c_{1}s\log(en/s)$ , and in the $L$ -subgaussian case, $c_{1}$ is a constant that depends only on $L$ . In Compressed Sensing literature, this isomorphic behaviour is called the Restricted Isometry property (RIP) (see, for example ): A matrix $\Gamma$ satisfies the RIP in $\Sigma_{s}$ with constant $0<\delta<1$ , if for every $t\in\Sigma_{s}$ ,

It is straightforward to show that if $\Gamma$ satisfies the RIP in $\Sigma_{2s}$ for a sufficiently small constant $\delta$ , then it has the exact reconstruction property of order $s$ (see, e.g. ).

Proving the RIP for subexponential ensembles is a much harder task than for subgaussian ensembles (see, e.g. ). Moreover, the RIP does not exhibit the same optimal quantitative behaviour as in the Gaussian case: it holds with high probability only when $N\geq c_{2}(L)s\log^{2}(en/s)$ , and this estimate cannot be improved, as can be seen when $X$ has independent, symmetric exponential random variables as coordinates .

Although the RIP need not be true for isotropic $L$ -subexponential ensemble using the optimal number of measurements, results in (see Theorem 7.3 there) and in show that exact reconstruction can still be achieved by such an ensemble and with the optimal number of measurements. This opens the door to an intriguing question: whether considerably weaker assumptions on the measurement vector may still lead to Exact Reconstruction even when the RIP fails.

The main result presented here does just that, using the small-ball method introduced in .

A random vector $X$ satisfies the small-ball condition in the set $\Sigma_{s}$ with constants $u,\beta>0$ if for every $t\in\Sigma_{s}$ ,

The small-ball condition is a rather minimal assumption on the measurement vector and is satisfied in fairly general situations for values of $u$ and $\beta$ that are suitable constants, independent of the dimension $n$ .

Under some normalization (like isotropicity), a small-ball condition is an immediate outcome of the Paley-Zygmund inequality (see, e.g. ) and moment equivalence. For example, in the following cases a small-ball condition holds with constants that depend only on $\kappa_{0}$ (and on $\varepsilon$ for the first case); the straightforward proof may be found in .

$\bullet$ $X$ is isotropic and for every $t\in\Sigma_{s}$ , $\left\|\bigl{<}X,t\bigr{>}\right\|_{L_{2+\varepsilon}}\leq\kappa_{0}\left\|\bigl{<}X,t\bigr{>}\right\|_{L_{2}}$ for some $\varepsilon>0$ ;

$\bullet$ $X$ is isotropic and for every $t\in\Sigma_{s}$ , $\left\|\bigl{<}X,t\bigr{>}\right\|_{L_{2}}\leq\kappa_{0}\left\|\bigl{<}X,t\bigr{>}\right\|_{L_{1}}$ .

Our first result shows that a combination of the small-ball condition and a weak moment assumption suffices to ensure the exact reconstruction property with the optimal number of measurements.

there are $\kappa_{1},\kappa_{2},w>1$ that satisfy that for every $1\leq j\leq n$ , $\|x_{j}\|_{L_{2}}=1$ and, for every $4\leq p\leq 2\kappa_{2}\log(wn)$ , $\|x_{j}\|_{L_{p}}\leq\kappa_{1}p^{\alpha}$ .

$X$ satisfies the small ball condition in $\Sigma_{s}$ with constants $u$ and $\beta$ .

and $X_{1},...,X_{N}$ are independent copies of $X$ , then, with probability at least

$\Gamma=N^{-1/2}\sum_{i=1}^{N}\bigl{<}X_{i},\cdot\bigr{>}f_{i}$ satisfies the exact reconstruction property in $\Sigma_{s_{1}}$ for $s_{1}=c_{2}u^{2}\beta s$ .

An immediate outcome of Theorem A is the following:

$\bullet$ Let $x$ be a centered random variable that has variance $1$ and for which $\|x\|_{L_{p}}\leq c\sqrt{p}$ for $1\leq p\leq 2\log n$ . If $X$ has independent coordinates distributed as $x$ , then the corresponding matrix $\Gamma$ with $N\geq c_{1}s\log(en/s)$ rows can be used as a measurement matrix and recover any $s$ -sparse vector with large probability.

It is relatively straightforward to derive many other results of a similar flavour, leading to random ensembles that satisfy the exact reconstruction property with the optimal number of measurements.

Our focus is on measurement matrices with independent rows, that satisfy conditions of a stochastic nature – they have i.i.d. rows. Other types of measurement matrices that have some structure have also been used in Compressed Sensing. One notable example is a random Fourier measurement matrix, obtained by randomly selecting rows from the discrete Fourier matrix (see, e.g. , or Chapter 12 in ).

One may wonder if the small-ball condition is satisfied for more structured matrices, as the argument we use here does not extend immediately to such cases. And, indeed, for structured ensembles one may encounter a different situation: a small-ball condition that is not uniform, in the sense that the constants $u$ and $\beta$ from Definition 1.4 are direction-dependent. Moreover, in some cases, the known estimates on these constants are far from what is expected.

Results of the same flavour of Theorem A may follow from a ‘good enough’ small-ball condition, even if it is not uniform, by slightly modifying the argument we use here. However, obtaining a satisfactory ‘non-uniform’ small-ball condition is a different story. For example, in the Fourier case, such an estimate is likely to require quantitative extensions of the Littlewood-Paley theory – a worthy challenge in its own right, and one which goes far beyond the goals of this article.

Just as noted for subexponential ensembles, Theorem A cannot be proved using an RIP-based argument. A key ingredient in the proof is the following observation:

for every $x\in\Sigma_{s}$ , $\left\|\Gamma x\right\|_{2}\geq c_{0}\left\|x\right\|_{2}$ , and

for every $j\in\{1,\ldots,n\}$ , $\left\|\Gamma e_{j}\right\|_{2}\leq c_{1}$ .

Setting $s_{1}=\big{\lfloor}(c_{0}^{2}(s-1))/(4c_{1}^{2})\big{\rfloor}-1$ , $\Gamma$ satisfies the exact reconstruction property in $\Sigma_{s_{1}}$ .

Compared with the RIP, conditions a) and b) in Theorem B are weaker, as it suffices to verify the right-hand side of (1.2) for $1$ -sparse vectors rather than for every $s$ -sparse vector. This happens to be a substantial difference: the assumption that for every $t\in\Sigma_{s}$ , $\left\|\Gamma t\right\|_{2}\leq(1+\delta)\left\|t\right\|_{2}$ is a costly one, and happens to be the reason for the gap between the RIP and the exact reconstruction property. Indeed, while the lower bound in the RIP holds for rather general ensembles (see and the next section for more details), and is guaranteed solely by the small-ball condition, the upper bound is almost equivalent to having the coordinates of $X$ exhibit a subgaussian behaviour of moments, at least up to some level. Even the fact that one has to verify the upper bound for $1$ -sparse vectors comes at a cost, namely, the moment assumption (1) in Theorem A.

The second goal of this note is to illustrate that while Exact Reconstruction is ‘cheaper’ than the RIP, it still comes at a cost – namely, that the moment condition (1) in Theorem A is truly needed.

A random matrix $\Gamma$ is generated by the random variable $x$ if $\Gamma=N^{-1/2}\sum_{i=1}^{N}\bigl{<}X_{i},\cdot\bigr{>}f_{i}$ and $X_{1},...,X_{N}$ are independent copies of the random vector $X=(x_{1},...,x_{n})^{\top}$ whose coordinates are independent copies of $x$ .

Theorem C. There exist absolute constants $c_{0},c_{1},c_{2}$ and $c_{3}$ for which the following holds. Given $n\geq c_{0}$ and $N\log N\leq c_{1}n$ , there exists a mean-zero, variance one random variable $x$ with the following properties:

$\bullet$ $\left\|x\right\|_{L_{p}}\leq c_{2}\sqrt{p}$ for $2<p\leq c_{3}(\log n)/(\log N)$ .

$\bullet$ If $(x_{j})_{j=1}^{n}$ are independent copies of $x$ then $X=(x_{1},...,x_{n})^{\top}$ satisfies the small-ball condition with constants $u$ and $\beta$ that depend only on $c_{2}$ .

$\bullet$ Denote by $\Gamma$ the $N\times n$ matrix generated by $x$ . For every $k\in\{1,\ldots,n\}$ , with probability larger than $1/2$ , $\operatorname*{argmin}\big{(}\left\|t\right\|_{1}:\Gamma t=\Gamma e_{k}\big{)}\neq\{e_{k}\}$ ; therefore, $e_{k}$ is not exactly reconstructed by Basis Pursuit and so $\Gamma$ does not satisfy the exact reconstruction property of order $1$ .

To put Theorem C in some perspective, note that if $\Gamma$ is generated by $x$ for which $\left\|x\right\|_{L_{2}}=1$ and $\left\|x\right\|_{L_{p}}\leq c_{4}\sqrt{p}$ for $2<p\leq c_{5}\log n$ , then $X=(x_{i})_{i=1}^{n}$ satisfies the small-ball condition with constants that depend only on $c_{4}$ , and by Theorem A, if $N\geq c_{6}\log n$ , $\Gamma$ satisfies ER( $1$ ) with high probability. On the other hand, the random ensemble from Theorem C is generated by $x$ that has almost identical properties – with one exception: its $L_{p}$ norm is well behaved only for $p\leq c_{7}(\log n)/\log\log n$ . This small gap in the number of moments has a significant impact: with probability at least $1/2$ , $\Gamma$ does not satisfy ER( $1$ ) when $N$ is of the order of $\log n$ .

The fact that the coordinates of $X$ do not have enough well behaved moments is the key feature that allows one to generate many ‘spiky’ columns in a typical $\Gamma$ .

An alternative formulation of Theorem C is the following: Theorem C′. There are absolute constants $c_{0},c_{1},c_{2}$ and $\kappa$ for which the following holds. If $n\geq c_{0}$ and $2<p<c_{1}\log n$ , there exists a mean-zero and variance $1$ random variable $x$ , for which $\|x\|_{L_{q}}\leq\kappa\sqrt{q}$ for $2<q\leq p$ , and if $N\leq c_{2}\sqrt{p}(n/\log n)^{1/p}$ and $\Gamma$ is the $N\times n$ matrix generated by $x$ , then with probability at least $1/2$ , $\Gamma$ does not satisfy the exact reconstruction property of order 1.

$X$ satisfies a weak small-ball condition in $\Sigma_{s}$ with constant $\beta$ if for every $t\in\Sigma_{s}$ ,

We end this introduction with a word about notation and the organization of the article. The proofs of Theorem A, Theorem B and Theorem D are presented in the next section, while the proofs of Theorem C and Theorem C′ may be found in Section 3. The final section is devoted to results in a natural ‘noisy’ extension of Compressed Sensing. In particular, we prove that both the Compatibility Condition and the Restricted Eigenvalue Condition hold under weak moment assumptions; we also study related properties of random polytopes.

As for notation, throughout, absolute constants or constants that depend on other parameters are denoted by $c$ , $C$ , $c_{1}$ , $c_{2}$ , etc., (and, of course, we will specify when a constant is absolute and when it depends on other parameters). The values of these constants may change from line to line. The notation $x\sim y$ (resp. $x\lesssim y$ ) means that there exist absolute constants $0<c<C$ for which $cy\leq x\leq Cy$ (resp. $x\leq Cy$ ). If $b>0$ is a parameter then $x\lesssim_{b}y$ means that $x\leq C(b)y$ for some constant $C(b)$ that depends only on $b$ .

Proof of Theorem A, B and D

The proof of Theorem A has several components, and although the first of which is rather standard, we present it for the sake of completeness.

Proof. Observe that if $x\in B_{1}^{n}$ and $\|x\|_{2}\geq r$ then $y=rx/\|x\|_{2}\in B_{1}^{n}\cap rS^{n-1}$ . Therefore, if $y\not\in{\rm ker}(\Gamma)$ , the same holds for $x$ ; thus

Let $s=\lfloor(2r)^{-2}\rfloor$ , fix $x_{0}\in\Sigma_{s}$ and put $I$ to be the set indices of coordinates on which $x_{0}$ is supported. Given a nonzero $h\in{\rm ker}(\Gamma)$ , let $h=h_{I}+h_{I^{c}}$ – the decomposition of $h$ to coordinates in $I$ and in $I^{c}$ . Since $h/\|h\|_{1}\in B_{1}^{n}\cap{\rm ker}(\Gamma)$ , it follows that $\|h\|_{2}<r\|h\|_{1}$ , and by the choice of $s$ , $2\sqrt{s}\|h\|_{2}<\|h\|_{1}$ . Therefore,

Hence, $\|x_{0}+h\|_{1}>\|x_{0}\|_{1}$ and $x_{0}$ is the unique minimizer of the basis pursuit algorithm.

The main ingredient in the proof of Theorem A is Lemma 2.3 below, which is based on the small-ball method introduced in . To formulate the lemma, one requires the notion of a VC class of sets.

Let ${\cal G}$ be a class of $\{0,1\}$ -valued functions defined on a set ${\cal X}$ . The set ${\cal G}$ is a VC-class if there exists an integer $V$ for which, given any $x_{1},...,x_{V+1}\in{\cal X}$ ,

The VC-dimension of ${\cal G}$ , denoted by $VC({\cal G})$ , is the smallest integer $V$ for which (2.1) holds.

The VC dimension is a combinatorial complexity measure that may be used to control the $L_{2}(\mu)$ -covering numbers of the class; indeed, set $N({\cal G},\varepsilon,L_{2}(\mu))$ to be the smallest number of open balls of radius $\varepsilon$ relative to the $L_{2}(\mu)$ norm that are needed to cover ${\cal G}$ . A well known result due to Dudley is that if $VC({\cal G})=V$ and $\mu$ is a probability measure on ${\cal X}$ then for every $0<\varepsilon<1$ ,

where $c_{1}$ and $c_{2}$ are absolute constants.

There exist absolute constants $c_{1}$ and $c_{2}$ for which the following holds. Let ${\cal F}$ be a class of functions and assume that there are $\beta>0$ and $u\geq 0$ for which

Note that $u=0$ is a ‘legal choice’ in Lemma 2.3, a fact that will be used in the proof of Theorem D.

Standard empirical processes arguments (symmetrization, the fact that Bernoulli processes are subgaussian and the entropy estimate (2.2) – see, for example, Chapters 2.2, 2.3 and 2.6 in ), show that since $VC({\cal G})\leq d$ ,

provided that $N\gtrsim d/\beta^{2}$ . Therefore, taking $t=N\beta^{2}/16c_{1}^{2}$ , it follows that with probability at least $1-\exp(-c_{3}\beta^{2}N)$ , for every $f\in{\cal F}$ ,

Therefore, on that event, $|\{i:|f(X_{i})|>u\}|\geq\beta N/2$ for every $f\in{\cal F}$ .

1. If there are $0<\beta\leq 1$ and $u\geq 0$ for which $P\big{(}|\bigl{<}t,X\bigr{>}|>u\big{)}\geq\beta$ for every $t\in S^{n-1}$ and if $N\geq c_{1}n/\beta^{2}$ , then with probability at least $1-\exp(-c_{2}N\beta^{2})$ ,

2. If there are $0<\beta\leq 1$ and $u\geq 0$ for which $P\big{(}|\bigl{<}t,X\bigr{>}|>u\big{)}\geq\beta$ for every $t\in\Sigma_{s}\cap S^{n-1}$ and if $N\geq c_{1}s\log(en/s)/\beta^{2}$ , then with probability at least $1-\exp(-c_{2}N\beta^{2})$ ,

Note that the first part of Corollary 2.5 gives an estimate on the smallest singular value of the random matrix $\Gamma=N^{-1/2}\sum_{i=1}^{N}\bigl{<}X_{i},\cdot\bigr{>}f_{i}$ . The proof follows the same path as in , but unlike the latter, no assumption on the covariance structure of $X$ , used both in and in , is required. In fact, Corollary 2.5 may be applied even if the covariance matrix does not exist. Thus, under a small-ball condition, the smallest singular value of $\Gamma$ is larger than $c(\beta,u)$ with high (exponential) probability.

is at most $c_{1}n$ for a suitable absolute constant $c_{1}$ (see, e.g., Chapter 2.6 in ). The claim now follows immediately from Lemma 2.3 because

Turning to the second part, note that $\Sigma_{s}\cap S^{n-1}$ is a union of $\binom{n}{s}$ spheres of dimension $s$ . Applying the first part to each one of those spheres, combined with the union bound, it follows that for $N\geq c_{2}\beta^{-2}s\log(en/s)$ , with probability at least $1-\exp(-c_{3}N\beta^{2})$ ,

Corollary 2.5 shows that the small-ball condition for linear functionals implies that $\Gamma$ ‘acts well’ on $s$ -sparse vectors. However, according to Lemma 2.1, exact recovery is possible if $\Gamma$ is well behaved on the set

for a well-chosen constant $\kappa_{0}$ . In the standard (RIP-based) argument, one proves exact reconstruction by first showing that the RIP holds in $\Sigma_{s}$ , and then the fact that each vector in $\sqrt{\kappa_{0}s}B_{1}^{n}\cap S^{n-1}$ is well approximated by vectors from $\Sigma_{s}$ (see, for instance, ) allows one to extend the RIP from $\Sigma_{s}$ to $\sqrt{\kappa_{0}s}B_{1}^{n}\cap S^{n-1}$ . Unfortunately, this extension requires both upper and lower estimates in the RIP.

Since the upper part of the RIP in $\Sigma_{s}$ forces severe restrictions on the random vector $X$ , one has to resort to a totally different argument if one wishes to extend the lower bound from $\Sigma_{s}$ (which only requires the small-ball condition) to $\sqrt{\kappa_{0}s}B_{1}^{n}\cap S^{n-1}$ .

The method presented below is based on Maurey’s empirical method and has been recently used in .

Let $Y_{1},...,Y_{s}$ be independent copies of $Y$ and set $Z=s^{-1}\sum_{k=1}^{s}Y_{k}$ . Note that $Z\in\Sigma_{s}$ for every realization of $Y_{1},...,Y_{s}$ ; thus $\|\Gamma Z\|_{2}^{2}\geq\lambda^{2}\|Z\|_{2}^{2}$ and

Therefore, setting $\mu_{j}=|y_{j}|/\|y\|_{1}$ and $W=\sum_{j=1}^{n}\left\|\Gamma e_{j}\right\|_{2}^{2}\mu_{j}$ ,

and using the same argument one may show that

Combining these two estimates with (2.4),

Proof of Theorem B: Assume that for every $x\in\Sigma_{s}$ , $\|\Gamma x\|_{2}\geq c_{0}\|x\|_{2}$ and that for every $1\leq i\leq n$ , $\|\Gamma e_{i}\|_{2}\leq c_{1}$ . It follows from Lemma 2.7 that if $s-1>c_{1}^{2}/(c_{0}^{2}r^{2})$ , then for every $y\in B_{1}^{n}\cap rS^{n-1}$ ,

which is an average of $N$ iid random variables (though $\left\|\Gamma e_{1}\right\|_{2},$ $\ldots,\left\|\Gamma e_{n}\right\|_{2}$ need not be independent).

Thanks to Theorem B and Corollary 2.5, the final component needed for the proof of Theorem A is information on the sum of iid random variables, which will be used to bound $\max_{1\leq j\leq n}\left\|\Gamma e_{j}\right\|_{2}^{2}$ from above.

There exists an absolute constant $c_{0}$ for which the following holds. Let $z$ be a mean-zero random variable and put $z_{1},\ldots,z_{N}$ to be $N$ independent copies of $z$ . Let $p_{0}\geq 2$ and assume that there exists $\kappa_{1}>0$ and $\alpha\geq 1/2$ for which $\|z\|_{L_{p}}\leq\kappa_{1}p^{\alpha}$ for every $2\leq p\leq p_{0}$ . If $N\geq p_{0}^{\max\{2\alpha-1,1\}}$ then for every $2\leq p\leq p_{0}$ ,

where $c_{1}(\alpha)=c_{0}\exp((2\alpha-1))$ .

Lemma 2.8 shows that even under a weak moment assumption, namely that $\|z\|_{L_{p}}\lesssim p^{\alpha}$ for $p\leq p_{0}$ and $\alpha\geq 1/2$ that can be large, a normalized sum of $N$ independent copies of $z$ exhibits a ‘subgaussian’ moment growth up to the same $p_{0}$ , as long as $N$ is sufficiently large.

The proof of Proposition 2.8 is based on the following fact due to Latała.

If $z$ is a mean-zero random variable and $z_{1},...,z_{N}$ are independent copies of $z$ , then for any $p\geq 2$ ,

Proof of Lemma 2.8. Let $2\leq p\leq p_{0}$ and $N\geq p$ . Since $\|z\|_{L_{s}}\leq\kappa_{1}s^{\alpha}$ for any $2\leq s\leq p$ , it follows from Theorem 2.9 that

It is straightforward to verify that the function $h(s)=(N/p)^{1/s}s^{-1+\alpha}$ is non-increasing when $\alpha\leq 1$ and attains its maximum in $s=\max\{2,p/N\}=2$ or in $s=p$ when $\alpha>1$ . Therefore, when $N\geq p$ ,

Finally, if $N\geq p^{2\alpha-1}$ then $e^{2\alpha-1}\sqrt{Np}\geq N^{1/p}p^{\alpha}$ , which completes the proof.

Proof of Theorem A. Consider $N\geq c_{1}s\log(en/s)/\beta^{2}$ . By Corollary 2.5, with probability at least $1-\exp(-c_{2}N\beta^{2})$ ,

Set $(X_{i})_{i=1}^{N}$ for which (2.5) holds and let $\Gamma=N^{-1/2}\sum_{i=1}^{N}\bigl{<}X_{i},\cdot\bigr{>}f_{i}$ . By Lemma 2.7 for $\lambda^{2}=u^{2}\beta/2$ , it follows that when $r\geq 1$ ,

Next, one has to obtain a high probability upper estimate on $\max_{1\leq j\leq n}\|\Gamma e_{j}\|_{2}^{2}$ . To that end, fix $w\geq 1$ and consider $z=x_{j}^{2}-1$ - where $x_{j}$ is the $j$ -th coordinate of $X$ . Observe that $z$ is a centered random variable and that $\left\|z\right\|_{L_{q}}\lesssim 4^{\alpha}\kappa_{1}^{2}q^{2\alpha}$ for every $1\leq q\leq\kappa_{2}\log(wn)$ . Thus, by Lemma 2.8 for $p=\kappa_{2}\log(wn)$ and $c_{3}(\alpha)\sim 4^{\alpha}\exp((4\alpha-1))$ ,

provided that $N\geq p^{\max\{4\alpha-1,1\}}=(\kappa_{2}\log(wn))^{\max\{4\alpha-1,1\}}$ . Hence, if $N\geq(c_{3}(\alpha)\kappa_{1}^{2})^{2}(\kappa_{2}\log(wn))^{\max\{4\alpha-1,1\}}$ , and setting $V_{j}=\left\|\Gamma e_{j}\right\|_{2}^{2}$ , one has

and $r\leq s\lambda^{2}/8e=su^{2}\beta/16e$ , then with probability at least $1-\exp(-c_{2}N\beta^{2})-1/(w^{\kappa_{2}}n^{\kappa_{2}-1})$ ,

Therefore, by Lemma 2.1, $\Gamma$ satisfies the exact reconstruction property for vectors that are $c_{4}u^{2}\beta s$ -sparse, as claimed.

Proof of Theorem C and Theorem C′

Proof. Let $w\in B_{1}^{J^{c}}$ for which $\Gamma v=\Gamma w$ and observe that $v\not=w$ (otherwise, $v\in B_{1}^{J}\cap B_{1}^{J^{c}}$ , implying that $v=0$ , which is impossible because $\left\|v\right\|_{1}=1$ ).

Set $x_{\cdot 1},\cdots,x_{\cdot n}$ to be the columns of $\Gamma$ . It immediately follows from Lemma 3.1 that if one wishes to prove that $\Gamma$ does not satisfy ER( $1$ ), it suffices to show that, for instance the first basis vector $e_{1}$ cannot be exactly reconstruct. This follows from

where ${\rm absconv}(S)$ is the convex hull of $S\cup-S$ . Therefore, if

for some absolute constant $c_{0}$ , then $\Gamma$ does not satisfy ER( $1$ ).

The proofs of Theorem C and of Theorem C′ follow from the construction of a random matrix ensemble for which (3.1) holds with probability larger than $1/2$ . We now turn on to such a construction.

Let $\eta$ be a selector (a $\{0,1\}$ -valued random variable) with mean $\delta$ to be named later, and let $\varepsilon$ be a symmetric $\{-1,1\}$ -valued random variable that is independent of $\eta$ . Fix $R>0$ and set

Observe that if $p\geq 2$ and $R\geq 1$ then

and the last equivalence holds when $R^{2}\delta\lesssim 1$ and $R^{p}\delta\gtrsim 1$ . Fix $2<p\leq 2\log(1/\delta)$ which will be specified later and set $R=\sqrt{p}(1/\delta)^{1/p}$ . Since the function $q\to\sqrt{q}/\delta^{1/q}$ is decreasing for $2\leq q\leq 2\log(1/\delta)$ one has that for $2\leq q\leq p$ and for $\delta$ that is small enough,

Note that $x=z/\left\|z\right\|_{L_{2}}$ is a mean-zero, variance one random variable that exhibits a ‘subgaussian’ moment behaviour only up to $p$ . Indeed, if $2\leq q\leq p$ , $\|z\|_{L_{q}}\lesssim\sqrt{q}\|z\|_{L_{2}}$ , and if $q>p$ , $\left\|z\right\|_{L_{q}}\sim\sqrt{p}\delta^{1/q-1/p}\left\|z\right\|_{L_{2}}$ , which may be far larger than $\sqrt{q}\left\|z\right\|_{L_{2}}$ if $\delta$ is sufficiently small.

Let $X=(x_{1},\ldots,x_{n})$ be a vector whose coordinates are independent, distributed as $x$ and let $\Gamma$ be the measurement matrix generated by $x$ . Note that up to the normalization factor of $\left\|z\right\|_{L_{2}}$ , which is of the order of a constant when $R^{2}\delta\lesssim 1$ , $\sqrt{N}\Gamma$ is a perturbation of a Rademacher matrix by a sparse matrix with few random spikes that are either $R$ or $-R$ .

the convex hull of $(\pm v_{j})_{j=2}^{n}$ .

We will show that with probability at least $1/2$ , $\sqrt{N}B_{2}^{N}\subset V$ and $\|v_{1}\|_{2}\leq\sqrt{N}$ , in three steps:

With probability at least $3/4$ , for every $1\leq i\leq N$ there is $y_{i}\in B_{\infty}^{N}$ for which $y_{i}+Rf_{i}\in V$ .

Hence, the claim follows by the union bound and integration with respect to the $(\varepsilon_{ij})$ .

Next, it is straightforward to verify that when $V$ contains such a perturbation of $RB_{1}^{N}$ (by vectors in $B_{\infty}^{N}$ ), it must also contain a large Euclidean ball, assuming that $R$ is large enough.

Let $R>N$ , and for every $1\leq i\leq N$ , set $y_{i}\in B_{\infty}^{N}$ and put $v_{i}=Rf_{i}+y_{i}$ . If $V$ is a convex, centrally symmetric set, and if $v_{i}\in V$ for every $1\leq i\leq N$ then $\big{(}R/\sqrt{N}-\sqrt{N}\big{)}B_{2}^{N}\subset V$ .

Proof. A separation argument shows that if $\sup_{v\in V}|\bigl{<}v,w\bigr{>}|\geq\rho$ for every $w\in S^{N-1}$ , then $\rho B_{2}^{N}\subset V$ (indeed, otherwise there would be some $x\in\rho B_{2}^{N}\backslash V$ ; but it is impossible to separate $x$ and the convex and centrally symmetric $V$ using any norm-one functional).

To complete the proof, observe that for every $w\in S^{N-1}$ ,

Applying Lemma 3.3, it follows that if $R\geq 2N$ then with probability at least $3/4$ , $\sqrt{N}B_{2}^{N}\subset V$ . Finally, if $\delta\lesssim 1/N$ then

and the same assertion holds for the normalized matrix $\Gamma$ , showing that it does not satisfy ER( $1$ ).

Of course, this assertion holds under several conditions on the parameters involved: namely, that $R=\sqrt{p}(1/\delta)^{1/p}\geq 2N$ ; that $(\log N)/n\lesssim\delta\lesssim\log\big{(}en/N\big{)}/N$ ; that $R^{4}\delta\lesssim 1$ ; that $p\leq 2\log(1/\delta)$ and that $\delta\lesssim 1/N$ .

For instance, one may select $\delta\sim(\log N)/n$ and $p\sim(\log n)/\log N$ , in which case all these conditions are met; hence, with probability at least $1/2$ , $\Gamma$ does not satisfy ER( $1$ ), proving Theorem C. A similar calculation leads to the proof of Theorem C′.

Note that the construction leads to a stronger, non-uniform result, namely, that for every basis vector $e_{k}$ , with probability at least $1/2$ , $e_{k}$ is not the unique solution of $\min(\|t\|_{1}:\ \Gamma t=\Gamma e_{k})$ . In particular, uniformity over all supports of size $1$ in the definition of ER( $1$ ) is not the reason why the moment assumption in Theorem A is required.

Results in the noisy measurements setup

In previous sections, we considered the idealized scenario, in which the data was noiseless. Here, we will study the noisy setup: one observes $N$ couples $(z_{i},X_{i})_{i=1}^{N}$ , and each $z_{i}$ is a noisy observation of $\bigl{<}X_{i},x_{0}\bigr{>}$ :

The goal is to obtain as much information as possible on the unknown vector $x_{0}$ with only the data $(z_{i},X_{i})_{i=1}^{N}$ at one’s disposal, and for the sake of simplicity, we will assume that the $g_{i}$ ’s are independent Gaussian random variables ${\cal N}(0,\sigma^{2})$ that are also independent of the $X_{i}$ ’s.

Unlike the noiseless case, there is no hope of reconstructing $x_{0}$ from the given data, and instead of exact reconstruction, there are three natural questions that one may consider:

These three problems are central in modern Statistics, and are featured in numerous statistical monographs, particularly in the context of the Gaussian regression model (Equation (4.1)).

Recently, all three problems have been recast in a ‘high-dimensional’ scenario, in which the number of observations $N$ may be much smaller than the ambient dimension $n$ . Unfortunately, such problems are often impossible to solve without additional assumptions, and just as in the noiseless case, the situation improves dramatically if $x_{0}$ has some low-dimensional structure, for example, if it is $s$ -sparse. The aim is therefore to design a procedure that performs as if the true dimension of the problem is $s$ rather than $n$ , despite the noisy data.

Both procedures may be implemented effectively, and their estimation and de-noising properties have been obtained under some assumptions on the measurement matrix (see, e.g. or Chapters 7 and 8 in ).

In this section, we shall focus on two such conditions on the measurement matrix. The first, called the Compatibility Condition, was introduced in (see also Definition 2.1 in ); the second, the Restricted Eigenvalue Condition, was introduced in .

Let $\Gamma$ be an ${N\times n}$ matrix. For $L>0$ and a set $S\subset\{1,\ldots,n\}$ , the compatibility constant associated with $L$ and $S$ is

where $\zeta_{S}$ (resp. $\zeta_{S^{c}}$ ) denotes a vector that is supported in $S$ (resp. $S^{c}$ ).

$\Gamma$ satisfies the Compatibility Condition for the set $S_{0}$ with constants $L>1$ and $c_{0}$ if $\phi(L,S_{0})\geq c_{0}$ ; it satisfies the uniform Compatibility Condition (CC) of order $s$ if $\min_{|S|\leq s}\phi(L,S)\geq c_{0}$ .

A typical result for the LASSO in the Gaussian model (4.1) and when $\Gamma$ satisfies the Compatibility Condition, is Theorem 6.1 in :

Even though the Compatibility Condition in $S_{0}$ suffices to show that the LASSO is an effective procedure, the fact remains that $S_{0}$ is not known. And while a non-uniform approach is still possible (e.g., if $\Gamma$ is a random matrix, one may try showing that with high probability it satisfies the Compatibility Condition for the fixed, but unknown $S_{0}$ ), the uniform Compatibility Condition is a safer requirement – and the one we shall explore below.

Let $\Gamma$ be an ${N\times n}$ matrix. Given $c_{0}\geq 1$ and an integer $1\leq s\leq m\leq n$ for which $m+s\leq n$ , the restricted eigenvalue constant is

The matrix $\Gamma$ satisfies the Restricted Eigenvalue Condition (REC) of order $s$ with a constant $c$ if $\kappa(s,s,3)\geq c$ .

Estimation and de-noising results follow from Theorem 6.1 (for the Dantzig selector) and Theorem 6.2 (for the LASSO) in , when the measurement matrix $\Gamma$ , normalized by having the diagonal elements of $\Gamma^{\top}\Gamma$ equal $1$ , satisfies the REC of an appropriate order and with a constant that is independent of the dimension. We also refer to Lemma 6.10 in for similar results that do not require normalization.

Because the two lead to bounds on the performance of the LASSO and the Dantzig selector, a question that comes to mind is whether there are matrices that satisfy the CC or the REC. And, as in Compressed Sensing, the only matrices that are known to satisfy those conditions for the optimal number of measurements (rows) are well-behaved random matrices (see for some examples).

Our aim in this final section is to extend our results to the noisy setup, by identifying almost necessary and sufficient moment assumptions for the CC and the REC. This turns out to be straightforward: on one hand, the proof of Theorem A actually provides a stronger quantitative version of the exact reconstruction property; on the other, the uniform compatibility condition can be viewed as a quantitative version of a geometric condition on the polytope $\Gamma B_{1}^{n}$ that characterizes Exact Reconstruction. A similar observation is true for the REC: it can be viewed as a quantitative version of the null space property (see and below) which is also equivalent to the exact reconstruction property.

It is well known that $\Gamma$ satisfies ER( $s$ ) if and only if $\Gamma B_{1}^{n}$ has $2n$ vertices and $\Gamma B_{1}^{n}$ is a centrally symmetric $s$ -neighbourly polytope. It turns out that this property is characterized by the uniform CC.

Let $\Gamma$ be an ${N\times n}$ matrix. The following are equivalent:

$\Gamma B_{1}^{n}$ has $2n$ vertices and is $s$ -neighbourly,

$\min\big{(}\phi(1,S):S\subset\{1,\ldots,n\},|S|\leq s\big{)}>0$ .

In particular, $\min_{|S|\leq s}\phi(L,S)$ for some $L\geq 1$ is a quantitative measure of the $s$ -neighbourly property of $\Gamma B_{1}^{n}$ : if $\Gamma B_{1}^{n}$ is $s$ -neighbourly and has $2n$ vertices then the two sets

are disjoint for every $|S|\leq s$ . However, $\min_{|S|\leq s}\phi(1,S)$ measures how far the two sets are from one another, uniformly over all subsets $S\subset\{1,\ldots,n\}$ of cardinality at most $s$ . Proof. Let $C_{1},\ldots,C_{n}$ be the $n$ columns of $\Gamma$ . It follows from Proposition 2.2.13 and Proposition 2.2.16 in that $\Gamma B_{1}^{n}$ has $2n$ vertices and is a centrally symmetric $s$ -neighbourly polytope if and only if for every $S\subset\{1,\ldots,n\}$ of cardinality $|S|\leq s$ and every choice of signs $(\varepsilon_{i})\in\{-1,1\}^{S}$ ,

As a consequence, (4.5) holds for every $S\subset\{1,\ldots,n\}$ of cardinality at most $s$ if and only if $\min\big{(}\phi(1,S):S\subset\{1,\ldots,n\},|S|\leq s\big{)}>0$ .

An observation of a similar nature is true for the REC: it can be viewed as a quantitative measure of the null space property.

Let $\Gamma$ be an ${N\times n}$ matrix. $\Gamma$ satisfies the null space property of order $s$ if it is invertible in the cone

In , the authors prove that $\Gamma$ satisfies ER( $s$ ) if and only if it has the null space property of order $s$ .

A natural way of quantifying the invertibility of $\Gamma$ in the cone (4.6) is to consider its smallest singular value, restricted to this cone, which is simply the REC $\kappa(s,n-s,1)$ . Unfortunately, statistical properties of the LASSO and of the Dantzig selector are not known under the assumption that $\kappa(s,n-s,1)$ is an absolute constant (though if $\kappa(s,s,3)$ is an absolute constant, LASSO is known to be optimal ).

The main result of this section is the following:

Theorem E. Let $L>0$ , $1\leq s\leq n$ and $c_{0}>0$ . Under the same assumptions as in Theorem A and with the same probability estimate, $\Gamma=N^{-1/2}\sum_{i=1}^{N}\bigl{<}X_{i},\cdot\bigr{>}f_{i}$ satisfies:

A uniform compatibility condition of order $c_{1}s$ , namely that

A restricted eigenvalue condition of order $c_{2}s$ , with

for any $1\leq m\leq n$ , as long as $(1+c_{0})^{2}c_{2}\leq u^{2}\beta/(16e)$ .

On the other hand, if $\Gamma$ is the matrix considered in Theorem C, then with probability at least $1/2$ , $\phi(1,\{e_{1}\})=0$ and $\kappa(1,m,1)=0$ for any $1\leq m\leq n$ .

Just like Theorem A and Theorem C, Theorem E shows that the requirement that the coordinates of the measurement vector have $\log n$ moments is almost a necessary and sufficient condition for the uniform Compatibility Condition and the Restricted Eigenvalue Condition to hold. Moreover, it shows the significance of the small-ball condition, even in the noisy setup.

It also follows from Theorem E that if $X$ satisfies the small-ball condition and its coordinates have $\log n$ well-behaved moments as in Theorem A, then $\Gamma B_{1}^{n}$ has $2n$ vertices and is $s$ -neighbourly with high probability for $N\sim s\log(en/s)$ . In particular, this improves Theorem 4.3 in by a logarithmic factor for matrices generated by subexponential variables.

Consider $\gamma=(\zeta_{S}-\zeta_{S^{c}})/\left\|\zeta_{S}-\zeta_{S^{c}}\right\|_{2}$ . Since

it follows that $\gamma\in\big{(}(1+L)\sqrt{|S|}\big{)}B_{1}^{n}\cap S^{n-1}$ .

Recall that by (2.7), if $r=(1+L)^{2}c_{1}s\leq su^{2}\beta/(16e)$ , then $\left\|\Gamma\gamma\right\|_{2}\geq(u^{2}\beta)/4$ . Therefore,

and thus $\min_{|S|\leq c_{1}s}\phi(L,S)\geq u^{2}\beta/4$ for $c_{1}=u^{2}\beta/\big{(}16e(1+L)^{2}\big{)}$ .

Turning to the REC, fix a constant $c_{2}$ to be named later. Consider $x$ in the cone and let $S_{0}\subset\{1,\ldots,n\}$ of cardinality $|S_{0}|\leq c_{2}s$ for which $\left\|x_{S_{0}^{c}}\right\|_{1}\leq c_{0}\left\|x_{S_{0}}\right\|_{1}$ . Let $S_{1}\subset\{1,\ldots,n\}$ be the set of indices of the $m$ largest coordinates of $(|x_{i}|)_{i=1}^{n}$ that are outside $S_{0}$ and put $S_{01}=S_{0}\cup S_{1}$ .

Observe that $\left\|x\right\|_{1}\leq(1+c_{0})\left\|x_{S_{0}}\right\|_{1}\leq(1+c_{0})\sqrt{|S_{0}|}\left\|x\right\|_{2}$ ; hence $x/\left\|x\right\|_{2}\in\big{(}(1+c_{0})\sqrt{|S_{0}|}\big{)}B_{1}^{n}\cap S^{n-1}$ . Applying (2.7) again, if $(1+c_{0})^{2}c_{2}s\leq su^{2}\beta/(16e)$ , then $\left\|\Gamma x\right\|_{2}\geq\big{(}(u^{2}\beta)/4\big{)}\left\|x\right\|_{2}$ . Thus,

and $\kappa(c_{2}s,m,c_{0})\geq u^{2}\beta/4$ for any $1\leq m\leq n$ , as long as $(1+c_{0})^{2}c_{2}\leq u^{2}\beta/(16e)$ .

The proof of the second part of Theorem E is an immediate corollary of the construction used in Theorem C. Recall that with probability at least $1/2$ , $\Gamma e_{1}\in{\rm absconv}(\Gamma e_{j}:j\in\{2,...,n\})$ . Setting $J=\{e_{2},...,e_{n}\}$ , there is $\zeta\in B_{1}^{J}$ for which $\left\|\Gamma e_{1}-\Gamma\zeta\right\|_{2}=0$ . Therefore, $\phi(1,\{e_{1}\})=0$ and $\kappa(1,m,1)=0$ for any $1\leq m\leq n$ , as claimed.

The results obtained in Theorem A and in parts (1) and (2) of Theorem E are also valid for the normalized (columns wise) measurement matrix:

The proof is almost identical to the one used for $\Gamma$ itself, even though $\Gamma_{1}$ does not have independent rows vectors, due to the normalization. For the sake of brevity, we will not present the straightforward proof of this observation.

Finally, the counterexample constructed in the proof of Theorem C and in which a typical $\Gamma$ does not satisfy ER( $1$ ), does not necessarily generate $\Gamma B_{1}^{n}$ that is not $s$ -neighbourly. Indeed, an inspection of the construction shows that the reason ER( $1$ ) fails is that $\Gamma B_{1}^{n}$ has less than $2n-2$ vertices, rather than that $\Gamma B_{1}^{n}$ is not $s$ -neighbourly. Thus, the question of whether a moment condition is necessary for the random polytope $\Gamma B_{1}^{n}$ to be $s$ -neighbourly with probability at least $1/2$ is still unresolved.

Introduction and main results

Proof of Theorem A, B and D

Proof of Theorem C and Theorem C′

Results in the noisy measurements setup

References