More data speeds up training time in learning halfspaces over sparse vectors

Amit Daniely, Nati Linial, Shai Shalev Shwartz

Introduction

In the modern digital period, we are facing a rapid growth of available datasets in science and technology. In most computing tasks (e.g. storing and searching in such datasets), large datasets are a burden and require more computation. However, for learning tasks the situation is radically different. A simple observation is that more data can never hinder you from performing a task. If you have more data than you need, just ignore it!

A basic question is how to learn from “big data”. The statistical learning literature classically studies questions like “how much data is needed to perform a learning task?” or “how does accuracy improve as the amount of data grows?” etc. In the modern, “data revolution era”, it is often the case that the amount of data available far exceeds the information theoretic requirements. We can wonder whether this, seemingly redundant data, can be used for other purposes. An intriguing question in this vein, studied recently by several researchers ((Decatur et al., 1998; Servedio., 2000; Shalev-Shwartz et al., 2012; Berthet and Rigollet, 2013; Chandrasekaran and Jordan, 2013)), is the following

Question 1: Are there any learning tasks in which more data, beyond the information theoretic barrier, can provably be leveraged to speed up computation time?

To prove this, we present a novel technique to establish computational-statistical tradeoffs in supervised learning problems. To the best of our knowledge, this is the first such a result that is not based on cryptographic primitives.

The natural learning problem we consider is the task of learning the class of halfspaces over $k$ -sparse vectors. Here, the instance space is the space of $k$ -sparse vectors,

and the hypothesis class is halfspaces over $k$ -sparse vectors, namely

In addition, we allow improper learning (a.k.a. representation independent learning), namely, the learning algorithm is not restricted to output a hypothesis from ${\mathcal{H}}_{n,k}$ , but only should output a hypothesis whose error is not much larger than the error of the best hypothesis in ${\mathcal{H}}_{n,k}$ . This gives the learner a lot of flexibility in choosing an appropriate representation of the problem. This additional freedom to the learner makes it much harder to prove lower bounds in this model. Concretely, it is not clear how to use standard reductions from NP hard problems in order to establish lower bounds for improper learning (moreover, Applebaum et al. (2008) give evidence that such simple reductions do not exist).

The classes ${\mathcal{H}}_{n,k}$ and similar classes have been studied by several authors (e.g. Long. and Servedio (2013)). They naturally arise in learning scenarios in which the set of all possible features is very large, but each example has only a small number of active features. For example:

Predicting an advertisement based on a search query: Here, the possible features of each instance are all English words, whereas the active features are only the set of words given in the query.

Learning Preferences (Hazan et al., 2012): Here, we have $n$ players. A ranking of the players is a permutation $\sigma:[n]\to[n]$ (think of $\sigma(i)$ as the rank of the $i$ ’th player). Each ranking induces a preference $h_{\sigma}$ over the ordered pairs, such that $h_{\sigma}(i,j)=1$ iff $i$ is ranked higher that $j$ . Namely,

We will show a positive answer to Question 1 for the class ${\mathcal{H}}_{n,3}$ . To do so, we showIn fact, similar results hold for every constant $k\geq 3$ . Indeed, since ${\mathcal{H}}_{n,3}\subset{\mathcal{H}}_{n,k}$ for every $k\geq 3$ , it is trivial that item $3$ below holds for every $k\geq 3$ . The upper bound given in item $1$ holds for every $k$ . For item 2, it is not hard to show that ${\mathcal{H}}_{n,k}$ can be learnt using a sample of $\Omega\left(\frac{n^{k}}{\epsilon^{2}}\right)$ examples by a naive improper learning algorithm, similar to the algorithm we describe in this section for $k=3$ . the following:

Ignoring computational issues, it is possible to learn the class ${\mathcal{H}}_{n,3}$ using $O\left(\frac{n}{\epsilon^{2}}\right)$ examples.

A graphical illustration of our main results is given below:

runtime $2^{O(n)}<math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>></mo><mi mathvariant="normal">poly</mi><mo>⁡</mo><mo stretchy="false">(</mo><mi>n</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">>\operatorname{poly}(n)</annotation></semantics></math>>poly(n)n^{O(1)}$ examples $n^{2}<math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msup><mi>n</mi><mn>1.5</mn></msup></mrow><annotation encoding="application/x-tex">n^{1.5}</annotation></semantics></math>n1.5n$ The proof of item 1 above is easy – simply note that $H_{n,3}$ has VC dimension $n+1$ .

Item 2 is proved in section 4, relying on the results of Hazan et al. (2012). We note, however, that a weaker result, that still suffices for answering Question 1 in the affirmative, can be proven using a naive improper learning algorithm. In particular, we show below how to learn ${\mathcal{H}}_{n,3}$ efficiently with a sample of $\Omega\left(\frac{n^{3}}{\epsilon^{2}}\right)$ examples. The idea is to replace the class ${\mathcal{H}}_{n,3}$ with the class $\{\pm 1\}^{C_{n,3}}$ containing all functions from $C_{n,3}$ to $\{\pm 1\}$ . Clearly, this class contains $H_{n,3}$ . In addition, we can efficiently find a function $f$ that minimizes the empirical training error over a training set $S$ as follows: For every $x\in C_{n,k}$ , if $x$ does not appear at all in the training set we will set $f(x)$ arbitrarily to $1$ . Otherwise, we will set $f(x)$ to be the majority of the labels in the training set that correspond to $x$ . Finally, note that the VC dimension of $\{\pm 1\}^{C_{n,3}}$ is smaller than $n^{3}$ (since $|C_{n,3}|<n^{3}$ ). Hence, standard generalization results (e.g. Vapnik (1995)) implies that a training set size of $\Omega\left(\frac{n^{3}}{\epsilon^{2}}\right)$ suffices for learning this class.

Item 3 is shown in section 3 by presenting a novel technique for establishing statistical-computational tradeoffs.

The class ${\mathcal{H}}_{n,2}$ . Our main result gives a positive answer to Question 1 for the task of improperly learning ${\mathcal{H}}_{n,k}$ for $k\geq 3$ . A natural question is what happens for $k=2$ and $k=1$ . Since $\operatorname*{VC}({\mathcal{H}}_{n,1})=\operatorname*{VC}({\mathcal{H}}_{n,2})=n+1$ , the information theoretic barrier for learning these classes is $\Theta\left(\frac{n}{\epsilon^{2}}\right)$ . In section 4, we prove that ${\mathcal{H}}_{n,2}$ (and, consequently, ${\mathcal{H}}_{n,1}\subset{\mathcal{H}}_{n,2}$ ) can be learnt using $O\left(\frac{n\log^{3}(n)}{\epsilon^{2}}\right)$ examples, indicating that significant computational-statistical tradeoffs start to manifest themselves only for $k\geq 3$ .

Recently, (Berthet and Rigollet, 2013) gave a positive answer to Question 1 in the context of unsupervised learning. Concretely, they studied the problem of sparse PCA, namely, finding a sparse vector that maximizes the variance of an unsupervised data. Conditioning on the hardness of the planted clique problem, they gave a positive answer to Question 1 for sparse PCA. Our work, as well as the previous work of Decatur et al. (1998); Servedio. (2000); Shalev-Shwartz et al. (2012), studies Question 1 in the supervised learning setup. We emphasize that unsupervised learning problems are radically different than supervised learning problems in the context of deriving lower bounds. The main reason for the difference is that in supervised learning problems, the learner is allowed to employ improper learning, which gives it a lot of power in choosing an adequate representation of the data. For example, the upper bound we have derived for the class of sparse halfspaces switched from representing hypotheses as halfspaces to representation of hypotheses as tables over $C_{n,3}$ , which made the learning problem easy from the computational perspective. The crux of the difficulty in constructing lower bounds is due to this freedom of the learner in choosing a convenient representation. This difficulty does not arise in the problem of sparse PCA detection, since there the learner must output a good sparse vector. Therefore, it is not clear whether the approach given in (Berthet and Rigollet, 2013) can be used to establish computational-statistical gaps in supervised learning problems.

Background and notation

For $h:C_{n,3}\to\{\pm 1\}$ and a distribution ${\mathcal{D}}$ on $C_{n,3}\times\{\pm 1\}$ we denote the error of $h$ w.r.t. ${\mathcal{D}}$ by $\operatorname{Err}_{\mathcal{D}}(h)=\Pr_{(x,y)\sim{\mathcal{D}}}\left(h(x)\neq y\right)$ . For ${\mathcal{H}}\subset\{\pm 1\}^{C_{n,3}}$ we denote the error of ${\mathcal{H}}$ w.r.t. ${\mathcal{D}}$ by $\operatorname{Err}_{{\mathcal{D}}}({\mathcal{H}})=\min_{h\in{\mathcal{H}}}\operatorname{Err}_{{\mathcal{D}}}(h)$ . For a sample $S\in\left(C_{n,3}\times\{\pm 1\}\right)^{m}$ we denote by $\operatorname{Err}_{S}(h)$ (resp. $\operatorname{Err}_{S}({\mathcal{H}})$ ) the error of $h$ (resp. ${\mathcal{H}}$ ) w.r.t. the empirical distribution induces by the sample $S$ .

A learning algorithm, $L$ , receives a sample $S\in\left(C_{n,3}\times\{\pm 1\}\right)^{m}$ and return a hypothesis $L(S):C_{n,3}\to\{\pm 1\}$ . We say that $L$ learns ${\mathcal{H}}_{n,3}$ using $m(n,\epsilon)$ examples if,For simplicity, we require the algorithm to succeed with probability of at least $9/10$ . This can be easily amplified to probability of at least $1-\delta$ , as in the usual definition of agnostic PAC learning, while increasing the sample complexity by a factor of $\log(1/\delta)$ . for every distribution ${\mathcal{D}}$ on $C_{n,3}\times\{\pm 1\}$ and a sample $S$ of more than $m(n,\epsilon)$ i.i.d. examples drawn from ${\mathcal{D}}$ ,

The algorithm $L$ is efficient if it runs in polynomial time in the sample size and returns a hypothesis that can be evaluated in polynomial time.

Soundness: If $\operatorname{Val}(\phi)\geq 1-\epsilon$ , then

In fact, for all we know, the following conjecture may be true for every $0\leq\mu\leq 0.5$ .

Note that Feige’s conjecture is equivalent to the -R3SAT hardness assumption.

Let $0\leq\mu\leq 0.5$ . If the $\mu$ -R3SAT hardness assumption (conjecture 2.2) is true, then there exists no efficient learning algorithm that learns the class ${\mathcal{H}}_{n,3}$ using $O\left(\frac{n^{1+\mu}}{\epsilon^{2}}\right)$ examples.

In the proof of Theorem 3.1 we rely on the validity of a conjecture, similar to conjecture 2.2 for $3$ -variables majority formulas. Following an argument from (Feige, 2002) (Theorem 3.2) the validity of the conjecture on which we rely for majority formulas follows the validity of conjecture 2.2.

An $n$ -variables 3MAJ clause is a boolean formula of the form

An $n$ -variables 3MAJ formula is a boolean formula of the form

where the $C_{i}$ ’s are 3MAJ clauses. By $\textrm{3MAJ}_{n,m}$ we denote the set of 3MAJ formulas with $n$ variables and $m$ clauses.

Let $0\leq\mu\leq 0.5$ . If the $\mu$ -R3SAT hardness assumption is true, then for every $\epsilon>0$ and for every large enough integer $\Delta>\Delta_{0}(\epsilon)$ there exists no efficient algorithm with the following properties.

If $\operatorname{Val}(\phi)\geq\frac{3}{4}-\epsilon$ , then

Next, we prove Theorem 3.1. In fact, we will prove a slightly stronger result. Namely, define the subclass ${\mathcal{H}}_{n,3}^{d}\subset{\mathcal{H}}_{n,3}$ , of homogenous halfspaces with binary weights, given by ${\mathcal{H}}_{n,3}^{d}=\left\{h_{w,0}\mid w\in\{\pm 1\}^{n}\right\}$ . As we show, under the $\mu$ -R3SAT hardness assumption, it is impossible to efficiently learn this subclass using only $O\left(\frac{n^{1+\mu}}{\epsilon^{2}}\right)$ examples.

The crux of the construction is that if $\phi$ is random, no algorithm (even improper and even inefficient) can return a hypothesis with a small error. The reason for that is that since the sample provided to the algorithm consists of only $\kappa\frac{n}{0.01^{2}}$ samples, the algorithm won’t see most of $\psi$ ’s clauses, and, consequently, the produced hypothesis $h$ will be independent of them. Since these clauses are random, $h$ is likely to err on about half of them, so that $\operatorname{Err}_{D_{\phi}}(h)$ will be close to half!

To summarize we constructed an efficient algorithm with the following properties: if $\phi$ is almost satisfiable, the algorithm will return a hypothesis with a small error, and then we will declare “exceptional”, while for random $\phi$ , the algorithm will return a hypothesis with a large error, and we will declare “typical”.

Our construction crucially relies on the restriction to learning algorithm with a small sample complexity. Indeed, if the learning algorithm obtains more than $n^{1+\mu}$ examples, then it will see most of $\psi$ ’s clauses, and therefore it might succeed in “learning” even when the source of the formula is random. Therefore, we will declare “exceptional” even when the source is random.

(of theorem 3.1) Assume by way of contradiction that the $\mu$ -R3SAT hardness assumption is true and yet there exists an efficient learning algorithm that learns the class ${\mathcal{H}}_{n,3}$ using $O\left(\frac{n^{1+\mu}}{\epsilon^{2}}\right)$ examples. Setting $\epsilon=\frac{1}{100}$ , we conclude that there exists an efficient algorithm $L$ and a constant $\kappa>0$ such that given a sample $S$ of more than $\kappa\cdot n^{1+\mu}$ examples drawn from a distribution ${\mathcal{D}}$ on $C_{n,3}\times\{\pm 1\}$ , returns a classifier $L(S):C_{n,3}\to\{\pm 1\}$ such that

W.p. $\geq\frac{3}{4}$ over the choice of $S$ , $\operatorname{Err}_{{\mathcal{D}}}(L(S))\leq\operatorname{Err}_{{\mathcal{D}}}({\mathcal{H}}_{n,3})+\frac{1}{100}$ .

Fix $\Delta$ large enough such that $\Delta>100\kappa$ and the conclusion of Theorem 3.2 holds with $\epsilon=\frac{1}{100}$ . We will construct an algorithm, $A$ , contradicting Theorem 3.2. On input $\phi\in\textrm{3MAJ}_{n,\Delta n^{1+\mu}}$ consisting of the 3MAJ clauses $C_{1},\ldots,C_{\Delta n^{1+\mu}}$ , the algorithm $A$ proceeds as follows

Generate a sample $S$ consisting of $\Delta n^{1+\mu}$ examples as follows. For every clause, $C_{k}=\textrm{MAJ}((-1)^{j_{1}}x_{i_{1}},(-1)^{j_{2}}x_{i_{2}},(-1)^{j_{3}}x_{i_{3}})$ , generate an example $(x_{k},y_{k})\in C_{n,3}\times\{\pm 1\}$ by choosing $b\in\{\pm 1\}$ at random and letting

For example, if $n=6$ , the clause is $\textrm{MAJ}(-x_{2},x_{3},x_{6})$ and $b=-1$ , we generate the example

Choose a sample $S_{1}$ consisting of $\frac{\Delta n^{1+\mu}}{100}\geq\kappa\cdot n^{1+\mu}$ examples by choosing at random (with repetitions) examples from $S$ .

We claim that $A$ contradicts Theorem 3.2. Clearly, $A$ runs in polynomial time. It remains to show that

If $\operatorname{Val}(\phi)\geq\frac{3}{4}-\frac{1}{100}$ , then

Assume first that $\phi\in\textrm{3MAJ}_{n,\Delta n^{1+\mu}}$ is chosen at random. Given the sample $S_{1}$ , the sample $S_{2}:=S\setminus S_{1}$ is a sample of $|S_{2}|$ i.i.d. examples which are independent from the sample $S_{1}$ , and hence also from $h=L(S_{1})$ . Moreover, for every example $(x_{k},y_{k})\in S_{2}$ , $y_{k}$ is a Bernoulli random variable with parameter $\frac{1}{2}$ which is independent of $x_{k}$ . To see that, note that an example whose instance is $x_{k}$ can be generated by exactly two clauses – one corresponds to $y_{k}=1$ , while the other corresponds to $y_{k}=-1$ (e.g., the instance $(1,-1,0,1)$ can be generated from the clause $\textrm{MAJ}(x_{1},-x_{2},x_{4})$ and $b=1$ or the clause $\textrm{MAJ}(-x_{1},x_{2},-x_{4})$ and $b=-1$ ). Thus, given the instance $x_{k}$ , the probability that $y_{k}=1$ is $\frac{1}{2}$ , independent of $x_{k}$ .

It follows that $\operatorname{Err}_{S_{2}}(h)$ is an average of at least $\left(1-\frac{1}{100}\right)\Delta n^{1+\mu}$ independent Bernoulli random variable. By Chernoff’s bound, with probability $\geq 1-o(1)$ , $\operatorname{Err}_{S_{2}}(h)>\frac{1}{2}-\frac{1}{100}$ . Thus,

Assume now that $\operatorname{Val}(\phi)\geq\frac{3}{4}-\frac{1}{100}$ and let $\psi\in\{\pm 1\}^{n}$ be an assignment that indicates that. Let $\Psi\in{\mathcal{H}}_{n,3}$ be the hypothesis $\Psi(x)=\textrm{sign}\left(\langle\psi,x\rangle\right)$ . It can be easily checked that $\Psi(x_{k})=y_{k}$ if and only if $\psi$ satisfies $C_{k}$ . Since $\operatorname{Val}(\phi)\geq\frac{3}{4}-\frac{1}{100}$ , it follows that

By the choice of $L$ , with probability $\geq 1-\frac{1}{4}=\frac{3}{4}$ ,

The following theorem derives upper bounds for learning ${\mathcal{H}}_{n,2}$ and ${\mathcal{H}}_{n,3}$ . Its proof relies on results from Hazan et al. (2012) about learning $\beta$ -decomposable matrices, and due to the lack of space is given in the appendix.

There exists an efficient algorithm that learns ${\mathcal{H}}_{n,2}$ using $O\left(\frac{n\log^{3}(n)}{\epsilon^{2}}\right)$ examples

There exists an efficient algorithm that learns ${\mathcal{H}}_{n,3}$ using $O\left(\frac{n^{2}\log^{3}(n)}{\epsilon^{2}}\right)$ examples

Discussion

We formally established a computational-sample complexity tradeoff for the task of (agnostically and improperly) PAC learning of halfspaces over $3$ -sparse vectors. Our proof of the lower bound relies on a novel, non cryptographic, technique for establishing such tradeoffs. We also derive a new non-trivial upper bound for this task.

Amit Daniely is a recipient of the Google Europe Fellowship in Learning Theory, and this research is supported in part by this Google Fellowship. Nati Linial is supported by grants from ISF, BSF and I-Core. Shai Shalev-Shwartz is supported by the Israeli Science Foundation grant number 590-10.

References

Appendix A Proof of Theorem 4.1

The proof of the theorem relies on results from Hazan et al. about learning $\beta$ -decomposable matrices. Let $W$ be an $n\times m$ matrix. We define the symmetrization of $W$ to be the $(n+m)\times(n+m)$ matrix

We say that $W$ is $\beta$ -decomposable if there exist positive semi-definite matrices $P,N$ for which

Each matrix in $\{\pm 1\}^{n\times m}$ can be naturally interpreted as a hypothesis on $[n]\times[m]$ .

We say that a learning algorithm $L$ learns a class ${\mathcal{H}}_{n}\subset\{\pm 1\}^{X_{n}}$ using $m(n,\epsilon,\delta)$ examples if, for every distribution ${\mathcal{D}}$ on $X_{n}\times\{\pm 1\}$ and a sample $S$ of more than $m(n,\epsilon,\delta)$ i.i.d. examples drawn from ${\mathcal{D}}$ ,

Hazan et al. have provedThe result of Hazan et al. is more general than what is stated here. Also, Hazan et al. considered the online scenario. The result for the statistical scenario, as stated here, can be derived by applying standard online-to-batch conversions (see for example Cesa-Bianchi et al. ). that

Hazan et al. The hypothesis class of $\beta$ -decomposable $n\times m$ matrices with $\pm 1$ entries ban be efficiently learnt using a sample of $O\left(\frac{\beta^{2}(n+m)\log(n+m)+\log(1/\delta)}{\epsilon^{2}}\right)$ examples.

We start with a generic reduction from a problem of learning a class ${\mathcal{G}}_{n}$ over an instance space $X_{n}\subset\{-1,1,0\}^{n}$ to the problem of learning $\beta(n)$ -decomposable matrices. We say that ${\mathcal{G}}_{n}$ is realized by $m_{n}\times m_{n}$ matrices that are $\beta(n)$ -decomposable if there exists a mapping $\psi_{n}:X_{n}\to[m_{n}]\times[m_{n}]$ such that for every $h\in{\mathcal{G}}_{n}$ there exists a $\beta(n)$ -decomposable $m_{n}\times m_{n}$ matrix $W$ for which $\forall x\in X_{n},\;h(x)=W_{\psi_{n}(x)}$ . The mapping $\psi_{n}$ is called a realization of ${\mathcal{G}}_{n}$ . In the case that the mapping $\psi_{n}$ can be computed in time polynomial in $n$ , we say that ${\mathcal{G}}_{n}$ is efficiently realized and $\psi_{n}$ is an efficient realization. It follows from Theorem A.1 that:

If ${\mathcal{G}}_{n}$ is efficiently realized by $m_{n}\times m_{n}$ matrices that are $\beta(n)$ -decomposable then ${\mathcal{G}}_{n}$ can be efficiently learnt using a sample of $O\left(\frac{\beta(n)^{2}m_{n}\log(m_{n})+\log(1/\delta)}{\epsilon^{2}}\right)$ examples.

We now turn to the proof of Theorem 4.1. We start with the first assertion, about learning ${\mathcal{H}}_{n,2}$ . The idea will be to partition the instance space into a disjoint union of subsets and show that the restriction of the hypothesis class to each subset can be efficiently realized by $\beta(n)$ -decomposable. Concretely, we decompose $C_{n,2}$ into a disjoint union of five sets

For every $-2\leq r\leq 2$ , ${\mathcal{H}}_{n,2}|_{A^{r}_{n}}$ can be efficiently realized by $n\times n$ matrices that are $O(\log(n))$ -decomposable.

To glue together the five restrictions, we will rely on the following Lemma, whose proof is given in section A.1.

Let $X_{1},...,X_{k}$ be partition of a domain $X$ and let $H$ be a hypothesis class over $X$ . Define $H_{i}=H|_{X_{i}}$ . Suppose the for every $H_{i}$ there exist a learning algorithm that learns $H_{i}$ using $\leq C(d+\log(1/\delta))/\epsilon^{2}$ examples, for some constant $C\geq 8$ . Consider the algorithm $A$ which receives an i.i.d. training set $S$ of $m$ examples from $X\times\{0,1\}$ and applies the learning algorithm for each $H_{i}$ on the examples in $S$ that belongs to $X_{i}$ . Then, $A$ learns $H$ using at most

The first part of Theorem 4.1 is therefore follows from Lemma A.3, Lemma A.4 and Corollary A.2.

Having the first part of Theorem 4.1 and Lemma A.4 at hand, it is not hard to prove the second part of Theorem 4.1:

For $1\leq i\leq n-2$ and $b\in\{\pm 1\}$ define

Let $\psi_{n}:C_{n,3}\to C_{n,2}$ be the mapping that zeros the first non zero coordinate. It is not hard to see that ${\mathcal{H}}_{n,3}|_{D_{n,i,b}}=\left\{h\circ\psi_{n}|_{D_{n,i,b}}\mid h\in{\mathcal{H}}_{n,2}\right\}$ . Therefore ${\mathcal{H}}_{n,3}|_{D_{n,i,b}}$ can be identified with ${\mathcal{H}}_{n,2}$ using the mapping $\psi_{n}$ , and therefore can efficiently learnt using $O\left(\frac{n\log^{3}(n)+\log(1/\delta)}{\epsilon^{2}}\right)$ examples (the dependency on $\delta$ does not appear in the statement, but can be easily inferred from the proof). The second part of Theorem 4.1 is therefore follows from the first part of the Theorem and Lemma A.4.

In the proof, we will rely on the following facts. The tensor product of two matrices $A\in M_{n\times m}$ and $B\in M_{k\times l}$ is defined as the $(n\cdot k)\times(m\cdot l)$ matrix

Let $W$ be a $\beta$ -decomposable matrix and let $A$ be a PSD matrix whose diagonal entries are upper bounded by $\alpha$ . Then $W\otimes A$ is $(\alpha\cdot\beta)$ -decomposable.

It is not hard to see that for every matrix $W$ and a symmetric matrix $A$ ,

Moreover, since the tensor product of two PSD matrices is PSD, if $\operatorname{sym}(W)=P-N$ is a $\beta$ -decomposition of $W$ , then

is a $(\alpha\cdot\beta)$ -decomposition of $W\otimes A$ . ∎

If $W$ is a $\beta$ -decomposable matrix, then so is every matrix obtained from $W$ by iteratively deleting rows and columns.

It is enough to show that deleting one row or column leaves $W$ $\beta$ -decomposable. Suppose that $W^{\prime}$ is obtained from $W\in M_{n\times m}$ by deleting the $i$ ’th row (the proof for deleting columns is similar). It is not hard to see that $\operatorname{sym}(W^{\prime})$ is the $i$ ’th principal minor of $\operatorname{sym}(W)$ . Therefore, since principal minors of PSD matrices are PSD matrices as well, if $\operatorname{sym}(W)=P-N$ is $\beta$ -decomposition of $W$ then $\operatorname{sym}(W^{\prime})=[P]_{i,i}-[N]_{i,i}$ is a $\beta$ -decomposition of $W^{\prime}$ . ∎

Hazan et al. Let $T_{n}$ be the upper triangular matrix whose all entries in the diagonal and above are $1$ , and whose all entries beneath the diagonal are $-1$ . Then $T_{n}$ is $O(\log(n))$ -decomposable.

Lastly, we will also need the following generalization of proposition A.7

Let $W$ be an $n\times n$ $\pm 1$ matrix. Assume that there exists a sequence $0\leq j(1),\ldots,j(n)\leq n$ such that

Since switching rows of a $\beta$ -decomposable matrix leaves a $\beta$ -decomposable matrix, we can assume without loss of generality that $j(1)\leq j(2)\leq\ldots\leq j(n)$ . Let $J$ be the $n\times n$ all ones matrix. It is not hard to see that $W$ can be obtained from $T_{n}\otimes J$ by iteratively deleting rows and columns. Combining propositions A.5, A.6 and A.7, we conclude that $W$ is $O(\log(n))$ -decomposable, as required. ∎

(of Lemma A.3) Denote ${\mathcal{A}}^{r}_{n}={\mathcal{H}}_{n,2}|_{A^{r}_{n}}$ . We split into cases.

Case 1, r=0: Note that $A_{n}^{0}=\{e_{i}-e_{j}\mid i,j\in[n]\}$ . Define $\psi_{n}:A_{n}^{0}\to[n]\times[n]$ by $\psi_{n}(e_{i}-e_{j})=(i,j)$ . We claim that $\psi_{n}$ is an efficient realization of ${\mathcal{A}}_{n}^{0}$ by $n\times n$ matrices that are $O(\log(n))$ decomposable. Indeed, let $h=h_{w,b}\in{\mathcal{A}}_{n}^{0}$ , and let $W$ be the $n\times n$ matrix $W_{ij}=W_{\psi_{n}(e_{i}-e_{j})}=h(e_{i}-e_{j})$ . It is enough to show that $W$ is $O(\log(n))$ -decomposable.

From equation (1), it is not hard to see that there exist numbers

The conclusion follows from Proposition A.8

Case 2, r=2 and r=-2: We confine ourselves to the case $r=2$ . The case $r=-2$ is similar. Note that $A_{n}^{2}=\{e_{i}+e_{j}\mid i\neq j\in[n]\}$ . Define $\psi_{n}:A_{n}^{2}\to[n]\times[n]$ by $\psi_{n}(e_{i}+e_{j})=(i,j)$ . We claim that $\psi_{n}$ is an efficient realization of ${\mathcal{A}}_{n}^{2}$ by $n\times n$ matrices that are $O(\log(n))$ decomposable. Indeed, let $h=h_{w,b}\in{\mathcal{A}}_{n}^{2}$ , and let $W$ be the $n\times n$ matrix $W_{ij}=W_{\psi_{n}(e_{i}+e_{j})}=h(e_{i}+e_{j})$ . It is enough to show that $W$ is $O(\log(n))$ -decomposable.

From equation (2), it is not hard to see that there exist numbers

The conclusion follows from Proposition A.8

Case 3, r=1 and r=-1: We confine ourselves to the case $r=1$ . The case $r=-1$ is similar. Note that $A_{n}^{1}=\{e_{i}\mid i\in[n]\}$ . Define $\psi_{n}:A_{n}^{0}\to[n]\times[n]$ by $\psi_{n}(e_{i})=(i,i)$ . We claim that $\psi_{n}$ is an efficient realization of ${\mathcal{A}}_{n}^{1}$ by $n\times n$ matrices that are $3$ -decomposable (let alone, $\log(n)$ -decomposable). Indeed, let $h=h_{w,b}\in{\mathcal{A}}_{n}^{1}$ , and let $W$ be the $n\times n$ matrix with $W_{ii}=W_{\psi_{n}(e_{i})}=h(e_{i})$ and $-1$ outside the diagonal. It is enough to show that $W$ is $3$ -decomposable. Since $J$ is $1$ -decomposable, it is enough to show that $W+J$ is $2$ -decomposable. However, it is not hard to see that every diagonal matrix $D$ is $(\max_{i}|D_{ii}|)$ -decomposable. ∎

(of Lemma A.4) Let $S=(x_{1},y_{1}),\ldots,(x_{m},y_{m})$ be a training set and let $\hat{m}_{i}$ be the number of examples in $S$ that belong to $X_{i}$ . Given that the values of the random variables $\hat{m}_{1},\ldots,\hat{m}_{i}$ is determined, we have that w.p. of at least $1-\delta$ ,

where $D_{i}$ is the induced distribution over $X_{i}$ , $h_{i}$ is the output of the $i$ ’th algorithm, and $h^{*}$ is the optimal hypothesis w.r.t. the original distribution $D$ . Define,

It follows from the above that we also have, w.p. at least $1-\delta$ , for every $i$ ,

Let $\alpha_{i}=D\{(x,y):x\in{\mathcal{X}}_{i}\}$ , and note that $\sum_{i}\alpha_{i}=1$ . Therefore,

Next note that if $\alpha_{i}m<C(d+\log(k/\delta))$ then $\alpha_{i}m/m_{i}\leq 1$ . Otherwise, using Chernoff’s inequality, for every $i$ we have

It follows that with probability of at least $1-\delta$ ,

All in all, we have shown that with probability of at least $1-2\delta$ it holds that

Therefore, the the algorithm learns ${\mathcal{H}}$ using