Representation Benefits of Deep Feedforward Networks

Matus Telgarsky

Overview

Let positive integer $k$ , number of layers $l$ , and number of nodes per layer $m$ be given with $m\leq 2^{(k-3)/l-1}$ . Then there exists a collection of $n:=2^{k}$ points $((x_{i},y_{i}))_{i=1}^{n}$ with $x_{i}\in$ and $y\in\{0,1\}$ such that

For example, approaching the error of the $2k$ -layer network (which has $\mathcal{O}(k)$ nodes and weights) with $2$ layers requires at least $2^{(k-3)/2-1}$ nodes, and with $\sqrt{k-3}$ layers needs at least $2^{\sqrt{k-3}-1}$ nodes.

The purpose of this note is to provide an elementary proof of Theorem 1.1 and its refinement Theorem 1.2, which amongst other improvements will use a recurrent neural network in the upper bound. Section 2 will present the proof, and Section 3 will tie these results to the literature on neural network expressive power and circuit complexity, which by contrast makes use of product nodes rather than standard feedforward networks when showing the benefits of depth.

There are three refinements to make: the classification problem will be specified, the perfect network will be an even simpler recurrent network, and $\sigma$ need not be $\sigma_{\textsc{r}}$ .

Let $n$ -ap (the $n$ -alternating-point problem) denote the set of $n$ uniformly spaced points within $[0,1-2^{-n}]$ with alternating labels, as depicted in Figure 1; that is, the points $((x_{i},y_{i}))_{i=1}^{n}$ with $x_{i}=i2^{-n}$ , and $y_{i}=0$ when $i$ is even, and otherwise $y_{i}=1$ . As the $x$ values pass from left to right, the labels change as often as possible; the key is that adding a constant number of nodes in a flat network only corrects predictions on a constant number of points, whereas adding a constant number of nodes in a deep network can correct predictions on a constant fraction of the points.

Let $\mathfrak{R}(\sigma;m,l;k)$ denote $k$ iterations of a recurrent network with $l$ layers of at most $m$ nodes each, defined as follows. Every $f\in\mathfrak{R}(\sigma;m,l;k)$ consists of some fixed network $g\in\mathfrak{N}(\sigma;m,l)$ applied $k$ times:

Consequently, $\mathfrak{R}(\sigma;m,l;k)\subseteq\mathfrak{N}(\sigma;m,lk)$ , but the former has $\mathcal{O}(ml)$ parameters whereas the latter has $\mathcal{O}(mlk)$ parameters.

This more refined result can thus say, for example, that on the $2^{k}$ -ap one needs exponentially (in $k$ ) many parameters when boosting decision stumps, linearly many parameters with a deep network, and constantly many parameters with a recurrent network.

Analysis

This section will first prove the lower bound via a counting argument, simply tracking the number of times a function within $\mathfrak{N}(\sigma;m,l)$ can cross 1/2. The upper bound will exhibit a network in $\mathfrak{N}(\sigma_{\textsc{r}};2,2)$ which can be composed with itself $k$ times to exactly fit the $n$ -ap. These bounds together prove Theorem 1.2, which in turn implies Theorem 1.1.

The lower bound is proved in two stages. First, composing and summing sawtooth functions must also yield a sawtooth function, thus elements of $\mathfrak{N}(\sigma;m,l)$ are sawtooth whenever $\sigma$ is. Secondly, a sawtooth function can not cross $1/2$ very often, meaning it can’t hope to match the quickly changing labels of the $n$ -ap.

To start, $\mathfrak{N}(\sigma;m,l)$ is sawtooth as follows.

The proof is straightforward and deferred momentarily. The key observation is that adding together sawtooth functions grows the number of regions very slowly, whereas composition grows the number very quickly, an early sign of the benefits of depth.

Given a sawtooth function, its classification error on the $n$ -ap may be lower bounded as follows.

To close, the proof of Section 2.1 proceeds as follows. First note how adding and composing sawtooths grows their complexity.

First consider $f+g$ , and moreover any intervals $U_{f}\in\mathcal{I}_{f}$ and $U_{g}\in\mathcal{I}_{g}$ . Necessarily, $f+g$ has a single slope along $U_{f}\cap U_{g}$ . Consequently, $f+g$ is $|\mathcal{I}|$ -sawtooth, where $\mathcal{I}$ is the set of all intersections of intervals from $\mathcal{I}_{f}$ and $\mathcal{I}_{g}$ , meaning $\mathcal{I}:=\{U_{f}\cap U_{g}:U_{f}\in\mathcal{I}_{f},U_{g}\in\mathcal{I}_{g}\}$ . By sorting the left endpoints of elements of $\mathcal{I}_{f}$ and $\mathcal{I}_{g}$ , it follows that $|\mathcal{I}|\leq k+l$ (the other intersections are empty).

Now consider $f\circ g$ , and in particular consider the image $f(g(U_{g}))$ for some interval $U_{g}\in\mathcal{I}_{g}$ . $g$ is affine with a single slope along $U_{g}$ , therefore $f$ is being considered along a single unbroken interval $g(U_{g})$ . However, nothing prevents $g(U_{g})$ from hitting all the elements of $\mathcal{I}_{f}$ ; since $U_{g}$ was arbitrary, it holds that $f\circ g$ is $(|\mathcal{I}_{f}|\cdot|\mathcal{I}_{g}|)$ -sawtooth. ∎

The proof of Section 2.1 follows by induction over layers of $\mathfrak{N}(\sigma;m,l)$ .

The proof proceeds by induction over layers, showing the output of each node in layer $i$ is $(tm)^{i}$ -sawtooth as a function of the neural network input. For the first layer, each node starts by computing $x\mapsto w_{0}+\left\langle w,x\right\rangle$ , which is itself affine and thus 1-sawtooth, so the full node computation $x\mapsto\sigma(w_{0}+\left\langle w,x\right\rangle)$ is $t$ -sawtooth by Section 2.1. Thereafter, the input to layer $i$ with $i>1$ is a collection of functions $(g_{1},\ldots,g_{m^{\prime}})$ with $m^{\prime}\leq m$ and $g_{j}$ being $(tm)^{i-1}$ -sawtooth by the inductive hypothesis; consequently, $x\mapsto w_{0}+\sum_{j}w_{j}g_{j}(x)$ is $m(tm)^{i-1}$ -sawtooth by Section 2.1, whereby applying $\sigma$ yields a $(tm)^{i}$ -sawtooth function (once again by Section 2.1). ∎

2 Upper bound

Note that $f_{\textup{m}}\in\mathfrak{N}(\sigma_{\textsc{r}};2,2)$ ; for instance, $f_{\textup{m}}(x)=\sigma_{\textsc{r}}(2\sigma_{\textsc{r}}(x)-4\sigma_{\textsc{r}}(x-1/2))$ . The upper bounds will use $f_{\textup{m}}^{k}\in\mathfrak{R}(\sigma_{\textsc{r}};2,2;k)\subseteq\mathfrak{N}(\sigma_{\textsc{r}};2,2k)$ .

These compositions may be written as follows.

Let real $x\in$ and positive integer $k$ be given, and choose the unique nonnegative integer $i_{k}\in\{0,\ldots,2^{k-1}\}$ and real $x_{k}\in[0,1)$ so that $x=(i_{k}+x_{k})2^{1-k}$ . Then

The proof proceeds by induction on the number of compositions $l$ . When $l=1$ , there is nothing to show. For the inductive step, the mirroring property of pre-composition with $f_{\textup{m}}$ combined with the symmetry of $f_{\textup{m}}^{l}$ (by the inductive hypothesis) implies that every $x\in[0,1/2]$ satisfies

Consequently, it suffices to consider $x\in[0,1/2]$ , which by the mirroring property means $(f_{\textup{m}}^{l}\circ f_{\textup{m}})(x)=f_{\textup{m}}^{l}(2x)$ . Since the unique nonnegative integer $i_{l+1}$ and real $x_{l+1}\in[0,1)$ satisfy $2x=2(i_{l+1}+x_{l+1})2^{-l-1}=(i_{l+1}+x_{l+1})2^{-l}$ , the inductive hypothesis applied to $2x$ grants

Before closing this subsection, it is interesting to view $f_{\textup{m}}^{k}$ in one more way, namely its effect on $((x_{i},y_{i}))_{i=1}^{n}$ provided by the $n$ -ap with $n:=2^{k}$ . Observe that $((f_{\textup{m}}(x_{i}),y_{i}))_{i=1}^{n}$ is an $(n/2)$ -ap with all points duplicated except $x_{1}=0$ , and an additional point with $x$ -coordinate $1$ .

3 Proof of Theorems 1.2 and 1.1

It suffices to prove Theorem 1.2, which yields Theorem 1.1 since $\sigma_{\textsc{r}}$ is 2-sawtooth, whereby the condition $m\leq 2^{(k-3)/l-1}$ implies

and the upper bound transfers since $\mathfrak{R}(\sigma_{\textsc{r}};2,2;k)\subseteq\mathfrak{N}(\sigma_{\textsc{r}};2,2k)$ .

Related work

The standard classical result on the representation power of neural networks is due to Cybenko (1989), who proved that neural networks can approximate continuous functions over $^{d}$ arbitrarily well. This result, however, is for flat networks.

An early result showing the benefits of depth is due to Håstad (1986), who established, via an incredible proof, that boolean circuits consisting only of and gates and or gates require exponential size in order to approximate the parity function well. These gates correspond to multiplication and addition over the boolean domain, and moreover the parity function is the Fourier basis over the boolean domain; as mentioned above, $f_{\textup{m}}^{k}$ as used here is a piecewise affine approximation of a Fourier basis, and it was suggested previously by Bengio and LeCun (2007) that Fourier transforms admit efficient representations with deep networks. Lastly, note that Håstad (1986)’s work has one of the same weaknesses as the present result, namely of only controlling a countable family of functions which is in no sense dense.

More generally, networks consisting of sum and product nodes, but now over the reals, have been studied in the machine learning literature, where it was showed by Bengio and Delalleau (2011) that again there is an exponential benefit to depth. While this result was again for a countable class of functions, more recent work by Cohen et al. (2015) aims to give a broader characterization.

Lastly, while this note was only concerned with finite sets of points, it is worthwhile to mention the relevance of representation power to statistical questions. Namely, by the seminal result of Anthony and Bartlett (1999, Theorem 8.14), the VC dimension of $\mathfrak{N}(\sigma_{\textsc{r}};m,l)$ is at most $\mathcal{O}(m^{8}l^{2})$ , indicating that these exponential representation benefits directly translate into statistical savings. Interestingly, note that $f_{\textup{m}}^{k}$ has an exponentially large Lipschitz constant (exactly $2^{k}$ ), and thus an elementary statistical analysis via Lipschitz constants and Rademacher complexity (Bartlett and Mendelson, 2002) can inadvertently erase the benefits of depth as presented here.

Overview

Analysis

2 Upper bound

3 Proof of Theorems 1.2 and 1.1

Related work

References