Deep learning with Elastic Averaging SGD

Sixin Zhang, Anna Choromanska, Yann LeCun

Introduction

One of the most challenging problems in large-scale machine learning is how to parallelize the training of large models that use a form of stochastic gradient descent (SGD) . There have been attempts to parallelize SGD-based training for large-scale deep learning models on large number of CPUs, including the Google’s Distbelief system . But practical image recognition systems consist of large-scale convolutional neural networks trained on few GPU cards sitting in a single computer . The main challenge is to devise parallel SGD algorithms to train large-scale deep learning models that yield a significant speedup when run on multiple GPU cards.

In this paper we introduce the Elastic Averaging SGD method (EASGD) and its variants. EASGD is motivated by quadratic penalty method , but is re-interpreted as a parallelized extension of the averaging SGD algorithm . The basic idea is to let each worker maintain its own local parameter, and the communication and coordination of work among the local workers is based on an elastic force which links the parameters they compute with a center variable stored by the master. The center variable is updated as a moving average where the average is taken in time and also in space over the parameters computed by local workers. The main contribution of this paper is a new algorithm that provides fast convergent minimization while outperforming DOWNPOUR method and other baseline approaches in practice. Simultaneously it reduces the communication overhead between the master and the local workers while at the same time it maintains high-quality performance measured by the test error. The new algorithm applies to deep learning settings such as parallelized training of convolutional neural networks.

The article is organized as follows. Section 2 explains the problem setting, Section 3 presents the synchronous EASGD algorithm and its asynchronous and momentum-based variants, Section 4 provides stability analysis of EASGD and ADMM in the round-robin scheme, Section 5 shows experimental results and Section 6 concludes. The Supplement contains additional material including additional theoretical analysis.

Problem setting

EASGD update rule

The communication period $\tau$ controls the frequency of the communication between every local worker and the master, and thus the trade-off between exploration and exploitation.

2 Momentum EASGD

The momentum EASGD (EAMSGD) is a variant of our Algorithm 1 and is captured in Algorithm 2. It is based on the Nesterov’s momentum scheme , where the update of the local worker of the form captured in Equation 3 is replaced by the following update

where $\delta$ is the momentum term. Note that when $\delta=0$ we recover the original EASGD algorithm.

As we are interested in reducing the communication overhead in the parallel computing environment where the parameter vector is very large, we will be exploring in the experimental section the asynchronous EASGD algorithm and its momentum-based variant in the relatively large $\tau$ regime (less frequent communication).

Stability analysis of EASGD and ADMM in the round-robin scheme

In this section we study the stability of the asynchronous EASGD and ADMM methods in the round-robin scheme . We first state the updates of both algorithms in this setting, and then we study their stability. We will show that in the one-dimensional quadratic case, ADMM algorithm can exhibit chaotic behavior, leading to exponential divergence. The analytic condition for the ADMM algorithm to be stable is still unknown, while for the EASGD algorithm it is very simpleThis condition resembles the stability condition for the synchronous EASGD algorithm (Condition 23 for $p=1$ ) in the analysis in the Supplement..

The analysis of the synchronous EASGD algorithm, including its convergence rate, and its averaging property, in the quadratic and strongly convex case, is deferred to the Supplement.

In our setting, the ADMM method involves solving the following minimax problemThe convergence analysis in is based on the assumption that “At any master iteration, updates from the workers have the same probability of arriving at the master.”, which is not satisfied in the round-robin scheme.,

where $\lambda^{i}$ ’s are the Lagrangian multipliers. The resulting updates of the ADMM algorithm in the round-robin scheme are given next. Let $t\geq 0$ be a global clock. At each $t$ , we linearize the function $F(x^{i})$ with $F(x^{i}_{t})+\scalprod{L2}{\nabla{F}(x^{i}_{t})}{x^{i}-x^{i}_{t}}+\frac{1}{2\eta}\norm{L2}{x^{i}-x^{i}_{t}}^{2}$ as in . The updates become

The EASGD algorithm in the round-robin scheme is defined similarly and is given below

At time $t$ , only the $i$ -th local worker (whose index $i-1$ equals $t$ modulo $p$ ) is activated, and performs the update in Equations 18 which is followed by the master update given in Equation 19.

For each of the $p$ linear maps, it’s possible to find a simple condition such that each map, where the $i^{\text{th}}$ map has the form $F^{i}_{3}\circ F^{i}_{2}\circ F^{i}_{1}$ , is stable (the absolute value of the eigenvalues of the map are smaller or equal to one). However, when these non-symmetric maps are composed one after another as follows $\mathcal{F}=F^{p}_{3}\circ F^{p}_{2}\circ F^{p}_{1}\circ\ldots\circ F^{1}_{3}\circ F^{1}_{2}\circ F^{1}_{1}$ , the resulting map $\mathcal{F}$ can become unstable! (more precisely, some eigenvalues of the map can sit outside the unit circle in the complex plane).

We now present the numerical conditions for which the ADMM algorithm becomes unstable in the round-robin scheme for $p=3$ and $p=8$ , by computing the largest absolute eigenvalue of the map $\mathcal{F}$ . Figure 1 summarizes the obtained result.

For the composite map $F^{p}\circ\ldots\circ F^{1}$ to be stable, the condition that needs to be satisfied is actually the same for each $i$ , and is furthermore independent of $p$ (since each linear map $F^{i}$ is symmetric). It essentially involves the stability of the $2\times 2$ matrix $\left(\begin{array}[]{cc}1-\eta-\alpha&\alpha\\ \alpha&1-\alpha\end{array}\right)$ , whose two (real) eigenvalues $\lambda$ satisfy $(1-\eta-\alpha-\lambda)(1-\alpha-\lambda)=\alpha^{2}$ . The resulting stability condition ( $|\lambda|\leq 1$ ) is simple and given as $0\leq\eta\leq 2,0\leq\alpha\leq\frac{4-2\eta}{4-\eta}$ .

Experiments

In this section we compare the performance of EASGD and EAMSGD with the parallel method DOWNPOUR and the sequential method SGD, as well as their averaging and momentum variants.

All the parallel comparator methods are listed belowWe have compared asynchronous ADMM with EASGD in our setting as well, the performance is nearly the same. However, ADMM’s momentum variant is not as stable for large communication periods.:

DOWNPOUR , the pseudo-code of the implementation of DOWNPOUR used in this paper is enclosed in the Supplement.

Momentum DOWNPOUR (MDOWNPOUR), where the Nesterov’s momentum scheme is applied to the master’s update (note it is unclear how to apply it to the local workers or for the case when $\tau>1$ ). The pseudo-code is in the Supplement.

All the sequential comparator methods ( $p=1$ ) are listed below:

Momentum SGD (MSGD) with constant momentum $\delta$ .

ASGD with moving rate $\alpha_{t+1}=\frac{1}{t+1}$ .

MVASGD with moving rate $\alpha$ set to a constant.

We perform experiments in a deep learning setting on two benchmark datasets: CIFAR-10 (we refer to it as CIFAR) Downloaded from http://www.cs.toronto.edu/~kriz/cifar.html. and ImageNet ILSVRC 2013 (we refer to it as ImageNet) Downloaded from http://image-net.org/challenges/LSVRC/2013.. We focus on the image classification task with deep convolutional neural networks. We next explain the experimental setup. The details of the data preprocessing and prefetching are deferred to the Supplement.

For all our experiments we use a GPU-cluster interconnected with InfiniBand. Each node has $4$ Titan GPU processors where each local worker corresponds to one GPU processor. The center variable of the master is stored and updated on the centralized parameter server Our implementation is available at https://github.com/sixin-zh/mpiT..

To describe the architecture of the convolutional neural network, we will first introduce a notation. Let $(c,y)$ denotes the size of the input image to each layer, where $c$ is the number of color channels and $y$ is both the horizontal and the vertical dimension of the input. Let $C$ denotes the fully-connected convolutional operator and let $P$ denotes the max pooling operator, $D$ denotes the linear operator with dropout rate equal to $0.5$ and $S$ denotes the linear operator with softmax output non-linearity. We use the cross-entropy loss and all inner layers use rectified linear units. For the ImageNet experiment we use the similar approach to with the following $11$ -layer convolutional neural network (3,221)C(96,108)P(96,36)C(256,32)P(256,16)C(384,14) C(384,13)C(256,12)P(256,6)D(4096,1)D(4096,1)S(1000,1). For the CIFAR experiment we use the similar approach to with the following $7$ -layer convolutional neural network (3,28)C(64,24)P(64,12)C(128,8)P(128,4)C(64,2)D(256,1)S(10,1).

In our experiments all the methods we run use the same initial parameter chosen randomly, except that we set all the biases to zero for CIFAR case and to 0.1 for ImageNet case. This parameter is used to initialize the master and all the local workersOn the contrary, initializing the local workers and the master with different random seeds ’traps’ the algorithm in the symmetry breaking phase.. We add $l_{2}$ -regularization $\frac{\lambda}{2}\norm{}{x}^{2}$ to the loss function $F(x)$ . For ImageNet we use $\lambda=10^{-5}$ and for CIFAR we use $\lambda=10^{-4}$ . We also compute the stochastic gradient using mini-batches of sample size $128$ .

2 Experimental results

For all experiments in this section we use EASGD with $\beta=0.9$ Intuitively the ’effective $\beta$ ’ is $\beta/\tau=p\alpha=p\eta\rho$ (thus $\rho=\frac{\beta}{\tau p\eta}$ ) in the asynchronous setting. , for all momentum-based methods we set the momentum term $\delta=0.99$ and finally for MVADOWNPOUR we set the moving rate to $\alpha=0.001$ . We start with the experiment on CIFAR dataset with $p=4$ local workers running on a single computing node. For all the methods, we examined the communication periods from the following set $\tau=\{1,4,16,64\}$ . For comparison we also report the performance of MSGD which outperformed SGD, ASGD and MVASGD as shown in Figure 6 in the Supplement. For each method we examined a wide range of learning rates (the learning rates explored in all experiments are summarized in Table 1, 2, 3 in the Supplement). The CIFAR experiment was run $3$ times independently from the same initialization and for each method we report its best performance measured by the smallest achievable test error.

From the results in Figure 2, we conclude that all DOWNPOUR-based methods achieve their best performance (test error) for small $\tau$ ( $\tau\in\{1,4\}$ ), and become highly unstable for $\tau\in\{16,64\}$ . While EAMSGD significantly outperforms comparator methods for all values of $\tau$ by having faster convergence. It also finds better-quality solution measured by the test error and this advantage becomes more significant for $\tau\in\{16,64\}$ . Note that the tendency to achieve better test performance with larger $\tau$ is also characteristic for the EASGD algorithm.

We next explore different number of local workers $p$ from the set $p=\{4,8,16\}$ for the CIFAR experiment, and $p=\{4,8\}$ for the ImageNet experimentFor the ImageNet experiment, the training loss is measured on a subset of the training data of size 50,000.. For the ImageNet experiment we report the results of one run with the best setting we have found. EASGD and EAMSGD were run with $\tau=10$ whereas DOWNPOUR and MDOWNPOUR were run with $\tau=1$ . The results are in Figure 3 and 4. For the CIFAR experiment, it’s noticeable that the lowest achievable test error by either EASGD or EAMSGD decreases with larger $p$ . This can potentially be explained by the fact that larger $p$ allows for more exploration of the parameter space. In the Supplement, we discuss further the trade-off between exploration and exploitation as a function of the learning rate (section 9.5) and the communication period (section 9.6). Finally, the results obtained for the ImageNet experiment also shows the advantage of EAMSGD over the competitor methods.

Conclusion

In this paper we describe a new algorithm called EASGD and its variants for training deep neural networks in the stochastic setting when the computations are parallelized over multiple GPUs. Experiments demonstrate that this new algorithm quickly achieves improvement in test error compared to more common baseline approaches such as DOWNPOUR and its variants. We show that our approach is very stable and plausible under communication constraints. We provide the stability analysis of the asynchronous EASGD in the round-robin scheme, and show the theoretical advantage of the method over ADMM. The different behavior of the EASGD algorithm from its momentum-based variant EAMSGD is intriguing and will be studied in future works.

The authors thank R. Power, J. Li for implementation guidance, J. Bruna, O. Henaff, C. Farabet, A. Szlam, Y. Bakhtin for helpful discussion, P. L. Combettes, S. Bengio and the referees for valuable feedback.

References

Additional theoretical results and proofs

We provide here the convergence analysis of the synchronous EASGD algorithm with constant learning rate. The analysis is focused on the convergence of the center variable to the local optimum. We discuss one-dimensional quadratic case first, then the generalization to multi-dimensional setting (Lemma 7.3) and finally to the strongly convex case (Theorem 7.1).

It follows from Lemma 7.1 that for the center variable to be stable the following has to hold

It can be verified that $\phi$ and $\gamma$ are the two zero-roots of the polynomial in $\lambda$ : $\lambda^{2}-(2-a)\lambda+(1-a+c^{2})$ . Recall that $\phi$ and $\lambda$ are the functions of $\eta$ and $\alpha$ . Thus (see proof in Section 7.1.2)

$\gamma<1$ iff $c^{2}>0$ (i.e. $\eta>0$ and $\alpha>0$ ).

$\phi>-1$ iff $(2-\eta h)(2-p\alpha)>2\alpha$ and $(2-\eta h)+(2-p\alpha)>\alpha$ .

$\phi=\gamma$ iff $a^{2}=4c^{2}$ (i.e. $\eta h=\alpha=0$ ).

The proof the above Lemma is based on the diagonalization of the linear gradient map (this map is symmetric due to the relation $\beta=p\alpha$ ). The stability analysis of the asynchronous EASGD algorithm in the round-robin scheme is similar due to this elastic symmetry.

Substituting the gradient from Equation 20 into the update rule used by each local worker in the synchronous EASGD algorithm (Equation 5 and 6) we obtain

where $\eta$ is the learning rate, and $\alpha$ is the moving rate. Recall that $\alpha=\eta\rho$ and $A=h$ .

and the (diffusion) vector $b_{t}=(\eta\xi^{1}_{t},\ldots,\eta\xi^{p}_{t},0)^{T}$ .

By combining Equation 25 and 26 as follows

where the last step results from the following relations: $\frac{p\alpha^{2}}{1-p\alpha-\phi}=1-\alpha-\eta h-\phi$ and $\phi+\gamma=1-\alpha-\eta h+1-p\alpha$ . Thus we obtained

Substituting $u_{0},u_{1},\dots,u_{t}$ , each given through Equation 28, into Equation 29 we obtain

To be more specific, the Equation 30 is obtained by integrating by parts,

Since the random variables $\xi_{l}$ are i.i.d, we may sum the variance term by term as follows

1.2 Condition in Equation 23

$\gamma<1$ iff $c^{2}>0$ (i.e. $\eta>0$ and $\beta>0$ ).

$\phi>-1$ iff $(2-\eta h)(2-\beta)>2\beta/p$ and $(2-\eta h)+(2-\beta)>\beta/p$ .

$\phi=\gamma$ iff $a^{2}=4c^{2}$ (i.e. $\eta h=\beta=0$ ).

Recall that $a=\eta h+(p+1)\alpha$ , $c^{2}=\eta hp\alpha$ , $\gamma=1-\frac{a-\sqrt{a^{2}-4c^{2}}}{2}$ , $\phi=1-\frac{a+\sqrt{a^{2}-4c^{2}}}{2}$ , and $\beta=p\alpha$ . We have

$\gamma<1\Leftrightarrow\frac{a-\sqrt{a^{2}-4c^{2}}}{2}>0\Leftrightarrow a>\sqrt{a^{2}-4c^{2}}\Leftrightarrow a^{2}>a^{2}-4c^{2}\Leftrightarrow c^{2}>0$ .

$\phi>-1\Leftrightarrow 2>\frac{a+\sqrt{a^{2}-4c^{2}}}{2}\Leftrightarrow 4-a>\sqrt{a^{2}-4c^{2}}\Leftrightarrow 4-a>0,(4-a)^{2}>a^{2}-4c^{2}\Leftrightarrow 4-a>0,4-2a+c^{2}>0\Leftrightarrow 4>\eta h+\beta+\alpha,4-2(\eta h+\beta+\alpha)+\eta h\beta>0$ .

$\phi=\gamma\Leftrightarrow\sqrt{a^{2}-4c^{2}}=0\Leftrightarrow a^{2}=4c^{2}$ .

The next corollary is a consequence of Lemma 7.1. As the number of workers $p$ grows, the averaging property of the EASGD can be characterized as follows

Let the Elastic Averaging relation $\beta=p\alpha$ and the condition 23 hold, then

Note that when $\beta$ is fixed, $\lim_{p\rightarrow\infty}a=\eta h+\beta$ and $c^{2}=\eta h\beta$ . Then $\lim_{p\rightarrow\infty}\phi=\min(1-\beta,1-\eta h)$ and $\lim_{p\rightarrow\infty}\gamma=\max(1-\beta,1-\eta h)$ . Also note that using Lemma 7.1 we obtain

Corollary 7.1 is obtained by plugining in the limiting values of $\phi$ and $\gamma$ . ∎

The crucial point of Corollary 7.1 is that the MSE in the limit $t\rightarrow\infty$ is in the order of $1/p$ which implies that as the number of processors $p$ grows, the MSE will decrease for the EASGD algorithm. Also note that the smaller the $\beta$ is (recall that $\beta=p\alpha=p\eta\rho$ ), the more exploration is allowed (small $\rho$ ) and simultaneously the smaller the MSE is.

2 Generalization to multidimensional case

The next lemma (Lemma 7.2) shows that EASGD algorithm achieves the highest possible rate of convergence when we consider the double averaging sequence (similarly to ) $\{z_{1},z_{2},\dots\}$ defined as below

If the condition in Equation 23 holds, then the normalized double averaging sequence defined in Equation 32 converges weakly to the normal distribution with zero mean and variance $\sigma^{2}/ph^{2}$ ,

Also recall that $\{\xi^{i}_{t}\}$ ’s are i.i.d. random variables (noise) with zero mean and the same covariance $\Sigma\succ 0$ . We are interested in the asymptotic behavior of the double averaging sequence $\{z_{1},z_{2},\dots\}$ defined as

Recall the Equation 30 from the proof of Lemma 7.1 (for the convenience it is provided below):

where $\xi_{t}=\sum_{i=1}^{p}\xi_{t}^{i}$ . Therefore

Therefore the expression in Equation 35 is asymptotically normal with zero mean and variance $\sigma^{2}/ph^{2}$ . ∎

The asymptotic variance in the Lemma 7.2 is optimal with any fixed $\eta$ and $\beta$ for which Equation 23 holds. The next lemma (Lemma 7.3) extends the result in Lemma 7.2 to the multi-dimensional setting.

Let $h$ denotes the largest eigenvalue of $A$ . If $(2-\eta h)(2-\beta)>2\beta/p$ , $(2-\eta h)+(2-\beta)>\beta/p$ , $\eta>0$ and $\beta>0$ , then the normalized double averaging sequence converges weakly to the normal distribution with zero mean and the covariance matrix $V=A^{-1}\Sigma(A^{-1})^{T}$ ,

Let the spatial average of the local parameters at time $t$ be denoted as $y_{t}$ where $y_{t}=\frac{1}{p}\sum_{i=1}^{p}x_{t}^{i}$ , and let the average noise be denoted as $\xi_{t}$ , where $\xi_{t}=\frac{1}{p}\sum_{i=1}^{p}\xi_{t}^{i}$ . Equations 24 and 25 can then be reduced to the following

From Equations 37 and 38 it follows that $U_{t+1}=MU_{t}+\Xi_{t}$ . Note that this linear system has a degenerate noise $\Xi_{t}$ which prevents us from directly applying results of . Expanding this recursive relation and summing by parts, we have

Note that the only non-vanishing term (in weak convergence) of $\frac{1}{\sqrt{t}}\sum_{k=0}^{t}U_{k}$ is $\frac{1}{\sqrt{t}}(\eta L)^{-1}\sum_{k=1}^{t}\Xi_{k-1}$ thus we have

The eigenvalue $\lambda$ of $M$ and the (non-zero) eigenvector $(y,z)$ of $M$ satisfy

Since $(y,z)$ is assumed to be non-zero, we can write $z=\beta y/(\lambda+\beta-1)$ . Then the Equation 50 can be reduced to

Thus $y$ is the eigenvector of $A$ . Let $\lambda_{A}$ be the eigenvalue of matrix $A$ such that $Ay=\lambda_{A}y$ . Thus based on Equation 51 it follows that

where $a=\eta\lambda_{A}+(p+1)\alpha$ , $c^{2}=\eta\lambda_{A}p\alpha$ . It follows from the condition in Equation 23 that $-1<\lambda<1$ iff $\eta>0$ , $\beta>0$ , $(2-\eta\lambda_{A})(2-\beta)>2\beta/p$ and $(2-\eta\lambda_{A})+(2-\beta)>\beta/p$ . Let $h$ denote the maximum eigenvalue of $A$ and note that $2-\eta\lambda_{A}\geq 2-\eta h$ . This implies that the condition of our lemma is sufficient. ∎

As in Lemma 7.2, the asymptotic covariance in the Lemma 7.3 is optimal, i.e. meets the Fisher information lower-bound. The fact that this asymptotic covariance matrix $V$ does not contain any term involving $\rho$ is quite remarkable, since the penalty term $\rho$ does have an impact on the condition number of the Hessian in Equation 2.

3 Strongly convex case

We have thus the update for the spatial average,

From Equation 54 the following relation holds,

By the cosine rule ( $2\scalprod{L2}{a-b}{c-d}=\norm{L2}{a-d}^{2}-\norm{L2}{a-c}^{2}+\norm{L2}{c-b}^{2}-\norm{L2}{d-b}^{2}$ ), we have

By the Cauchy-Schwarz inequality, we have

Combining the above estimates in Equations 57, 58, 59, 60, we obtain

Now we apply similar idea to estimate $\norm{L2}{y_{t}-x^{\ast}}^{2}$ . From Equation 56 the following relation holds,

By $\scalprod{L2}{\frac{1}{p}\sum_{i=1}^{p}a_{i}}{\frac{1}{p}\sum_{j=1}^{p}b_{j}}=\frac{1}{p}\sum_{i=1}^{p}\scalprod{L2}{a_{i}}{b_{i}}-\frac{1}{p^{2}}\sum_{i>j}\scalprod{L2}{a_{i}-a_{j}}{b_{i}-b_{j}}$ , we have

Denote $\xi_{t}=\frac{1}{p}\sum_{i=1}^{p}\xi^{i}_{t}$ , we can rewrite Equation 65 as

By combining the above Equations 66, 67 with 68, we obtain

Thus it follows from Equation 57 and 70 that

Recall $y_{t}=\frac{1}{p}\sum_{i=1}^{p}x^{i}_{t}$ , we have the following bias-variance relation,

By the Cauchy-Schwarz inequality, we have

Combining the above estimates in Equations 71, 72, 73, we obtain

Since $\frac{2\sqrt{\mu L}}{\mu+L}\leq 1$ , we need also bound the non-linear term $\scalprod{L2}{\nabla f^{i}_{t}-\nabla f^{j}_{t}}{x_{t}^{i}-x_{t}^{j}}\leq L\norm{L2}{x_{t}^{i}-x_{t}^{j}}^{2}$ . Recall the bias-variance relation $\frac{1}{p}\sum_{i=1}^{p}\norm{L2}{x^{i}_{t}-x^{\ast}}^{2}=\frac{1}{p^{2}}\sum_{i>j}\norm{L2}{x^{i}_{t}-x^{j}_{t}}^{2}+\norm{L2}{y_{t}-x^{\ast}}^{2}$ . The key observation is that if $\frac{1}{p}\sum_{i=1}^{p}\norm{L2}{x^{i}_{t}-x^{\ast}}^{2}$ remains bounded, then larger variance $\sum_{i>j}\norm{L2}{x^{i}_{t}-x^{j}_{t}}^{2}$ implies smaller bias $\norm{L2}{y_{t}-x^{\ast}}^{2}$ . Thus this non-linear term can be compensated.

Again choose $\eta$ small enough such that $\eta^{2}+\frac{\eta^{2}\alpha}{1-\alpha}-\frac{2\eta}{\mu+L}\leq 0$ and take expectation in Equation 75,

As for the center variable in Equation 55, we apply simply the convexity of the norm $\norm{L2}{\cdot}^{2}$ to obtain

as long as $0\leq\beta\leq 1$ , $0\leq\alpha<1$ and $\eta^{2}+\frac{\eta^{2}\alpha}{1-\alpha}-\frac{2\eta}{\mu+L}\leq 0$ , i.e. $0\leq\eta\leq\frac{2}{\mu+L}(1-\alpha).$ ∎

Additional pseudo-codes of the algorithms

Algorithm 3 captures the pseudo-code of the implementation of the DOWNPOUR used in this paper.

2 MDOWNPOUR pseudo-code

Algorithms 4 and 5 capture the pseudo-codes of the implementation of momentum DOWNPOUR (MDOWNPOUR) used in this paper. Algorithm 4 shows the behavior of each local worker and Algorithm 5 shows the behavior of the master.

Experiments - additional material

For the ImageNet experiment, we re-size each RGB image so that the smallest dimension is $256$ pixels. We also re-scale each pixel value to the interval $ $. We then extract random crops (and their horizontal flips) of size$ 3\times 221\times 221 $pixels and present these to the network in mini-batches of size$ 128$.

For the CIFAR experiment, we use the original RGB image of size $3\times 32\times 32$ . As before, we re-scale each pixel value to the interval $ $. We then extract random crops (and their horizontal flips) of size$ 3\times 28\times 28 $pixels and present these to the network in mini-batches of size$ 128$.

The training and test loss and the test error are only computed from the center patch ( $3\times 28\times 28$ ) for the CIFAR experiment and the center patch ( $3\times 221\times 221$ ) for the ImageNet experiment.

2 Data prefetching (Sampling the dataset by the local workers)

We will now explain precisely how the dataset is sampled by each local worker as uniformly and efficiently as possible. The general parallel data loading scheme on a single machine is as follows: we use $k$ CPUs, where $k=8$ , to load the data in parallel. Each data loader reads from the memory-mapped (mmap) file a chunk of $c$ raw images (preprocessing was described in the previous subsection) and their labels (for CIFAR $c=512$ and for ImageNet $c=64$ ). For the CIFAR, the mmap file of each data loader contains the entire dataset whereas for ImageNet, each mmap file of each data loader contains different $1/k$ fractions of the entire dataset. A chunk of data is always sent by one of the data loaders to the first worker who requests the data. The next worker requesting the data from the same data loader will get the next chunk. Each worker requests in total $k$ data chunks from $k$ different data loaders and then process them before asking for new data chunks. Notice that each data loader cycles through the data in the mmap file, sending consecutive chunks to the workers in order in which it receives requests from them. When the data loader reaches the end of the mmap file, it selects the address in memory uniformly at random from the interval $[0,s]$ , where $s=(\textsf{number of images in the mmap file}\text{\>\>modulo\>\>}\textsf{mini-batch size})$ , and uses this address to start cycling again through the data in the mmap file. After the local worker receives the $k$ data chunks from the data loaders, it shuffles them and divides it into mini-batches of size $128$ .

3 Learning rates

In Table 1 we summarize the learning rates $\eta$ (we used constant learning rates) explored for each method shown in Figure 2. For all values of $\tau$ the same set of learning rates was explored for each method.

In Table 2 we summarize the learning rates $\eta$ (we used constant learning rates) explored for each method shown in Figure 3. For all values of $p$ the same set of learning rates was explored for each method.

In Table 3 we summarize the initial learning rates $\eta$ we use for each method shown in Figure 4. For all values of $p$ the same set of learning rates was explored for each method. We also used the rule of the thumb to decrease the initial learning rate twice, first time we divided it by $5$ and the second time by $2$ , when we observed that the decrease of the online predictive (training) loss saturates.

4 Comparison of SGD, ASGD, MVASGD and MSGD

Figure 6 shows the convergence of the training and test loss (negative log-likelihood) and the test error computed for the center variable as a function of wallclock time for SGD, ASGD, MVASGD and MSGD ( $p=1$ ) on the CIFAR experiment. For all CIFAR experiments we always start the averaging for the $ADOWNPOUR$ and $ASGD$ methods from the very beginning of each experiment. For all ImageNet experiments we start the averaging for the $ASGD$ at the same time when we first reduce the initial learning rate.

Figure 7 shows the convergence of the training and test loss (negative log-likelihood) and the test error computed for the center variable as a function of wallclock time for SGD, ASGD, MVASGD and MSGD ( $p=1$ ) on the ImageNet experiment.

5 Dependence of the learning rate

This section discusses the dependence of the trade-off between exploration and exploitation on the learning rate. We compare the performance of respectively EAMSGD and EASGD for different learning rates $\eta$ when $p=16$ and $\tau=10$ on the CIFAR experiment. We observe in Figure 8 that higher learning rates $\eta$ lead to better test performance for the EAMSGD algorithm which potentially can be justified by the fact that they sustain higher fluctuations of the local workers. We conjecture that higher fluctuations lead to more exploration and simultaneously they also impose higher regularization. This picture however seems to be opposite for the EASGD algorithm for which larger learning rates hurt the performance of the method and lead to overfitting. Interestingly in this experiment for both EASGD and EAMSGD algorithm, the learning rate for which the best training performance was achieved simultaneously led to the worst test performance.

6 Dependence of the communication period

This section discusses the dependence of the trade-off between exploration and exploitation on the communication period. We have observed from the CIFAR experiment that EASGD algorithm exhibits very similar convergence behavior when $\tau=1$ up to even $\tau=1000$ , whereas EAMSGD can get trapped at worse energy (loss) level for $\tau=100$ . This behavior of EAMSGD is most likely due to the non-convexity of the objective function. Luckily, it can be avoided by gradually decreasing the learning rate, i.e. increasing the penalty term $\rho$ (recall $\alpha=\eta\rho$ ), as shown in Figure 9. In contrast, the EASGD algorithm does not seem to get trapped at all along its trajectory. The performance of EASGD is less sensitive to increasing the communication period compared to EAMSGD, whereas for the EAMSGD the careful choice of the learning rate for large communication periods seems crucial.

Compared to all earlier results, the experiment in this section is re-run three times with a new randomTo clarify, the random initialization we use is by default in Torch’s implementation. seed and with faster cuDNNhttps://developer.nvidia.com/cuDNN packagehttps://github.com/soumith/cudnn.torch. All our methods are implemented in Torchhttp://torch.ch. The Message Passing Interface implementation MVAPICH2http://mvapich.cse.ohio-state.edu is used for the GPU-CPU communication.

7 Breakdown of the wallclock time

In addition, we report in Table 4 the breakdown of the total running time for EASGD when $\tau=10$ (the time breakdown for EAMSGD is almost identical) and DOWNPOUR when $\tau=1$ into computation time, data loading time and parameter communication time. For the CIFAR experiment the reported time corresponds to processing $400\times 128$ data samples whereas for the ImageNet experiment it corresponds to processing $1024\times 128$ data samples. For $\tau=1$ and $p\in\{8,16\}$ we observe that the communication time accounts for significant portion of the total running time whereas for $\tau=10$ the communication time becomes negligible compared to the total running time (recall that based on previous results EASGD and EAMSGD achieve best performance with larger $\tau$ which is ideal in the setting when communication is time-consuming).

8 Time speed-up

In Figure 10 and 11, we summarize the wall clock time needed to achieve the same level of the test error for all the methods in the CIFAR and ImageNet experiment as a function of the number of local workers $p$ . For the CIFAR (Figure 10) we examined the following levels: $\{21\%,20\%,19\%,18\%\}$ and for the ImageNet (Figure 11) we examined: $\{49\%,47\%,45\%,43\%\}$ . If some method does not appear on the figure for a given test error level, it indicates that this method never achieved this level. For the CIFAR experiment we observe that from among EASGD, DOWNPOUR and MDOWNPOUR methods, the EASGD method needs less time to achieve a particular level of test error. We observe that with higher $p$ each of these methods does not necessarily need less time to achieve the same level of test error. This seems counter intuitive though recall that the learning rate for the methods is selected based on the smallest achievable test error. For larger $p$ smaller learning rates were selected than for smaller $p$ which explains our results. Meanwhile, the EAMSGD method achieves significant speed-up over other methods for all the test error levels. For the ImageNet experiment we observe that all methods outperform MSGD and furthermore with $p=4$ or $p=8$ each of these methods requires less time to achieve the same level of test error. The EAMSGD consistently needs less time than any other method, in particular DOWNPOUR, to achieve any of the test error levels.