On Optimal Probabilities in Stochastic Coordinate Descent Methods

Peter Richtárik, Martin Takáč

Introduction

In this work we consider the optimization problem

where $\phi$ is strongly convex and smooth. We propose a new algorithm, and call it ‘NSync (Nonuniform SYNchronous Coordinate descent).

In ‘NSync, we first assign a probability $p_{S}\geq 0$ to every subset $S$ of $[n]:=\{1,\dots,n\}$ , with $\sum_{S}p_{S}=1$ , and pick stepsize parameters $w_{i}>0$ , $i=1,2,\dots,n$ . At every iteration, a random set $\hat{S}$ is generated, independently from previous iterations, following the law $\mathbf{Prob}(\hat{S}=S)=p_{S}$ , and then coordinates $i\in\hat{S}$ are updated in parallel by moving in the direction of the negative partial derivative with stepsize $1/w_{i}$ . The updates are synchronized: no processor/thread is allowed to proceed before all updates are applied, generating the new iterate $x^{k+1}$ . We specifically study samplings $\hat{S}$ which are non-uniform in the sense that $p_{i}:=\mathbf{Prob}(i\in\hat{S})=\sum_{S:i\in S}p_{S}$ is allowed to vary with $i$ . By $\nabla_{i}\phi(x)$ we mean $\langle\nabla\phi(x),e^{i}\rangle$ , where $e^{i}\in\mathbf{R}^{n}$ is the $i$ -th unit coordinate vector.

Serial stochastic coordinate descent methods were proposed and analyzed in , and more recently in various settings in . Parallel methods were considered in , and more recently in . A memory distributed method scaling to big data problems was recently developed in . A nonuniform coordinate descent method updating a single coordinate at a time was proposed in , and one updating two coordinates at a time in . To the best of our knowledge, ‘NSync is the first nonuniform parallel coordinate descent method.

Analysis

Our analysis of ‘NSync is based on two assumptions. The first assumption generalizes the ESO concept introduced in and later used in to nonuniform samplings. The second assumption requires that $\phi$ be strongly convex.

Notation: For $x,y,u\in\mathbf{R}^{n}$ we write $\|x\|_{u}^{2}:=\sum_{i}u_{i}x_{i}^{2}$ , $\langle x,y\rangle_{u}:=\sum_{i=1}^{n}u_{i}y_{i}x_{i}$ , $x\bullet y:=(x_{1}y_{1},\dots,x_{n}y_{n})$ and $u^{-1}:=(1/u_{1},\dots,1/u_{n})$ . For $S\subseteq[n]$ and $h\in\mathbf{R}^{n}$ , let $h_{[S]}:=\sum_{i\in S}h_{i}e^{i}$ .

Assume $p=(p_{1},\dots,p_{n})^{T}>0$ and that for some positive vector $w\in\mathbf{R}^{n}$ and all $x,h\in\mathbf{R}^{n}$ ,

Inequalities of type (2), in the uniform case ( $p_{i}=p_{j}$ for all $i,j$ ), were studied in .

We assume that $\phi$ is $\gamma$ -strongly convex with respect to the norm $\|\cdot\|_{v}$ , where $v=(v_{1},\dots,v_{n})^{T}>0$ and $\gamma>0$ . That is, we require that for all $x,h\in\mathbf{R}^{n}$ ,

We can now establish a bound on the number of iterations sufficient for ‘NSync to approximately solve (1) with high probability.

Let Assumptions 1 and 2 be satisfied. Choose $x^{0}\in\mathbf{R}^{n}$ , $0<\epsilon<\phi(x^{0})-\phi^{*}$ and $0<\rho<1$ , where $\phi^{*}:=\min_{x}\phi(x)$ . Let

If $\{x^{k}\}$ are the random iterates generated by ‘NSync, then

Moreover, we have the lower bound $\Lambda\geq(\sum_{i}\tfrac{w_{i}}{v_{i}})/\mathbf{E}[|\hat{S}|]$ .

We first claim that $\phi$ is $\mu$ -strongly convex with respect to the norm $\|\cdot\|_{w\bullet p^{-1}}$ , i.e.,

where $\mu:=\gamma/\Lambda$ . Indeed, this follows by comparing (3) and (6) in the light of (4). Let $x^{*}$ be such that $\phi(x^{*})=\phi^{*}$ . Using (6) with $h=x^{*}-x$ ,

Let $h^{k}:=-(\operatorname*{Diag}(w))^{-1}\nabla\phi(x^{k})$ . Then $x^{k+1}=x^{k}+(h^{k})_{[\hat{S}]}$ , and utilizing Assumption 1, we get

Taking expectations in the last inequality and rearranging the terms, we obtain $\mathbf{E}[\phi(x^{k+1})-\phi^{*}]\leq(1-\mu)\mathbf{E}[\phi(x^{k})-\phi^{*}]\leq(1-\mu)^{k+1}(\phi(x^{0})-\phi^{*})$ . Using this, Markov inequality, and the definition of $K$ , we finally get $\mathbf{Prob}(\phi(x^{K})-\phi^{*}\geq\epsilon)\leq\mathbf{E}[\phi(x^{K})-\phi^{*}]/\epsilon\leq(1-\mu)^{K}(\phi(x^{0})-\phi^{*})/\epsilon\leq\rho$ . Let us now establish the last claim. First, note that (see [16, Sec 3.2] for more results of this type),

Letting $\Delta:=\{p^{\prime}\in\mathbf{R}^{n}:p^{\prime}\geq 0,\sum_{i}p_{i}^{\prime}=\mathbf{E}[|\hat{S}|]\}$ , we have

where the last equality follows since optimal $p_{i}^{\prime}$ is proportional to $w_{i}/v_{i}$ . ∎

Theorem 3 is generic in the sense that we do not say when Assumptions 1 and 2 are satisfied, how should one go about to choose the stepsizes $w$ and probabilities $\{p_{S}\}$ . In the next section we address these issues. On the other hand, this abstract setting allowed us to write a brief complexity proof.

Change of variables. Consider the change of variables $y=\operatorname*{Diag}(d)x$ , where $d>0$ . Defining $\phi^{d}(y):=\phi(x)$ , we get $\nabla\phi^{d}(y)=(\operatorname*{Diag}(d))^{-1}\nabla\phi(x)$ . It can be seen that (2), (3) can equivalently be written in terms of $\phi^{d}$ , with $w$ replaced by $w^{d}:=w\bullet d^{-2}$ and $v$ replaced by $v^{d}:=v\bullet d^{-2}$ . By choosing $d_{i}=\sqrt{v_{i}}$ , we obtain $v^{d}_{i}=1$ for all $i$ , recovering standard strong convexity.

Nonuniform samplings and ESO

Consider now problem (1) with $\phi$ of the form

where $v>0$ . Note that Assumption 2 is satisfied. We further make the following two assumptions.

$f$ has Lipschitz gradient with respect to the coordinates, with positive constants $L_{1},\dots,L_{n}$ . That is, $|\nabla_{i}f(x)-\nabla_{i}f(x+te_{i})|\leq L_{i}|t|$ for all $x\in\mathbf{R}^{n}$ and $t\in\mathbf{R}$ .

$f(x)=\sum_{J\in\mathcal{J}}f_{J}(x)$ , where $\mathcal{J}$ is a finite collection of nonempty subsets of $[n]$ and $f_{J}$ are differentiable convex functions such that $f_{J}$ depends on coordinates $i\in J$ only. Let $\omega:=\max_{J}|J|$ . We say that $f$ is separable of degree $\omega$ .

Uniform parallel coordinate descent methods for regularized problems with $f$ of the above structure were analyzed in .

Let $f(x)=\tfrac{1}{2}\|Ax-b\|_{2}^{2}$ , where $A\in\mathbf{R}^{m\times n}$ . Then $L_{i}=\|A_{:i}\|_{2}^{2}$ and $f(x)=\tfrac{1}{2}\sum_{j=1}^{m}(A_{j:}x-b_{j})^{2}$ , whence $\omega$ is the maximum # of nonzeros in a row of $A$ .

Instead of considering the general case of arbitrary $p_{S}$ assigned to all subsets of $[n]$ , here we consider a special kind of sampling having two advantages: i) sets can be generated easily, ii) it leads to larger stepsizes $1/w_{i}$ and hence improved convergence rate. Fix $\tau\in[n]$ and $c\geq 1$ and let $S_{1},\dots,S_{c}$ be a collection of (possibly overlapping) subsets of $[n]$ such that $|S_{j}|\geq\tau$ for all $i$ and $\cup_{j=1}^{c}S_{j}=[n]$ . Moreover, let $q=(q_{1},\dots,q_{c})>0$ be a probability vector. Let $\hat{S}_{j}$ be $\tau$ -nice sampling from $S_{j}$ ; that is, $\hat{S}_{j}$ picks subsets of $S_{j}$ having cardinality $\tau$ , uniformly at random. We assume these samplings are independent. Now, $\hat{S}$ is generated as follows. We first pick $j\in\{1,\dots,c\}$ with probability $q_{j}$ , and then draw $\hat{S}_{j}$ . Note that we do not need to compute the quantities $p_{S}$ , $S\subseteq[n]$ , to execute ‘NSync. In fact, it is much easier to implement the sampling via the two-tier procedure explained above. Sampling $\hat{S}$ is a nonuniform variant of the $\tau$ -nice sampling studied in , which here arises as a special case for $c=1$ . Note that

where $\delta_{ij}=1$ if $i\in S_{j}$ , and otherwise.

Let Assumptions 4 and 5 be satisfied, and let $\hat{S}$ be the sampling described above. Then Assumption 1 is satisfied with $p$ given by (12) and any $w=(w_{1},\dots,w_{n})^{T}$ for which

where $\omega_{j}:=\max_{J\in\mathcal{J}}|J\cap S_{j}|\leq\omega$ .

Since $f$ is separable of degree $\omega$ , so is $\phi$ (because $\frac{1}{2}\|x\|_{v}^{2}$ is separable). Now,

where the last inequality follows from the ESO for $\tau$ -nice samplings established in [16, Theorem 15]. The claim now follows by comparing the above expression and (2). ∎

Optimal probabilities

Observe that formula (13) can be used to design a sampling (characterized by the sets $S_{j}$ and probabilities $q_{j}$ ) that minimizes $\Lambda$ , which in view of Theorem 3 optimizes the convergence rate of the method.

Serial setting. Consider the serial version of ‘NSync ( $\mathbf{Prob}(|\hat{S}|=1)=1$ ). We can model this via $c=n$ , with $S_{i}=\{i\}$ and $p_{i}=q_{i}$ for all $i\in[n]$ . In this case, using (12) and (13), we get $w_{i}=w_{i}^{*}=L_{i}+v_{i}$ . Minimizing $\Lambda$ in (4) over the probability vector $p$ gives the optimal probabilities (we refer to this as the optimal serial method) and optimal complexity

respectively. Note that the uniform sampling, $p_{i}=1/n$ for all $i$ , leads to $\Lambda_{US}:=n+n\max_{j}L_{j}/v_{j}$ (we call this the uniform serial method), which can be much larger than $\Lambda_{OS}$ . Moreover, under the change of variables $y=\operatorname*{Diag}(d)x$ , the gradient of $f^{d}(y):=f(\operatorname*{Diag}(d^{-1})y)$ has coordinate Lipschitz constants $L_{i}^{d}=L_{i}/d_{i}^{2}$ , while the weights in (11) change to $v_{i}^{d}=v_{i}/d_{i}^{2}$ . Hence, the condition numbers $L_{i}/v_{i}$ can not be improved via such a change of variables.

Optimal serial method can be faster than the fully parallel method. To model the fully parallel setting (i.e., the variant of ‘NSync updating all coordinates at every iteration), we can set $c=1$ and $\tau=n$ , which yields $\Lambda_{FP}=\omega+\omega\max_{j}L_{j}/v_{j}$ . Since $\omega\leq n$ , it is clear that $\Lambda_{US}\geq\Lambda_{FP}$ . However, for large enough $\omega$ it will be the case that $\Lambda_{FP}\geq\Lambda_{OS}$ , implying, surprisingly, that the optimal serial method can be faster than the fully parallel method.

Parallel setting. Fix $\tau$ and sets $S_{j}$ , $j=1,2,\dots,c$ , and define $\theta:=\max_{j}\left(1+\tfrac{(\tau-1)(\omega_{j}-1)}{\max\{1,|S_{j}|-1\}}\right)$ . Consider running ‘NSync with stepsizes $w_{i}=\theta(L_{i}+v_{i})$ (note that $w_{i}\geq w_{i}^{*}$ , so we are fine). From (4), (12) and (13) we see that the complexity of ‘NSync is determined by

The probability vector $q$ minimizing this quantity can be computed by solving a linear program with $c+1$ variables ( $q_{1},\dots,q_{c},\alpha$ ), $2n$ linear inequality constraints and a single linear equality constraint:

where $b^{i}\in\mathbf{R}^{c}$ , $i\in[n]$ , are given by $b^{i}_{j}=\tfrac{v_{i}}{(L_{i}+v_{i})}\tfrac{\delta_{ij}}{|S_{j}|}$ .

Experiments

We now conduct 2 preliminary small scale experiments to illustrate the theory; the results are depicted below. All experiments are with problems of the form (11) with $f$ chosen as in Example 6.

In the left plot we chose $A\in\mathbf{R}^{2\times 30}$ , $\gamma=1$ , $v_{1}=0.05$ , $v_{i}=1$ for $i\neq 1$ and $L_{i}=1$ for all $i$ . We compare the US method ( $p_{i}=1/n$ , blue) with the OS method ( $p_{i}$ given by (16), red). The dashed lines show 95% confidence intervals (we run the methods 100 times, the line in the middle is the average behavior). While OS can be faster, it is sensitive to over/under-estimation of the constants $L_{i},v_{i}$ . In the right plot we show that a nonuniform serial (NS) method can be faster than the fully parallel (FP) variant (we have chosen $m=8$ , $n=10$ and 3 values of $\omega$ ). On the horizontal axis we display the number of epochs, where 1 epoch corresponds to updating $n$ coordinates (for FP this is a single iteration, whereas for NS it corresponds to $n$ iterations).