Algorithms of Robust Stochastic Optimization Based on Mirror Descent Method

Anatoli Juditsky, Alexander Nazin, Arkadi Nemirovsky, Alexandre Tsybakov

Introduction

In this paper, we consider the problem of convex composite stochastic optimization:

where $X$ is a compact convex subset of a finite-dimensional real vector space $E$ with norm $\|\cdot\|$ , $\omega$ is a random variable on a probability space $\Omega$ with distribution $P$ , function $\psi$ is convex and continuous, and function $\Phi:\;X\times\Omega\to{\bf R}$ . Suppose that the expectation

is finite for all $x\in X$ , and is a convex and differentiable function of $x$ . Under these assumptions, the problem (1) has a solution with optimal value $F_{*}=\min_{x\in{X}}F(x)$ .

Assume that there is an oracle, which for any input $(x,\omega)\in X\times\Omega$ returns a stochastic gradient that is a vector $G(x,\omega)$ satisfying

where $\|\cdot\|_{*}$ is conjugate norm to $\|\cdot\|$ , and $\sigma>0$ is a constant. The aim of this paper is to construct $(1-\alpha)$ -reliable approximate solutions of the problem (1), i.e., solutions ${\widehat{x}}_{N}$ , based on $N$ queries of the oracle and satisfying the condition

with as small as possible $\delta_{N}(\alpha)>0$ .

Note that stochastic optimization problems of the form (1) arise in the context of penalized risk minimization, where the confidence bounds (3) are directly converted into confidence bounds for the accuracy of the obtained estimators. In this paper, the bounds (3) are derived with $\delta_{N}(\alpha)$ of order $\sqrt{\ln(1/\alpha)/N}$ . Such bounds are often called sub-Gaussian confidence bounds. Standard results on sub-Gaussian confidence bounds for stochastic optimization algorithms assume boundedness of exponential or subexponential moments of the stochastic noise of the oracle $G(x,\omega)-\nabla\phi(x)$ (cf. ). In the present paper, we propose robust stochastic algorithms that satisfy sub-Gaussian bounds of type (3) under a significantly less restrictive condition (2).

Recall that the notion of robustness of statistical decision procedures was introduced by J. Tukey and P. Huber in the $1960$ ies, which led to the subsequent development of robust stochastic approximation algorithms. In particular, in the 1970ies–1980ies, algorithms that are robust for wide classes of noise distributions were proposed for problems of stochastic optimization and parametric identification. Their asymptotic properties when the sample size increases have been well studied, see, for example, and references therein. An important contribution to the development of the robust approach was made by Ya.Z. Tsypkin. Thus, a significant place in the monographs is devoted to the study of iterative robust identification algorithms.

The interest in robust estimation resumed in the 2010ies due to the need to develop statistical procedures that are resistant to noise with heavy tails in high-dimensional problems. Some recent work develops the method of median of means for constructing estimates that satisfy sub-Gaussian confidence bounds for noise with heavy tails. Thus, in the median of means approach was used to construct an $(1-\alpha)$ -reliable version of stochastic approximation with averaging (“batch” algorithm) in a stochastic optimization setting similar to (1). Other original approaches were developed in , in particular, the geometric median techniques for robust estimation of signals and covariance matrices with sub-Gaussian guarantees . Also there was a renewal of interest in robust iterative algorithms. Thus, it was shown that robustness of stochastic approximation algorithms can be enhanced by using the geometric median of stochastic gradients . Another variant of the stochastic approximation procedure for calculating the geometric median was studied in , where a specific property of the problem (boundedness of the stochastic gradients) allowed the authors to construct $(1-\alpha)$ -reliable bounds under a very weak assumption about the tails of the noise distribution.

This paper discusses an approach to the construction of robust stochastic algorithms based on truncation of the stochastic gradients. It is shown that this method satisfies sub-Gaussian confidence bounds. In Sections 2 and 3, we define the main components of the optimization problem under consideration. In Section 4, we define the robust stochastic mirror descent algorithm and establish confidence bounds for it. Section 5 is devoted to robust accuracy estimates for general stochastic algorithms. Finally, Section 6 establishes robust confidence bounds for problems, in which $F$ has a quadratic growth. The Appendix contains the proofs of the results of the paper.

Notation and Definitions

Let $E$ be a finite-dimensional real vector space with norm $\|\cdot\|$ and let $E^{*}$ be the conjugate space to $E$ . Denote by $\langle s,x\rangle$ the value of linear function $s\in E^{*}$ at point $x\in E$ and by $\|\cdot\|_{*}$ the conjugate to norm $\|\cdot\|$ on $E^{*}$ , i.e.,

we consider a continuous convex function $\theta:B\to\textbf{R}$ with the following property:

where $\theta^{\prime}(\cdot)$ is a continuous in $B^{o}=\{x\in B:\,\partial\theta(x)\neq\emptyset\}$ version of the subgradient of $\theta(\cdot)$ and $\partial\theta(x)$ denotes the subdifferential of function $\theta(\cdot)$ at point $x$ , i.e., the set of all subgradients at this point. In other words, function $\theta(\cdot)$ is strongly convex on $B$ with coefficient 1 with respect to the norm $\|\cdot\|$ . We will call $\theta(\cdot)$ the normalized proxy function. Examples of such functions are:

$\theta(x)=\mbox{\small$ \frac{1}{2} $}\|x\|_{2}^{2}$ for $\left(E,\|\cdot\|\right)=\left(\textbf{R}^{n},\|\cdot\|_{2}\right)$ ;

$\theta(x)={2\rm{e}(\ln n)}\|x\|_{p}^{p}$ with $p=p(n):=1+{1\over 2\ln n}$ for $\left(E,\|\cdot\|\right)=\left(\textbf{R}^{n},\|\cdot\|_{1}\right)$ ;

$\theta(x)=4\rm{e}(\ln n)\sum_{i=1}^{n}|\lambda_{i}(x)|^{p}$ with $p=p(n)$ for $E=S_{n}$ , where $S_{n}$ is the space of symmetric $n\times n$ matrices equipped with the nuclear norm $\|x\|=\sum_{i=1}^{n}|\lambda_{i}(x)|$ and $\lambda_{i}(x)$ are eigenvalues of matrix $x$ .

Now, let $X$ be a convex compact subset in $E$ and let $x_{0}\in X$ and $R>0$ be such that $\max_{x\in X}\|x-x_{0}\|\leq R$ . We equip $X$ with a proxy function

Note that $\vartheta(\cdot)$ is strongly convex with coefficient 1 and

Let $D:=\max_{x,x^{\prime}\in X}\|x-x^{\prime}\|$ be the diameter of the set $X$ . Then $D\leq 2R$ .

In the following, we denote by $C$ and $C^{\prime}$ positive numerical constants, not necessarily the same in different cases.

Assumptions

Consider a convex composite stochastic optimisation problem (1) on a convex compact set $X\subset E$ . Assume in the following that the function

is convex on $X$ , differentiable at each point of the set $X$ and its gradient satisfies the Lipschitz condition

Assume also that function $\psi$ is convex and continuous. In what follows, we assume that we have at our disposal a stochastic oracle, which for any input $(x,\omega)\in X\times\Omega$ , returns a random vector $G(x,\omega)$ , satisfying the conditions (2). In addition, it is assumed that for any $a\in E^{*}$ and $\beta>0$ an exact solution of the minimization problem

with a constant $\upsilon\geq 0$ . This assumption is motivated as follows.

First, if we a priori know that the global minimum of function $\phi$ is attained at an interior point $x_{\phi}$ of the set $X$ (what is common in statistical applications of stochastic approximation), we have $\nabla\phi(x_{\phi})=0$ . Therefore, choosing $\bar{x}=x_{\phi}$ , one can put $g(\bar{x})=0$ and assumption (6) holds automatically with $\upsilon=0$ .

Second, in general, one can choose $\bar{x}$ as any point of the set $X$ and $g(\bar{x})$ as a geometric median of stochastic gradients $G(\bar{x},\omega_{i})$ , $i=1,\dots,m$ , over $m$ oracle queries. It follows from that if $m$ is of order $\ln\left(\varepsilon^{-1}\right)$ with some sufficiently small $\varepsilon>0$ , then

Thus, the confidence bounds obtained below will remain valid up to an $\varepsilon$ -correction in the probability of deviations.

Accuracy bounds for Algorithm RSMD

In what follows, we consider that the assumptions of Section 3 are fulfilled. Introduce a composite proximal transform

For $i=1,2,\dots$ , define the algorithm of Robust Stochastic Mirror Descent (RSMD) by the recursion

Here $\beta_{i}>0$ , $i=0,1,\dots$ , and $\lambda>0$ are tuning parameters that will be defined below, and $\omega_{1},\omega_{2},\dots$ are independent identically distributed (i.i.d.) realizations of a random variable $\omega$ , corresponding to the oracle queries at each step of the algorithm.

The approximate solution of problem (1) after $N$ iterations is defined as the weighted average

If the global minimum of function $\phi$ is attained at an interior point of the set $X$ and $\upsilon=0$ , then definition (12) is simplified. In this case, replacing $\|\bar{x}-x_{i-1}\|$ by the upper bound $D$ and putting $\upsilon=0$ and $g(\bar{x})=0$ in (12), we define the truncated stochastic gradient by the formula

The next result describes some useful properties of mirror descent recursion (9). Define

Let $\beta_{i}\geq 2L$ for all $i=0,1,...$ , and let $\widehat{x}_{N}$ be defined in (13), where $x_{i}$ are iterations (9) for any values $y_{i}$ , not necessarily given by (12). Then for any $z\in X$ we have

where $z_{i}$ is a random vector with values in $X$ depending only on $x_{0},\xi_{1},\dots,\xi_{i}$ .

Using Proposition 1 we obtain the following bounds on the expected error $F({\widehat{x}}_{N})-F_{*}$ of the approximate solution of problem (1) based on the RSMD algorithm. In what follows, we denote by ${\bf E}\{\cdot\}$ the expectation with respect to the distribution of $\omega^{N}=(\omega_{1},...,\omega_{N})\in\Omega^{\otimes N}$ .

Set $M=LR$ . Assume that $\lambda\geq\max\{M,\sigma\sqrt{N}\}+\upsilon\sigma$ and $\beta_{i}\geq 2L$ for all $i=0,1,...$ . Let ${\widehat{x}}_{N}$ be the approximate solution (13), where $x_{i}$ are the iterations of the RSMD algorithm defined by relations (9) and (12). Then

In particular, if $\beta_{i}=\bar{\beta}$ for all $i=0,1,...$ , where

Moreover, in this case we have the following inequality with explicit constants:

This result shows that if the truncation threshold $\lambda$ is large enough, then the expected error of the proposed algorithm is bounded similarly to the expected error of the standard mirror descent algorithm with averaging, i.e., the algorithm in which stochastic gradients are taken without truncation: $y_{i}=G(x_{i-1},\omega_{i})$ .

The following theorem gives confidence bounds for the proposed algorithm.

Let $\beta_{i}=\bar{\beta}\geq 2L$ for all $i=0,1,...$ , and let $1\leq\tau\leq N/\upsilon^{2}$ ,

Let ${\widehat{x}}_{N}$ be the approximate solution (13), where $x_{i}$ are the RSMD iterations defined by relations (9) and (12). Then there is a random event ${\cal A}_{N}\subset\Omega^{\otimes N}$ of probability at least $1-2e^{-\tau}$ such that for all $\omega^{N}\in{\cal A}_{N}$ the following inequalities hold:

In paticular, chosing $\bar{\beta}$ as in formula (18) we have, for all $\omega^{N}\in{\cal A}_{N}$ ,

where $C_{1}>0$ and $C_{2}>0$ are numerical constants.

The values of the numerical constants $C_{1}$ and $C_{2}$ in (21) can be obtained from the proof of the theorem, cf. the bound in (51).

Confidence bound (21) in Theorem 1 contains two terms corresponding to the deterministic error and to the stochastic error. Unlike the case of noise with a “light tail” (see, for example, ) and the bound in expectation (19), the deterministic error $LR^{2}[\tau\vee\Theta]/N$ depends on $\tau$ . Note also that Theorem 1 gives a sub-Gaussian confidence bound (the order of the stochastic error is $\sigma R\sqrt{[\tau\vee\Theta]/N}$ ). However, the truncation threshold $\lambda$ depends on the confidence level $\tau$ . This can be inconvenient for the implementation of the algorithms. Some simple but coarser confidence bounds can be obtained by using a universal threshold independent of $\tau$ , which is $\lambda=\max\{\sigma\sqrt{N},M\}+\upsilon\sigma$ . In particular, we have the following result.

Let $\beta_{i}=\bar{\beta}\geq 2L$ for all $i=0,1,...$ , and let $N\geq\upsilon^{2}$ . Set

Let ${\widehat{x}}_{N}=N^{-1}\sum_{i=1}^{N}x_{i}$ , where $x_{i}$ are the iterations of the RSMD algorithm defined by relations (9) and (12). Then there is a random event ${\cal A}_{N}\subset\Omega^{\otimes N}$ of probability at least $1-2e^{-\tau}$ such that for all $\omega^{N}\in{\cal A}_{N}$ the following inequalities hold:

In particular, choosing $\bar{\beta}$ as in formula (18) we have

The values of the numerical constants $C$ in Theorem 2 can be obtained from the proof, cf. the bound in (51).

Robust Confidence Bounds for Stochastic Optimization Methods

Consider an arbitrary algorithm for solving the problem (1) based on $N$ queries of the stochastic oracle. Assume that we have a sequence $\big{(}x_{i},G(x_{i},\omega_{i+1})\big{)},\;i=0,...,N$ , where $x_{i}\in X$ are the search points of some stochastic algorithm and $G(x_{i},\omega_{i+1})$ are the corresponding observations of the stochastic gradient. It is assumed that $x_{i}$ depends only on $\{(x_{j-1},\omega_{j}),j=1,\dots,i\}$ . The approximate solution of the problem (1) is defined in the form:

Our goal is to construct a confidence interval with sub-Gaussian accuracy for $F({\widehat{x}}_{N})-F_{*}$ . To do this, we use the following fact. Note that for any $t\geq L$ the value

is an upper bound on the accuracy of the approximate solution ${\widehat{x}}_{N}$ :

(see Lemma 1 in Appendix). This fact is true for any sequence of points $x_{0},\dots,x_{N}$ in $X$ , regardless of how they are obtained. However, since the function $\nabla\phi(\cdot)$ is not known, the estimate (24) cannot be used in practice. Replacing the gradients $\nabla\phi(x_{i-1})$ in (23) with their truncated estimates $y_{i}$ defined in (12) we get an implementable analogue of $\epsilon_{N}(t)$ :

Note that computing ${\widehat{\epsilon}}_{N}(t)$ reduces to solving a problem of the form (4) with $\beta=0$ . Thus, it is computationally not more complex than, for example, one step of the RSMD algorithm. Replacing $\nabla\phi(x_{i-1})$ with $y_{i}$ introduces a random error. In order to get a reliable upper bound for $\epsilon_{N}(t)$ , we need to compensate this error by slightly increasing ${\widehat{\epsilon}}_{N}(t)$ . Specifically, we add to ${\widehat{\epsilon}}_{N}(t)$ the value

Let $\big{(}x_{i},G(x_{i},\omega_{i+1})\big{)}_{i=0}^{N}$ be the trajectory of a stochastic algorithm for which $x_{i}$ depends only on $\{(x_{j-1},\omega_{j}),j=1,\dots,i\}$ . Let $0<\tau\leq N/\upsilon^{2}$ and let $y_{i}=y_{i}(\tau)$ be truncated stochastic gradients defined in (12), where the threshold $\lambda=\lambda(\tau)$ is chosen in the form (20). Then for any $t\geq L$ the value

Since $\Delta_{N}(\tau,t)$ monotonically increases in $t$ it suffices to use this bound for $t=L$ when $L$ is known. Note that, although $\Delta_{N}(\tau,t)$ gives an upper bound for $\epsilon_{N}(t)$ , Proposition 2 does not guarantee that $\Delta_{N}(\tau,t)$ is sufficiently close to $\epsilon_{N}(t)$ . However, this property holds for the RSMD algorithm with a constant step, as follows from the next result.

Under the conditions of Proposition 2, let the vectors $x_{0},\dots,x_{N}$ be given by the RSMD recursion (9)–(12), where $\beta_{i}=\bar{\beta}\geq 2L$ , $i=0,...,N-1$ . Then

Moreover, if $\bar{\beta}\geq\max\left\{2L,{\sigma\sqrt{N}\over R\sqrt{\Theta}}\right\}$ then

where $C_{3}>0$ and $C_{4}>0$ are numerical constants.

The values of the numerical constants $C_{3}$ and $C_{4}$ can be derived from the proof of this corollary.

Robust Confidence Bounds for Quadratic Growth Problems

In this section, it is assumed that $F$ is a function with quadratic growth on $X$ in the following sense (cf. ). Let $F$ be a continuous function on $X$ and let $X_{*}\subset X$ be the set of its minimizers on $X$ . Then $F$ is called a function with quadratic growth on $X$ if there is a constant $\kappa>0$ such that for any $x\in X$ there exists $\bar{x}(x)\in X_{*}$ such that the following inequality holds:

The RSMD algorithm for quadratically growing functions will be defined in stages. At each stage, for specially selected $r>0$ and $y\in X$ it solves an auxiliary problem

We initialize the algorithm by choosing arbitrary $y_{0}=x_{0}\in X$ and $r_{0}\geq\max_{z\in X}\|z-x_{0}\|$ . We set $r^{2}_{k}=2^{-k}r^{2}_{0}$ , $k=1,2,...$ . Let $C_{1}$ and $C_{2}$ be the numerical constants in the bound (21) of Theorem 1. For a given parameter $\tau>0$ , and $k=1,2,\dots$ we define the values

Here $\rfloor t\lfloor$ denotes the smallest integer greater than or equal to $t$ . Set

Now, let $k\in\{1,2,\dots,m(N)\}$ . At the $k$ -th stage of the algorithm, we solve the problem of minimization of $F$ on the ball $X_{r_{k-1}}(y_{k-1})$ , we find its approximate solution ${\widehat{x}}_{N_{k}}$ according to (9)–(13), where we replace $x_{0}$ by $y_{k-1}$ , $X$ by $X_{r_{k-1}}(y_{k-1})$ , $R$ by $r_{k-1}$ , $N$ by $N_{k}$ , and set

It is assumed that, at each stage $k$ of the algorithm, an exact solution of the minimization problem

is available for any $a\in E$ and $\beta>0$ . At the output of the $k$ -th stage of the algorithm, we obtain $y_{k}:={\widehat{x}}_{N_{k}}$ .

Assume that $m(N)\geq 1$ , i.e. at least one stage of the algorithm described above is completed. Then there is a random event ${\cal B}_{N}\subset\Omega^{\otimes N}$ of probability at least $1-2m(N)e^{-\tau}$ such that for $\omega^{N}\in{\cal B}_{N}$ the approximate solution $y_{m(N)}$ after $m(N)$ stages of the algorithm satisfies the inequality

Theorem 3 shows that, for functions with quadratic growth, the deterministic error component can be significantly reduced – it becomes exponentially decreasing in $N$ . The stochastic error component is also significantly reduced. Note that the factor $m(N)$ is of logarithmic order and has little effect on the probability of deviations. Indeed, it follows from (29) that $m(N)\leq C\ln\left(\frac{C^{\prime}\kappa^{2}r_{0}^{2}N}{\sigma^{2}(\tau\vee\Theta)}\right)$ . Neglecting this factor in the probability of deviations and considering the stochastic component of the error, we see that the confidence bound of Theorem 3 is approximately sub-exponential rather than sub-Gaussian.

Conclusion

We have considered algorithms of smooth stochastic optimization when the distribution of noise in observations has heavy tails. It is shown that by truncating the observed gradients with a suitable threshold one can construct confidence sets for the approximate solutions that are similar to those in the case of “light tails”. It should be noted that the order of the deterministic error in the obtained bounds is suboptimal — it is substantially greater than the optimal rates achieved by the accelerated algorithms , namely, $O(LR^{2}N^{-2})$ in the case of convex objective function and $O(\exp(-N\sqrt{\kappa/L}))$ in the strongly convex case. On the other hand, the proposed approach cannot be used to obtain robust versions of the accelerated algorithms since applying it to such algorithms leads to accumulation of the bias caused by the truncation of the gradients. The problem of constructing accelerated robust stochastic algorithms with optimal guarantees remains open.

A.1. Preliminary remarks. We start with the following known result.

Assume that $\phi$ and $\psi$ satisfy the assumptions of Section 3, and let $x_{0},\dots,x_{N}$ be some points of the set $X$ . Define

Then for any $z\in X$ the following inequality holds:

In addition, for $\widehat{x}_{N}={1\over N}\sum_{i=1}^{N}x_{i}$ we have

Proof Using the property $V_{x}(z)\geq\mbox{\small$ \frac{1}{2} $}\|x-z\|^{2}$ , the convexity of functions $\phi$ and $\psi$ and the Lipschitz condition on $\nabla\phi$ we get that, for any $z\in X$ ,

Summing up over $i$ from 0 to $N-1$ and using the convexity of $F$ we obtain the second result of the lemma. $\square$

In what follows, we denote by ${\bf E}_{x_{i}}\{\cdot\}$ the conditional expectation for fixed $x_{i}$ .

Let the assumptions of Section 3 be fulfilled and let $x_{i}$ and $y_{i}$ satisfy the RSMD recursion, cf. (9) and (12). Then

Proof Set $\chi_{i}=1_{\|G(x_{i-1},\omega_{i})-g(\bar{x})\|_{*}>L\|x_{i-1}-\bar{x}\|+\lambda+\upsilon\sigma}$ . Note that by construction $\chi_{i}\leq\eta_{i}:=1_{\|G(x_{i-1},\omega_{i})-\nabla f(x_{i-1}\|_{*}>\lambda}$ . We have

Moreover, since ${\bf E}_{x_{i-1}}\{G(x_{i-1},\omega_{i})\}=\nabla\phi(x_{i-1})$ we have

The following lemma gives bounds for the deviations of the sums $\sum_{i}\langle\xi_{i},x_{i-1}-z\rangle$ and $\sum_{i}\|\xi_{i}\|_{*}^{2}$ .

Let the assumptions of Section 3 be fulfilled and let $x_{i}$ and $y_{i}$ satisfy the recursion of RSMD, cf. (9) and (12).

(i) If $\tau\leq{N/\upsilon^{2}}$ and $\lambda=\max\left\{\sigma\sqrt{N\over\tau},M\right\}+\upsilon\sigma$ then, for any $z\in X$ ,

(ii) If $N\geq\upsilon^{2}$ and $\lambda=\max\left\{\sigma\sqrt{N},M\right\}+\upsilon\sigma$ then, for any $z\in X$ ,

Proof Set $\zeta_{i}=\langle\xi_{i},z-x_{i-1}\rangle$ and $\varsigma_{i}=\|\xi_{i}\|_{*}^{2}$ , $i=1,2,\dots$ Using Lemma 2 it is easy to check that the following inequalities are fulfilled

In what follows, we apply several times the Bernstein inequality, and each time we will use the same notation $r$ , $A$ , $s$ for the values that are, respectively, the uniform upper bound of the expectation, the maximum absolute value, and the standard deviation of a random variable.

1o. We first prove the statement $(i)$ . We start with the case $M\leq\sigma\sqrt{N\over\tau}$ . It follows from (39) that in this case

Using (47) and Bernstein’s inequality for martingales (see, for example, ) we get

for all $\tau>0$ satisfying the condition $\tau\leq{16N/(9\upsilon^{2})}$ . On the other hand, in the case under consideration, the following inequalities hold (cf. (43) and (47))

for $0<\tau\leq{N/\upsilon^{2}}$ . Applying again the Bernstein inequality, we get

for all $\tau>0$ satisfying the condition $\tau\leq N/\upsilon^{2}$ .

2o. Assume now that $M>\sigma\sqrt{N\over\tau}$ , so that $\lambda=M+\upsilon\sigma$ and $\sigma^{2}\leq{M^{2}\tau/N}$ . Then

and applying again the Bernstein inequality we get

for all $\tau>0$ , satisfying the condition $\tau\leq 16N/(9\upsilon^{2})$ . Next, in this case

for $\tau\leq{N/\upsilon^{2}}$ . Applying once again the Bernstein inequality we get

for all $\tau>0$ satisfying the condition $\tau\leq N/\upsilon^{2}$ .

3o. Now, consider the case $\lambda=\max\{M,\sigma\sqrt{N}\}+\sigma\upsilon$ . Let $M\leq\sigma\sqrt{N}$ , so that $\lambda=\sigma(\sqrt{N}+\upsilon)$ . We argue in the same way as in the proof of $(i)$ . By virtue of (39) we have

Hence, using the Bernstein inequality we get

Now, applying again the Bernstein inequality we get

Proofs of the bounds (34) and (35) in the case $M>\sigma\sqrt{N}$ and $\lambda=M+\sigma\upsilon$ follow the same lines. $\square$

A.2. Proof of Proposition 1. We first prove inequality (15). In view of (4), the optimality condition for (9) has the form

where the last equality follows from the following remarkable identity (see, for example, ): for any $u,u^{\prime}$ and $w\in X$

Since, by definition, $\xi_{i}=y_{i}-\nabla\phi(x_{i-1})$ we get

It follows from Lemma 1 and the condition $\beta_{i}\geq 2L$ that

Together with (48), this inequality implies

On the other hand, due to the strong convexity of $V_{x}(\cdot)$ we have

for all $z\in X$ . Dividing (49) by $\beta_{i}$ and taking the sum over $i$ from to $N-1$ we obtain (15).

We now prove the bound (16). Applying Lemma 6.1 of with $z_{0}=x_{0}$ we get

where $z_{i}={\rm arg\,min}_{z\in X}\big{\{}\mu_{i-1}\langle\xi_{i},z\rangle+V_{z_{i-1}}(z)\big{\}}$ depends only on $z_{0},\xi_{1},\dots,\xi_{i}$ . Further,

Combining this inequality with (15), we get (16). $\square$

A.3. Proof of Corollary 1. Note that (17) is an immediate consequence of (15) and of the bounds for the moments of $\|\xi_{i}\|_{*}$ given in Lemma 2. Indeed, (31)(b) implies that, under the conditions of Corollary 1,

Taking the expectation of both sides of (15) and using the last two inequalities we get (17). The bound (19) is proved in a similar way, with the only difference that instead of inequality (15) we use (16). $\square$

A.4. Proof of Theorem 1. By virtue of part (i) of Lemma 3, under the condition $\tau\leq N/\upsilon^{2}$ we have that, with probability of at least $1-2e^{-\tau}$ ,

Plugging these bounds in (16) we obtain that, with probability at least $1-2e^{-\tau}$ , the following holds:

Next, taking $\bar{\beta}=\max\big{\{}2L,{\sigma\over R}\sqrt{N\over\Theta}\big{\}}$ we get

for $1\leq\tau\leq N/\upsilon^{2}$ . This implies (21). $\square$

A.5. Proof of Theorem 2. We act in the same way as in the proof of Theorem 1 with the only difference that instead of part (i) of Lemma 3 we use part (ii) of that lemma, which implies that if $N\geq\upsilon^{2}$ then with probability at least $1-2e^{-\tau}$ the following inequalities hold:

The proposition is a direct consequence of the following result.

Then, for $0<\tau\leq N/\upsilon^{2}$ and $t\geq L$ the following inequalities hold

Proof of Lemma . Let us prove the first inequality in (55). Recall that $\xi_{i}=y_{i}-\nabla\phi(x_{i-1})$ , $i=1,...,N$ . Due to the strong convexity of $V_{x}(\cdot)$ , for any $z\in X$ and $\mu>0$ we have

Recalling that $V_{x_{0}}(z)\leq R^{2}\Theta$ , we conclude that $\sum_{i=1}^{N}\langle\xi_{i},z-x_{i}\rangle\leq\rho_{N}(\tau;\mu,\nu)$ for all $z\in X$ and all $\omega^{N}\in{\cal A}_{N}$ . Therefore, for $\omega^{N}\in{\cal A}_{N}$ we have

which proves the first inequality in (55). The proof of the second inequality in (55) is similar and therefore it is omitted. $\square$

A.7. Proof of Corollary 2. From the definition of $\epsilon_{N}(\cdot)$ we deduce that

and we get (26) by taking $\mu=1/\bar{\beta}$ . On the other hand, one can check that for $\bar{\beta}\geq\max\left\{2L,{\sigma\sqrt{N}\over R\sqrt{\Theta}}\right\}$ the following inequalities hold:

with the same probability. This implies (27). $\square$

1o. We first show that for each $k=1,\dots,m=m(N)$ , the following is true.

Fact ${I_{k}}$ . There is a random event ${\cal B}_{k}\subseteq\Omega^{\otimes N}$ of probability at least $1-2ke^{-\tau}$ such that for all $\omega^{N}\in{\cal B}_{k}$ the following inequalities hold:

The proof of Fact ${I_{k}}$ is carried out by induction. Note that (58)(a) holds with probability 1 for $k=0$ . Set ${\cal B}_{0}=\Omega^{\otimes N}$ . Assume that (58)(a) holds for some $k\in\{0,\dots,m-1\}$ with probability at least $1-2ke^{-\tau}$ , and let us show that then Fact ${I_{k+1}}$ is true.

Define $F_{*}^{k}=\min_{x\in X_{r_{k}}(y_{k})}F(x)$ and let $X_{*}^{k}$ be the set of all minimizers of function $F$ on $X_{r_{k}}(y_{k})$ . By Theorem 1 and the definition of $N_{k}$ (cf. (29)), there is an event ${\cal A}_{k}$ of probability at least $1-2e^{-\tau}$ such that for $\omega^{N}\in{\cal A}_{k}$ after the $(k+1)$ -th stage of the algorithm we have

where $\bar{x}_{k}(y_{k+1})$ is the projection of $y_{k+1}$ onto $X_{*}^{k}$ . Set ${\cal B}_{k+1}={\cal B}_{k}\cap{\cal A}_{k}$ . Then

In addition, due to the assumption of induction, on the set ${\cal B}_{k}$ (and, therefore, on ${\cal B}_{k+1}$ ) we have

i.e., the distance between $y_{k}$ and the set $X_{*}$ of global minimizers does not exceed $r_{k}$ . Therefore, the set $X_{r_{k}}(y_{k})$ has a non-empty intersection with $X_{*}$ . Thus, $X^{k}_{*}\subseteq X_{*}$ , the point $\bar{x}_{k}(y_{k+1})$ is contained in $X_{*}$ and $F_{*}^{k}$ coincides with the optimal value $F_{*}$ of the initial problem. We conclude that

2o. We now prove the theorem in the case ${\overline{N}}_{1}\geq 1$ . This condition is equivalent to the fact that ${\overline{N}}_{k}\geq 1$ for all $k=1,\dots,m(N)$ , since ${\overline{N}}_{1}\leq{\overline{N}}_{2}\leq\cdots\leq{\overline{N}}_{m(N)}$ by construction. Assume that $\omega^{N}\in{\cal B}_{m(N)}$ , so that (58) holds with $k=m(N)$ . Since ${\overline{N}}_{1}\geq 1$ we have $N_{k}\leq 2{\overline{N}}_{k}$ . In addition, ${\overline{N}}_{k+1}\leq 2{\overline{N}}_{k}$ . Using these remarks and the definition of $m(N)$ we get

Thus, using the definition of ${\overline{N}}_{k}$ (cf. (29)) we obtain

Two cases are possible: $S_{1}\geq N/2$ or $S_{2}\geq N/2$ . If $S_{1}\geq N/2$ , then

so that if $\omega^{N}\in{\cal B}_{m(N)}$ then

If $S_{2}\geq N/2$ the following inequalities hold:

Therefore, in this case for $\omega^{N}\in{\cal B}_{m(N)}$ we have

Let $k_{*}\geq 2$ be the smallest integer $k$ such that ${\overline{N}}_{k}\geq 1$ . If $k_{*}>N/4$ it is not difficult to see that $m(N)\geq N/4$ and therefore for $\omega^{N}\in{\cal B}_{m(N)}$ we have

If $2\leq k_{*}\leq N/4$ we have the following chain of inequalities:

where the first inequality uses the fact that ${\overline{N}}_{m(N)+1}\leq 2{\overline{N}}_{m(N)}$ and the last inequality follows from the definition of $m(N)$ . Based on this remark and on the fact that ${\overline{N}}_{k}/2^{k}\leq{\overline{N}}_{k_{*}}/2^{k_{*}}$ for $k\geq k_{*}$ we obtain

where the last inequality follows by noticing that ${\overline{N}}_{k_{*}}\leq 2{\overline{N}}_{k_{*}-1}<2$ . Hence, taking into account (58)(b) we get that, for $\omega^{N}\in{\cal B}_{m(N)}$ ,

Combining this bound with (60), (61) and (63) we get (30). $\square$