Halpern Iteration for Near-Optimal and Parameter-Free Monotone Inclusion and Strong Solutions to Variational Inequalities

Jelena Diakonikolas

Introduction

the monotone inclusion problem consists in finding a point $\mathbf{u}^{*}$ that satisfies:

Monotone inclusion is a fundamental problem in continuous optimization that is closely related to variational inequalities (VIs) with monotone operators, which model a plethora of problems in mathematical programming, game theory, engineering, and finance (Facchinei and Pang, 2003, Section 1.4). Within machine learning, VIs with monotone operators and associated monotone inclusion problems arise, for example, as an abstraction of convex-concave min-max optimization problems, which naturally model adversarial training (Madry et al., 2018; Arjovsky et al., 2017; Arjovsky and Bottou, 2017; Goodfellow et al., 2014).

On the other hand, approximate monotone inclusion is well-defined even for unbounded feasible sets. In the context of min-max optimization, it corresponds to guarantees in terms of stationarity. Specifically, in the unconstrained setting, solving monotone inclusion corresponds to minimizing the norm of the gradient of $\Phi.$ Note that even in the special setting of convex optimization, convergence in norm of the gradient is much less understood than convergence in optimality gap (Nesterov, 2012; Kim and Fessler, 2018). Further, unlike classical results for VIs that provide convergence guarantees for approximating weak solutions (Nemirovski, 2004; Nesterov, 2007), approximations to monotone inclusion lead to approximations to strong solutions (see Section 1.2 for definitions of weak and strong solutions and their relationship to monotone inclusion).

We leverage the connections between nonexpansive maps, structured monotone operators, and proximal maps to obtain near-optimal algorithms for solving monotone inclusion over different classes of problems with Lipschitz-continuous operators. In particular, we make use of the classical Halpern iteration, which is defined by (Halpern, 1967):

In addition to its simplicity, Halpern iteration is particularly relevant to machine learning applications, as it is an implicitly regularized method with the following property: if the set of fixed points of $T$ is non-empty, then Halpern iteration (Hal) started at a point $\mathbf{u}_{0}$ and applied with any choice of step sizes $\{\lambda_{k}\}_{k\geq 1}$ that satisfy all of the following conditions:

A special case of what is now known as the Halpern iteration (Hal) was introduced and its asymptotic convergence properties were analyzed by Halpern (1967) in the setting of $\mathbf{u}_{0}=\textbf{0}$ and $T:\mathcal{B}_{2}\to\mathcal{B}_{2},$ where $\mathcal{B}_{2}$ is the unit Euclidean ball. Using the proof-theoretic techniques of Kohlenbach (2008), Leustean (2007) extracted from the asymptotic convergence result of Wittmann (1992) the rate at which Halpern iteration converges to a fixed point. The results obtained by Leustean (2007) are rather loose and provide guarantees of the form $\|T(\mathbf{u}_{k})-\mathbf{u}_{k}\|=O(\frac{M}{\log(k)})$ in the best case (obtained for $\lambda_{k}=\Theta(\frac{1}{k})$ ), where $M\geq\|\mathbf{u}_{0}\|+\|T(\mathbf{u}_{0})\|+\|\mathbf{u}_{k}\|,$ $\forall k.$ A tighter result that shows that $\|T(\mathbf{u}_{k})-\mathbf{u}_{k}\|$ decreases at rate that is at least as good as $1/\sqrt{k}$ was obtained by Kohlenbach (2011). The results of Leustean (2007) and Kohlenbach (2011) apply to general normed spaces. The work of Kohlenbach (2011) also provided an explicit rate of metastability that characterizes the convergence of the sequence of iterates $\{\mathbf{u}_{k}\}$ in Hilbert spaces.

More recently, Lieder (2017) proved that under the standard assumption that $T$ has a fixed point $\mathbf{u}^{*}$ and for the step size $\lambda_{k}=\frac{1}{k+1},$ Halpern iteration converges to a fixed point as $\|T(\mathbf{u}_{k})-\mathbf{u}_{k}\|=\frac{2\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{k+1}.$ A similar result but for an alternative algorithm was recently obtained by Kim (2019). These two results (as well as all the results from this paper) only apply to Hilbert spaces. Unlike Halpern iteration, the algorithm introduced by Kim (2019) is not known to possess the implicit regularization property discussed earlier in this paper. The results of Lieder (2017) and Kim (2019) can be used to obtain the same $1/k$ convergence rate for monotone inclusion with a cocoercive operator but only if the cocoercivity parameter is known, which is rarely the case in practice. Similarly, those results can also be extended to more general monotone Lipschitz operators but only if the proximal map (or resolvent) of $F$ can be computed exactly, an assumption that can rarely be met (see Section 1.2 for definitions of cocoercive operators and proximal maps). We also note that the results of Lieder (2017) and Kim (2019) were obtained using the performance estimation (PEP) framework of Drori and Teboulle (2014). The convergence proofs resulting from the use of PEP are computer-assisted: they are generated as solutions to large semidefinite programs, which typically makes them hard to interpret and generalize.

Our approach is arguably simpler, as it relies on the use of a potential function, which allows us to remove the assumptions about the knowledge of the problem parameters and availability of exact proximal maps. Our main contributions are summarized as follows:

We introduce a new, potential-based, proof of convergence of Halpern iteration that applies to more general step sizes $\lambda_{k}$ than handled by the analysis of Lieder (2017) (Section 2). The proof is simple and only requires elementary algebra. Further, the proof is derived for cocoercive operators and leads to a parameter-free algorithm for monotone inclusion. We also extend this parameter-free method to the constrained setting using the concept of gradient mapping generalized to monotone operators (Section 2.1). To the best of our knowledge, this is the first work to obtain the $1/k$ convergence rate with a parameter-free method.

Results for monotone Lipschitz operators.

Up to a logarithmic factor, we obtain the same $1/k$ convergence rate for the parameter-free setting of the more general monotone Lipschitz operators (Section 2.2). The best known convergence rate established by previous work for the same setting was of the order $1/\sqrt{k}$ (Dang and Lan, 2015; Ryu et al., 2019). We obtain the improved convergence rate through the use of the Halpern iteration with inexact proximal maps that can be implemented efficiently. The idea of coupling inexact proximal maps with another method is similar in spirit to the Catalyst framework (Lin et al., 2017) and other instantiations of the inexact proximal-point method, such as, e.g., in the work of Davis and Drusvyatskiy (2019); Asi and Duchi (2019); Lin et al. (2018). However, we note that, unlike in the previous work, the coupling used here is with a method (Halpern iteration) whose convergence properties were not well-understood and for which no simple potential-based convergence proof existed prior to our work.

Results for strongly monotone Lipschitz operators.

We show that a simple restarting-based approach applied to our method for operators that are only monotone and Lipschitz (described above) leads to a parameter-free method for strongly monotone and Lipschitz operators (Section 2.3). Under mild assumptions about the problem parameters and up to a poly-logarithmic factor, the resulting algorithm is iteration-complexity-optimal. To the best of our knowledge, this is the first near-optimal parameter-free method for the setting of strongly monotone Lipschitz operators and any of the associated problems – monotone inclusion, VIs, or convex-concave min-max optimization.

Lower bounds.

To certify near-optimality of the analyzed methods, we provide lower bounds that rely on algorithmic reductions between different problem classes and highlight connections between them (Section 3). The lower bounds are derived by leveraging the recent lower bound of Ouyang and Xu (2019) for approximating the optimality gap in convex-concave min-max optimization.

2 Notation and Preliminaries

Let $\mathcal{U}\subseteq E$ be closed and convex, and let $F:E\rightarrow E$ be an $L$ -Lipschitz-continuous operator defined on $\mathcal{U}.$ Namely, we assume that:

The definition of monotonicity was already provided in Eq. (1.1), and easily specializes to monotonicity on the set $\mathcal{U}$ by restricting $\mathbf{u},\mathbf{v}$ to be from $\mathcal{U}.$ Further, $F$ is said to be:

strongly monotone (or coercive) on $\mathcal{U}$ with parameter $m$ , if:

cocoercive on $\mathcal{U}$ with parameter $\gamma$ , if:

It is immediate from the definition of cocoercivity that every $\gamma$ -cocoercive operator is monotone and $1/\gamma$ -Lipschitz. The latter follows by applying the Cauchy-Schwarz inequality to the left-hand side of Eq. (1.5) and then dividing both sides by $\gamma\|F(\mathbf{u})-F(\mathbf{v})\|$ .

Examples of monotone operators include the gradient of a convex function and appropriately modified gradient of a convex-concave function. Namely, if a function $\Phi(\mathbf{x},\mathbf{y})$ is convex in $\mathbf{x}$ and concave in $\mathbf{y},$ then $F([\vbox{\Let@\restore@math@cr\default@tag\halign{\hfil$ \m@th\scriptstyle# $&$ \m@th\scriptstyle{}# $\hfil\cr\mathbf{x}\\ \mathbf{y}\crcr}}])=[\vbox{\Let@\restore@math@cr\default@tag\halign{\hfil$ \m@th\scriptstyle# $&$ \m@th\scriptstyle{}# $\hfil\cr\nabla_{\mathbf{x}}\Phi(\mathbf{x},\mathbf{y})\\ -\nabla_{\mathbf{y}}\Phi(\mathbf{x},\mathbf{y})\crcr}}]$ is monotone.

The Stampacchia Variational Inequality (SVI) problem consists in finding $\mathbf{u}^{*}\in\mathcal{U}$ such that:

In this case, $\mathbf{u}^{*}$ is also referred to as a strong solution to the variational inequality (VI) corresponding to $F$ and $\mathcal{U}$ . The Minty Variational Inequality (MVI) problem consists in finding $\mathbf{u}^{*}$ such that:

in which case $\mathbf{u}^{*}$ is referred to as a weak solution to the variational inequality corresponding to $F$ and $\mathcal{U}$ . In general, if $F$ is continuous, then the solutions to (MVI) are a subset of the solutions to (SVI). If we assume that $F$ is monotone, then (1.1) implies that every solution to (SVI) is also a solution to (MVI), and thus the two solution sets are equivalent. The solution set to monotone inclusion is the same as the solution set to (SVI).

Approximate versions of variational inequality problems (SVI) and (MVI) are defined as follows: Given $\epsilon>0,$ find an $\epsilon$ -approximate solution $\mathbf{u}^{*}_{\epsilon}\in\mathcal{U},$ which is a solution that satisfies:

Clearly, when $F$ is monotone, an $\epsilon$ -approximate solution to (SVI) is also an $\epsilon$ -approximate solution to (MVI); the reverse does not hold in general.

Similarly, $\epsilon$ -approximate monotone inclusion can be defined as fidning $\mathbf{u}^{*}_{\epsilon}$ that satisfies:

where $\mathcal{B}(\epsilon)$ is the ball w.r.t. $\|\cdot\|$ , centered at 0 and of radius $\epsilon.$ We will sometimes write Eq. (1.6) in the equivalent form $-F(\mathbf{u}^{*}_{\epsilon})\in\partial I_{\mathcal{U}}(\mathbf{u}^{*}_{\epsilon})+\mathcal{B}(\epsilon).$ The following fact is immediate from Eq. (1.6).

Given $F$ and $\mathcal{U},$ let $\mathbf{u}^{*}_{\epsilon}$ satisfy Eq. (1.6). Then:

where $\mathcal{B}_{\mathbf{u}^{*}_{\epsilon}}$ denotes the unit ball w.r.t. $\|\cdot\|,$ centered at $\mathcal{B}_{\mathbf{u}^{*}_{\epsilon}}.$

Further, if the diameter of $\mathcal{U}$ , $D=\sup_{\mathbf{u},\mathbf{v}\in\mathcal{U}}\|\mathbf{u}-\mathbf{v}\|$ , is bounded, then:

Thus, when the diameter $D$ is bounded, any $\frac{\epsilon}{D}$ -approximate solution to monotone inclusion is an $\epsilon$ -approximate solution to (SVI) (and thus also to (MVI)); the converse does not hold in general. Recall that when $D$ is unbounded, neither (SVI) nor (MVI) can be approximated.

Nonexpansive Maps.

Let $T:E\to E$ . We say that $T$ is nonexpansive on $\mathcal{U}\subseteq E$ , if $\forall\mathbf{u},\mathbf{v}\in\mathcal{U}:$

Nonexpansive maps are closely related to cocoercive operators, and here we summarize some of the basic properties that are used in our analysis. More information can be found in, e.g., the book by Bauschke and Combettes (2011).

$T$ is said to be firmly nonexpansive or averaged, if $\forall\mathbf{u},\mathbf{v}\in\mathcal{U}:$

Useful properties of firmly nonexpansive maps are summarized in the following fact.

Halpern Iteration for Monotone Inclusion and Variational Inequalities

Halpern iteration is typically stated for nonexpansive maps $T$ as in (Hal). Because our interest is in cocoercive operators $F$ with the unknown parameter $1/L,$ we instead work with the following version of the Halpern iteration:

where $L_{k}\in(0,\infty),\,\forall k.$ If $L$ was known, we could simply set $L_{k+1}=L,$ in which case (H) would be equivalent to the standard Halpern iteration, due to Fact 1.2. We assume throughout that $\lambda_{1}=\frac{1}{2}.$

We start with the assumption that the setting is unconstrained: $\mathcal{U}\equiv E.$ We will see in Section 2.1 how the result can be extended to the constrained case. Section 2.2 will consider the case of operators that are monotone and Lipschitz, while Section 2.3 will deal with the strongly monotone and Lipschitz case. Some of the proofs are omitted and are instead provided in Appendix A.

To analyze the convergence of (H) for the appropriate choices of sequences $\{\lambda_{i}\}_{i\geq 1}$ and $\{L_{i}\}_{i\geq 1},$ we make use of the following potential function:

Let us first show that if $A_{k}\mathcal{C}_{k}$ is non-increasing with $k$ for an appropriately chosen sequence of positive numbers $\{A_{k}\}_{k\geq 1},$ then we can deduce a property that, under suitable conditions on $\{\lambda_{i}\}_{i\geq 1}$ and $\{L_{i}\}_{i\geq 1},$ implies a convergence rate for (H).

Let $\mathcal{C}_{k}$ be defined as in Eq. (2.1) and let $\mathbf{u}^{*}$ be the solution to (MI) that minimizes $\|\mathbf{u}_{0}-\mathbf{u}^{*}\|$ . Assume further that $\left\langle F(\mathbf{u}_{1})-F(\mathbf{u}_{0}),\mathbf{u}_{1}-\mathbf{u}_{0}\right\rangle\geq\frac{1}{L_{1}}\|F(\mathbf{u}_{1})-F(\mathbf{u}_{0})\|^{2}.$ If $A_{k+1}\mathcal{C}_{k+1}\leq A_{k}\mathcal{C}_{k},$ $\forall k\geq 1,$ where $\{A_{i}\}_{i\geq 1}$ is a sequence of positive numbers that satisfies $A_{1}=1$ , then:

Using Lemma 2.1, our goal is now to show that we can choose $L_{k}=O(L)$ and $\lambda_{k}=O(\frac{1}{k}),$ which in turn would imply the desired $1/k$ convergence rate: $\|F(\mathbf{u}_{k})\|=O(\frac{L\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{k}).$ The following lemma provides sufficient conditions for $\{A_{i}\}_{i\geq 1},$ $\{\lambda_{i}\}_{i\geq 1}$ , and $\{L_{i}\}_{i\geq 1}$ to ensure that $A_{k+1}\mathcal{C}_{k+1}\leq A_{k}\mathcal{C}_{k},$ $\forall k\geq 1,$ so that Lemma 2.1 applies.

Let $\mathcal{C}_{k}$ be defined as in Eq. (2.1). Let $\{A_{i}\}_{i\geq 1}$ be defined recursively as $A_{1}=1$ and $A_{k+1}=A_{k}\frac{\lambda_{k}}{(1-\lambda_{k})\lambda_{k+1}}$ for $k\geq 1.$ Assume that $\{\lambda_{i}\}_{i\geq 1}$ is chosen so that $\lambda_{1}=\frac{1}{2}$ and for $k\geq 1:$ $\frac{\lambda_{k+1}}{1-2\lambda_{k+1}}\geq\frac{\lambda_{k}L_{k}}{(1-\lambda_{k})L_{k+1}}$ . Finally, assume that $L_{k}\in(0,\infty)$ and $\left\langle F(\mathbf{u}_{k})-F(\mathbf{u}_{k-1}),\mathbf{u}_{k}-\mathbf{u}_{k-1}\right\rangle\geq\frac{1}{L_{k}}\|F(\mathbf{u}_{k})-F(\mathbf{u}_{k-1})\|^{2}$ , $\forall k.$ Then,

Observe first the following. If we knew $L$ and set $L_{k}=L,$ $\lambda_{k}=\frac{1}{k+1},$ and $A_{k}=k(k+1)/2,$ then all of the conditions from Lemma 2.2 would be satisfied, and Lemma 2.1 would then imply $\|F(\mathbf{u}_{k})\|\leq\frac{L\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{k},$ which recovers the result of Lieder (2017). The choice $\lambda_{k}=\frac{1}{k+1}$ is also the tightest possible that satisfies the conditions Lemma 2.2 – the inequality relating $\lambda_{k+1}$ and $\lambda_{k}$ is satisfied with equality. This result is in line with the numerical observations made by Lieder (2017), who observed that the convergence of Halpern iteration is fastest for $\lambda_{k}=\frac{1}{k+1}$ .

To construct a parameter-free method, we use that $F$ is $L$ -cocoercive; namely, that there exists a constant $L<\infty$ such that $F$ satisfies Eq. (1.5) with $\gamma=1/L$ . The idea is to start to with a “guess” of $L$ (e.g., $L_{0}=1$ ) and double the guess $L_{k}$ as long as $\left\langle F(\mathbf{u}_{k})-F(\mathbf{u}_{k-1}),\mathbf{u}_{k}-\mathbf{u}_{k-1}\right\rangle<\frac{1}{L_{k}}\|F(\mathbf{u}_{k})-F(\mathbf{u}_{k-1})\|^{2}.$ The total number of times that the guess can be doubled is bounded above by $\max\{0,\log_{2}(2L/L_{0})\}.$ Parameter $\lambda_{k}$ is simply chosen to satisfy the condition from Lemma 2.2. The algorithm pseudocode is stated in Algorithm 1 for a given accuracy specified at the input.

We now prove the first of our main results. Note that the total number of arithmetic operations in Algorithm 1 is of the order of the number of oracle queries to $F$ multiplied by the complexity of evaluating $F$ at a point. The same will be true for all the algorithms stated in this paper, except that the complexity of evaluating $F$ may be replaced by the complexity of projections onto $\mathcal{U}$ .

Given $\mathbf{u}_{0}\in\mathcal{U}$ and an operator $F$ that is $\frac{1}{L}$ -cocoercive on $E,$ Algorithm 1 returns a point $\mathbf{u}_{k}$ such that $\|F(\mathbf{u}_{k})\|\leq\epsilon$ after at most $\frac{\max\{2L,L_{0}\}\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}+\max\{0,\log_{2}(2L/L_{0})\}$ oracle queries to $F$ .

As $F$ is $\frac{1}{L}$ -cocoercive, $L_{k}\leq\max\{2L,L_{0}\}$ and the total number of times that the algorithm enters the inner while loop is at most $\max\{0,\log_{2}(2L/L_{0})\}.$ The parameters satisfy the assumptions of Lemmas 2.1 and 2.2, and, thus, $\|F(\mathbf{u}_{k})\|\leq L_{k}\frac{\lambda_{k}}{1-\lambda_{k}}\|\mathbf{u}_{0}-\mathbf{u}^{*}\|.$ Hence, we only need to show that $\lambda_{k}$ decreases sufficiently fast with $k.$ As $L_{k}$ can only be increased in any iteration, we have that

Hence, the total number of outer iterations is at most $\frac{\max\{2L,L_{0}\}\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}$ . Combining with the maximum total number of inner iterations from the beginning of the proof, the result follows. ∎

Assume now that $\mathcal{U}\subseteq E.$ We will make use of a counterpart to gradient mapping (Nesterov, 2018, Chapter 2) that we refer to as the operator mapping, defined as:

where $\Pi_{\mathcal{U}}\big{(}\mathbf{u}-\frac{1}{\eta}F(\mathbf{u})\big{)}$ is the projection operator, namely:

Operator mapping generalizes a cocoercive operator to the constrained case: when $\mathcal{U}\equiv E,$ $G_{\eta}\equiv F.$

It is a well-known fact that the projection operator is firmly-nonexpansive (Bauschke and Combettes, 2011, Proposition 4.16). Thus, Fact 1.3 can be used to show that, if $F$ is $\frac{1}{L}$ -cocoercive and $\eta\geq L,$ then $G_{\eta}$ is $\frac{1}{2\eta}$ -cocoercive. This is shown in the following (simple) proposition.

Let $F$ be an $\frac{1}{L}$ -cocoercive operator and let $G_{\eta}$ be defined as in Eq. (1.1), where $\eta\geq L.$ Then $G_{\eta}$ is $\frac{1}{2\eta}$ -cocoercive.

As $G_{\eta}$ is $\frac{1}{2\eta}$ -cocoercive, applying results from the beginning of the section to $G_{\eta}$ , it is now immediate that Algorithm 2 (provided for completeness) produces $\mathbf{u}_{k}$ with $\|G_{L_{k}}(\mathbf{u}_{k})\|\leq\epsilon$ after at most $\frac{\max\{4L,L_{0}\}\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}+\max\{0,\log_{2}(4L/L_{0})\}$ oracle queries to $F$ (as each computation of $G_{\eta}$ requires one oracle query to $F$ ).

To complete this subsection, it remains to show that $G_{\eta}$ is a good surrogate for approximating (MI) (and (SVI)). This is indeed the case and it follows as a suitable generalization of Lemma 3 from Ghadimi and Lan (2016), which is provided here for completeness.

Let $G_{\eta}$ be defined as in Eq. (2.2). Denote $\bar{\mathbf{u}}=\Pi_{\mathcal{U}}(\mathbf{u}-F(\mathbf{u})/\eta),$ so that $G_{\eta}(\mathbf{u})=\eta(\mathbf{u}-\bar{\mathbf{u}}).$ If, for some $\mathbf{u}\in\mathcal{U},$ $\|G_{\eta}(\mathbf{u})\|\leq\epsilon,$ then

Lemma 2.5 implies that when the operator mapping is small in norm $\|\cdot\|,$ then $\bar{\mathbf{u}}=\Pi_{\mathcal{U}}(\mathbf{u}-F(\mathbf{u})/\eta)$ is an approximate solution to (MI) corresponding to $F$ on $\mathcal{U}.$ We can now formally bound the number of oracle queries to $F$ needed to approximate (MI) and (SVI).

Given $\mathbf{u}_{0}\in\mathcal{U}$ and a $\frac{1}{L}$ -cocoercive operator $F$ , Algorithm 2 returns $\bar{\mathbf{u}}_{k}\in\mathcal{U}$ such that

$\|G_{L_{k}}(\bar{\mathbf{u}}_{k})\|\leq\frac{\epsilon}{2}$ , $\max_{\mathbf{v}\in\{\mathcal{U}\cap{\cal B}_{\bar{\mathbf{u}}_{k}}\}}\left\langle F(\bar{\mathbf{u}}_{k}),\bar{\mathbf{u}}_{k}-\mathbf{v}\right\rangle\leq\epsilon$ after at most

$\max_{\mathbf{v}\in\mathcal{U}}\left\langle F(\bar{\mathbf{u}}),\bar{\mathbf{u}}-\mathbf{v}\right\rangle\leq\epsilon$ after at most

Further, every point $\mathbf{u}_{k}$ that Algorithm 2 constructs is from the feasible set: $\mathbf{u}_{k}\in\mathcal{U},$ $\forall k\geq 0$ , and a simple modification to the algorithm takes at most $\frac{\max\{4L,L_{0}\}\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}+\max\{0,\log_{2}(4L/L_{0})\}$ oracle queries to $F$ to construct a point such that $\|G_{L_{k}}(\mathbf{u}_{k})\|\leq\epsilon$ .

By the definition of $G_{\eta},$ if $\mathbf{u}_{0}\in\mathcal{U},$ then $\mathbf{u}_{k}\in\mathcal{U},$ for all $k.$ This follows simply as:

Observe that, due to Line 2 of Algorithm 2, $L_{k}\geq\bar{L}_{k}.$ The rest of the proof follows using Lemma 2.5, Fact 1.1, and the same reasoning as in the proof of Theorem 2.3. Observe that if the goal is to only output a point $\mathbf{u}_{k}$ such that $\|G_{L_{k}}(\mathbf{u}_{k})\|\leq\epsilon$ , then computing $\bar{\mathbf{u}}_{k}$ and $F(\bar{\mathbf{u}}_{k})$ is not needed, and the algorithm can instead use $\|G_{L_{k}}(\mathbf{u}_{k})\|>\epsilon$ as the exit condition in the outer while loop. ∎

2 Setups with non-Cocoercive Lipschitz Operators

Finding a point $\mathbf{u}\in\mathcal{U}$ such that $\|P(\mathbf{u})\|\leq\epsilon$ is sufficient for approximating monotone inclusion (and (SVI)). This is shown in the following simple proposition, provided here for completeness.

As $\|P(\mathbf{u})\|\leq\epsilon,$ the result follows. ∎

Let $\bar{\mathbf{u}}_{k}^{*}=J_{F+I_{\mathcal{U}}}(\mathbf{u}_{k}),$ where $\mathbf{u}_{k}\in\mathcal{U}$ and $F$ is $L$ -Lipschitz. Then, there exists a parameter-free algorithm that queries $F$ at most $O((L+1)\log(\frac{L\|\mathbf{u}_{k}-\bar{\mathbf{u}}_{k}^{*}\|}{\epsilon}))$ times and outputs a point $\bar{\mathbf{u}}_{k}$ such that $\|\bar{\mathbf{u}}_{k}-\bar{\mathbf{u}}^{*}_{k}\|\leq\epsilon.$

To obtain the desired result, we need to prove the convergence of a Halpern iteration with inexact evaluations of the cocoercive operator $P$ . Note that here we do know the cocoercivity parameter of $P$ – it is equal to $1/2$ . The resulting inexact version of Halpern’s iteration for $P$ is:

To analyze the convergence of (2.3), we again use the potential function $\mathcal{C}_{k}$ from Eq. (2.1), with $P$ as the operator. For simplicity of exposition, we take the best choice of $\lambda_{i}=\frac{1}{i+1}$ that can be obtained from Lemma 2.1 for $L_{i}=L=2,$ $\forall i.$ The key result for this setting is provided in the following lemma, whose proof is deferred to the appendix.

Let $\mathcal{C}_{k}$ be defined as in Eq. (2.1) with $P$ as the $\frac{1}{2}$ -cocoercive operator, and let $L_{k}=2,$ $\lambda_{k}=\frac{1}{k+1},$ and $A_{k}=\frac{k(k+1)}{2},$ $\forall k\geq 1$ . If the iterates $\mathbf{u}_{k}$ evolve according to (2.3) for an arbitrary initial point $\mathbf{u}_{0}\in\mathcal{U},$ then:

Further, if, $\forall k\geq 1,$ $\|\mathbf{e}_{k-1}\|\leq\frac{\epsilon}{4k(k+1)},$ then $\|P(\mathbf{u}_{K})\|\leq\epsilon$ after at most $K=\frac{4\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}$ iterations.

We are now ready to state the algorithm and prove the main theorem for this subsection.

Let $F$ be a monotone and $L$ -Lipschitz operator and let $\mathbf{u}_{0}\in\mathcal{U}$ be an arbitrary initial point. For any $\epsilon>0,$ Algorithm 3 outputs a point with $\|P(\mathbf{u}_{k})\|\leq\epsilon$ after at most $\frac{8\|\mathbf{u}^{*}-\mathbf{u}_{0}\|}{\epsilon}$ iterations, where each iteration can be implemented with $O((L+1)\log(\frac{(L+1)\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon})$ oracle queries to $F.$ Hence, the total number of oracle queries to $F$ is: $O\big{(}\frac{(L+1)\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}\log\big{(}\frac{(L+1)\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}\big{)}\big{)}.$

Similarly as before, $\|P(\mathbf{u}_{k})\|\leq\epsilon$ implies an $\epsilon$ -approximate solution to (MI), by Proposition 2.7. When the diameter $D$ is bounded, $\|P(\mathbf{u}_{k})\|\leq\frac{\epsilon}{D}$ implies an $\epsilon$ -approximate solution to (SVI).

3 Setups with Strongly Monotone and Lipschitz Operators

We now show that by restarting Algorithm 3, we can obtain a parameter-free method with near-optimal oracle complexity. To simplify the exposition, we assume w.l.o.g. that $L=\Omega(1).$

Given $F$ that is $L$ -Lipschitz and $m$ -strongly monotone, consider running the following algorithm $\mathcal{A}$ , starting with $\mathbf{u}_{0}\in\mathcal{U}$ :

Then, $\mathcal{A}$ outputs $\mathbf{u}_{k}\in\mathcal{U}$ with $\|P(\mathbf{u}_{k})\|\leq\epsilon$ after at most $1+\log_{2}(\frac{\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon})$ iterations, for any $\epsilon\in(0,\frac{1}{2}]$ . The total number of queries to $F$ until $\|P(\mathbf{u}_{k})\|\leq\epsilon$ is $O\big{(}(L+\frac{L}{m})\log(\frac{\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon})\log(L+\frac{L}{m})\big{)}.$

The first part is immediate, as each call to Algorithm 3 ensures, due to Theorem 2.10, that

and $\|P(\mathbf{u}_{0})\|\leq 2\|\mathbf{u}_{0}-\mathbf{u}^{*}\|$ as $P$ is 2-Lipschitz (because it is $\frac{1}{2}$ -cocoercive) and $P(\mathbf{u}^{*})=\textbf{0}.$

On the other hand, as $F$ is $m$ -strongly monotone and $\mathbf{u}^{*}$ is an (MVI) solution,

Hence: $\|\bar{\mathbf{u}}^{*}_{k-1}-\mathbf{u}^{*}\|\leq\frac{1}{m}\|P(\mathbf{u}_{k-1})\|.$ It remains to use the triangle inequality and $P(\mathbf{u}_{k-1})=\mathbf{u}_{k-1}-\bar{\mathbf{u}}^{*}_{k-1}$ to obtain: $\|\mathbf{u}_{k-1}-\mathbf{u}^{*}\|\leq\big{(}1+\frac{1}{m}\big{)}\|P(\mathbf{u}_{k-1})\|.$ ∎

Lower Bound Reductions

In this section, we only state the lower bounds, while more details about the oracle model and the proof are deferred to Appendix A.

For any deterministic algorithm working in the operator oracle model and any $L,D>0$ , there exists an $L$ -Lipschitz-continuous operator $F$ and a closed convex feasible set $\mathcal{U}$ with diameter $D$ such that:

For all $\epsilon>0$ such that $k=\frac{LD^{2}}{\epsilon}=O(d)$ , $\max_{\mathbf{u}\in\mathcal{U}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=\Omega(\epsilon)$ ;

For all $\epsilon>0$ such that $k=\frac{LD}{\epsilon}=O(d)$ , $\max_{\mathbf{u}\in\{\mathcal{U}\cap\mathcal{B}_{\mathbf{u}_{k}}\}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=\Omega(\epsilon)$ ;

If $F$ is $\frac{1}{L}$ -cocoercive, then for all $\epsilon>0$ such that $k=\frac{LD}{\epsilon\log(D/\epsilon)}=O(d)$ , it holds that

If $F$ is $m$ -strongly monotone, then for all $\epsilon>0$ such that $k=\frac{L}{m}=O(d)$ , it holds that

Parts (a) and (b) of Lemma 3.1 certify that Algorithm 3 is optimal up to a logarithmic factor, due to Theorem 2.10. This is true because we can run Algorithm 3 with accuracy $\frac{\epsilon}{D}$ to obtain $\max_{\mathbf{u}\in\mathcal{U}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=O(\epsilon)$ in $k=O(\frac{LD^{2}}{\epsilon}\log(\frac{LD}{\epsilon}))$ iterations, or with accuracy $\epsilon$ to obtain $\max_{\mathbf{u}\in\{\mathcal{U}\cap\mathcal{B}_{\mathbf{u}_{k}}\}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=O(\epsilon)$ in $k=O(\frac{LD}{\epsilon}\log(\frac{LD}{\epsilon}))$ iterations (see Proposition 2.7).

Part (c) of Lemma 3.1 certifies that Algorithm 2 is optimal up to a $\log(D/\epsilon)$ factor, due to Theorem 2.6. Part (d) certifies that the restarting algorithm from Theorem 2.12 is optimal up to a factor $\log(D/\epsilon)\log(L/m)$ whenever $L=\Omega(L/m).$ Note that $L=\Omega(L/m)$ can be ensured by a proper scaling of the problem instance, as any such scaling would leave the condition number $L/m$ unaffected and would only impact the target error $\epsilon,$ which only appears under a logarithm.

Conclusion

We showed that variants of Halpern iteration can be used to obtain near-optimal methods for solving different classes of monotone inclusion problems with Lipschitz operators. The results highlight connections between monotone inclusion, variational inequalities, fixed points of nonexpansive maps, and proximal-point-type algorithms. Some interesting questions that merit further investigation remain. In particular, one open question that arises is to close the gap between the upper and lower bounds provided here. We conjecture that the optimal complexity of monotone inclusion is: (i) $\Theta(\frac{LD}{\epsilon})$ when the operator is either $L$ -Lipschitz or $\frac{1}{L}$ -cocoercive, and (ii) $\Theta(\frac{L}{m}\log(\frac{LD}{\epsilon}))$ when the operator is $L$ -Lipschitz and $m$ -strongly monotone.

Acknowledgements

We thank Prof. Ulrich Kohlenbach for useful comments and pointers to the literature. We also thank Howard Heaton for pointing out a typo in the proof of Lemma 2.1 in a previous version of this paper.

References

Appendix A Omitted Proofs

The statement holds trivially if $\|F(\mathbf{u}_{k})\|=0,$ so assume that $\|F(\mathbf{u}_{k})\|>0.$ Under the assumption of the lemma, we have that $A_{k}\mathcal{C}_{k}\leq\mathcal{C}_{1},$ $\forall k\geq 1.$ From (H) and $\lambda_{1}=\frac{1}{2}$ , $\mathbf{u}_{1}=\mathbf{u}_{0}-\frac{1}{L_{1}}F(\mathbf{u}_{0}),$ and thus: $\mathcal{C}_{1}=\frac{1}{L_{1}}\|F(\mathbf{u}_{1})\|^{2}-\frac{1}{L_{1}}\left\langle F(\mathbf{u}_{1}),F(\mathbf{u}_{0})\right\rangle.$

Let $\mathbf{u}^{*}$ be an arbitrary solution to (MI) (and thus also to (MVI)). As $\left\langle F(\mathbf{u}_{1})-F(\mathbf{u}_{0}),\mathbf{u}_{1}-\mathbf{u}_{0}\right\rangle\geq\frac{1}{L_{1}}\|F(\mathbf{u}_{1})-F(\mathbf{u}_{0})\|^{2}$ and $\mathbf{u}_{1}=\mathbf{u}_{0}-\frac{1}{L_{1}}F(\mathbf{u}_{0}),$ it follows that $\|F(\mathbf{u}_{1})\|^{2}\leq\left\langle F(\mathbf{u}_{0}),F(\mathbf{u}_{1})\right\rangle,$ and, thus $\mathcal{C}_{1}\leq 0.$ Further, as $A_{k}>0,$ we also have $\mathcal{C}_{k}\leq 0,$ and, hence:

where the last line is by $\mathbf{u}^{*}$ being a solution to (MVI) and by the Cauchy-Schwarz inequality. The conclusion of the lemma now follows by dividing both sides of $\|F(\mathbf{u}_{k})\|^{2}\leq L_{k}\frac{\lambda_{k}}{1-\lambda_{k}}\|F(\mathbf{u}_{k})\|\cdot\|\mathbf{u}_{0}-\mathbf{u}^{*}\|$ by $\|F(\mathbf{u}_{k})\|$ and observing that the statement holds for an arbitrary solution $\mathbf{u}^{*}$ to (MI), and thus, it also holds for the one that minimizes the distance to $\mathbf{u}_{0}.$ ∎

which, after expanding the left-hand side, can be equivalently written as:

From (H), we have that $\mathbf{u}_{k+1}-\mathbf{u}_{k}=\frac{\lambda_{k+1}}{1-\lambda_{k+1}}(\mathbf{u}_{0}-\mathbf{u}_{k+1})-\frac{2}{L_{k+1}}F(\mathbf{u}_{k})$ and $\mathbf{u}_{k+1}-\mathbf{u}_{k}=\lambda_{k+1}(\mathbf{u}_{0}-\mathbf{u}_{k})-\frac{2(1-\lambda_{k+1})}{L_{k+1}}F(\mathbf{u}_{k}).$ Hence:

Rearranging the last inequality and multiplying both sides by $A_{k+1},$ we have:

The left-hand side of the last inequality if precisely $A_{k+1}\mathcal{C}_{k+1}.$ The right-hand side is $\leq A_{k}\mathcal{C}_{k},$ by the choice of sequences $\{A_{i}\}_{i\geq 1},$ $\{\lambda_{i}\}_{i\geq 1}.$ ∎

A.2 Operator Mapping

Let $F$ be an $\frac{1}{L}$ -cocoercive operator and let $G_{\eta}$ be defined as in Eq. (1.1), where $\eta\geq L.$ Then $G_{\eta}$ is $\frac{1}{2\eta}$ -cocoercive.

As $\eta\geq L$ and $F$ is $\frac{1}{L}$ -cocoercive, $\frac{1}{\eta}\left\langle F(\mathbf{u})-F(\mathbf{v}),\mathbf{u}-\mathbf{v}\right\rangle\geq\frac{1}{\eta^{2}}\|F(\mathbf{u})-F(\mathbf{v})\|^{2}.$ It remains to apply Young’s inequality, which implies $\left\langle G_{\eta}(\mathbf{u})-G_{\eta}(\mathbf{v}),F(\mathbf{u})-F(\mathbf{v})\right\rangle\leq\frac{1}{2}\|G_{\eta}(\mathbf{u})-G_{\eta}(\mathbf{v})\|^{2}+\frac{1}{2}\|F(\mathbf{u})-F(\mathbf{v})\|^{2}.$ ∎

A.3 Approximating the Resolvent

Let us start by proving the convergence of a version of the Extragradient method of Korpelevich that does not require the knowledge of the Lipschitz constant $L$ (but does require knowledge of the strong monotonicity parameter $m$ ; when computing the resolvent we have $m=1$ ). The algorithm is summarized in Algorithm 4.

Observe that the update step for $\mathbf{u}_{k}$ from Lines 4 and 4 can be written in the form of a projection onto $\mathcal{U};$ we chose to write it in the current form as it is more convenient for the analysis.

We now bound the convergence of Algorithm 4.

Let $a_{0}>0$ and let $F$ be $m$ -strongly monotone and $L$ -Lipschitz. Then, Algorithm 4 outputs a point $\mathbf{u}_{k}$ with $\|\mathbf{u}_{k}-\mathbf{u}^{*}\|\leq\epsilon$ after at most $k=O\big{(}\frac{L}{m}\log(\frac{L\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{m\epsilon}\big{)})$ oracle queries to $F,$ where $\mathbf{u}^{*}$ solves (SVI).

Define $A_{k}=\sum_{i=0}^{k}a_{i}.$ To prove the lemma, we will use the following gap (or merit) functions:

As $F$ is strongly monotone, $f_{k}\geq 0,\,\forall k.$ By convention, we take $f_{-1}=0$ and $A_{-1}=0$ , so that $A_{k}f_{k}-A_{k-1}f_{k-1}=a_{k}\Big{(}\left\langle F(\bar{\mathbf{u}}_{k}),\bar{\mathbf{u}}_{k}-\mathbf{u}^{*}\right\rangle-\frac{m}{2}\|\bar{\mathbf{u}}_{k}-\mathbf{u}^{*}\|^{2}\Big{)}.$ Let us now bound $A_{k}f_{k}-A_{k-1}f_{k-1}$ , and observe that $A_{k}f_{k}-A_{k-1}f_{k-1}\geq 0$ . First, write

By the first-order optimality of $\mathbf{u}_{k+1}$ in its definition, we have, $\forall\mathbf{u}:$

By the standard three-point identity (which can also be verified directly):

Thus, setting $\mathbf{u}=\mathbf{u}^{*}:$

By the condition of the while loop in Line 4 of Algorithm 4, and because $A_{k}f_{k}-A_{k-1}f_{k-1}\geq 0$ ,

The condition of the while loop in Line 4 of Algorithm 4 is satisfied for any $a_{k}\leq\frac{1}{2L},$ as

where we have used the Cauchy-Schwarz inequality, the fact that $F$ is $L$ -Lipschitz, and the Young inequality. Thus, in any iteration, $a_{k}>\frac{1}{4L},$ and the total number of times the while loop from Line 4 is entered is at most $\log_{2}(4L/a_{0}).$

From Eq. (A.5), $\|\mathbf{u}^{*}-\mathbf{u}_{k+1}\|^{2}\leq\frac{1}{1+m/(4L)}\|\mathbf{u}^{*}-\mathbf{u}_{k}\|^{2}\leq(1-\frac{m}{8L})\|\mathbf{u}^{*}-\mathbf{u}_{k}\|^{2}.$ Thus, for any $\delta>0,$ $\|\mathbf{u}^{*}-\mathbf{u}_{k}\|\leq\delta$ for $k\geq\frac{16L}{m}\log(\frac{\|\mathbf{u}^{*}-\mathbf{u}_{0}\|}{\delta}).$ Consequently, from Eq. (A.5), $\|\bar{\mathbf{u}}_{k}-\mathbf{u}_{k}\|\leq\sqrt{2}\delta$ whenever $\|\mathbf{u}^{*}-\mathbf{u}_{k}\|\leq\delta.$ In particular, for $\delta=\frac{a_{k}m\epsilon}{5\sqrt{2}}\geq\frac{m\epsilon}{20\sqrt{2}L},$ $\|\bar{\mathbf{u}}_{k}-\mathbf{u}_{k}\|\leq\sqrt{2}\delta=\frac{a_{k}m}{5}\epsilon$ after at most $k=\frac{16L}{m}\log(\frac{20\sqrt{2}L\|\mathbf{u}^{*}-\mathbf{u}_{0}\|}{m\epsilon}$ (outer loop) iterations.

On the other hand, as $F$ is $m$ -strongly monotone, we also have $\left\langle F(\bar{\mathbf{u}}_{k}),\bar{\mathbf{u}}_{k}-\mathbf{u}^{*}\right\rangle\geq\frac{m}{2}\|\bar{\mathbf{u}}_{k}-\mathbf{u}^{*}\|^{2}.$ Hence, $\|\bar{\mathbf{u}}_{k}-\mathbf{u}^{*}\|\leq\frac{4\epsilon}{5}.$ Finally, applying the triangle inequality and as $a+k\leq 1/m:$

Note that we have already bounded the total number of inner and outer loop iterations. Observing that each inner iteration makes 2 oracle queries to $F$ and each outer iteration makes $2$ oracle queries to $F$ outside of the inner iteration, the bound on the total number of oracle queries to $F$ follows. ∎

Let $\bar{\mathbf{u}}_{k}^{*}=J_{F+I_{\mathcal{U}}}(\mathbf{u}_{k}),$ where $\mathbf{u}_{k}\in\mathcal{U}$ and $F$ is $L$ -Lipschitz. Then, there exists a parameter-free algorithm that queries $F$ at most $O((L+1)\log(\frac{(L+1)\|\mathbf{u}_{k}-\bar{\mathbf{u}}_{k}^{*}\|}{\epsilon}))$ times and outputs a point $\bar{\mathbf{u}}_{k}$ such that $\|\bar{\mathbf{u}}_{k}-\bar{\mathbf{u}}^{*}_{k}\|\leq\epsilon.$

Observe first that $\bar{\mathbf{u}}^{*}_{k}$ solves (SVI) for operator $\bar{F}(\mathbf{u})=F(\mathbf{u})+\mathbf{u}-\mathbf{u}_{k}$ over the set $\mathcal{U}.$ This follows from the definition of the resolvent, which implies:

Equivalently: $\textbf{0}\in\bar{F}(\bar{\mathbf{u}}_{k}^{*})+\partial I_{\mathcal{U}}(\bar{\mathbf{u}}^{*}_{k})$ .

The rest of the proof follows by applying Lemma A.1 to $\bar{F},$ which is $(L+1)$ -Lipschitz and $1$ -strongly monotone. ∎

A.4 Inexact Halpern Iteration

We start by first proving the following auxiliary result.

Given an initial point $\mathbf{u}_{0}\in\mathcal{U},$ let $\mathbf{u}_{k}$ evolve according to Eq. (2.3), where $\lambda_{k}=\frac{1}{k+1}$ . Then,

where $\mathbf{u}^{*}$ is such that $\|P(\mathbf{u}^{*})\|=0.$

where we have used the triangle inequality and nonexpansivity of $T.$ The result follows by recursively applying the last inequality and observing that $\prod_{j=i}^{k}(1-\lambda_{j})=\frac{i}{k+1}.$ ∎

Using this proposition, we can now prove the following lemma.

By the same arguments as in the proof of Lemma 2.1:

Plugging $\lambda_{k+1}=\frac{1}{k+2}$ in the last inequality and using the definition of $\mathcal{C}_{k}$ and the choice of $A_{k}$ from the statement of the lemma completes the proof of the first part.

Using the same arguments as in the proof of Lemma 2.2, we can conclude from $\quad A_{k+1}\mathcal{C}_{k+1}\leq A_{k}\mathcal{C}_{k}+A_{k+1}\left\langle\mathbf{e}_{k},(1-\lambda_{k+1})P(\mathbf{u}_{k})-P(\mathbf{u}_{k+1})\right\rangle,$ $\forall k\geq 1$ that:

Let us now bound each $\left\langle\mathbf{e}_{i-1},\frac{i}{i+1}P(\mathbf{u}_{i-1})-P(\mathbf{u}_{i})\right\rangle$ term. Recall that $P(\mathbf{u}^{*})=\textbf{0}$ and $P$ is 2-Lipschitz (as discussed in Section 1.2, this follows from $P$ being $\frac{1}{2}$ -cocoercive). Thus, we have:

where we have used Proposition A.2 in the last inequality. In particular, if $\|\mathbf{e}_{i-1}\|\leq\frac{\epsilon}{4i(i+1)}$ , then, $\forall i\geq 1$ :

Observe that if $\|\mathbf{u}_{0}-\mathbf{u}^{*}\|\leq\epsilon/2,$ as $P$ is 2-Lipschitz and $P(\mathbf{u}^{*})=\textbf{0},$ we would have $\|P(\mathbf{u}_{0})\|\leq\epsilon,$ and the statement of the second part of the lemma would hold trivially. Assume from now on that $\|\mathbf{u}_{0}-\mathbf{u}^{*}\|>\epsilon/2.$ Suppose that $\|P(\mathbf{u}_{k})\|>\epsilon$ and $k\geq\frac{4\|\mathbf{u}_{0}-\mathbf{u}^{*}\|}{\epsilon}.$ Then, dividing both sides of Eq. (A.7) by $\|P(\mathbf{u}_{k})\|/2$ and using that $\|P(\mathbf{u}_{k})\|>\epsilon$ and $\|\mathbf{u}_{0}-\mathbf{u}^{*}\|>\epsilon/2$ , we get:

contradicting the assumption that $\|P(\mathbf{u}_{k})\|>\epsilon$ and completing the proof. ∎

A.5 Strongly Monotone Lipschitz Operators

Given $F$ that is $L$ -Lipschitz and $m$ -strongly monotone, consider running the following algorithm $\mathcal{A}$ , starting with $\bar{\mathbf{u}}_{0}\in\mathcal{U}$ :

Then, $\mathcal{A}$ outputs a point $\mathbf{u}_{k}\in\mathcal{U}$ with $\|P(\mathbf{u}_{k})\|\leq\epsilon$ after at most $\log_{2}(\|\mathbf{u}_{0}-\mathbf{u}^{*}\|/\epsilon)$ iterations, where w.l.o.g. $\epsilon\leq\frac{1}{2}$ . The total number of oracle queries to $F$ until this happens is $O\big{(}(L+\frac{L}{m})\log(\|\mathbf{u}_{0}-\mathbf{u}^{*}\|/\epsilon)\log(L+\frac{L}{m})\big{)}.$

The first part of the theorem is immediate, as each call to Algorithm 3 ensures, due to Theorem 2.10, that

and $\|P(\mathbf{u}_{0})\|\leq 2\|\mathbf{u}_{0}-\mathbf{u}^{*}\|$ as $P$ is 2-Lipschitz (because it is $\frac{1}{2}$ -cocoercive) and $P(\mathbf{u}^{*})=\textbf{0}.$

On the other hand, as $F$ is $m$ -strongly monotone and $\mathbf{u}^{*}$ is an (MVI) solution,

A.6 Lower Bounds

We make use of the lower bound from Ouyang and Xu and the algorithmic reductions between the problems considered in previous sections to derive (near-tight) lower bounds for all of the problems considered in this paper.

The lower bounds are for deterministic algorithms working in a (first-order) oracle model. For convex-concave saddle-point problems with the objective $\Phi(\mathbf{x},\mathbf{y})$ and closed convex feasible set $\mathcal{X}\times\mathcal{Y},$ any such algorithm $\mathcal{A}$ can be described as follows: in each iteration $k$ , $\mathcal{A}$ queries a pair of points $(\bar{\mathbf{x}}_{k},\bar{\mathbf{y}}_{k})\in\mathcal{X}\times\mathcal{Y}$ to obtain $(\nabla_{\mathbf{x}}\Phi(\bar{\mathbf{x}}_{k},\bar{\mathbf{y}}_{k}),\,\nabla_{\mathbf{y}}\Phi(\bar{\mathbf{x}}_{k},\bar{\mathbf{y}}_{k})),$ and outputs a candidate solution pair $(\mathbf{x}_{k},\mathbf{y}_{k})\in\mathcal{X}\times\mathcal{Y}.$ Both the query points pair $(\bar{\mathbf{x}}_{k},\bar{\mathbf{y}}_{k})$ and the candidate solution pair $(\mathbf{x}_{k},\mathbf{y}_{k})$ can only depend on (i) global problem parameters (such as the Lipschitz constant of $\Phi$ ’s gradients or the feasible sets $\mathcal{X},\mathcal{Y}$ ) and (ii) oracle queries and answers up to iteration $k:$

We start by summarizing the result from [Ouyang and Xu, 2019, Theorem 9].

where $(\mathbf{x}_{k},\mathbf{y}_{k})\in\mathcal{X}\times\mathcal{Y}$ is the algorithm output after $k$ iterations and $R_{\mathcal{X}},\,R_{\mathcal{Y}}$ denote the diameters of the feasible sets $\mathcal{X},\,\mathcal{Y},$ respectively, and where both $\mathcal{X},\,\mathcal{Y},$ are closed and convex.

The assumption of the theorem that $k=O(d)$ means that the lower bound applies in the high-dimensional regime $d=\Omega(\frac{L({R_{\mathcal{X}}}^{2}+R_{\mathcal{X}}R_{\mathcal{Y}})}{\epsilon}),$ which is standard and generally unavoidable.

In the setting of VIs, we consider a related model in which an algorithm has oracle access to $F$ and refer to it as the operator oracle model. Similarly as for the saddle-point problems, we consider deterministic algorithms that on a given problem instance described by $(F,\mathcal{U})$ operate as follows: in each iteration $k$ the algorithm queries a point $\bar{\mathbf{u}}_{k}\in\mathcal{U}$ , receives $F(\bar{\mathbf{u}}_{k}),$ and outputs a solution candidate $\mathbf{u}_{k}\in\mathcal{U}$ . Both $\mathbf{u}_{k}$ and $\bar{\mathbf{u}}_{k}$ can only depend on (i) global problem parameters (such as the feasible set $\mathcal{U}$ and the Lipschitz parameter of $F$ ), and (ii) oracle queries and answers up to iteration $k:$ $\{\bar{\mathbf{u}}_{i},F(\bar{\mathbf{u}}_{i})\}_{i=0}^{k-1}.$ Note that all methods described in this paper and most of the commonly used methods for solving VIs, such as, e.g., the mirror-prox method of Nemirovski and dual extrapolation method of Nesterov , work in this oracle model.

For any deterministic algorithm working in the operator oracle model described above and any $L,D>0$ , there exists a VI described by an $L$ -Lipschitz-continuous operator $F$ and a closed convex feasible set $\mathcal{U}$ with diameter $D$ such that:

For all $\epsilon>0$ such that $k=\frac{LD^{2}}{\epsilon}=O(d)$ , $\max_{\mathbf{u}\in\mathcal{U}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=\Omega(\epsilon)$ ;

If $F$ is $\frac{1}{L}$ -cocoercive, then for all $\epsilon>0$ such that $k=\frac{LD}{\epsilon\log(D/\epsilon)}=O(d)$ , it holds that

If $F$ is $m$ -strongly monotone, then for all $\epsilon>0$ such that $k=\frac{L}{m}=O(d)$ , it holds that

Proof of (a): Suppose that this claim was not true. Then we would be able to solve any instance with $L$ -Lipschitz $F$ and $\mathcal{U}$ with diameter bounded by $D$ and obtain $\mathbf{u}_{k}$ with $\max_{\mathbf{u}\in\mathcal{U}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle\leq\epsilon$ in $o(\frac{LD^{2}}{\epsilon})$ iterations, assuming the appropriate high-dimensional regime. In particular, given any fixed convex-concave $\Phi(\mathbf{x},\mathbf{y})$ with $L$ -Lipschitz gradients and feasible sets $\mathcal{X},\mathcal{Y}$ whose diameter is $D/2,$ let $\mathbf{u}=[\vbox{\Let@\restore@math@cr\default@tag\halign{\hfil$ \m@th\scriptstyle# $&$ \m@th\scriptstyle{}# $\hfil\cr\mathbf{x}\\ \mathbf{y}\crcr}}],$ $F(\mathbf{u})=[\vbox{\Let@\restore@math@cr\default@tag\halign{\hfil$ \m@th\scriptstyle# $&$ \m@th\scriptstyle{}# $\hfil\cr\nabla_{\mathbf{x}}\Phi(\mathbf{x},\mathbf{y})\\ -\nabla_{\mathbf{y}}\Phi(\mathbf{x},\mathbf{y})\crcr}}]$ , $\mathcal{U}=\mathcal{X}\times\mathcal{Y}.$ Then, it is not hard to verify that $F$ is monotone and $L$ -Lipschitz (see, e.g., Nemirovski , Facchinei and Pang ) and the diameter of $\mathcal{U}$ is $D.$ Thus, by assumption, we would be able to construct a point $\mathbf{u}_{k}=[\vbox{\Let@\restore@math@cr\default@tag\halign{\hfil$ \m@th\scriptstyle# $&$ \m@th\scriptstyle{}# $\hfil\cr\mathbf{x}_{k}\\ \mathbf{y}_{k}\crcr}}]$ for which $\max_{\mathbf{u}\in\mathcal{U}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle\leq\epsilon$ in $o(\frac{LD^{2}}{\epsilon})$ iterations. But then, because $\Phi$ is convex-concave, we would also have, for any $\mathbf{x}\in\mathcal{X},\mathbf{y}\in\mathcal{Y}$ :

Because we obtained this bound for an arbitrary $L$ -Lipschitz convex-concave $\Phi$ and arbitrary feasible sets $\mathcal{X},\mathcal{Y}$ with diameters $D/2,$ Theorem A.3 leads to a contradiction.

Proof of (b): If (b) was not true, then we would be able to obtain a point $\mathbf{u}_{k}$ with

in $k=\frac{LD^{2}}{\epsilon}$ iterations. But the same point would satisfy $\max_{\mathbf{u}\in\mathcal{U}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=o(\epsilon),$ which is a contradiction, due to (a).

Proof of (c): We prove the claim for $L=2.$ This is w.l.o.g., due to the standard rescaling argument: if $F$ is $\frac{1}{L}$ -cocoercive, then $\bar{F}=F/(2L)$ is $\frac{1}{2}$ -cocoercive. Further, if, for some $\mathbf{u}_{k}\in\mathcal{U},$

then $\max_{\mathbf{u}\in\{\mathcal{U}\cap\mathcal{B}_{\mathbf{u}_{k}}\}}\left\langle F(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=\Omega(L\epsilon).$

Suppose that the claim was not true for a $\frac{1}{2}$ -cocoercive operator $F.$ Then for any $M$ -Lipschitz monotone operator $G,$ we would be able to use the strategy from Section 2.2 to obtain a point $\mathbf{u}_{k}$ with

in $k=\frac{MD}{\epsilon}$ iterations. This is a contradiction, due to (b).

Proof of (d): Suppose that the claim was not true, i.e., that there existed an algorithm that, for any $m,L>0,$ could output $\mathbf{u}_{k}$ with $\max_{\mathbf{u}\in\{\mathcal{U}\cap\mathcal{B}_{\mathbf{u}_{k}}\}}\left\langle\bar{F}(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=\epsilon/2$ in $k=o(L/m)$ iterations, for any $m$ -strongly monotone and $L$ -Lipschitz operator. Then for any $L$ -Lipschitz monotone operator $F$ , we could apply that algorithm to $\bar{F}(\cdot)=F(\cdot)+\frac{\epsilon}{2D}(\cdot-\mathbf{u}_{0})$ to obtain a point $\mathbf{u}_{k}$ with $\max_{\mathbf{u}\in\{\mathcal{U}\cap\mathcal{B}_{\mathbf{u}_{k}}\}}\left\langle\bar{F}(\mathbf{u}_{k}),\mathbf{u}_{k}-\mathbf{u}\right\rangle=\epsilon/2$ in $k=o(LD/\epsilon)$ iterations. But then we would also have: