Randomized Composable Core-sets for Distributed Submodular Maximization

Vahab Mirrokni, Morteza Zadimoghaddam

Introduction

An effective way of processing massive data is to first extract a compact representation of the data and then perform further processing only on the representation itself. This approach significantly reduces the cost of processing, communicating and storing the data, as the representation size can be much smaller than the size of the original data set. Typically, the representation provides a smooth tradeoff between its size and the representation accuracy. Examples of this approach include techniques such as sampling, sketching, (composable) core-sets and mergeable summaries. Among these techniques, the concept of composable core-sets has been employed in several distributed optimization models such as nearest neighbor search , and the streaming and MapReduce models . Roughly speaking, the main idea behind this technique is as follows: First partition the data into smaller parts. Then compute a representative solution, referred to as a core-set, from each part. Finally, obtain a solution by solving the optimization problem over the union of core-sets for all parts. While this technique has been successfully applied to diversity maximization and clustering problems , for coverage and submodular maximization problems, impossibility bounds are known for this technique .

In this paper, we focus on efficient construction of a randomized variant of composable core-sets where the above idea is applied on a random clustering of the data. We employ this technique for the coverage, monotone and non-monotone submodular problems. Our results significantly improve upon the hardness results for non-randomized core-sets, and imply improved results for submodular maximization in a distributed and streaming settings. The effectiveness of this technique has been confirmed empirically for several machine learning applications , and our proof provides a theoretical foundation to this idea. Let us first define this concept, and then discuss its applications, and our results.

Here, we discuss the formal problem definition, and the distributed model motivating it.

Given an integer size constraint $k$ , we let $f_{k}$ be

where the expectation is taken over the random choice of $\{T_{1},T_{2},\ldots,T_{m}\}$ . For brevity, instead of saying that $\mathsf{ALG}$ implements a composable core-set, we say that $\mathsf{ALG}$ is an $\alpha$ -approximate randomized composable core-set.

For ease of notation, when it is clear from the context, we may drop the term composable, and refer to composable core-sets as core-sets. Throughout this paper, we discuss randomized composable core-sets for the submodular maximization problem with a cardinality constraint $k$ .

Distributed Approximation Algorithm. Note that we can use a randomized $\alpha$ -approximate composable core-set algorithm $\mathsf{ALG}$ to design the following simple distributed $(1-{1\over e})\alpha$ -approximation algorithm for monotone submodular maximization:

Each machine $i$ computes a randomized composable core-set $S_{i}\subseteq T_{i}$ of size $k^{\prime}$ , i.e., $S_{i}=\mathsf{ALG}(T_{i})$ for each $1\leq i\leq m$ .

In the second phase, first collect the union of all core-sets, $U=\cup_{1\leq i\leq m}S_{i}$ , on one machine. Then apply a post-processing $(1-{1\over e})$ -approximation algorithm (e.g., algorithm $\mathsf{Greedy}$ ) to compute a solution $S$ to the submodular maximization problem over the set $U$ . Output $S$ .

It follows from the definition of the $\alpha$ -approximate randomized composable core-set that the above algorithm is a distributed $(1-{1\over e})\alpha$ -approximation algorithm for submodular maximization problem. We refer to this two-phase algorithmic approach as the distributed algorithm, and the overall approximation factor of the distributed algorithm as the distributed approximation factor. For all our algorithms in this paper, in addition to presenting an algorithm that achieves an approximation factor $\alpha$ as a randomized composable core-set, we propose a post-processing algorithm for the second phase, and present an improved analysis that achieves much better than $(1-{1\over e})\alpha$ -approximation as the distributed approximation factor.

Note that the above algorithm can be implemented in a distributed manner only if $k^{\prime}$ is small enough such that $mk^{\prime}$ items can be processed on one machine. In all our results the size of the composable core-set, $k^{\prime}$ , is a function of the cardinality constraint, $k$ : In particular, in Section 2, we apply a composable core-set of size $k^{\prime}=k$ . In Section 4, we apply a composable core-set of size $k^{\prime}<4k$ , and as a result, achieve a better approximation factor. We call a core-set, a small-size core-set, if its size $k^{\prime}$ is less than $k$ (See Section 5). As we will see, the hardness results for small-size core-sets are much stronger than that of core-sets of size $k$ or larger.

Non-randomized Composable Core-sets. The above definition for randomized composable core-sets is introduced in this paper. Prior work define a non-randomized variant of composable core-sets where the above property holds for any (arbitrary) partitioning $\{T_{1},T_{2},\ldots,T_{m}\}$ of data into $m$ parts It is not hard to see that for non-randomized composable core-set, the multiplicity parameter $C$ is not relevant., i.e., an algorithm $\mathsf{ALG}$ as described above is a $\alpha$ -approximate (non-randomized) composable core-set of size $k^{\prime}$ for $f$ , if for any cardinality constraint $k$ , and any arbitrary partitioning $\{T_{1},T_{2},\ldots,T_{m}\}$ of the items into $m$ sets, we have $f_{k}(\mathsf{ALG}(T_{1})\cup\ldots\cup\mathsf{ALG}(T_{m}))\geq{\alpha}\cdot f_{k}(T_{1}\cup\ldots\cup T_{m})$ .

2 Applications and Motivations

An $\alpha$ -approximate randomized composable core-set of size $k^{\prime}=O(k)$ for a problem can be applied in three types of applications These results assume $k\leq n^{1-\epsilon}$ for a constant $\epsilon$ .: (i) in distributed computation , where it implies an $\alpha$ -approximation in one or two rounds of MapReduces using the total communication complexity of $O(n)$ , (ii) in the random-order streaming model, where it implies an $\alpha$ -approximation algorithm in one pass using sublinear memory, (iii) in a class of approximate nearest neighbor search problems, where it implies an $\alpha$ -approximation algorithm based on the locally sensitive hashing (under an assumption). Here, we discuss the application for the MapReduce and Streaming framework, and for details of the approximate nearest neighbor application, we refer to .

We first show how to use a randomized composable core-set of size $O(k)$ to design a distributed algorithm in one or two rounds of MapReduces The straightforward way of applying the ideas will result in two rounds of MapReduce. However, if we assume that the data is originally sharded randomly and each part is in a single shard, and the memory for each machine is more than the size of each shard, then it can be implemented via one round of MapReduce computation. using linear total communication complexity: Let $m=\sqrt{n/k}$ , and let $(T_{1},\ldots,T_{m})$ be a random partitioning where $T_{i}$ has $\sqrt{kn}$ items. In the distributed algorithm, we assume that the random partitioning is produced in one round of MapReduce where each of $m$ reducers receives $T_{i}$ as input, and produces a core-set $S_{i}$ for the next round. Alternatively, we may assume that the data (or the items) are distributed uniformly at random among machines, or similarity each of $m$ mappers receives $T_{i}$ as input, and produces a core-set $S_{i}$ for the reducer. In either case, the produced core-sets are passed to a single reducer in the first or the second round. The total input to the reducer, i.e., the union of the core-sets, is of size at most $mk^{\prime}=O(k)\sqrt{n/k}=O(\sqrt{kn})$ . The solution computed by the reducer for the union of the core-sets is, by definition, a good approximation to the original problem. It is easy to see that the total communication complexity of this algorithm is $O(n)$ , and this computation can be performed in one or two rounds as formally defined in the MapReduce computation model .

Next, we elaborate on the application for a streaming computation model: In the random-order data stream model, a random sequence of $n$ data points needs to be processed “on-the-fly” while using only limited storage. An algorithm for a randomized composable core-set can be easily used to obtain an algorithm for this setting The paper introduced this approach for the special case of $k$ -median clustering. More general formulation of this method with other applications appeared in .. In particular, if a randomized composable core-set for a given problem has size $k$ , we start by dividing the random stream of data into $\sqrt{n/k}$ blocks of size $s=\sqrt{nk}$ . This way, each block will be a random subset of items. The algorithm then proceeds block by block. Each block is read and stored in the main memory, its core-set is computed and stored, and the block is deleted. At the end, the algorithm solves the problem for the union of the core-sets. The whole algorithm takes only $O(\sqrt{kn})$ space. The storage can be reduced further by utilizing more than one level of compression, at the cost of increasing the approximation factor.

Variants of the composable core-set technique have been applied for optimization under MapReduce framework . However, none of these previous results formally study the difference between randomized and non-randomized variants and in most cases, they employ non-randomized composable core-sets. Indyk et al. observed that the idea of non-randomized composable core-sets cannot be applied to the coverage maximization (or more generally submodular maximization) problems. In fact, all our hardness results also apply to a class of submodular maximization problems known as the maximum $k$ -coverage problems, i.e., given a number $k$ , and a family of subsets ${\cal A}\subset 2^{X}$ , find a subfamily of $k$ subsets $A_{1},\ldots,A_{k}$ whose union $\cup_{j=1}^{k}A_{j}$ is maximized. Solving max $k$ -coverage and submodular maximization in a distributed manner have attracted a significant amount of research over the last few years . Other than the importance of these problems, one reason for the popularity of this problem in this context is the fact that its approximation algorithm is algorithm $\mathsf{Greedy}$ which is naturally sequential and it is hard to parallelize or implement in a distributed manner.

3 Our Contributions

Our results are summarized in Table 1. As our first result, we prove that a family of efficient algorithms including a variant of algorithm $\mathsf{Greedy}$ with a consistent tie-breaking rule leads to an almost $1/3$ -approximate randomized composable core-set of size $k$ for any monotone submodular function and cardinality constraint $k$ with multiplicity of $1$ (see Section 2). This is in contrast to a known $O({\log k\over\sqrt{k}})$ hardness result for any (non-randomized) composable core-set , and shows the advantage of using the randomization here. Furthermore, by constructing this randomized core-set and applying algorithm $\mathsf{Greedy}$ afterwards, we show a $0.27$ distributed approximation factor for the monotone submodular maximization problem in one or two rounds of MapReduces with a linear communication complexity. Previous results lead to algorithms with either much larger number of rounds of MapReduce , and/or larger communication complexity . This improvement is important, since the number of rounds of MapReduce computation and communication complexity are the most important factors in determining the performance of a MapReduced-based algorithm . The effectiveness of using this technique has been confirmed empirically by Mirzasoleiman et al who studied a similar algorithm on a subclass of submodular maximization problems. However, they only provide provable guarantees for a subclass of submodular functions satisfying a certain Lipchitz condition . Our result not only works for monotone submodular functions, but also extends to non-monotone (non-negative) submodular functions, and leads to the first constant-round MapReduce-based constant-factor approximation algorithm for non-monotone submodular maximization (with $O(n)$ total communication complexity and approximation factor of $0.18$ ). It also leads to the first constant-factor approximation algorithm for non-monotone submodular maximization in a random-order streaming model in one pass with sublinear memory.

Our next goal is to improve the approximation factor of the above algorithm for monotone submodular functions. To this end, we first observe that one cannot achieve a better than the $1/2$ factor via core-sets of size $k$ using algorithm $\mathsf{Greedy}$ or any algorithm in a family of local search algorithms. In Section 4, we show how to go beyond the $1/2$ -approximation by applying core-sets of size higher than $k$ but still of size $O(k)$ , and prove that algorithm $\mathsf{Greedy}$ with a consistent tie-breaking rule provides a $0.585$ -approximate randomized composable core-set of size $k^{\prime}<4k$ for our problem. We then present algorithm $\mathsf{PseudoGreedy}$ that can be applied as a post-processing step to design a distributed $0.545$ -approximation algorithm in one or two rounds of MapReduces, and with linear total communication complexity. For monotone submodular maximization, this result implies the first distributed approximation algorithm with approximation factor better than $1/2$ that runs in a constant number of rounds. We achieve this approximation factor using one or two rounds of MapReduces and with the total communication complexity of $O(n)$ . In addition, this result implies the first approximation algorithm beating the $1/2$ factor for the random-order streaming model with constant number of passes on the data and sublinear memory. To complement this result, we first show that our analysis for algorithm $\mathsf{Greedy}$ is tight. Moreover, we show that it is information theoretically impossible to achieve an approximation factor better than $1-{1/e}$ using a core-set with size polynomial in $k$ .

Finally, we consider the construction of small-size core-sets, i.e., a core-set of size $k^{\prime}<k$ . Studying such core-sets is important particularly for cases with large parameter $k$ , e.g., $k=\Omega(n)$ or $k={n\over\log n}$ For such large $k$ , a core-set of size $k$ may not be as useful since outputting the whole core-set may be impossible. For example, in the formal MapReduce model , outputting a core-set of size $k$ for $k=\Omega(n)$ is not feasible.. For our problem, we first observe a hardness bound of $O({k^{\prime}\over k})$ for non-randomized core-sets. On the other hand, in Subsection 5.2, we show an $\Omega(\sqrt{k^{\prime}\over k})$ -approximate randomized composable core-set for this problem, and accompany this result by a matching hardness bound of $O(\sqrt{k^{\prime}\over k})$ for randomized composable core-set. The hardness result is presented in Subsection 5.1.

4 Other Related Work.

Submodular Maximization in Streaming and MapReduce: Solving max $k$ -coverage and submodular maximization in a distributed manner have attracted a significant amount of research over the last few years . From theoretical point of view, for the coverage maximization problem, Chierchetti et al. present a $(1-1/e)$ -approximation algorithm in polylogarithmic number of MapReduce rounds, and Belloch et al improved this result and achieved $\log^{2}n$ number of rounds. Recently, Kumar et al. present a $(1-1/e)$ -approximation algorithm using a logarithmic number of rounds of MapReduces. They also derive $(1/2-{\epsilon})$ -approximation algorithm that runs in $O({1\over\delta})$ number of rounds of MapReduce (for a constant $\delta$ ), but this algorithm needs a $\log n$ blowup in the communication complexity. As observed in various empirical studies , the communication complexity and the number of MapReduce rounds are important factors in determining the performance of a MapReduce-based algorithm and a $\log n$ blowup in the communication complexity can play a crucial role in applicability of the algorithm in practice. Our algorithm on the other hand runs only in (one or) two rounds, and can run on any number of machines as long as they can store the data, i.e, it needs $m$ machines each with memory proportional to ${1\over m}$ of the size of the input. One previous attempt to apply the idea of core-sets for submodular maximization is by Indyk et al. who rule out the applicability of non-randomized core-sets by showing a hardness bound of $O({\log k\over\sqrt{k}})$ for non-randomized core-sets. The most relevant previous attempt in applying randomized core-sets to submodular maximization is by Mirzasoleiman et al , where the authors study a class of algorithms similar to the ones discussed here, and show the effectiveness of applying algorithm $\mathsf{Greedy}$ over a random partitioning empirically for several machine learning applications. The authors also prove theoretical guarantees for algorithm $\mathsf{Greedy}$ for special classes of submodular functions satisfying a certain Lipschitz condition . Here, on the other hand, we present guaranteed approximation results for all monotone and non-monotone submodular functions. In fact, Ashwinkumar and Karbasi observed that if one applies the greedy algorithm without a consistent tie-breaking rule, the approximation factor of the algorithm is not bounded for coverage functions. In this paper, we prove that a class of algorithms including greedy with a consistent tie-breaking rule provide a guaranteed $0.27$ -approximation algorithm for all monotone submodular functions. We also show how to achieve an improved $0.54$ -approximation using slightly larger core-sets of size $O(k)$ . Finally, there is a recent paper in the streaming model in which the authors present a streaming $1/2$ -approximation algorithm with one pass and linear memory. Our results lead to improved results for random-order streaming model for monotone and non-monotone submodular maximization.

Core-sets: The notion of core-sets has been introduced in . In this paper, we use the term core-sets to refer to “composable core-sets” which was formally defined in a recent paper . This notion has been also implicitly used in Section 5 of Agarwal et al. where the authors specify composability properties of $\epsilon$ -kernels (a variant of core-sets). The notion of (composable) core-sets are also related to the concept of mergeable summaries that have been studied in the literature . As discussed before, the idea of using core-sets has been applied either explicitly or implicitly in the streaming model and in the MapReduce framework . Moreover, notions similar to randomized core-sets have been studied for random-order streaming models . Finally, the idea of random projection in performing algebraic projections can be viewed as a related topic but it does not discuss the concept of composability over a random partitioning.

5 More Notation

Consider a submodular set function $f$ . Let $\mathsf{ALG}$ be an algorithm that given any subset $T$ returns subset ${\mathsf{ALG}}(T)\subseteq T$ with size at most $k^{\prime}$ . We say that this Algorithm $\mathsf{ALG}$ is a $\beta$ -nice algorithm for function $f$ and some parameter $\beta$ iff for any set $T$ and any item $x\in T\setminus{\mathsf{ALG}}(T)$ (item $x$ is in set $T$ but is not selected in the output of algorithm $\mathsf{ALG}$ ), then the following two properties hold:

Set ${\mathsf{ALG}}(T\setminus\{x\})$ is equal to ${\mathsf{ALG}}(T)$ , i.e., intuitively the output of the algorithm should not depend on the items it does not select, and

$\Delta(x,{\mathsf{ALG}}(T))$ is at most $\beta\frac{f({\mathsf{ALG}}(T))}{k^{\prime}}$ . In other words, the marginal $f$ value of any not-selected item cannot be more than $\beta$ times the average contribution of selected items.

Randomized Core-sets for Submodular Maximization

In this section, we show that a family of $\beta$ -nice algorithms, introduced in Section 1.5, leads to a constant-factor approximate randomized composable core-set, and a constant-factor distributed approximation algorithm for monotone and non-monotone submodular maximization problems with cardinality constraints. Later, in Subsection 2.2, we show that several efficient algorithms in the literature of submodular maximization are $\beta$ -nice for some $\beta\in[1,1+\epsilon]$ (for $\epsilon=o(1)$ ) including some variant of algorithm $\mathsf{Greedy}$ with a consistent tie-breaking rule, and also an almost linear-time algorithm in . Before stating the theorem, we emphasize that in this section, we apply a composable core-set of multiplicity 1 which corresponds to a random partitioning of items into $m$ disjoint pieces.

For any $\beta>0$ , any $\beta$ -nice algorithm $\mathsf{ALG}$ is a $\frac{1}{2+\beta}$ -approximate randomized composable core-set of multiplicity 1 and size $k$ for the monotone, and $\frac{1-\frac{1}{m}}{2+\beta}$ -approximate for non-monotone submodular maximization problems with cardinality constraint $k$ .

We want to show that there exists a subset of $\cup_{i=1}^{m}S_{i}$ with size at most $k$ , and at least an expected $f$ value of $\frac{f(\textsc{OPT}{})}{2+\beta}$ for monotone $f$ , and $\frac{(1-1/m)f(\textsc{OPT}{})}{2+\beta}$ for non-monotone $f$ . Toward this goal, we take the maximum of $\max_{1\leq i\leq m}f(S_{i})$ and $f(\textsc{OPT}{}^{S})$ as a candidate solution. We define some notation to simplify the rest of the proof. Consider an arbitrary permutation $\pi$ on items of OPT, and for each item $x\in\textsc{OPT}{}$ let $\textsc{OPT}{}^{x}$ be the set of items in $\pi$ that appear before $x$ . We will first lower bound $f(\textsc{OPT}{}^{S})$ using the submodularity property in Lemma 2.2.

For any set of selected items, $f(\textsc{OPT}{}^{S})\geq f(\textsc{OPT}{})-\sum_{x\in\textsc{OPT}{}\setminus(\cup_{i=1}^{m}S_{i})}\Delta(x,\textsc{OPT}{}^{x})$ for a monotone or non-monotone submodular function $f$ .

First we note that $f(\textsc{OPT}{})-f(\textsc{OPT}{}^{S})=\sum_{x\in\textsc{OPT}{}\setminus\textsc{OPT}{}^{S}}\Delta(x,\textsc{OPT}{}^{x}\cup\textsc{OPT}{}^{S})$ . Using submodularity property, we know that $\Delta(x,\textsc{OPT}{}^{x}\cup\textsc{OPT}{}^{S})$ is at most $\Delta(x,\textsc{OPT}{}^{x})$ because $\textsc{OPT}{}^{x}$ is a subset of $\textsc{OPT}{}^{x}\cup\textsc{OPT}{}^{S}$ . Therefore $f(\textsc{OPT}{})-f(\textsc{OPT}{}\cap A)\leq\sum_{x\in\textsc{OPT}{}\setminus\textsc{OPT}{}^{S}}\Delta(x,\textsc{OPT}{}^{x})$ which is equal to $\sum_{x\in\textsc{OPT}{}\setminus(\cup_{i=1}^{m}S_{i})}\Delta(x,\textsc{OPT}{}^{x})$ by definition of $\textsc{OPT}{}^{S}$ . This concludes the proof. ∎

Lemma 2.2 suggests that we should upper bound $\sum_{x\in\textsc{OPT}{}\setminus(\cup_{i=1}^{m}S_{i})}\Delta(x,\textsc{OPT}{}^{x})$ which is done in the next lemma.

The sum $\sum_{i=1}^{m}\sum_{x\in\textsc{OPT}{}\cap T_{i}\setminus S_{i}}\Delta(x,\textsc{OPT}{}^{x})$ is at most $\beta\left(\max_{1\leq i\leq m}f(S_{i})\right)+$ $\sum_{i=1}^{m}\sum_{x\in\textsc{OPT}{}\cap T_{i}\setminus S_{i}}$ $\left(\Delta(x,\textsc{OPT}{}^{x})-\Delta(x,\textsc{OPT}{}^{x}\cup S_{i})\right)$ for a monotone or non-monotone submodular function $f$ .

The sum $\sum_{i=1}^{m}\sum_{x\in\textsc{OPT}{}\cap T_{i}\setminus S_{i}}\Delta(x,\textsc{OPT}{}^{x})$ can be written as:

The first term in the sum is upper bounded by $\beta\frac{f(S_{i})}{k}$ using the second property of $\beta$ -nice algorithms. To conclude the proof, we apply inequality $f(S_{i})\leq\max_{1\leq i^{\prime}\leq m}f(S_{i^{\prime}})$ , and use the fact that there are at most $k$ items in $\textsc{OPT}{}\setminus\textsc{OPT}{}^{S}=\cup_{i=1}^{m}(\textsc{OPT}{}\cap T_{i}\setminus S_{i})$ . ∎

At this stage of the analysis, we use randomness of partition $\{T_{i}\}_{i=1}^{m}$ to upper bound the expected value of these differences in $\Delta$ values with the expected value of average of $f(S_{i})$ . This is stated in the following lemma:

Before proving the lemmas, we observe that putting these three lemmas together, we can finish the proof of the theorem. In particular, for a monotone or non-monotone submodular function $f$ , we have that:

Proof of Lemma 2.4. The main part of the proof is to show that the sum of the $\Delta$ differences in the statement of the lemma is in expectation at most $\frac{1}{m}$ fraction of sum of $\Delta$ differences for a larger set of pairs $(i,x)$ . In particular, we show that

In Theorem 2.1, we prove that if on each part (set $T_{i}$ ) of the partitioning, we run a $\beta$ -nice algorithm $\mathsf{ALG}$ , the union of output sets of $\mathsf{ALG}$ will contain a set of size at most $k$ that preserves at least $\frac{1}{2+\beta}$ fraction of value of optimum set. If we run algorithm $\mathsf{Greedy}$ on the union of output sets $\cup_{i=1}^{m}S_{i}$ , using the classic analysis of algorithm $\mathsf{Greedy}$ , we can easily claim that the overall value of the output set at the end is at least $\frac{1-1/e}{2+\beta}$ fraction of $f(\textsc{OPT}{})$ . In Theorem 2.5, we show an improved distributed approximation factor.

Let $S$ be the output of algorithm $\mathsf{Greedy}$ over $\cup_{i=1}^{m}S_{i}$ , i.e. $S=\mathsf{Greedy}(\cup_{i=1}^{m}S_{i})$ for a monotone submodular function $f$ . Also let $S$ be the output of non-monotone submodular maximization algorithm of Buchbinder et al. on $\cup_{i=1}^{m}S_{i}$ when $f$ is a non-monotone submodular function. The expected value of $\max\{f(S),\max_{i=1}^{m}\{f(S_{i})\}\}$ is at least $\frac{(1-1/e)f(\textsc{OPT}{})}{1+(1-1/e)(1+\beta)}$ for monotone and $\frac{(1-\frac{1}{m})f(\textsc{OPT}{})/e}{1+(1+\beta)/e}$ for non-monotone submodular $f$ . In particular, for $\beta=1$ , the distributed approximation factors are $\geq 0.27$ , and $\geq 0.21-\frac{1}{4m}$ for monotone and non-monotone $f$ respectively.

By applying lemmas 2.2, 2.3, and 2.4, we have that: $f(\textsc{OPT}{}^{S})\geq f(\textsc{OPT}{})-\beta\max_{i=1}^{m}f(S_{i})-\frac{\sum_{i=1}^{m}f(S_{i})}{m}\geq f(\textsc{OPT}{})-(1+\beta)\max_{i=1}^{m}f(S_{i})$ for a monotone submodular function $f$ . Using the classic analysis of $\mathsf{Greedy}$ on submodular maximization in , one can prove that $f(S)\geq(1-1/e)f(\textsc{OPT}{}^{S})$ when $f$ is monotone. By taking the expectation of the two sides of this inequality, and using Lemma 2.4, we have that:

If $f$ is non-monotone, we get a weaker inequality $f(S)\geq\frac{f(\textsc{OPT}{}^{S})}{e}$ by applying the algorithm of Buchbinder et al. . Using lemmas 2.2, 2.3, and 2.4, we have that: $f(\textsc{OPT}{}^{S})\geq f(\textsc{OPT}{})-\beta\max_{i=1}^{m}f(S_{i})-\frac{\sum_{i=1}^{m}f(S_{i})}{m}-\frac{f(\textsc{OPT}{})}{m}\geq(1-\frac{1}{m})f(\textsc{OPT}{})-(1+\beta)\max_{i=1}^{m}f(S_{i})$ . Similarly, we can claim that $\rho\geq\frac{(1-\frac{1}{m})/e}{1+(1+\beta)/e}$ . This implies the desired lower bound of $\frac{(1-\frac{1}{m})/e}{1+(1+\beta)/e}$ on $\rho$ . ∎

2 Examples of β𝛽\beta-Nice Algorithms

In this section, we show that several existing algorithms for submodular maximization in the literature belong to the family of $\beta$ -nice algorithms.

Algorithm $\mathsf{Greedy}$ with Consistent Tie-breaking: First, we observe that algorithm $\mathsf{Greedy}$ is $1$ -nice if it has a consistent tie breaking rule: while selecting among the items with the same marginal value, $\mathsf{Greedy}$ can have a fixed strict total ordering ( $\Pi$ ) of the items, and among the set of items with the maximum marginal value chooses the one highest rank in $\Pi$ . The consistency of the tie breaking rule implies the first property of nice algorithms. To see the second property, first observe that (i) $\mathsf{Greedy}$ always adds an item with the maximum marginal $f$ value, and (ii) using submodularity of $f$ , the marginal $f$ values are decreasing as more items are added to the selected items. Therefore, after $k$ iterations, the marginal value of adding any other item, is less than each of the $k$ marginal $f$ values we achieved while adding the first $k$ items. This implies the 2nd property, and concludes that $\mathsf{Greedy}$ with a consistent tie-breaking rule is a $1$ -nice algorithm.

An almost linear-time $(1+\epsilon)$ -nice Algorithm: Badanidiyuru and Vondrak present an almost linear-time $(1-\frac{1}{e}-\epsilon)$ -approximation algorithm for monotone submodular maximization with a cardinality constraint. We observe that this algorithm is $(1+2\epsilon)$ -nice. The algorithm is a relaxed version of $\mathsf{Greedy}$ where in each iteration, it adds an item with almost maximum marginal value (with at least $1-\epsilon$ fraction of the maximum marginal). As a result, similar to the proof for $\mathsf{Greedy}$ , one can show that this linear-time algorithm $\frac{1}{1-\epsilon}$ -nice and consequently $(1+2\epsilon)$ -nice for $\epsilon\leq 0.5$ .

Hardness Results for Randomized Core-sets

In Section 2, we showed that a family of $\beta$ -nice algorithms are $\frac{1}{2+\beta}$ -approximate randomized core sets (e.g., $\frac{1}{3}$ -approximate for algorithm $\mathsf{Greedy}$ ). Here we show what kinds of randomized core-sets are not achievable. In particular, we prove, in Theorem 3.1 that if we restrict our attention to core-sets of size $k$ , algorithm $\mathsf{Greedy}$ or any local search algorithm does not achieve an approximation factor better than $\frac{1}{2}$ even if each item is sent to multiple machines (up to multiplicity $C=o(\sqrt{m})$ ). This leads to the following question: does increasing the output size of core-sets, $k^{\prime}$ , help with the approximation factor? In other words, can we get a better than $1/2$ approximation factor if we allow the algorithm to select more than $k$ items on each machine? To answer this question, we first prove, in Theorem 3.3 that it is not possible to achieve a randomized composable core-set of size $k^{\prime}=o(\frac{n}{Cm})$ with approximation better than $1-\frac{1}{e}$ even when we allow for multiplicity $C=o(\sqrt{\frac{m}{k}})$ . We then show in Section 4 that although it is not possible to beat the $1-\frac{1}{e}$ barrier, we can slightly increase the output sizes, apply algorithm $\mathsf{Greedy}$ to achieve an approximation factor $\approx 2-\sqrt{2}>\frac{1}{2}$ with a constant multiplicity.

Following we show a limitation on core-sets of size $k$ . In particular, we introduce a family of instances for which algorithm $\mathsf{Greedy}$ and any algorithm that returns a locally optimum solution of size at most $k$ do not achieve a better than $\frac{1}{2}+\epsilon$ -approximate core-set for any $\epsilon>0$ . This lower-bound result applies to a coverage valuation (and therefore submodular) function $f$ and it holds even if we send each item to multiple machines.

For any $\epsilon>0$ , assuming each item is sent to at least one random machine (to set $T_{i}$ for a random $1\leq i\leq n$ ), and at most $C\leq\sqrt{\frac{\epsilon m}{2}}$ random machines, and the number of items an algorithm is allowed to return is at most $k$ , there exists a family of instances for which algorithm $\mathsf{Greedy}$ and any other local search algorithm returns an at most $(\frac{1}{2}+\frac{1}{k}+\epsilon)$ -approximate composable core-set.

We say a machine is a good machine if for each $1\leq i\leq k$ , it receives at least a set $B_{i,j}$ for some $1\leq j\leq L$ , and we call it is a bad machine otherwise. At first, we show that the output of algorithm $\mathsf{Greedy}$ or any local search algorithm on a good machine that has not received set $A_{1}$ is exactly one set $B_{i,j}$ for each $1\leq i\leq k$ , and nothing else. In other words, these algorithms do not return any of the sets $A_{2},A_{3},\cdots,A_{k}$ unless they have a set $A_{1}$ as part of their input, or they are running on a bad machine. It is not hard to see that if $A_{1}$ is not part of the input each set $B_{i,j}$ has marginal value $k$ if no other set $B_{i,j^{\prime}}$ for $j^{\prime}\neq j$ has been selected before. On the other hand, the marginal value of each of the sets $\{A_{i}\}_{i=2}^{k}$ is $k-1$ . So $\mathsf{Greedy}$ or any local search algorithm does not select any of the sets $\{A_{i}\}_{i=2}^{k}$ unless some set $B_{i,j}$ has been selected for each $1\leq i\leq k$ . The fact that the output sizes are limited to $k$ implies that sets $\{A_{i}\}_{i=2}^{k}$ are not selected.

Now it suffices to prove that most machines are good, and most sets in $\{A_{i}\}_{i=2}^{k}$ are not sent to a machine that has set $A_{1}$ as well. To prove this, we show that each machine is good with probability at least $1-\frac{\epsilon}{2C}$ . To see this, note that for each $i$ , there are $L$ identical sets $B_{i,j}$ , and the probability that a machine does not receive any of these $L$ copies is at most $(1-\frac{1}{m})^{L}\leq e^{-L/m}\leq\frac{\epsilon}{2Ck}$ . So the probability that a machine is good is at least $(1-\frac{\epsilon}{2Ck})^{k}\geq 1-\frac{\epsilon}{2C}$ . Since each set is sent to at most $C$ machines, for each $1\leq i\leq k$ , we know that set $A_{i}$ is sent to only good machines with probability at least $(1-\frac{\epsilon}{2C})^{C}\geq 1-\frac{\epsilon}{2}$ . We also know that the probability that for each $2\leq i\leq k$ , the probability of set $A_{i}$ sharing a machine with set $A_{1}$ is at most $\frac{C^{2}}{m}\leq\frac{\epsilon}{2}$ since each set is sent to at most $C$ random machines. As a result, at most $\frac{\epsilon}{2}+\frac{\epsilon}{2}=\epsilon$ fraction of sets $\{A_{i}\}_{i=2}^{k}$ are selected by at least one machine in expectation, and therefore in expectation the size of the union of all selected sets (not only the best $k$ of them) is at most $k^{2}+\epsilon(k-1)^{2}$ which is less than $\frac{1}{2}+\frac{1}{k}+\epsilon$ fraction of the value of the optimum $k$ sets ( $\{A_{i}\}_{i=1}^{k}$ ). ∎

Following, we prove that it is not possible to achieve a better than $1-\frac{1}{e}+\epsilon$ approximation factor for submodular maximization subject to a cardinality constraint even if each item is sent to at most $C$ machines where $C\leq\sqrt{\epsilon m}$ , and each machine is allowed to return $k^{\prime}=\frac{\epsilon(1-\epsilon)n}{8Cm}$ items. We note that in this hardness result, the size of the output sets can be arbitrarily large in terms of $k$ , i.e. for instance $k^{\prime}$ could be $\Omega(2^{k})$ for some values of $n,m,C$ , and $k$ . This is an information theoretic hardness result that does not use any complexity theoretic assumption. In fact, the instance itself can be optimally solved on a single machine, but distributing the items among several machines makes it hard to preserve the optimum solution. Before presenting the hardness result, we state the following version of Chernoff bound (which we use in the proof) as given on page 267, Corollary $A.1.10$ and Theorem $A.1.13$ in :

Suppose $X_{1},X_{2},\cdots,X_{n}$ are $0-1$ random variables such that $Pr[X_{i}=1]=p_{i}$ , and let $\mu=\sum_{i=1}^{n}p_{i}$ , and $X=\sum_{i=1}^{n}X_{i}$ . Then for any $a>0$

For any $\epsilon>0$ , $k\geq\frac{8}{\epsilon}$ and $C\leq\sqrt{\frac{\epsilon m}{4k}}$ , assuming each item is sent to at most $C$ machines randomly, and each machine can output at most $k^{\prime}=\frac{\epsilon(1-\epsilon)}{8C}\times\frac{n}{m}$ items, there exists a family of instances for which no algorithm can guarantee a core-set of expected value more than $(1-\frac{1}{e}+\epsilon)$ fraction of the optimum solution.

So among the optimum solution items that are alone in their machines, at most $\frac{\epsilon}{4C}$ fraction of them will be selected. We note that each item is sent to at most $C$ machines, therefore in total $\frac{\epsilon k}{4}$ optimum items will be selected in this category in expectation.

On the other hand, the total number of optimum items that share a machine with some other optimum item is at most $\frac{(Ck)^{2}}{m}\leq\frac{\epsilon k}{4}$ (an upper bound on the expected number of collisions). We conclude that in total at most $\frac{\epsilon k}{2}$ optimum items will be selected in expectation. Therefore the total value of the selected optimum items does not exceed $\frac{\epsilon}{2}$ fraction of the optimum solution.

Better Randomized Core-sets for Monotone Submodular Maximization

In this section, we prove that although it is not possible to beat the $1-\frac{1}{e}$ barrier, we can slightly increase the output sizes (to $k^{\prime}=(\sqrt{2}+1)k$ ), and apply algorithm $\mathsf{Greedy}$ to achieve an approximation factor $\approx 2-\sqrt{2}>\frac{1}{2}$ with a constant multiplicity. Furthermore, we show in Theorem 4.8 that our analysis is tight for algorithm $\mathsf{Greedy}$ even if we increase the core-set sizes significantly. Finally, we present in Subsection 4.1 a post-processing algorithm $\mathsf{PseudoGreedy}$ that achieves an overall distributed approximation factor better than $1/2$ . In particular, after the first phase, we show how to find a size $k$ subset of the union of selected items with expected value at least $(0.545-o(1))f(\textsc{OPT}{})$ . Since in this section, we are dealing with a monotone submodular function $f$ , we can assume WLOG that $f(\emptyset)=0$ .

For any integer $C\geq 1$ , any cardinality constraint $k=o(m)$ , algorithm $\mathsf{Greedy}$ is a $\left(2-\sqrt{2}-O\left(\frac{1}{k}+\frac{\ln(C)}{C}\right)\right)$ -approximate randomized composable core-set of multiplicity $C$ and size $k^{\prime}=(2\sqrt{2}+1)k$ for any monotone submodular function $f$ . By letting $C={1\over\epsilon}$ , this leads to a randomized composable core-set of approximation factor $0.5857$ .

Let $D\stackrel{{\scriptstyle\text{def}}}{{=}}2\sqrt{2}+1$ , and $k^{\prime}=Dk$ . Following our notation from Section 1.5, let $S_{i}\stackrel{{\scriptstyle\text{def}}}{{=}}\mathsf{Greedy}(T_{i})$ for $1\leq i\leq m$ , where $T_{i}$ is the set of items sent to the machine $i$ . Note that we can let $|S_{i}|=Dk$ , since if there are less than $Dk$ items in $T_{i}$ , WLOG we can assume the algorithm returns some extra dummy items just for the sake of analysis.

Consider an item $x\in\textsc{OPT}{}$ . We say that $x$ survives from machine $i$ , if, when we send $x$ to machine $i$ in addition to items of $T_{i}$ , algorithm $\mathsf{Greedy}$ would choose this item $x$ in its output of size $k^{\prime}$ , i.e., if $x\in\mathsf{Greedy}(T_{i}\cup\{x\})\}$ . For the sake of analysis, we partition the optimum solution into two sets as follows: let $\textsc{OPT}{}_{1}$ be the set of items in the optimum solution that would survive the first machine, i.e., $\textsc{OPT}{}_{1}\stackrel{{\scriptstyle\text{def}}}{{=}}\{x|x\in\textsc{OPT}{}\cap\mathsf{Greedy}(T_{1}\cup\{x\})\}$ . Let $\textsc{OPT}{}_{2}\stackrel{{\scriptstyle\text{def}}}{{=}}\textsc{OPT}{}\setminus\textsc{OPT}{}_{1}$ , and $k_{1}\stackrel{{\scriptstyle\text{def}}}{{=}}|\textsc{OPT}{}_{1}|$ , and $k_{2}\stackrel{{\scriptstyle\text{def}}}{{=}}|\textsc{OPT}{}_{2}|$ (note that $k_{1}+k_{2}=k$ ). We also define $\textsc{OPT}{}_{1}^{\prime}\stackrel{{\scriptstyle\text{def}}}{{=}}\textsc{OPT}{}_{1}\cap\textsc{OPT}{}^{S}$ , and $\textsc{OPT}{}^{\prime}\stackrel{{\scriptstyle\text{def}}}{{=}}\textsc{OPT}{}_{1}^{\prime}\cup\textsc{OPT}{}_{2}$ where $\textsc{OPT}{}^{S}$ is defined in Subsection 1.5.

The optimum value $f(\textsc{OPT}{})$ is equal to $\sum_{x\in\textsc{OPT}{}}\Delta(x,\pi^{x})$ , and the term $f(\textsc{OPT}{})-f(\textsc{OPT}{}^{\prime})$ is at most $\sum_{x\in\textsc{OPT}{}_{1}\setminus\textsc{OPT}{}^{S}}\Delta(x,\pi^{x})$ .

By definition of $\Delta$ values, we have that: $\sum_{x\in\textsc{OPT}{}}\Delta(x,\pi^{x})=f(\textsc{OPT}{})-f(\emptyset)=f(\textsc{OPT}{})$ . Similarly, we have that $f(\textsc{OPT}{}^{\prime})$ is equal to $\sum_{x\in\textsc{OPT}{}^{\prime}}\Delta(x,\pi^{x}\cap\textsc{OPT}{}^{\prime})$ . By submodularity, we know that $\Delta(x,\pi^{x}\cap\textsc{OPT}{}^{\prime})$ is at least $\Delta(x,\pi^{x})$ because $\pi^{x}\cap\textsc{OPT}{}^{\prime}$ is a subset of $\pi^{x}$ . Therefore, we have:

where the last equality holds by definition of $\textsc{OPT}{}^{\prime}$ . ∎

We note that $\Delta(x,\pi^{x})$ is a fixed (non-random) term and therefore we can take it out of the expectation. Since $f(\textsc{OPT}{})$ is equal to $\sum_{x\in\textsc{OPT}{}}\Delta(x,\pi^{x})$ , we just need to prove that $Pr[x\in\textsc{OPT}{}_{1}\setminus\textsc{OPT}{}^{S}]$ is at most $O\left(\frac{\ln(C)}{C}\right)$ for any item $x\in\textsc{OPT}{}$ .

We note that the first machine is just a random machine, and the distribution of set of items sent to it, $T_{1}$ , is the same as any other set $T_{i}$ for any $2\leq i\leq m$ . We consider two cases for an item $x\in\textsc{OPT}{}$ :

The probability of $x$ being chosen when added to a random machine is at most $\frac{\ln(C)}{C}$ , i.e. $Pr[x\in\mathsf{Greedy}(T_{1}\cup\{x\})]=\frac{\ln(C)}{C}$ .

The probability $Pr[x\in\mathsf{Greedy}(T_{1}\cup\{x\})]$ is at least $\frac{\ln(C)}{C}$ .

In the first case, we know that $Pr[x\in\textsc{OPT}{}_{1}]\leq\frac{\ln(C)}{C}$ , and therefore $Pr[x\in\textsc{OPT}{}_{1}\setminus\textsc{OPT}{}^{S}]\leq\frac{\ln(C)}{C}$ which concludes the proof.

Using Lemma 4.2, it is sufficient prove that ratio $\frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ is at least $2-\sqrt{2}-O\left(\frac{1}{k}\right)$ . In order to lower bound the ratio $\frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ , we write the following factor-revealing linear program $LP^{k,k_{2}}$ , and prove in Lemma 4.4 that the solution to this LP is a lower bound on the aforementioned ratio.

For any integer $k>0$ , the ratio $\frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ is lower bounded by the optimum solution of linear program $LP^{k,k_{2}}$ for some integer $1\leq k_{2}\leq k$ .

We want to prove that ratio $\frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ is lower bounded by the solution of minimization linear program $LP^{k,k_{2}}$ . It suffices to construct one feasible solution with objective value $\beta$ equal to $\frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ . At first, we construct this solution for every instance of the problem, and then prove its feasibility in Claim 4.5.

We remind that $\textsc{OPT}{}^{\prime}$ is the union of two disjoint sets $\textsc{OPT}{}_{1}^{\prime}$ , and $\textsc{OPT}{}_{2}$ . Fix a permutation $\pi$ on the items of $\textsc{OPT}{}^{\prime}$ such that every item of $\textsc{OPT}{}_{1}^{\prime}$ appears before every item of $\textsc{OPT}{}_{2}$ in $\pi$ . In other words, $\pi$ is an arbitrary permutation on items of $\textsc{OPT}{}_{1}^{\prime}$ followed by an arbitrary permutation on items of $\textsc{OPT}{}_{2}$ . For any item $x$ in $\textsc{OPT}{}^{\prime}$ , define $\pi^{x}$ to be the set of items in $\textsc{OPT}{}^{\prime}$ that appear prior to $x$ in permutation $\pi$ . For any $1\leq j\leq Dk$ , we define set $S^{j}$ to be the first $j$ items of $S_{1}$ . We set the linear program variables as follows:

The above assignment forms a feasible solution of $LP^{k,k_{2}}$ , and its solution is equal to $\frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ .

The claim on the objective value of the solution is evident by definition of $\beta$ . To prove that the constraints hold, we show some simple and useful facts about the marginal values of items in OPT. We note that $\sum_{x\in\textsc{OPT}{}^{\prime}}\Delta(x,\pi^{x})=f(\textsc{OPT}{}^{\prime})$ by definition of $\Delta$ values and the fact that $f(\emptyset)=0$ . Similarly for any $1\leq j\leq Dk$ , we have that:

We are ready to prove that all constraints $(1),(2),\cdots,(5)$ one by one. We start with constraint $(1)$ . For any set $J\subset[DK]$ with size $|J|=k_{2}$ , we define $S(J)$ to be $\{y_{j}|j\in J\}$ where $y_{j}$ is the $j$ th item selected by algorithm $\mathsf{Greedy}$ in $S_{1}$ . We also define $S^{\prime}(J)$ to be $\textsc{OPT}{}_{1}^{\prime}\cup S(J)$ . Set $S^{\prime}(J)$ is a subset of $\textsc{OPT}{}_{1}^{\prime}\cup S_{1}$ with size at most $k_{2}+k_{1}=k$ . Therefore $f(S^{\prime}(J))$ is a lower bound on $f_{k}(\textsc{OPT}{}_{1}^{\prime}\cup S_{1})$ . We can also lower bound $f(S^{\prime}(J))$ as follows:

The first equality holds by definition of $\Delta$ . The first inequality holds by submodularity of $f$ , and knowing that $S^{j-1}\cap S(J)\subseteq S^{j-1}$ . The second equality holds by definition of $\alpha$ , and the last equality holds by definition of $c_{j}$ . We claim that $\Big{(}f(\textsc{OPT}{}_{1}^{\prime}\cup S^{j})-f(S^{j})\Big{)}-\Big{(}f(\textsc{OPT}{}_{1}^{\prime}\cup S^{j-1})-f(S^{j-1})\Big{)}$ (which is part of the right hand side of the last equality) is equal to $-b_{j}f(\textsc{OPT}{}^{\prime})$ . We note that $\Big{(}f(\textsc{OPT}{}_{1}^{\prime}\cup S^{j})-f(S^{j})\Big{)}$ is equal to $\sum_{x\in\textsc{OPT}{}_{1}^{\prime}}\Delta(x,\pi^{x}\cup S^{j})$ , and similarly $\Big{(}f(\textsc{OPT}{}_{1}^{\prime}\cup S^{j-1})-f(S^{j-1})\Big{)}$ is equal to $\sum_{x\in\textsc{OPT}{}_{1}^{\prime}}\Delta(x,\pi^{x}\cup S^{j-1})$ . By taking the difference of them, we have:

which is (by definition of $b_{j}$ ) equal to $-b_{j}f(\textsc{OPT}{}^{\prime})$ . We conclude that $f(S^{\prime})$ , and consequently $f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})$ are both at least $(1-\alpha+\sum_{j\in J}a_{j}+c_{j})f(\textsc{OPT}{}^{\prime})$ which concludes the proof of constraint $(1)$ .

We prove constraint $(2)$ using the fact that algorithm $\mathsf{Greedy}$ selects the item with maximum marginal value in each step. We note that the right hand side of constraint $(2)$ is $a_{j}+b_{j}+c_{j}$ which is by definition the marginal gain of item $y_{j}$ divided by $f(\textsc{OPT}{}^{\prime})$ , i.e. $\frac{\Delta(y_{j},S^{j-1})}{f(\textsc{OPT}{}^{\prime})}$ . We know that any item $x\in\textsc{OPT}{}_{2}$ will not be selected by algorithm $\mathsf{Greedy}$ if it is part of the input set which means that the marginal gain $(a_{j}+b_{j}+c_{j})f(\textsc{OPT}{}^{\prime})=\Delta(y_{j},S^{j-1})$ is at least the marginal gain $\Delta(x,S^{j-1})$ for any $x\in\textsc{OPT}{}_{2}$ , and it is also greater than the average of these marginal gains. In other words, we have:

To finish the proof of constraint $(2)$ , it suffices to prove the following inequality:

By definition of $\alpha$ , and $a$ values, we have that:

which completes the proof of Equation 2, and consequently constraint $(2)$ .

Now, we prove that constraint $(3)$ holds. By definition of $\alpha$ , and the fact that $f(\textsc{OPT}{}^{\prime})=$ $\sum_{x\in\textsc{OPT}{}^{\prime}}$ $\Delta(x,\pi^{x})$ , we know that $1-\alpha$ is equal to $\frac{\sum_{x\in\textsc{OPT}{}_{1}^{\prime}}\Delta(x,\pi^{x})}{f(\textsc{OPT}{}^{\prime})}$ . We also know that:

where the inequality holds because valuation function $f$ is monotone, and therefore all $\Delta$ values are non-negative. This proves that constraint $(3)$ holds.

To prove constraint $(4)$ , we should show that variables $a_{j},b_{j}$ , $c_{j}$ , and $\alpha$ are all in range $ $for any$ 1\leq j\leq DK $. We know that$ f(\textsc{OPT}{}^{\prime}) $is equal to$ \sum_{x\in\textsc{OPT}{}^{\prime}}\Delta(x,\pi^{x}) $. Therefore by definition,$ \alpha $is equal to$ \frac{\sum_{x\in\textsc{OPT}{}_{2}}\Delta(x,\pi^{x})}{\sum_{x\in\textsc{OPT}{}^{\prime}}\Delta(x,\pi^{x})} $. Since$ \textsc{OPT}{}_{2} $is a subset of$ \textsc{OPT}{}^{\prime} $, we imply that$ \alpha $is at most$ 1 $. We also know that$ \Delta $values are all non-negative, and therefore$ \alpha $is non-negative. Now, we should prove that variables$ a_{j},b_{j} $, and$ c_{j} $are all in range$ $. By definition,$ a_{j}+b_{j}+c_{j} $is equal to$ \frac{\Delta(y_{j},S^{j-1})}{f(\textsc{OPT}{}^{\prime})}\leq\frac{f(\{y_{j}\})-f(\emptyset)}{f(\textsc{OPT}{}^{\prime})}=\frac{f(\{y_{j}\})}{f(\textsc{OPT}{}^{\prime})} $where the inequality and equality are implied by the submodularity of$ f $, and the fact$ f(\emptyset)=0 $respectively. If$ f(\{y_{j}\}) $is at least$ f(\textsc{OPT}{}^{\prime}) $, the proof of Lemma 4.4 can be completed as follows. We know that$ f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})\geq f(\{y_{j}\}) $, and therefore the ratio$ \frac{f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})}{f(\textsc{OPT}{}^{\prime})} $is at least$ 1 $. On the other hand, there exists a very simple solution for$ LP^{k_{2},k} $with objective value$ \beta=1 $by just setting all variables to zero, and$ \beta $equal to one which completes the proof in the case$ f(\{y_{j}\})\geq f(\textsc{OPT}{}^{\prime}) $. So we can focus on the case,$ f(\{y_{j}\})\leq f(\textsc{OPT}{}^{\prime}) $in which we have$ a_{j}+b_{j}+c_{j}\leq\frac{f(\{y_{j}\})}{f(\textsc{OPT}{}^{\prime})}\leq 1 $. So it suffices to prove that these three variables are non-negative to prove that constraint$ (4) $holds. Variables$ a_{j} $and$ b_{j} $are non-negative because$ f $is submodular, and$ S^{j-1} $is a subset of$ S^{j} $. We use Equation 1 to prove non-negativity of$ c_{j} $. By definition,$ a_{j}+b_{j} $is equal to$ \frac{\sum_{x\in\textsc{OPT}{}^{\prime}}\Delta(x,\pi^{x}\cup S^{j-1})-\Delta(x,\pi^{x}\cup S^{j})}{f(\textsc{OPT}{}^{\prime})}$. By applying Equation 1, we have:

where the last inequality holds because of monotonicity of $f$ .

We prove constraint $(5)$ as follows. At first, we show that the right hand side of constraint $(5)$ is simply equal to $\frac{f(S^{k})}{f(\textsc{OPT}{}^{\prime})}$ . We know that $a_{j}+b_{j}+c_{j}=\frac{f(S^{j})-f(S^{j-1})}{f(\textsc{OPT}{}^{\prime})}$ for each $1\leq j\leq k$ . By a telescopic summation, we have that the right hand side of constraint $(5)$ , $\sum_{j=1}^{k}a_{j}+b_{j}+c_{j}$ , is equal to $\frac{f(S^{k})-f(\emptyset)}{f(\textsc{OPT}{}^{\prime})}=\frac{f(S^{k})}{f(\textsc{OPT}{}^{\prime})}$ . By definition of $\beta$ , and the fact that $f_{k}(S_{1}\cup\textsc{OPT}{}_{1}^{\prime})$ is at least $f(S^{k})$ , we conclude that constraint $(5)$ holds.

To prove constraint $(6)$ , we note that by definition, $a_{j}+b_{j}+c_{j}$ is $\Delta(y_{j},S^{j-1})$ . Since algorithm $\mathsf{Greedy}$ chooses the item with maximum marginal value at each step, we have $\Delta(y_{j},S^{j-1})\geq\Delta(y_{j+1},S^{j-1})$ . By submodularity, we have $\Delta(y_{j+1},S^{j-1})\geq\Delta(y_{j+1},S^{j})=a_{j+1}+b_{j+1}+c_{j+1}$ . We conclude that constraint $(6)$ holds. Therefore the proofs of Claim 4.5, and Lemma 4.4 are also complete. ∎

Finally, we show that the solution of $LP^{k,k_{2}}$ is at least $2-\sqrt{2}-O(\frac{1}{k})$ for any possible value of $k_{2}$ which concludes the proof of Theorem 4.1.

The optimum solution of linear program $LP^{k,k_{2}}$ is at least $2-\sqrt{2}-O(\frac{1}{{k}})$ for any $0\leq k_{2}\leq k$ .

We consider two cases: a) $k_{2}\leq\frac{k}{10}$ , and b) $k_{2}>\frac{k}{10}$ . We first consider the former case which is easier to prove, and then focus on the latter case. If $k_{2}$ is at most $\frac{k}{10}$ , we prove that the objective function of linear program $LP^{k,k_{2}}$ (which is $\beta$ ) cannot be less than $0.6$ which concludes the proof of this lemma. Since all variables are non-negative, we can apply constraint $(2)$ for each $j$ in range $[1,k]$ , and imply that $\sum_{j=1}^{k}a_{j}+b_{j}+c_{j}$ is at least $\left(1-\left(1-\frac{1}{k_{2}}\right)^{k}\right)\alpha$ . Using constraint $(5)$ , we know that $\sum_{j=1}^{k}a_{j}+b_{j}+c_{j}$ is a lower bound for $\beta$ . We are also considering the case $k_{2}\leq\frac{k}{10}$ , therefore $\beta$ is at least $\left(1-\left(1-\frac{1}{k_{2}}\right)^{k}\right)\alpha\geq\left(1-e^{-10}\right)\alpha\geq 0.9999\alpha$ . If $\alpha$ is at least $0.586$ , the claim of Lemma 4.6 is proved. So we assume $\alpha$ is at most $0.586$ . We define three sets of indices: $J_{1}\stackrel{{\scriptstyle\text{def}}}{{=}}\{1,2,\cdots,k_{2}\}$ , $J_{2}\stackrel{{\scriptstyle\text{def}}}{{=}}\{k_{2}+1,k_{2}+2,\cdots,2k_{2}\}$ , and $J_{3}\stackrel{{\scriptstyle\text{def}}}{{=}}\{2k_{2}+1,2k_{2}+2,\cdots,3k_{2}\}$ . We note that these sets have size $k_{2}$ , and therefore constraint $(1)$ should hold for them. If there exists some set $J\subset Dk$ with size $k_{2}$ such that $\sum_{j\in J}a_{j}+c_{j}$ is at least $0.3\alpha$ , we can use constraint $(1)$ to lower bound $\beta$ with $1-\alpha+0.3\alpha=1-0.7\alpha>2-\sqrt{2}$ . Therefore we can assume that $\sum_{j\in J_{i}}a_{j}+c_{j}$ is at most $0.3\alpha$ for $i\in\{1,2,3\}$ . Using constraint $(2)$ , we imply that:

We also know that $J_{1}\cup J_{2}\cup J_{3}\subset\{1,2,\cdots,k\}$ . We apply constraint $(5)$ to imply that $\beta$ is at least $(0.7+0.4+0.1)\alpha=1.2\alpha$ . This yields a stronger upper bound on $\alpha$ . If $\alpha$ is at least $\frac{2-\sqrt{2}}{1.2}<0.49$ , the claim of Lemma 4.6 is proved. So we assume $\alpha\leq 0.49$ . We follow the above argument one more time, and the proof is complete. If for some $i\in\{1,2,3\}$ , the sum $\sum_{j\in J_{i}}a_{j}+c_{j}$ is at least $0.16\alpha$ , using constraint $(1)$ , we can lower bound $\beta$ with $1-0.84\alpha>2-\sqrt{2}$ . Therefore we have $\sum_{j\in J_{i}}a_{j}+c_{j}\leq 0.16\alpha$ for each $i\in\{1,2,3\}$ which yields the following stronger inequalities:

By applying constraint $(5)$ , we have $\beta\geq(0.84+0.68+0.52)\alpha>2\alpha$ . We can also use constraint $(1)$ , and conclude that $\beta\geq\max\{2\alpha,1-\alpha\}>2-\sqrt{2}$ which completes the proof for the former case $k_{2}\leq\frac{k}{10}$ .

In the rest of the proof, we consider the latter case $k_{2}>\frac{k}{10}$ . The structure of the proof is as follows: In Claim 4.7, we first show that without loss of generality, one can assume a special structure in an optimum solution of $LP^{k,k_{2}}$ , and then exploit this structure to show that any solution of $LP^{k,k_{2}}$ is lower bounded by a simple system of two equations with some $O(\frac{1}{k_{2}})=O(\frac{1}{k})$ error. We can explicitly analyze this system of equations, and achieve a lower bound of $2-\sqrt{2}$ on $\beta$ (the key variable of the system of equations, and also the objective function of linear program $LP^{k,k_{2}}$ ).

There exists an optimum solution for linear program $LP^{k,k_{2}}$ with the following three properties:

constraint $(2)$ is tight for all $1\leq j\leq Dk$

It suffices to show that every optimum solution of $LP^{k,k_{2}}$ without changing the objective $\beta$ can be transformed to a feasible solution with the above three properties. Consider an optimum solution $\left(\beta^{*},\alpha^{*},\{a^{*}_{j},b^{*}_{j},c^{*}_{j}\}_{j=1}^{Dk}\right)$ . We start by showing how $c^{*}_{j}$ s can be set to zero. Suppose $c^{*}_{j}>0$ for some $1\leq j\leq Dk$ . We can increase the value of $a^{*}_{j}$ by $c^{*}_{j}$ , and then set $c^{*}_{j}$ to zero. This update keeps $a^{*}_{j}+c^{*}_{j}$ intact, and therefore does not change anything in constraints $(1)$ , $(3)$ , $(5)$ , and $(6)$ . It also makes it easier to satisfy constraint $(2)$ since it (possibly) reduces the right hand side, and keeps the left hand side intact. Constraint $(4)$ remains satisfied since $a^{*}_{j}+c^{*}_{j}\leq 1$ (otherwise $\beta$ is also at least $1$ which proves the claim of Lemma 4.6 directly). Therefore we can assume $c$ variables are equal to zero, and exclude them to have a simpler linear program:

Now we prove how to make $a^{*}$ variables monotone decreasing. Suppose for some $j_{1}<j_{2}$ , we have $a^{*}_{j_{1}}=a^{*}_{j_{2}}-2\delta$ for some positive $\delta$ . We set both variables $a^{*}_{j_{1}}$ and $a^{*}_{j_{2}}$ to their average, i.e. increase $a^{*}_{j_{1}}$ by $\delta$ , and decrease $a^{*}_{j_{2}}$ by $\delta$ .

We also decrease $b^{*}_{j_{1}}$ and increase $b^{*}_{j_{2}}$ by $\delta$ :

Now, we show that all constraints holds one by one. We note that it suffices to consider constraint $(1)$ only for set $J$ with maximum $\sum_{j\in J}a^{*}_{j}$ . We can assume that this set $J$ with maximum $\sum_{j\in J}a^{*}_{j}$ cannot contain $j_{1}$ without having $j_{2}$ (either before of after the update) because $a^{*}_{j_{1}}$ is at most $a^{*}_{j_{2}}$ in both cases. Therefore, the right hand side of constraint $(1)$ is intact, and it still holds.

Constraints $(2)$ , $(5)$ , and $(6)$ are all intact because $a^{*}_{j}+b^{*}_{j}$ is invariant in this operation for any $1\leq j\leq Dk$ .

Constraint $(3)$ holds because sum of $b^{*}$ variables remain the same. To prove constraint $(4)$ holds, it suffices to show that all variables stay in range $ $. It is evident for$ a^{*} $values since we are setting them to their average. For$ b^{*} $values, we first prove that$ b^{*}_{j_{1}} $stays non-negative. We note that$ a^{*}_{j_{1}}+b^{*}_{j_{1}}\geq a^{*}_{j_{2}}+b^{*}_{j_{2}} $using constraint$ (6) $. We also have$ a^{*}_{j_{1}}=a^{*}_{j_{2}} $, and$ b^{*}_{j_{2}}\geq 0 $. Therefore,$ b^{*}_{j_{1}} $cannot be negative. To prove that$ b^{*}_{j_{2}} $is (still) at most$ 1 $, it suffices to note that the sum of all$ b^{*} $values is (still) at most$ 1-\alpha\leq 1 $, and$ b^{*}$ variables are all non-negative. Therefore all constraints are still valid after this operation.

We prove that after a finite number of times (at most ${DK\choose 2}$ ) of applying this operation, we reach a feasible solution with monotone non-increasing sequence of $a^{*}$ values. If we start with $j_{1}=1$ , and do this operation for any pair $(j_{1},j_{2})$ with $a^{*}_{j_{1}}<a^{*}_{j_{2}}$ , after at most $Dk-1$ steps, we reach a solution in which $a^{*}_{1}\geq a^{*}_{j}$ for any $1<j\leq Dk$ . We continue the same process by increasing $j_{1}$ one by one, and after at most ${DK\choose 2}$ updates, we reach a sorted sequence of $a^{*}$ values. By monotonicity of $a^{*}$ values, we can simplify the linear program even further:

Now we prove that we can assume that constraint $(2)$ is tight for all $1\leq j\leq Dk$ . At first, we prove by contradiction that the right hand side of constraint $(2)$ is non-negative. Let $j$ be the minimum index for which the right hand side of constraint $(2)$ is negative. We can set all $a^{*}$ and $b^{*}$ values to zero for any index greater than $j$ , and also reduce $a^{*}_{j-1}$ by some amount to make this right hand side zero. All constraints hold, and we will have a solution in which the right hand side of constraint $(2)$ is always non-negative. Now we make constraint $(2)$ always tight as follows. Let $j_{1}$ be the maximum index in range $[1,Dk]$ , for which constraint $(2)$ is loose by some $\delta>0$ . We update the variables as follows. If $b^{*}_{j_{1}}$ is positive, we reduce it to $\max\{b^{*}_{j_{1}}-\delta,0\}$ , and do not change any other variable. We note that in this case, all constraints still hold, and constraint $(2)$ for all indices $j_{1},j_{1}+1,\cdots,Dk$ is tight.

If $b^{*}_{j_{1}}$ is zero, we decrease $a^{*}_{j_{1}}$ by $\delta$ , and for any $j_{2}>j_{1}$ , we increase $a^{*}_{j_{2}}$ by $\frac{\delta}{k_{2}}(1-\frac{1}{k_{2}})^{j_{2}-j_{1}-1}$ . We prove that constraint $(2)$ for all indices $j_{1},j_{1}+1,\cdots,Dk$ is tight, and all other constraints still hold after this update. By definition of $\delta$ , constraint $(2)$ is tight for index $j_{1}$ after this update. Constraint $(2)$ was tight before the update for $j_{2}>j_{1}$ because of the special choice of $j_{1}$ . We prove that the right and left hand sides of constraint $(2)$ increased by the same amount for each $j_{2}>j_{1}$ . The left hand side increased by $\frac{\delta}{k_{2}}(1-\frac{1}{k_{2}})^{j_{2}-j_{1}-1}$ . We also know that the right hand side changed by:

Therefore constraint $(2)$ is tight for all $j_{2}>j_{1}$ after the update. Now we prove feasibility of the new solution. For constraint $(1)$ , we note that all increments of $a^{*}$ variables is less than $\frac{\delta}{k_{2}}\sum_{r=0}^{\infty}(1-\frac{1}{k_{2}})^{r}=\delta$ . We also note that $a^{*}_{j_{1}}$ is decreased by $\delta$ . So the right hand side of constraint $(1)$ is decreased, and it remains feasible. We just showed that for $j_{2}\geq j_{1}$ , constraint $(2)$ is tight and therefore valid, and it is intact for smaller indices. Constraint $(3)$ is also intact. To prove constraint $(4)$ , we should show that the new $a^{*}$ values are in range $ $. Since constraint$ (2) $is tight, and$ \alpha $is at most$ 1 $, these new$ a^{*} $values are all at most$ 1 $. Non-negativity of the right hand side of constraint$ (2) $implies that these new values are all non-negative. Constraint$ (5) $holds since its right hand side is only decreased. To prove constraint$ (6) $, we note that the right hand sides of constraint$ (2) $is decreasing in$ j $, and they are all tight for indices$ \geq j_{1} $. Therefore constraint$ (6) $holds for$ j\geq j_{1} $. For$ j=j_{1}-1 $, it clearly holds since we are decreasing$ a^{*}_{j_{1}} $, and consequently its right hand side. To prove constraint$ (7) $, we note that the increments in$ a^{*} $values is decreasing as$ j_{2}>j_{1} $increases. So constraint$ (7) $remains feasible for$ j>j_{1} $. For$ j=j_{1} $, we note that$ b^{*}_{j_{1}} $is zero. So using constraint$ (6) $, we can prove constraint$ (7) $holds for$ j=j_{1} $which completes the feasibility proof. Each time, we make these updates, the index$ j_{1} $(the maximum index for which constraint$ (2) $is loose) reduces by at least$ (1) $. Therefore after at most$ Dk$ operations, we have an optimum solution with all three properties of Claim 4.7. ∎

Using Claim 4.7, we can assume that the solution of $LP^{k,k_{2}}$ is lower bounded by the next $LP^{new,k,k_{2}}$ . We focus on lower bounding the solution of $LP^{new,k,k_{2}}$ in the rest of the proof.

We note that we eliminated one of the lower bounds on $\beta$ . This only reduces the optimum solution of the linear program which is consistent with our approach. We also used Claim 4.7 to replace the inequality constraint $(2)$ with an equality constraint. We also removed the $c$ variables, and added the monotonicity constraint $(6)$ . In the rest of the proof, we introduce some notation, and show some extra structure in the optimum solution of $LP^{new,k,k_{2}}$ . This will help us lower bound the optimum solution by analyzing a system of two equations explicitly.

We start with proving the extra structure. Let $\tau\stackrel{{\scriptstyle\text{def}}}{{=}}a_{k_{2}}$ . We show that for any pair of indices $1\leq j_{1}<j_{2}\leq k_{2}$ , either $b_{j_{1}}$ is zero or $a_{j_{2}}$ is equal to $\tau$ . Let $j_{1}$ be the minimum index with $b_{j_{1}}>0$ . If $j_{1}$ is at least $k_{2}$ , the claim holds clearly. So we consider $j_{1}<k_{2}$ . If $a_{j_{1}+1}$ is equal to $\tau$ , by monotonicity of $a$ values, the claim is proved. So we define $\delta>0$ to be $\min\{b_{j_{1}},a_{j_{1}+1}-\tau\}$ . We increase $a_{j_{1}}$ , and $b_{j_{1}+1}$ by $\delta$ , and $\delta(1-\frac{1}{k_{2}})$ respectively. We also decrease both of $a_{j_{1}+1}$ , and $b_{j_{1}}$ by $\delta$ . After these changes, constraints $(1)$ , $(2)$ , $(3)$ , $(4)$ , and $(5)$ in $LP^{new,k,k_{2}}$ still hold. In particular, we made the changes in this special way to make sure that constraint $(2)$ still holds. Constraint $(5)$ also holds since the right hand side of constraint $(2)$ is decreasing in $j$ . But the monotonicity constraint $(6)$ may be violated for $j=j_{1}+1$ . This happens if the new $a_{j_{1}+1}$ is less than $a_{j_{1}+2}$ . In this case, we swap the variables $a_{j_{1}+1}$ and $a_{j_{1}+2}$ . We also change $b_{j_{1}+1}$ , and $b_{j_{1}+2}$ in a way that constraint $(2)$ holds for both $j=j_{1}+1$ , and $j=j_{1}+2$ . Similarly, we have that all constraints $(1),(2),\cdots,(5)$ hold. But constraint $(6)$ may be violated for $j=j_{1}+2$ . We continue doing this swap operation until constraint $(6)$ holds as well. This will happen after at most $k_{2}$ swap operations, since the new $a$ values are all at least $\tau$ . Finally, we reach a feasible solution for $LP^{new,k,k_{2}}$ in which either one more $a$ variable is equal to $\tau$ or one more $b$ variable is set to zero. Therefore after at most $2k_{2}$ updates, for any pair of indices $1\leq j_{1}<j_{2}\leq k_{2}$ , we have that either $b_{j_{1}}$ is zero or $a_{j_{2}}$ is equal to $\tau$ .

We also claim that for any $j>k_{2}$ , we can assume either $a_{j}=\tau$ , or $b_{j}=0$ . Otherwise, suppose $j_{1}>k_{2}$ is the smallest index for which $a_{j_{1}}<\tau$ , and $b_{j_{1}}>0$ . We can increase $a_{j_{1}}$ by $\delta=\min\{\tau-a_{j_{1}},b_{j_{1}}\}$ , and decrease $b_{j_{1}}$ by $\delta$ . We note that since $a_{j_{1}}+b_{j_{1}}$ is invariant, the monotonicity constraints $(5)$ still holds. We need to prove constraint $(6)$ for $j=j_{1}-1$ . It can be violated only if $a_{j_{1}-1}$ is less than $\tau$ , and $b_{j_{1}-1}$ is non-negative which contradicts the choice of $j_{1}$ . To prove feasibility, we only need to prove constraint $(2)$ for $j>j_{1}$ . We make it hold by the following adjustments. We start by $j=j_{1}+1$ , and increase it one by one. If constraint $(2)$ is loose by some $\epsilon$ for index $j$ , we decrease $b_{j}$ by $\min\{b_{j},\epsilon\}$ . We also decrease $a_{j}$ by $\max\{\epsilon-b_{j},0\}$ . The sum $a_{j}+b_{j}$ is reduced by $\epsilon$ , and constraint $(2)$ now holds for $j$ . It is also clear that sum of $b$ variables do not increase, and all other constraints still hold. With these adjustments for each $j>j_{1}$ in the increasing order, we know the solution is feasible. We also have all the ingredients to characterize the optimum solution of $LP^{new,k,k_{2}}$ , and lower bound it.

We can formalize this optimum solution in terms of a few parameters including $\tau,\beta,k,$ and $k_{2}$ . At this final stage of the proof, we conclude with two lower bounds (system of two equations) on $\beta$ and $1-\alpha$ in terms of these few parameters. Let $t$ be the smallest index in range $1\leq t\leq k_{2}$ with $b_{t}\neq 0$ . If such an index does not exist, define $t$ to be $k_{2}$ . Using Claim 4.7, we know that constraint $(2)$ is tight, and $b_{j}=0$ for any $j<t$ , we can inductively prove that $a_{j}=\frac{\alpha}{k_{2}}(\frac{k_{2}-1}{k_{2}})^{j-1}$ for any $j<t$ . Consequently, we have that $\sum_{j=1}^{t-1}a_{j}$ is equal to $\alpha(1-(\frac{k_{2}-1}{k_{2}})^{t-1})\geq\alpha(1-e^{-r})$ where $r$ is defined to be $\frac{t-1}{k_{2}}$ . Therefore, constraint $(1)$ implies that: $\beta\geq 1-\alpha+(1-e^{-r})\alpha+(1-r)k_{2}\tau$ . This lower bound on $\beta$ is the first inequality we wanted to prove. To achieve the second inequality (lower bound on $1-\alpha$ ), we start by upper bounding the sum of $a$ variables.

We show that $\sum_{j=1}^{t}a_{j}\leq\alpha\left(1-e^{-r}+\frac{2}{k_{2}}\right)$ as follows. Since $a$ variables are monotone and, constraint $(2)$ is tight for $j=1$ , we have $a_{t}\leq a_{1}=\frac{\alpha}{k_{2}}$ . We also have that $\sum_{j=1}^{t-1}a_{j}$ is equal to $\alpha\left(1-\left(1-\frac{1}{k_{2}}\right)^{t-1}\right)$ . We can upper bound $\left(1-\frac{1}{k_{2}}\right)^{k_{2}-1}$ by $e^{-1}$ as follows. We prove that $\left(1-\frac{1}{k_{2}}\right)^{k_{2}-1}$ is a monotone decreasing sequence for $k_{2}=1,2,\cdots$ .

We also know that $\lim_{k_{2}\to\infty}\left(1-\frac{1}{k_{2}}\right)^{k_{2}-1}=e^{-1}$ . Therefore each term $\left(1-\frac{1}{k_{2}}\right)^{k_{2}-1}$ is at least $e^{-1}$ . Therefore, we have that

which yields the desired upper bound on $\sum_{j=1}^{t}a_{j}$ .

We conclude that if $\beta^{*}$ is the solution of linear program $LP^{new,k,k_{2}}$ , the following system of equations should have a solution with $\beta=\beta^{*}$ :

where $\lambda$ is defined to be $k_{2}\tau$ . We note that $\alpha,\lambda,$ and $r$ are the variables of the above system of two equations, and they should be in range $ $. We also note that$ k_{2} $is another variable which can be any positive integer. To simplify, we solve the following system of equations to eliminate$ k_{2}$:

It is easy to see that if system of equations 3 has a solution $(\beta_{1},\alpha_{1},\lambda_{1},r_{1},k_{2})$ , system of equations 4 has the following solution: $(\beta=\beta_{1}+\frac{2}{k_{2}},\alpha=\alpha_{1},\lambda=\lambda_{1}+\frac{2}{k_{2}},r=r_{1})$ . Therefore it suffices to lower bound $\beta$ in system of equations 4. Because the same lower bound plus the term $\frac{2}{k_{2}}=O(\frac{1}{k})$ holds for $\beta$ in system of equations 3.

By computing the partial derivatives, and considering boundary values, one can find the minimum $\beta$ for which the system of equations 4 has a valid solution. Its minimum occurs when $r$ is zero, and the second inequality is tight. Therefore we have $\alpha=\sqrt{\lambda(2-\lambda)}$ . We conclude that $\beta$ is the minimum of $1-\sqrt{\lambda(2-\lambda)}+\lambda$ which is equal to $1-\sqrt{(1-\sqrt{\frac{1}{2}})(1+\sqrt{\frac{1}{2}})}+(1-\sqrt{\frac{1}{2}})=2-2\sqrt{\frac{1}{2}}=2-\sqrt{2}\approx 0.5857$ and occurs at $\lambda=1-\sqrt{\frac{1}{2}}\approx 0.2928$ .

So we have a slightly different set of two inequalities in this case to lower bound $\beta$ .

It is evident that both inequalities should be tight to minimize $\beta$ , and therefore $\alpha$ is equal to $\frac{1+\lambda}{1+D^{\prime}e^{-r}}$ where $D^{\prime}$ is $\frac{D-1}{2}$ . So we can write $\beta$ as a function of just $\lambda$ and $r$ :

To minimize $\beta$ , either $\lambda$ should be at one of its boundary values $\{0,1\}$ , or the partial derivative $\frac{\partial\beta}{\partial\lambda}$ should be zero. For $\lambda=1$ , $\beta$ cannot be less than $2-r-e^{-r}\geq 1-\frac{1}{e}>2-\sqrt{2}$ . For $\lambda=0$ , we have $1-\alpha\geq D^{\prime}\alpha^{\prime}\geq D^{\prime}\alpha$ , so $\alpha$ is at most $\frac{1}{1+D^{\prime}}=\sqrt{2}-1$ , and therefore $\beta$ is at least $1-\alpha\geq 2-\sqrt{2}$ . The only case to consider is when $\frac{\partial\beta}{\partial\lambda}=0$ which means $\frac{e^{-r}}{1+D^{\prime}e^{-r}}$ should be equal to $1-r$ with a unique solution $r^{*}=0.71\pm 0.001$ . Therefore $\beta$ is equal to $1-(1-r^{*})(1+\lambda-\lambda)=r^{*}>2-\sqrt{2}$ . We conclude that any feasible solution of linear program $LP^{k,k_{2}}$ has $\beta$ at least $2-\sqrt{2}-O(\frac{1}{k})$ which completes the proof. ∎

We show in the following Theorem that the $(2-\sqrt{2})\approx 0.585$ lower bound on the approximation ratio of the core-sets that $\mathsf{Greedy}$ finds is tight even if we allow the core-sets to be significantly large.

For any $\epsilon>0$ , and any core-set size $k^{\prime}\geq k$ , there are instances of monotone submodular maximization problem with cardinality constraint $k$ for which $\mathsf{Greedy}$ is at most a $(2-\sqrt{2}+O(\epsilon))$ -approximate randomized composable core-set even if each item is sent to $C\leq\sqrt{\epsilon m}$ machines.

We also note that for any $1\leq i\leq k_{2}$ , with probability at least $1-\epsilon$ , none of the at most $C$ copies of set $R_{i}$ shares a machine with one copy of $B^{\prime}$ . There are at most $C$ copies of $B^{\prime}$ , and $C$ copies of $R_{i}$ , and the probability that two sets are sent to the same machine is $\frac{1}{m}$ . So with probability at most $\frac{C^{2}}{m}\leq\epsilon$ , one copy of $R_{i}$ and one copy of $B^{\prime}$ are sent to the same machine.

where the last inequality holds for $k=\Omega(\frac{1}{\epsilon})$ . On the other hand, the optimum solution consists of $k_{2}=k-1$ row sets $R_{i}$ s, and set $B^{\prime}$ with value at least:

We conclude that the expected value of maximum value size $k$ subset of selected sets is upper bounded by $\lambda+1-\alpha+O(\epsilon)=2-\sqrt{2}+O(\epsilon)$ times the optimum solution $f(\textsc{OPT}{})$ . ∎

We remind that in the first phase, each machine $1\leq i\leq m$ runs algorithm $\mathsf{Greedy}$ on set $T_{i}$ with $k^{\prime}=Dk$ (where $D$ is $2\sqrt{2}+1$ ).By Theorem 4.1, there exists a size $k$ subset of selected items $\cup_{i=1}^{m}S_{i}$ with expected value at least $0.585f(\textsc{OPT}{})$ , but we do not know how to find this set efficiently. If we apply algorithm $\mathsf{Greedy}$ again on $\cup_{i=1}^{m}S_{i}$ to select $k$ items in total, we achieve a distributed approximation factor of $(1-\frac{1}{e})\times 0.585\approx 0.37$ . In the following, we present a post-processing algorithm $\mathsf{PseudoGreedy}$ that achieves an overall distributed approximation factor better than $1/2$ . In particular, after the first phase, we show how to find a size $k$ subset of the union of selected items $\cup_{i=1}^{m}S_{i}$ with expected value at least $(0.545-o(1))f(\textsc{OPT}{})$ .

Algorithm $\mathsf{PseudoGreedy}$ proceeds as follows: it first computes a family of candidate solutions of size $k+O(1)$ , and keeps the one candidate solution $V$ with the maximum value. It then lets $S$ to be a random size $k$ subset of $V$ , and returns $S$ as the solution. These candidate solutions, denoted by $S_{k^{\prime}_{2},I}$ (for $1\leq k^{\prime}_{2}\leq k$ and $4I\subseteq\{1,\cdots,8\}$ ) in Algorithm 1, are computed as follows: We first enumerate all $k$ possible values of $1\leq k^{\prime}_{2}\leq k$ (this notation is used to be consistent with the proof of Theorem 4.1). Then, by letting $k^{\prime}_{1}=k-k^{\prime}_{2}$ , and $k_{3}=32\lceil\frac{k^{\prime}_{2}}{128}\rceil$ , we partition the first $8k_{3}$ items in set $S_{1}$ We choose any machine instead of machine 1, and since the clustering is done at random, the analysis goes through. into eight subsets $\{A_{i^{\prime}}\}_{i^{\prime}=1}^{8}$ , each of size $k_{3}$ . The next step of the algorithm proceeds as follows: for any $I\subseteq\{1,2,\cdots,8\}$ with $\lvert I\rvert\leq 4+\frac{k^{\prime}_{1}}{k_{3}}$ , we initialize set $S_{k^{\prime}_{2},I}$ with the union of all sets $A_{i^{\prime}}$ where $i^{\prime}\in I$ . Then for $k^{\prime}_{1}+(4-\lvert I\rvert)k_{3}$ iterations, we search all items in $\cup_{i=1}^{m}S_{i}$ , and insert the item with the maximum marginal value to $S_{k_{2},I}$ . Roughly speaking, by starting from all these subsets, we ensure that the selected set hits enough number of items in OPT. The upper bound we enforce on $\lvert I\rvert$ is to make sure that the number of iterations at this step is non-negative, i.e. $k^{\prime}_{1}+(4-\lvert I\rvert)k_{3}\geq 0$ . Finally we define $V$ to be the set $S_{k^{\prime}_{2},I}$ with the maximum $f$ value, and return a random subset of size $k$ of $V$ as the output set $S$ .

1 Input: A collection of $m$ subsets $\{S_{1},\ldots,S_{m}\}$ . 2[0.6ex] Output: Set $S\subset\cup_{i=1}^{m}S_{i}$ with $|S|\leq k$ . 3[0.6ex] $V\leftarrow\emptyset$ ; 4 forall the $1\leq k^{\prime}_{2}\leq k$ do 5 $k_{3}\leftarrow 32\lceil\frac{k^{\prime}_{2}}{128}\rceil$ ; 6 $k^{\prime}_{1}\leftarrow k-k^{\prime}_{2}$ ; 7 Partition the first $8k_{3}$ items of $S_{1}$ into $8$ sets $\{A_{i^{\prime}}\}_{i^{\prime}=1}^{8}$ each of size $k_{3}$ ; 8 forall the $I\subseteq\{1,2,\cdots,8\}$ with $|I|\leq 4+\frac{k^{\prime}_{1}}{k_{3}}$ do 9 $S_{k^{\prime}_{2},I}\leftarrow\cup_{i^{\prime}\in I}A_{i^{\prime}}$ ; 10 for $k^{\prime}_{1}+(4-|I|)k_{3}$ times do 11 Find $\operatorname*{arg\,max}_{x\in\cup_{i=1}^{m}S_{i}}\Delta(x,S_{k^{\prime}_{2},I})$ , and insert it to $S_{k^{\prime}_{2},I}$ ; 12 13 end for 14 if $f(S_{k_{2},I})>f(V)$ then $V\leftarrow S_{k_{2},I}$ ; 15 16 end forall 17 18 end forall 19 $S\leftarrow$ a random size $k$ subset of $V$ ; Algorithm 1 Algorithm $\mathsf{PseudoGreedy}$

Algorithm $\mathsf{PseudoGreedy}$ returns a subset $S$ with size at most $k$ , and expected value at least $(0.545-O(\frac{1}{k}+\frac{\ln(C)}{C}))f(\textsc{OPT}{})$ .

Since algorithm $\mathsf{PseudoGreedy}$ enumerates all $k$ possible values of $k^{\prime}_{2}$ , in one of these trials, $k^{\prime}_{2}$ is equal to $k_{2}=|\textsc{OPT}{}_{2}|$ . From now on, we focus on this specific value of $k^{\prime}_{2}$ . We remind that $k_{3}=32\lceil\frac{k^{\prime}_{2}}{128}\rceil$ . Let $I^{\prime}$ be the set of indices that maximizes $S_{k_{2},I}$ , i.e., $I^{\prime}\stackrel{{\scriptstyle\text{def}}}{{=}}\operatorname*{arg\,max}_{I\subseteq\{1,2,\cdots,8\}\&|I|\leq 4+\frac{k_{1}}{k_{3}}}f(S_{k_{2},I})$ . Define $V^{\prime}$ be the set $S_{k_{2},I^{\prime}}$ . By definition, we have $f(V^{\prime})\leq f(V)$ . So it suffices to show that $f(V^{\prime})$ is at least $f(\textsc{OPT}{}^{\prime})$ times the solution of $LP^{r}$ for some $0\leq r\leq 1$ as follows. We define $r$ to be $\frac{4k_{3}}{4k_{3}+k_{1}}$ .

At first, we remind the feasible solution of $LP^{k,k_{2}}$ constructed in the proof of Lemma 4.4, and then construct a feasible solution for $LP^{r}$ with $\beta$ equal to $\frac{f(V^{\prime})}{f(\textsc{OPT}{}^{\prime})}$ as follows. Fix a permutation $\pi$ on the items of $\textsc{OPT}{}^{\prime}$ such that every item of $\textsc{OPT}{}_{1}^{\prime}$ appears before every item of $\textsc{OPT}{}_{2}$ in $\pi$ . In other words, $\pi$ is an arbitrary permutation on items of $\textsc{OPT}{}_{1}^{\prime}$ followed by an arbitrary permutation on items of $\textsc{OPT}{}_{2}$ . For any item $x$ in $\textsc{OPT}{}^{\prime}$ , define $\pi^{x}$ to be the set of items in $\textsc{OPT}{}^{\prime}$ that appear prior to $x$ in permutation $\pi$ . We set $\alpha$ to be $\frac{\sum_{x\in\textsc{OPT}{}_{2}}\Delta(x,\pi^{x})}{f(\textsc{OPT}{}^{\prime})}$ . For any $1\leq j\leq 8k_{3}$ , we define set $S^{j}$ to be the first $j$ items of $S_{1}$ . Now we are ready to define variables $\{a_{j},b_{j},c_{j}\}_{j=1}^{8k_{3}}$ as follows.

We let the variable $c_{j}$ be the marginal gain of item $y_{j}$ divided by $f(\textsc{OPT}{}^{\prime})$ minus $a_{j}+b_{j}$ where $y_{j}$ is the $j$ th item selected in $S_{1}$ , i.e., $c_{j}:=\frac{\Delta(y_{j},S^{j-1})}{f(\textsc{OPT}{}^{\prime})}-a_{j}-b_{j}$ .

We note that the last inequality holds assuming the numerator is non-negative. In case, the numerator is negative, the constraint $(2)$ holds using non-negativity of $a^{\prime},b^{\prime},$ and $c^{\prime}$ variables. So we just need to prove constraint $(1)$ (which is the most important constraint) of $LP^{r}$ . We remind that $\beta$ is defined to be $\frac{f(V^{\prime})}{f(\textsc{OPT}{}^{\prime})}=\max_{I\subseteq\{1,2,\cdots,8\}\&|I|\leq 4+\frac{k_{1}}{k_{3}}}\frac{f(S_{k_{2},I})}{f(\textsc{OPT}{}^{\prime})}$ . So for every $I\subseteq\{1,2,\cdots,8\}$ with size at most $4+\frac{k_{1}}{k_{3}}$ , it suffices to prove that:

We note that $S_{k_{2},I}$ consists of two sets of items:

Set $S(I)$ with $|I|k_{3}$ items corresponding to sets $\{B_{i^{\prime}}\}_{i^{\prime}\in I}$ which are added in the first phase. In other words, $S(I)$ is $\cup_{i^{\prime}\in I}B_{i^{\prime}}=\cup_{i^{\prime}\in I}\cup_{j=(i-1)k_{3}+1}^{ik_{3}}\{y_{j}\}$ where $y_{j}$ is the $j$ th item of $S_{1}$ .

$k_{1}+(4-|I|)k_{3}$ items added greedily in the second phase.

Similar to proof of Claim 4.5, we define $S^{\prime}(I)$ to be $\textsc{OPT}{}_{1}^{\prime}\cup S(I)$ . To be consistent in notation, we define $J$ to be the indices of items in $S(I)$ , i.e. $J\stackrel{{\scriptstyle\text{def}}}{{=}}\{j|y_{j}\in S(I)\}$ . Using the same argument in proof of Claim 4.5, we know that:

This means that there are $|\textsc{OPT}{}_{1}^{\prime}|\leq k_{1}$ items that can be added to $S(I)$ in the second phase, and increase its value to $f(S^{\prime}(I))$ . Since we are adding $k_{1}+(4-|I|)k_{3}$ items greedily in the second phase, our final value $f(S_{k_{2},I})$ is at least:

where the equality holds by definition of $r$ . We can also lower bound $f(S(I))$ as follows.

where the inequality holds by submodularity of $f$ , and equalities hold by definition. By combining inequalities 6, 7, and 8, we conclude that:

Small-size Core-sets for Submodular Maximization

We start by presenting the hardness results for non-randomized core-sets.

For any $1\leq k^{\prime}\leq k$ , the output of any algorithm is at most an $O(\frac{k^{\prime}}{k})$ -approximate composable non-randomized core-set for the submodular maximization problem.

Now we construct the following family of instances to achieve hardness results for randomized core-sets in the submodular maximization problem.

We define instance $I^{k,k^{\prime}}$ of the submodular maximization problem as follows. Define $\Gamma$ to be $\lfloor\sqrt{k^{\prime}k}\rfloor$ . For each $1\leq i\leq k-\Gamma$ , we add an item that represents the set $\{i\}$ . For each $k-\Gamma<i\leq k$ , we add $\lfloor\frac{k}{k^{\prime}}\rfloor$ identical items all representing the same set $\{i\}$ . The value of a subset $S$ of items, $f(S)$ , is equal to the cardinality of the union of all sets the items in $S$ represent. This is a coverage valuation function and subsequently monotone submodular.

For any $1\leq k^{\prime}\leq k$ , with $m=\Theta(\frac{k}{k^{\prime}})$ machines, the output of any algorithm is at most an $O(\sqrt{\frac{k^{\prime}}{k}})$ -approximate randomized composable core-set for the submodular maximization problem.

In Theorem 5.3, we proved that it is not possible to achieve better than $O(\sqrt{\frac{k^{\prime}}{k}})$ -approximate randomized composable core-sets for the submodular maximization problem. Here we show that this bound is tight. We prove this by applying a randomized algorithm (which is different from algorithm $\mathsf{Greedy}$ ):

For any $1\leq k^{\prime}\leq k$ , and $m\geq\frac{k}{k^{\prime}}$ machines, the union of output sets by the above algorithm form an $\Omega(\sqrt{\frac{k^{\prime}}{k}})$ -approximate randomized core-set for the submodular maximization problem.

Similar to the proofs of Lemmas 2.3, and 2.4, we also have that:

Conclusion

The concept of composable core-sets has been introduced recently in the context of distributed and streaming algorithms and have been applied to several problems . In this paper, we introduced the concept of randomized composable core-sets and showed its effectiveness in maximizing submodular functions in a distributed manner. While we mainly discuss the cardinality constraint in this paper, we expect that the ideas and the proof techniques be applicable to more general packing constraints such as matroid constraints. There are several research problems that are interesting to explore in this line of research.

For the submodular maximization problem, it remains to find a randomized composable core-set of approximation factor $1-{1\over e}$ , or rule out the possibility of constructing such a core-set.

We discussed how the size and multiplicity of the composable core-set can help improve the approximation factor of the algorithms. It would be nice to get tight bounds on the approximation factor for each range of the size of composable core-sets. Moreover, it would be interesting to study the impact of increasing the multiplicity of the core-set on the achievable approximation factors.

While we provided a tight result for the small-size composable core-set problem, the achievable approximation factor is not satisfactory. A natural way to improve this factor is to apply the composable core-set idea iteratively, and achieve an improve approximation factor. Even for composable core-sets of size $k$ and above, it might be possible to improve the approximation factor by applying such a core-set multiple times. This approach leads to several interesting follow-up questions.

While randomized composable core-sets are applicable to random-order streaming models, applying the proof techniques in this paper may result in stronger approximation factors in pure random-order streaming models (compared to the ones presented here). We leave this problem to future research.

Finally, it would be nice to explore applicability of these ideas on other optimization and graph theoretic problems.