Achieving Exact Cluster Recovery Threshold via Semidefinite Programming: Extensions

Bruce Hajek, Yihong Wu, Jiaming Xu

Introduction

The stochastic block model (SBM) , also known as the planted partition model , is a popular statistical model for studying the community detection and graph partitioning problem (see, e.g., and the references therein). In its simple form, it assumes that out of a total of $n$ vertices, $(K_{1}+\cdots+K_{r})$ of them are partitioned into $r$ clusters with sizes $K_{1},\ldots,K_{r},$ and the remaining $n-(K_{1}+\cdots+K_{r})$ vertices do not belong to any clusters (called outlier vertices); a random graph $G$ is then generated based on the cluster structure, where each pair of vertices is connected independently with probability $p$ if they are in the same cluster or $q$ otherwise. In this paper, we focus on the problem of exactly recovering the clusters (up to a permutation of cluster indices) based on the graph $G$ .

In the setting of two equal-sized clusters or a single cluster plus outlier vertices, recently it has been shown in that the semidefinite programming (SDP) relaxation of the maximum likelihood (ML) estimator achieves the optimal recovery threshold with high probability, in the asymptotic regime of $p=a\log n/n$ and $q=b\log n/n$ for fixed constants $a,b$ and cluster sizes growing linearly in $n$ as $n\to\infty$ . The result for two equal-sized clusters was originally conjectured in and another resolution was recently given in independently.

In this paper, we extend the optimality of SDP to the following three cases, while still assuming $p=a\log n/n$ and $q=b\log n/n$ with $a>b>0$ :

Stochastic block model with two asymmetric clusters: the first cluster consists of $K$ vertices and the second cluster consists of $n-K$ vertices with $K=\lfloor\rho n\rfloor$ for some $\rho\in[0,1/2].$ The value of $\rho$ may be known or unknown to the recovery procedure.

Stochastic block model with $r$ clusters of equal size $K$ : $r\geq 2$ is a fixed integer and $n=rK$ .

Censored block model with two clusters: given an Erdős-Rényi random graph $G\sim{\mathcal{G}}(n,p)$ , each edge $(i,j)$ has a label $L_{ij}\in\{\pm 1\}$ independently drawn according to the distribution:

where $\sigma^{\ast}_{i}=1$ if vertex $i$ is in the first cluster and $\sigma^{\ast}_{i}=-1$ otherwise; $\epsilon\in[0,1/2]$ is a fixed constant.Under the censored block model, the graph itself does not contain any information about the underlying clusters and we are interested in recovering the clusters by observing the graph and edge labels.

In all three cases, we show that a necessary condition for the maximum likelihood (ML) estimator to succeed coincides with a sufficient condition for the correctness of the SDP procedure, thereby establishing both the optimal recovery threshold and the optimality of the SDP relaxation. The proof techniques in this paper are similar to those in ; however, the construction and validation of dual certificates for the success of SDP are more challenging especially in the multiple-cluster case. Notably, we resolve an open problem raised in [1, Section 6] about the optimal recovery threshold in the censored block model and show that the optimal recovery threshold can be achieved in polynomial-time via SDP.

To further investigate the applicability of SDP procedures for community detection, we explored two cases for which the algorithm is adaptive to the unknown cluster sizes. First, we found that for two clusters, the conditions for exact recovery are the strongest in the equal-sized case. This suggests, and it is shown in Section 2.2, that if the cluster size constraint is replaced by an appropriate Lagrangian term not depending on the cluster size, exact recovery is achieved for all cluster sizes under the condition required for two equal-sized clusters. Secondly, we examined the general community detection problem with a fixed number of unequal-sized clusters and with outlier vertices, and identified a sufficient condition for the SDP procedure to achieve exact recovery with knowledge of only the smallest cluster size and the parameters $a,b.$ (See Section 5.)

The optimality result of SDP has recently been extended to the cases with $o(\log n)$ number of equal-sized clusters in and a fixed number of clusters with unequal sizes in .

For the case of $r$ equal-sized clusters, it is independently shown in that the optimal recovery threshold can be obtained in polynomial-time. Their clustering algorithm is a two-step procedure similar to , where the partial recovery is achieved via a simple spectral algorithm. For the case with two unequal-sized clusters, a sufficient (but not tight) recovery condition is also derived in .

Further literature on SDP for cluster recovery

There has been a recent surge of interest in analyzing the semidefinite programming relaxation approach for cluster recovery; some of the latest development are summarized below. For different recovery approaches such as spectral methods, we refer the reader to for details.

The SDP approach is mostly analyzed in the regime where the average degrees scale as $\log n$ , with the objective of exact cluster recovery. In this setting, the analysis often relies on the standard technique of dual witnesses, which amounts to constructing the dual variables so that the desired KKT conditions are satisfied for the primal variable corresponding to the true clusters. The SDP has been applied to recover cliques or densest subgraphs in . For the stochastic block model with possibly unbounded number of clusters, a sufficient condition for an SDP procedure to achieve exact recovery is obtained in , which improves the sufficient conditions in in terms of scaling. Various formulations of SDP for cluster recovery are discussed in . The robustness of the SDP has been investigated in for minimum bisection in the semirandom model with monotone adversary and, more recently, in for generalized SBM with arbitrary outlier vertices. The SDP machinery has also been applied to recover clusters with partially observed graphs and binary matrices . In the converse direction, necessary conditions for the success of particular SDPs are obtained in . In contrast to the previous work mentioned above where the constants are often loose, the recent line of work initiated by , and followed by and the current paper, focus on establishing necessary and sufficient conditions in the special case of a fixed number of clusters with sharp constants, attained via SDP relaxations.

In the sparse graph case with bounded average degree, exact recovery is provably impossible and instead the goal is to achieve partial recovery, namely, to correctly cluster all but a small fraction of vertices. Using Grothendieck’s inequality, a sufficient condition for SDP to achieve partial recovery is obtained in ; the technique is extended to the labeled stochastic block model in . In , an SDP-based test is applied to distinguish the binary symmetric stochastic block model versus the Erdős-Rényi random graph and shown to attain the optimal detection threshold.

Notation

Denote the identity matrix by $\mathbf{I}$ , the all-one matrix by $\mathbf{J}$ and the all-one vector by $\mathbf{1}$ . We write $X\succeq 0$ if $X$ is symmetric and positive semidefinite and $X\geq 0$ if all the entries of $X$ are non-negative. Let ${\mathcal{S}}^{n}$ denote the set of all $n\times n$ symmetric matrices. For $X\in{\mathcal{S}}^{n}$ , let $\lambda_{2}(X)$ denote its second smallest eigenvalue. For any matrix $Y$ , let $\|Y\|$ denote its spectral norm. For any positive integer $n$ , let $[n]=\{1,\ldots,n\}$ . For any set $T\subset[n]$ , let $|T|$ denote its cardinality and $T^{c}$ denote its complement. For $\rho\in$ , let $\bar{\rho}=1-\rho$ . We use standard big $O$ notations, e.g., for any sequences $\{a_{n}\}$ and $\{b_{n}\}$ , $a_{n}=\Theta(b_{n})$ if there is an absolute constant $c>0$ such that $1/c\leq a_{n}/b_{n}\leq c$ ; $a_{n}=\Omega(b_{n})$ or $b_{n}=O(a_{n})$ if there exists an absolute constant $c>0$ such that $a_{n}/b_{n}\geq c$ . Let ${\rm Bern}(p)$ denote the Bernoulli distribution with mean $p$ and ${\rm Binom}(n,p)$ denote the binomial distribution with $n$ trials and success probability $p$ . All logarithms are natural and we use the convention $0\log 0=0$ .

Binary asymmetric SBM

Let $A$ denote the adjacency matrix of the graph, and $(C^{\ast}_{1},C^{\ast}_{2})$ denote the underlying true partition, where the clusters $C^{\ast}_{1}$ and $C^{\ast}_{2}$ have cardinalities $K$ and $n-K$ , respectively, and we consider the asymptotic regime $K=\lceil n\rho\rceil$ as $n\to\infty$ for $\rho\in[0,\frac{1}{2}]$ fixed. In this subsection we assume that $\rho$ is known to the recovery procedure and the goal is to obtain the $\rho$ -dependent optimal recovery threshold attained by SDP relaxations.

The cluster structure under the binary stochastic block model can be represented by a vector $\sigma\in\{\pm 1\}^{n}$ such that $\sigma_{i}=1$ if vertex $i$ is in the first cluster and $\sigma_{i}=-1$ otherwise. Let $\sigma^{\ast}$ correspond to the true clusters. Then the ML estimator of $\sigma^{\ast}$ for the case $a>b$ can be simply stated as

which maximizes the number of in-cluster edges minus the number of out-cluster edges subject to the cluster size constraint. If $K=n/2$ , (1) reduces to the minimum graph bisection problem which is NP-hard in the worst case. Due to the computational intractability of the ML estimator, next we turn to its convex relaxation. Let $Y=\sigma\sigma^{\top}$ . Then $Y_{ii}=1$ is equivalent to $\sigma_{i}=\pm 1$ , and $\sigma^{\top}\mathbf{1}=\pm(2K-n)$ if and only if $\langle Y,\mathbf{J}\rangle=(2K-n)^{2}$ . Therefore, (1) can be recast asHenceforth, all matrix variables in the optimization are symmetric.

Notice that any feasible solution is a rank-one positive semidefinite matrix. Relaxing this condition by dropping the rank-one restriction, we obtain the following convex relaxation of (2), which is a semidefinite program:

We note that the only model parameter needed by the estimator (3) is the cluster size $K$ .

Let $Y^{\ast}=\sigma^{\ast}(\sigma^{\ast})^{\top}$ correspond to the true partition and ${\mathcal{Y}}_{n}\triangleq\{\sigma\sigma^{\top}:\sigma\in\{\pm 1\}^{n},\sigma^{\top}\mathbf{1}=2K-n\}$ denote the set of all admissible partitions. The following result establishes the optimality of the SDP procedure.

with $\bar{\rho}=1-\rho$ , $\tau=\frac{a-b}{\log a-\log b}$ , $\gamma=\sqrt{(\bar{\rho}-\rho)^{2}\tau^{2}+4\rho\bar{\rho}ab}$ , and $\eta(0,a,b)=\lim_{\rho\to 0}\eta(\rho,a,b)=\frac{a+b}{2}-\tau\log\frac{{\rm e}\sqrt{ab}}{\tau}.$

The proof of Theorem 1 is similar in outline to the proof given in , but a considerable detour is needed to handle the imbalance. Notice that by definition, $\eta(\rho,a,b)=\eta(\bar{\rho},a,b)$ , and $\eta(1/2,a,b)=\frac{1}{2}(\sqrt{a}-\sqrt{b})^{2}$ . The threshold function $\eta(\rho,a,b)$ turns out to be the error exponent in the following large deviation events. For vertex $i$ , let $e(i,C^{*}_{1})$ denotes the number of edges between vertex $i$ and vertices in $C^{*}_{1}$ , and define $e(i,C^{*}_{2})$ similarly. Then,

Next we prove a converse for Theorem 1 which shows that the recovery threshold achieved by the SDP relaxation is in fact optimal.

In the special case with two equal-sized clusters, we have $K=n/2$ and $\eta(1/2,a,b)=\frac{1}{2}(\sqrt{a}-\sqrt{b})^{2}$ . The corresponding threshold $(\sqrt{a}-\sqrt{b})^{2}>2$ has been established in , and the achievability by SDP has been shown in and independently by later.

A recent work also studies the exact recovery problem in the unbalanced case and provides the sufficient (but not tight) recovery condition for a polynomial-time two-step procedure based on the spectral method.

2 Unknown cluster size

Theorem 4 shows that if one knows the relative cluster size $\rho$ , the SDP relaxation (3) achieves the size-dependent optimal threshold $\eta(\rho,a,b)>1$ . For fixed $a$ and $b$ , $\eta(\rho,a,b)$ is minimized at $\rho=\frac{1}{2}$ (see Appendix A for a proof). This suggests that for two communities the equal-sized case is the most difficult to cluster. Indeed, the next result proves that if there is no constraint on the cluster size, then the optimal recovery threshold coincides with that in the balanced case, i.e., $(\sqrt{a}-\sqrt{b})^{2}>2$ , which can be achieved by a penalized SDP.

Theorem 3 holds for all cluster sizes $K,$ including the extreme case where the entire network forms a single cluster ( $K=0$ ), in which case the SDP (5) outputs $Y^{*}=\mathbf{J}$ with high probability. The downside is that the penalization parameter $\lambda^{*}$ depends on the parameters $a$ and $b$ . Nevertheless, there exists a fully data-driven choice of $\lambda^{*}$ based on the degree distribution of the network, so that Theorem 3 continues to hold whenever the cluster sizes scale linearly, i.e., $K/n\to\rho\in(0,\frac{1}{2}]$ ; the price to pay for adaptivity is that the probability of error vanishes polylogarithmically instead of polynomially as $n\to\infty$ . See Appendix B for details.

SBM with multiple equal-sized clusters

The cluster structure under the stochastic block model with $r$ clusters of equal size $K$ can be represented by $r$ binary vectors $\xi_{1},\ldots,\xi_{r}\in\{0,1\}^{n}$ , where $\xi_{k}$ is the indicator function of the cluster $k$ , such that $\xi_{k}(i)=1$ if vertex $i$ is in cluster $k$ and $\xi_{k}(i)=0$ otherwise. Let $\xi^{\ast}_{1},\ldots,\xi^{\ast}_{r}$ correspond to the true clusters and let $A$ denote the adjacency matrix. Then the maximum likelihood (ML) estimator of $\xi^{\ast}$ for the case $a>b$ can be simply stated as

We remark that we could as well have worked with the constraint $\langle Y,\mathbf{J}\rangle=0$ , which, for $Y\succeq 0$ , is equivalent to the last constraint in (8). Letting $Z=\frac{r-1}{r}Y+\frac{1}{r}\mathbf{J}$ , we can also equivalently rewrite (8) as

The only model parameter needed by the estimator (9) is the cluster size $K$ . Let $Z^{*}=\sum_{k=1}^{r}\xi_{k}^{*}(\xi_{k}^{*})^{\top}$ correspond to the true clusters and define

The sufficient condition for the success of SDP in (9) is given as follows.

The following result establishes the optimality of the SDP procedure.

The optimal recovery threshold $\sqrt{a}-\sqrt{b}=\sqrt{r}$ is also obtained by two parallel independent works via a polynomial-time two-step procedure, consisting of a partial recovery algorithm followed by a cleanup stage. The previous work studies the stochastic block model in a much more general setting with $r$ clusters of equal size $K$ plus outlier vertices, where $r,K$ and the edge probabilities $p,q$ may scale with $n$ arbitrarily as long as $rK\leq n$ ; it is shown that an SDP achieves exact recovery with high probability provided that

for some universal constant $C$ . In the special setting where the network consists of a fixed number of clusters without outliers and $p=a\log n/n>q=b\log n/n$ , the sufficient condition (10) simplifies to $\sqrt{a}-\sqrt{b}\geq C^{\prime}\sqrt{r}$ for some absolute constant $C^{\prime}$ , which is off by a constant factor compared to the sharp sufficient condition $\sqrt{a}-\sqrt{b}>\sqrt{r}$ given by Theorem 4.

It is straightforward to extend the current proof of Theorem 5 to the regime where $r=\gamma\log^{s}n$ , $p=\frac{a\log^{s+1}n}{n}>q=\frac{b\log^{s+1}n}{n}$ for any fixed $\gamma,a,b>0$ and $s\in[0,1)$ , showing that SDP achieves the optimal recovery threshold $\sqrt{a}-\sqrt{b}=\sqrt{\gamma}$ . Indeed, the preprint shows similar optimality results of SDP for $r=o(\log n)$ number of equal-sized clusters. Conversely, it has been recently proved in that SDP relaxations cease to be optimal for logarithmically many communities in the sense that SDP is constantwise suboptimal when $r\geq C\log n$ for a large enough constant $C$ and orderwise suboptimal when $r=\omega(\log n)$ .

Binary censored block model

Under the binary censored block model, with possibly unequal cluster sizes, the cluster structure can be represented by a vector $\sigma\in\{\pm 1\}^{n}$ such that $\sigma_{i}=1$ if vertex $i$ is in the first cluster and $\sigma_{i}=-1$ if vertex $i$ is in the second cluster. Let $\sigma^{\ast}\in\{\pm 1\}^{n}$ correspond to the true clusters. Let $A$ denote the weighted adjacency matrix such that $A_{ij}=0$ if $i,j$ are not connected by an edge; $A_{ij}=1$ if $i,j$ are connected by an edge with label $+1$ ; $A_{ij}=-1$ if $i,j$ are connected by an edge with label $-1$ . Then the ML estimator of $\sigma^{\ast}$ can be simply stated as

which maximizes the number of in-cluster $+1$ edges minus that of in-cluster $-1$ edges, or equivalently, maximizes the number of cross-cluster $-1$ edges minus that of cross-cluster $+1$ edges. The NP-hard max-cut problem can be reduced to (11) by simply labeling all the edges in the input graph as $-1$ edges, and thus (11) is computationally intractable in the worst case. Instead, we consider the SDP studied in obtained by convex relaxation. Let $Y=\sigma\sigma^{\top}$ . Then $Y_{ii}=1$ is equivalent to $\sigma_{i}=\pm 1$ . Therefore, (6) can be recast as

Replacing the rank-one constraint by positive semidefiniteness, we obtain the following convex relaxation of (12), which is an SDP:

We remark that (13) does not rely on any knowledge of the model parameters. Let $Y^{\ast}=\sigma^{\ast}(\sigma^{\ast})^{\top}$ and ${\mathcal{Y}}_{n}\triangleq\{\sigma\sigma^{\top}\colon\sigma\in\{\pm 1\}^{n}\}$ . The following result establishes the success condition of the SDP procedure in the scaling regime $p=a\log n/n$ for a fixed constant $a$ :

Next we prove a converse for Theorem 6 which shows that the recovery threshold achieved by the SDP relaxation is in fact optimal.

Theorem 7 still holds if the cluster sizes are proportional to $n$ and known to the estimators, i.e., the prior distribution of $\sigma^{\ast}$ is uniform over $\{\sigma\in\{\pm 1\}^{n}:\sigma^{\top}\mathbf{1}=2K-n\}$ for $K=\lfloor\rho n\rfloor$ with $\rho\in(0,1/2]$ .

Denote by $a^{*}(\epsilon)$ the optimal recovery threshold, namely, the infimum of $a>0$ such that exact cluster recovery is possible with probability converging to one as $n\to\infty$ . Our results show that for all $\epsilon\in[0,1/2]$ , the optimal recovery threshold is given by

and can be achieved by the SDP relaxations. The optimal recovery threshold is insensitive to $\rho$ , which is in contrast to what we have seen for the binary stochastic block model.

Exact cluster recovery in the censored block model is previously studied in and it is shown that if $\epsilon\to 1/2$ , the maximum likelihood estimator achieves the optimal recovery threshold $a(1-2\epsilon)^{2}>2+o(1)$ , while an SDP relaxation of the ML estimator succeeds if $a(1-2\epsilon)^{2}>4+o(1)$ . The optimal recovery threshold for any fixed $\epsilon\in(0,1/2)$ and whether it can be achieved in polynomial-time were previously unknown. Theorem 6 and Theorem 7 together show that the SDP relaxation achieves the optimal recovery threshold $a(\sqrt{1-\epsilon}-\sqrt{\epsilon})^{2}>1$ for any fixed constant $\epsilon\in[0,1/2]$ . Notice that $(\sqrt{1-\epsilon}-\sqrt{\epsilon})^{2}=\frac{1}{2}(1-2\epsilon)^{2}+o((1-2\epsilon)^{2})$ when $\epsilon\to 1/2$ . For the censored block model with the background graph being random regular graph, it is further shown in that the SDP relaxations also achieve the optimal exact recovery threshold.

The above exact recovery threshold in the regime $p=a\log n/n$ shall be contrasted with the positively correlated recovery threshold in the sparse regime $p=a/n$ for constant $a$ . In this sparse regime, there exists at least a constant fraction of vertices with no neighbors and exactly recovering the clusters is hopeless; instead, the goal is to find an estimator $\widehat{\sigma}$ positively correlated with $\sigma^{\ast}$ up to a global flip of signs. It was conjectured in that the positively correlated recovery is possible if and only if $a(1-2\epsilon)^{2}>1$ ; the converse part is shown in and recently it is proved in that spectral algorithms achieve the sharp threshold in polynomial-time.

An SDP for general cluster structure

In this section we consider SDPs for the general case of multiple clusters and outliers. We assume there are $r$ clusters with sizes $K_{1},\ldots,K_{r},$ and $n-(K_{1}+\cdots+K_{r})$ outlier vertices. Vertices in the same cluster are connected with probability $p$ , while other pairs of vertices are connected between them with probability $q.$ We consider the asymptotic regime $p=\frac{a\log n}{n},$ $q=\frac{b\log n}{n}$ and $K_{k}=\rho_{k}n$ as $n\rightarrow\infty$ for $a,b,\rho_{0},\ldots,\rho_{r}$ fixed, with $\rho_{1}\geq\ldots\geq\rho_{r}>0.$ Let $\rho_{\min}=\rho_{r}.$ We derive sufficient conditions for exact recovery by SDPs. While the conditions are not the tightest possible for specific cases, we would like to identify an algorithm that recovers the cluster matrix exactly without knowing the details of the cluster structure. As in Section 3, the true cluster matrix can be expressed as $Z^{*}=\sum_{k=1}^{r}\xi_{k}^{*}(\xi_{k}^{*})^{\top},$ where $\xi_{k}^{*}$ is the indicator function of the $k^{{\rm th}}$ cluster. Denote by ${\mathcal{Z}}_{n}$ the collection of all such cluster matrices.

Implementing the SDP (15) requires no knowledge of the density parameters $a$ and $b$ , the number of clusters $r,$ or the sizes of the individual clusters; but it does require the exact knowledge of the sum as well as the sum of squares of the cluster sizes, which, in practical applications, may be unrealistic to assume. Therefore, similar to (5), we also consider the following penalized SDP, obtained by removing the constraints for those two quantities while augmenting the objective function:

Here the penalization parameters $\eta^{*}$ and $\lambda^{*}$ must be specified.

Clearly the above two SDPs are different and need not have the same solutions; nevertheless, they are similar enough so that in the following theorem we state a sufficient condition for either of the SDPs to exactly recover $Z^{*}$ with high probability. Define

For $\mu>0$ fixed, $I(\mu,d)$ is a strictly convex, nonnegative function in $d$ which is zero if and only if $d=\mu.$

Suppose there exists $\psi_{1}>0$ and $\psi_{2}>0$ with $b+\psi_{1}+\psi_{2}<a$ such that

We examine two simpler sufficient conditions for recovery, assuming we have enough information to implement one of the two SDPs, and we also have a lower bound $\rho_{\min},$ on the $\rho_{k}$ ’s, but we don’t know how many clusters there are nor whether there are outlier vertices. The conditions of Theorem 8 are most stringent when there are two clusters of the smallest possible size $\rho_{\min},$ and in that case we get the tightest result from the theorem by selecting $\psi_{1}=\psi_{2}=\psi,$ yielding the following corollary:

There is no simple expression for $\psi$ in Corollary 1. If instead we consider the equation $I(a,b+2\psi)=I(b,b+2\psi),$ we have the smaller but explicit solution $\psi=\frac{\tau-b}{2},$ where $\tau=\frac{a-b}{\log(a/b)}.$ Using this $\psi$ in the test $I(b,b+\psi)>1/\rho_{\min},$ we obtain the following weaker but more explicit recovery condition, which, nevertheless, is within a factor of eight of the necessary condition (see Remark 32 below):

Let us compare the sufficient condition provided by Corollary 2 with necessary conditions for recovery. In the presence of outliers, $I(b,\tau)>1/\rho_{\min}$ is a necessary condition as shown in [24, Theorem 4], for otherwise we can swap a vertex in the smallest cluster with an outlier vertex to increase the number of in-cluster edges. Also, with at least two clusters,

is necessary, because we could have two smallest clusters of sizes $\rho_{\min}n,$ and even if a genie were to reveal all the other clusters, we would still need (32) to recover the two smallest ones, as shown by [2, Theorem 1]. By Lemma 11, $I(b,\tau)\leq(\sqrt{a}-\sqrt{b})^{2}\leq 2I(b,\tau)$ ; so with or without outliers, $2I(b,\tau)>1/\rho_{\min}$ is necessary. By Lemma 12, $I(b,\tau)\leq 4I(b,\frac{\tau+b}{2})$ . Therefore we conclude that the sufficient condition of Corollary 2 is within a factor of four (resp. eight) of the necessary condition in the presence (resp. absence) of outliers.

Conclusions

This paper shows that the SDP procedure works for recovering community structure at the asymptotically optimal threshold in various important settings beyond the case of two equal-sized clusters or that of a single cluster and outliers considered in . In particular, SDP relaxations works asymptotically optimally for two unequal clusters (with or without knowing the cluster size), or $r$ equal clusters, or the binary censored block model with the background graph being Erdős-Rényi. These results demonstrate the versatility of SDP relaxation as a simple, general purpose, computationally feasible methodology for community detection.

The picture is less impressive when these cases are combined to have a general case with $r$ clusters of various sizes plus outliers. Still, we found that an SDP procedure can achieve exact recovery even without the knowledge of the cluster sizes; the sufficient condition for recovery is within a factor of eight of the necessary information-theoretic bound. An interesting open problem is whether the SDP relaxation can achieve the optimal recovery threshold in this general case. The preprint addresses this problem, showing that the SDP relaxation still achieves the optimal threshold for recovering a fixed number of clusters with unequal sizes.

Proofs

with $\gamma=\sqrt{\alpha^{2}+4\rho_{1}\rho_{2}ab}.$

We first prove the upper tail bound in (36) using Chernoff’s bound. In particular,

with $\gamma_{n}=\sqrt{\alpha_{n}^{2}+4\rho_{1,n}\rho_{2,n}ab}$ . Thus, using the inequality that $\log(1-x)\leq-x$ , we have

If $\rho_{1,n}=\rho_{1}+o(1)$ , $\rho_{2,n}=\rho_{2}+o(1)$ , and $\alpha_{n}=\alpha+o(1)$ , then we let

with $\gamma=\sqrt{\alpha^{2}+4\rho_{1}\rho_{2}ab}$ . It follows that

and thus the upper tail bound in (35) holds in view of (37). Next, we prove the lower tail bound in (35).

Case 1: $\rho_{1},\rho_{2}>0$ . For any choice of the constant $\alpha^{\prime}$ with $\alpha^{\prime}>|\alpha|,$

Setting $\alpha^{\prime}=\sqrt{\alpha^{2}+4\rho_{1}\rho_{2}ab}$ in the last displayed equation yields

where the last inequality follows from Lemma 1. ∎

The following lemma provides a deterministic sufficient condition for the success of SDP (3) in the case $a>b$ .

Then $\widehat{Y}_{{\rm SDP}}=Y^{\ast}$ is the unique solution to (3).

where $(a)$ holds because $\langle S^{\ast},Y\rangle\geq 0$ ; $(b)$ holds because $\langle Y^{\ast},S^{\ast}\rangle=(\sigma^{\ast})^{\top}S^{\ast}\sigma^{\ast}=0$ by (38). Hence, $Y^{\ast}$ is an optimal solution. It remains to establish its uniqueness. To this end, suppose ${\widetilde{Y}}$ is an optimal solution. Then,

where $(a)$ holds because $\langle\mathbf{J},{\widetilde{Y}}\rangle=\langle\mathbf{J},Y^{\ast}\rangle$ , $\langle A,{\widetilde{Y}}\rangle=\langle A,Y^{\ast}\rangle$ , and ${\widetilde{Y}}_{ii}=Y^{*}_{ii}=1$ for all $i\in[n]$ . In view of (38), since ${\widetilde{Y}}\succeq 0$ , $S^{\ast}\succeq 0$ with $\lambda_{2}(S^{*})>0$ , ${\widetilde{Y}}$ must be a multiple of $Y^{*}=\sigma^{\ast}(\sigma^{\ast})^{\top}$ . Because ${\widetilde{Y}}_{ii}=1$ for all $i\in[n]$ , ${\widetilde{Y}}=Y^{\ast}$ . ∎

Let $D^{\ast}=\mathsf{diag}\left\{{d^{\ast}_{i}}\right\}$ with

and choose $\lambda^{*}=\tau\log n/n$ , where $\tau=\frac{a-b}{\log a-\log b}$ . It suffices to show that $S^{*}=D^{\ast}-A+\lambda^{\ast}\mathbf{J}$ satisfies the conditions in Lemma 3 with high probability.

By definition, $d^{\ast}_{i}\sigma_{i}^{\ast}=\sum_{j}A_{ij}\sigma^{\ast}_{j}-\lambda^{\ast}(2K-n)$ for all $i\in[n]$ , i.e., $D^{\ast}\sigma^{\ast}=A\sigma^{\ast}-\lambda^{\ast}(2K-n)\mathbf{1}$ . Since $\mathbf{J}\sigma^{\ast}=(2K-n)\mathbf{1}$ , it follows that the desired (38) holds, that is, $S^{*}\sigma^{*}=0$ . It remains to verify that $S^{\ast}\succeq 0$ and $\lambda_{2}(S^{\ast})>0$ with high probability, which amounts to showing that

where $(a)$ holds because $\left\langle x,\sigma^{\ast}\right\rangle=0$ . It follows from (41) that for any $x\perp\sigma^{\ast}$ and $\|x\|_{2}=1$ , $x^{T}S^{*}x=t_{1}(x)+t_{2}(x)$ where

We next bound $\inf_{x\perp\sigma^{*},\|x\|_{2}=1}t_{1}(x)$ from the below. Consider the specific vector $\check{x}$ that maximizes $x^{\top}\mathbf{J}x$ subject to the unit norm constraint and $\langle x,\sigma^{*}\rangle=0.$ It has coordinates $\sqrt{\frac{n-K}{nK}}$ for the $K$ vertices of the first cluster and coordinates $\sqrt{\frac{K}{n(n-K)}}$ for the $n-K$ vertices of the other cluster. Let $E_{2}=\mbox{span}(\sigma^{*},\check{x});$ $E_{2}$ is the set of vectors that are constant over each cluster. Then

Notice that for any vector $x$ with $x\perp E_{2},$ $\mathbf{J}x=0.$ It follows that

We bound the three terms in the parenthesis separately in the sequel.

Lower bound on $t_{1}(\check{x})$ : Notice that $\check{x}^{\top}\mathbf{J}\check{x}=4K(n-K)/n$ and thus

Since $\check{x}^{\top}D^{\ast}\check{x}=\langle A,B\rangle-\lambda^{\ast}(2K-n)\sum_{i=1}^{n}\check{x}_{i}^{2}\sigma_{i}^{\ast}$ , where $B_{ij}=\sigma_{i}\sigma_{j}\check{x}_{i}^{2}$ , it follows that $\check{x}^{\top}D\check{x}$ is Lipschitz continuous in $A$ with Lipschitz constant $\|B\|_{\rm F}=\sqrt{(1-\rho)^{2}/{\rho}+\rho^{2}/(1-\rho)}+o(1)$ . Moreover, $A_{ij}$ is $ $-valued. It follows from the Talagrand’s concentration inequality for Lipschitz convex functions (see, e.g., [40, Theorem 2.1.13]) that for any$ c>0 $, there exists$ c^{\prime}>0 $only depending on$ \rho$, such that

Hence, with probability at least $1-n^{-c}$ ,

Lower bound on $\inf_{\|x\|_{2}=1,x\perp E_{2}}x^{\top}D^{*}\check{x}$ : Note that $E[D^{*}]\check{x}\in E_{2}.$ So for any vector $x$ with $x\perp E_{2},$ $x^{\top}E[D^{*}]\check{x}=0$ . Hence,

where the last inequality follows from the Cauchy-Schwartz inequality. It follows from the Talagrand’s concentration inequality for Lipschitz convex functions that for any $c>0$ , there exists $c^{\prime}>0$ such that

Lower bound on $\inf_{\|x\|_{2}=1,x\perp E_{2}}x^{\top}D^{*}x$ : Notice that for $\|x\|_{2}=1,$ $x^{\top}D^{*}x\geq\min_{i}d^{\ast}_{i}$ , so it suffices to bound $\min_{i}d^{\ast}_{i}$ from the below. For $i\in C_{1}$ , $A_{ij}\sigma_{i}\sigma_{j}$ is equal in distribution to $X-R$ , where $X\sim{\rm Binom}(K-1,\frac{a\log n}{n})$ and $R\sim{\rm Binom}(n-K,\frac{b\log n}{n})$ . It follows from Lemma 2 that

For $i\in C_{2}$ , $A_{ij}\sigma_{i}\sigma_{j}$ is equal in distribution to $X-R$ , where $X\sim{\rm Binom}(n-K-1,\frac{a\log n}{n})$ and $R\sim{\rm Binom}(K,\frac{b\log n}{n})$ . It follows from Lemma 2 that

It follows from the definition of $d^{\ast}_{i}$ that

where the last inequality follows from the assumption that $\eta(\rho,a,b)>1$ .

Combing all the three lower bounds together, we get that with high probability,

Notice that we have shown that with high probability $\inf_{x\perp\sigma^{*},\|x\|_{2}=1}t_{2}(x)\geq p-c^{\prime}\sqrt{\log n}$ . It follows from (44) that with high probability,

Notice that $a>b>0$ and thus $\tau>b$ . Therefore, the desired (40) holds and the theorem follows from Lemma 3. ∎

By symmetry, we can condition on $C_{1}^{*}$ being the first $K$ vertices. Let $T$ denote the set of first $\lfloor\frac{n}{\log^{2}n}\rfloor$ vertex. Then

Suppose that the true clusters are $C^{\ast}_{1}$ and $C^{\ast}_{2}$ of cardinality $K_{n}$ and $n-K_{n}$ , respectively. One can easily check that Lemma 3 still holds with $\lambda^{\ast}=\tau\log n/n$ , where $\tau=\frac{a-b}{\log a-\log b}$ . Choose the same $d_{i}^{\ast}$ in (39) as in the proof of Theorem 1. It suffices to show for any $0\leq K_{n}\leq n$ ,

First, consider the case $K_{n}=0$ or $n$ where $Y^{*}=\mathbf{J}$ and the graph is simply ${\mathcal{G}}(n,p)$ . Then for $i\in[n]$ , $\sum_{j}A_{ij}\sigma^{\ast}_{i}\sigma^{\ast}_{j}\sim{\rm Binom}(n-1,\frac{a\log n}{n})$ . Recall that $\tau=\frac{a-b}{\log a-\log b}$ and notice that in this case, $d_{i}^{\ast}=\sum_{j}A_{ij}\sigma^{\ast}_{i}\sigma^{\ast}_{j}-\tau\log n$ . It follows from Lemma 1 that

where $\eta(0,a,b)=a-\tau\log({\rm e}a/\tau)$ and the last inequality follows from Lemma 13 in Appendix A. By the union bound,

where the last inequality holds because $\eta(1/2,a,b)=\frac{1}{2}(\sqrt{a}-\sqrt{b})^{2}>1$ by assumption. Moreover, since $\sigma^{*}=\pm\mathbf{1}$ , any $x$ such that $x\perp\sigma^{\ast}$ satisfies $x^{\top}\mathbf{J}x=0$ . It follows from (41) that

Next, we consider the case $1\leq K_{n}\leq n-1$ . For $i\in C_{1}$ , $\sum_{j}A_{ij}\sigma_{i}\sigma_{j}$ is stochastically larger than $X-R-1$ , where $X\sim{\rm Binom}(K_{n},\frac{a\log n}{n})$ and $R\sim{\rm Binom}(n-K_{n},\frac{b\log n}{n})$ . Let $\rho_{n}=\frac{K_{n}}{n}\in(0,1)$ and $t_{n}=\tau(1-2\rho_{n})\log n-\frac{\log n}{\log\log n}-1$ . Applying the non-asymptotic upper bound in Lemma 2 yields

We proceed to show that $g(\rho_{n},1-\rho_{n},a,b,-\frac{t_{n}}{\log n})\geq\eta(1/2,a,b)+o(1)$ . First note that

and $-\frac{t_{n}}{\log n}=\tau(\rho_{n}-\bar{\rho}_{n})-\epsilon_{n}$ , where $\epsilon_{n}=\frac{1}{\log\log n}+\frac{1}{\log n}$ and $\bar{\rho}_{n}\triangleq 1-\rho_{n}$ . Furthermore, for any fixed $a,b>0$ ,

for some function $F(a,b)$ independent of $n$ and $\bar{\rho}\triangleq 1-\rho$ . To see this, let $t=\tau(\rho-\bar{\rho})-\delta$ , where $0\leq\delta\leq\epsilon_{n}$ . First consider the case of $t<0$ . Then $\rho\leq\frac{1}{2}+o(1)\leq 2/3$ . Hence $1/3\leq\bar{\rho}\leq 1$ . Then

Since $\sqrt{ab}<\tau<\frac{a+b}{2}$ whenever $a\neq b$ and $\bar{\rho}-\rho\in$ , both the numerator and denominator in (50) are bounded away from zero and infinity uniformly in $\rho$ . The case of $t>0$ follows analogously. Therefore

where (51) is due to (49), (52) is by definition of $\eta$ , and (53) follows from Lemma 13.

Similarly, for $i\in C_{2}$ , $A_{ij}\sigma_{i}\sigma_{j}$ is stochastically larger than $X-R-1$ , where $X\sim{\rm Binom}(n-K_{n},\frac{a\log n}{n})$ and $R\sim{\rm Binom}(K_{n},\frac{b\log n}{n})$ . Let $k^{\prime}_{n}=\tau(1-2\rho_{n})\log n+\frac{\log n}{\log\log n}+1$ . It follows from Lemma 2 that

where the last inequality follows from the same steps as in (51) – (53). It follows from the definition of $d^{\ast}_{i}$ that

where the last inequality follows from the assumption that $\eta(1/2,a,b)=\frac{1}{2}(\sqrt{a}-\sqrt{b})^{2}>1$ . Furthermore one can verify that $\inf_{x\perp\sigma^{\ast}}t_{2}(x)\geq p-O(\sqrt{\log n})$ with high probability, where the functions $t_{1}$ and $t_{2}$ are defined in (42)–(43). We divide the remaining analysis into the two cases:

Case 1: $K_{n}\leq n/\sqrt{\log n}$ or $n-K_{n}\leq n/\sqrt{\log n}$ . Notice that $\tau\leq\frac{a+b}{2}$ and recall the definition of $\check{x}$ in the proof of Theorem 1. Then

Thus the desired (48) follows by the same argument used in the proof of Theorem 1.

Then the desired (48) follows by the same argument used in the proof of Theorem 1.

2 Proofs for Section 3: Multiple equal-sized clusters

Theorem 4 is proved after three lemmas are given. For $k\in[r]$ , denote by $C_{k}\subset[n]$ the support of the $k^{\rm th}$ cluster. For a set $T$ of vertices, let $e(i,T)\triangleq\sum_{j\in T}A_{ij}$ and $e(T^{\prime},T)=\sum_{i\in T^{\prime}}e(i,T).$ Let $k(i)$ denote the index of the cluster containing vertex $i.$ Denote the number of neighbors of $i$ in its own cluster by $s_{i}=e(i,C_{k(i)})$ and the maximum number of neighbors of $i$ in other clusters by $r_{i}=\max_{k^{\prime}\neq k(i)}e(i,C_{k^{\prime}}).$

Notice that $s_{i}\sim{\rm Binom}(K,p)$ and for $k^{\prime}\neq k(i)$ , $e(i,C_{k^{\prime}})\sim{\rm Binom}(K,q)$ . It is shown in that

Applying the union bound over all possible vertices, we complete the proof. ∎

There exists a constant $c>0$ depending only on $b$ and $r$ such that

where $(a)$ follows from Bernstein’s inequality. Furthermore,

It follows from McDiarmid’s inequality that

Thus, with probability at most $n^{-2}$ , $\frac{1}{K}\sum_{i\in C_{k}}r_{i}\geq Kq+O(\sqrt{\log n})$ . The lemma follows in view of the union bound. ∎

The following lemma provides a deterministic sufficient condition for the success of SDP (9) in the case $a>b$ .

Then $\widehat{Z}_{SDP}=Z^{*}$ is the unique solution to (9).

Let $H=Z-Z^{*},$ where $Z$ is an arbitrary feasible matrix for the SDP (9). Since $Z$ and $Z^{*}$ are both feasible, $\langle D^{*},H\rangle=\langle\lambda^{*}\mathbf{1}^{\top},H\rangle=\langle\mathbf{1}(\lambda^{\ast})^{\top},H\rangle=0.$ Since $A=D^{*}-B^{*}-S^{*}+\lambda^{*}\mathbf{1}^{\top}+\mathbf{1}(\lambda^{\ast})^{\top},$

$\langle B^{*},H\rangle\geq 0,$ with equality if and only if $\langle B^{*},Z\rangle=0.$ That is because $B^{*}\geq 0,$ $Z\geq 0,$ and $\langle B^{*},Z^{*}\rangle=0.$

$\langle S^{*},H\rangle\geq 0,$ with equality if and only if $\langle S^{*},Z\rangle=0.$ That is because $\langle S^{*},Z\rangle\geq 0$ (because $S^{*},Z\succeq 0$ ) and $\langle S^{*},Z^{*}\rangle=0$ (because $Z^{*}=\sum_{k=1}^{r}\xi^{*}_{k}(\xi^{*}_{k})^{\top}$ and $S^{*}\xi^{*}_{k}=0$ for all $k\in[r]$ ).

Thus, $\langle A,H\rangle\leq 0,$ so that $Z^{*}$ is a solution to the SDP.

To prove that $Z^{*}$ is the unique solution, restrict attention to the case that $Z$ is another solution to the SDP. We need to show $Z=Z^{*}.$ Since both $Z$ and $Z^{*}$ are solutions, $\langle A,H\rangle=0,$ so that $\langle B^{*},H\rangle=\langle S^{*},H\rangle=0.$ Therefore, by the above two points: $\langle B^{*},Z\rangle=\langle S^{*},Z\rangle=0.$ For each $i$ , $B^{*}_{i,j}=0$ if and only if vertices $i$ and $j$ are in the same cluster. Also, the fact $Z\succeq 0$ and $Z_{ii}\leq 1$ for all $i$ implies $Z_{ij}\leq 1$ for all $i,j.$ Thus, the only way $Z$ can meet the constraint $Z\mathbf{1}=K\mathbf{1}$ is that $Z_{ij}=1$ whenever $i$ and $j$ are in the same cluster. Therefore $Z=Z^{*}$ and hence $Z^{*}$ is the unique solution. ∎

We now begin the proof of Theorem 4. Let $E$ denote the subspace spanned by vectors $\{\xi^{\ast}_{k}\}_{k\in[r]}$ , i.e., $E=\text{span}(\xi^{\ast}_{k}:k\in[r]).$ Ultimately, we will show that

Since $B^{*}$ is assumed to be symmetric, (59) is equivalent to requiring that

for some $y_{kk^{\prime}}^{\ast}$ and $z_{kk^{\prime}}^{\ast}$ . Next we ensure that $S^{*}\xi^{\ast}_{k}=0$ for $k\in[r]$ . Equivalently, we want to ensure that for any distinct $k,k^{\prime}\in[r]$ and any $i\in C_{k}$ ,

Requiring (62) for all distinct $k,k^{\prime}\in[r]$ and all $i\in C_{k}$ is equivalent to requiring

for all distinct $k,k^{\prime}\in[r]$ and all $j\in C_{k^{\prime}}$ (by swapping $i$ for $j$ and $k$ for $k^{\prime}$ ). Moreover, it is equivalent to checking both (62) and (63) under the additional assumption that $k<k^{\prime}.$ Substituting (60) into (62) and (63) gives that for all $k,k^{\prime}\in[r]$ with $k<k^{\prime},$

For $k<k^{\prime}$ and $i\in C_{k},j\in C_{k^{\prime}}$ , set

where $u_{kk^{\prime}}$ and $\alpha_{k}$ are to be determined. Equations (64) and (65) both reduce to:

(which must hold whenever $k<k^{\prime}$ ) and (61) becomes

In view of Lemma 4 and the assumption that $\sqrt{a}-\sqrt{b}>\sqrt{r}$ , $\min_{i}(s_{i}-r_{i})\geq\log n/\log\log n$ with high probability. By Lemma 5, $\max_{k\in[r]}\frac{1}{K}\sum_{i\in C_{k}}r_{i}\leq Kq+O(\sqrt{\log n})$ with high probability. Finally set

Thus, the desired (57) holds in view of (69) and (58). Also, note that $e(C_{k},C_{k^{\prime}})\sim{\rm Binom}(K^{2},q)$ . For $X\sim{\rm Binom}(n,p_{0})$ with $p_{0}\in$ , Chernoff’s bound yields

Applying the union bound, we have that with high probability, $e(C_{k},C_{k^{\prime}})>K^{2}q-K\sqrt{\log n}$ for all $1\leq k<k^{\prime}\leq r$ . Hence $y^{\ast}_{kk^{\prime}}(i)>0$ and $z^{\ast}_{kk^{\prime}}(j)>0$ for all $1\leq k<k^{\prime}\leq r$ and $i\in C_{k}$ and $j\in C_{k^{\prime}}$ so that $B^{\ast}_{ij}>0$ for all $i,j$ in distinct clusters as desired. ∎

3 Proofs for Section 4: Binary censored block model

Our analysis of the SDP relies on two key ingredients: the spectrum of labeled Erdős-Rényi random graph and the tail bounds for the binomial distributions, which we first present.

Let $E=(E_{ij})$ denote an $n\times n$ matrix with independent entries drawn from $\widehat{\mu}\triangleq\frac{p}{2}\delta_{1}+\frac{p}{2}\delta_{-1}+(1-p)\delta_{0}$ , which is the distribution of a Rademacher random variable multiplied with an independent Bernoulli with bias $p$ . Define $E^{\prime}$ as $E^{\prime}_{ii}=E_{ii}$ and $E^{\prime}_{ij}=-E_{ji}$ for all $i\neq j$ . Let $A^{\prime}$ be an independent copy of $A$ . Let $D$ be a zero-diagonal symmetric matrix whose entries are drawn from $\widehat{\mu}$ and $D^{\prime}$ be an independent copy of $D$ . Let $M=(M_{ij})$ denote an $n\times n$ zero-diagonal symmetric matrix whose entries are Rademacher and independent from $C$ and $C^{\prime}$ . We apply the usual symmetrization arguments:

where $(a),(d)$ follow from the Jensen’s inequality; $(b)$ follows because $A-A^{\prime}$ has the same distribution as $(A-A^{\prime})\circ M$ , where $\circ$ denotes the element-wise product; $(c),(f)$ follow from the triangle inequality; $(e)$ follows from the fact that $D-D^{\prime}$ has the same distribution as $E-E^{\prime}$ . In particular, first, the diagonal entries of $D-D^{\prime}$ and $E-E^{\prime}$ are all equal to zero. Second, both $D-D^{\prime}$ and $E-E^{\prime}$ are symmetric matrices with independent upper triangular entries. Third, $D_{ij}-D^{\prime}_{ij}$ is equal in distribution to $E_{ij}-E^{\prime}_{ij}$ for all $i<j$ by definition.

Then, we apply the result of Seginer which characterized the expected spectral norm of i.i.d. random matrices within universal constant factors. Let $X_{j}\triangleq\sum_{i=1}^{n}E_{ij}^{2}$ , which are independent ${\rm Binom}(n,p)$ . Since $\widehat{\mu}$ is symmetric, [39, Theorem 1.1] and Jensen’s inequality yield

for some universal constant $\kappa$ . In view of the following Chernoff bound for the binomial distribution [31, Theorem 4.4]:

for all $t\geq 6np$ , setting $t_{0}=6\max\{np/\log n,1\}$ and applying the union bound, we have

where the last inequality follows from $np\geq c_{0}\log n$ . Assembling (70) – (72), we obtain

Assume that $k_{n}\in[m]$ and $k_{n}=(1+o(1))\frac{\log n}{\log\log n}$ . Then

If $\epsilon=0$ , then $\sum_{i=1}^{m}X_{i}\sim{\rm Binom}(m,p)$ and the lemma follows from Lemma 1. Next we focus on the case $\epsilon>0$ . It follows from the Chernoff bound that

Hence, by setting $x=k_{n}/m$ , we get $\lambda^{\ast}=\frac{1}{2}\log\frac{1-\epsilon}{\epsilon}+o(1)$ and thus

where the last equality holds due to the Taylor expansion of $\log(1-x)$ at $x=0$ and $p=a\log n/n$ . Combining the last displayed equation with (75) gives the desired (74). ∎

The following lemma establishes a lower tail bound for $\sum_{i=1}^{m}X_{i}$ .

Let $k^{\ast}\triangleq\lfloor 2a\sqrt{\epsilon(1-\epsilon)}\log n\rfloor$ . Notice that $\sum_{i=1}^{m}X^{2}_{i}\sim{\rm Binom}(m,p)$ . Let $Z_{1},Z_{2},\ldots,Z_{n}{\stackrel{{\scriptstyle\text{i.i.d.}}}{{\sim}}}(1-\epsilon)\delta_{+1}+\epsilon\delta_{-1}.$ Then

We use the following non-asymptotic bound on the binomial tail probability [9, Lemma 4.7.2]: For $U\sim{\rm Binom}(n,p)$ ,

where $\lambda=\frac{k}{n}\in(0,1)\geq p$ and $D(\lambda\|p)=\lambda\log\frac{\lambda}{p}+(1-\lambda)\log\frac{1-\lambda}{1-p}$ is the binary divergence function. Let $W\sim{\rm Binom}(k^{\ast},\epsilon)$ . Then,

Moreover, using the following bound on binomial coefficients [9, Lemma 4.7.1]:

where $\lambda=\frac{k}{n}\in(0,1)$ and $h(\lambda)=-\lambda\log\lambda-(1-\lambda)\log(1-\lambda)$ is the binary entropy function, we have that

Observe that by the definition of $k^{\ast}$ , $\log\frac{a\log n}{k^{\ast}}=D(1/2\|\epsilon)+o(1)$ and it follows from (76) that

The following lemma provides a deterministic sufficient condition for the success of SDP (13) in the case of $a>b$ .

Suppose there exist $D^{\ast}=\mathsf{diag}\left\{{d^{\ast}_{i}}\right\}$ such that $S^{*}\triangleq D^{\ast}-A$ satisfies $S^{\ast}\succeq 0$ , $\lambda_{2}(S^{\ast})>0$ and

Then $\widehat{Y}_{{\rm SDP}}=Y^{\ast}$ is the unique solution to (13).

where the Lagrangian multipliers are $S\succeq 0$ and $D=\mathsf{diag}\left\{{d_{i}}\right\}$ . Then for any $Y$ satisfying the constraints in (13),

where $(a)$ holds because $\langle S^{\ast},Y\rangle\geq 0$ ; $(b)$ holds because $\langle Y^{\ast},S^{\ast}\rangle=(\sigma^{\ast})^{\top}S^{\ast}\sigma^{\ast}=0$ by (77). Hence, $Y^{\ast}$ is an optimal solution. It remains to establish its uniqueness. To this end, suppose ${\widetilde{Y}}$ is an optimal solution. Then,

where $(a)$ holds because $\langle A,{\widetilde{Y}}\rangle=\langle A,Y^{\ast}\rangle$ and ${\widetilde{Y}}_{ii}=Y^{*}_{ii}=1$ for all $i\in[n]$ . In view of (77), since ${\widetilde{Y}}\succeq 0$ , $S^{\ast}\succeq 0$ with $\lambda_{2}(S^{*})>0$ , ${\widetilde{Y}}$ must be a multiple of $Y^{*}=\sigma^{\ast}(\sigma^{\ast})^{\top}$ . Because ${\widetilde{Y}}_{ii}=1$ for all $i\in[n]$ , ${\widetilde{Y}}=Y^{\ast}$ . ∎

Let $D^{\ast}=\mathsf{diag}\left\{{d^{\ast}_{i}}\right\}$ with

It suffices to show that $S^{*}=D^{\ast}-A$ satisfies the conditions in Lemma 9 with high probability.

By definition, $d^{\ast}_{i}\sigma_{i}^{\ast}=\sum_{j}A_{ij}\sigma^{\ast}_{j}$ for all $i$ , i.e., $D^{\ast}\sigma^{\ast}=A\sigma^{\ast}$ . Thus (77) holds, that is, $S^{*}\sigma^{*}=0$ . It remains to verify that $S^{\ast}\succeq 0$ and $\lambda_{2}(S^{\ast})>0$ with probability converging to one, which amounts to showing that

Applying the union bound implies that $\min_{i\in[n]}d^{\ast}_{i}\geq\frac{\log n}{\log\log n}$ holds with probability at least $1-n^{1-a(\sqrt{1-\epsilon}-\sqrt{\epsilon})^{2}+o(1)}$ . It follows from the assumption $a(\sqrt{1-\epsilon}-\sqrt{\epsilon})^{2}>1$ and (80) that the desired (79) holds, completing the proof. ∎

The prior distribution of $\sigma^{\ast}$ is uniform over $\{\pm 1\}^{n}$ . First consider the case of $\epsilon=0$ . If $a<1$ , then the number of isolated vertices tends to infinity in probability . Notice that for isolated vertices $i$ , vertex $\sigma^{\ast}_{i}$ is equally likely to be $+1$ or $-1$ conditional on the graph. Hence, the probability of exact recovery converges to .

Let $T$ denote the set of first $\lfloor\frac{n}{\log^{2}n}\rfloor$ vertices and $T^{c}=[n]\backslash T$ . Let $s^{\prime}_{i}=\sum_{j\in T^{c}:\sigma^{\ast}_{j}=\sigma^{\ast}_{i}}A_{ij}$ and $r^{\prime}_{i}=\sum_{j\in T^{c}:\sigma^{\ast}_{j}\neq\sigma^{\ast}_{i}}A_{ij}$ . Then

4 Proofs for Section 5: General cluster structure

We first present a dual certificate lemma which is useful for the proof of Theorem 8. Recall that $\xi_{k}^{*}$ denotes the indicator vector of cluster $k$ for $k\in[r]$ and $Z^{*}=\sum_{k}\xi_{k}^{\top}\xi_{k}.$

(If the penalized SDP (22) is used, the same $\eta^{*}$ and $\lambda^{*}$ should be used in the SDP and in this lemma.) Then $Z^{*}$ is the unique solution to both SDP (15) and (22) (i.e., $\widehat{Z}_{SDP}$ produced by either SDP is equal to $Z^{*}$ ).

Let $H=Z-Z^{*},$ where $Z$ is either an arbitrary feasible matrix for the SDP (15) or an arbitrary feasible matrix for the SDP (22). Since $A-\eta^{*}\mathbf{I}-\lambda^{*}\mathbf{J}=D^{*}-B^{*}-S^{*},$

$\langle D^{*},H\rangle\leq 0,$ with equality if and only if $Z_{ii}=1$ for all inlier s $i.$ That is because for inliers $i$ , $d^{*}_{i}>0$ , $Z_{ii}\leq 1=Z^{*}_{ii},$ and for outliers $i,$ $d^{*}_{i}=0.$

$\langle B^{*},H\rangle\geq 0,$ with equality if and only if $\langle B^{*},Z\rangle=0.$ That is because $B^{*}\geq 0,$ $Z\geq 0,$ and $\langle B^{*},Z^{*}\rangle=0.$

$\langle S^{*},H\rangle\geq 0,$ with equality if and only if $\langle S^{*},Z\rangle=0.$ That is because $\langle S^{*},Z\rangle\geq 0$ (because $S^{*},Z\succeq 0$ ) and $\langle S^{*},Z^{*}\rangle=0$ (because $Z^{*}$ is a sum of matrices of the form $\xi_{k}\xi_{k}^{\top}$ and $S^{*}\xi_{k}=0$ for all $k.$ )

Thus, $\langle A,H\rangle-\eta^{*}\langle\mathbf{I},H\rangle-\lambda^{*}\langle\mathbf{J},H\rangle\leq 0.$ Therefore, $Z^{*}$ is a solution to SDP (22). If $Z$ is a feasible solution for the SDP (15), (as $Z^{*}$ is), then $\langle\mathbf{I},H\rangle=\langle\mathbf{J},H\rangle=0,$ so we conclude that $\langle A,H\rangle\leq 0,$ so $Z^{*}$ is also a solution to SDP (15).

To prove that $Z^{*}$ is the unique solution, restrict attention to the case that $Z$ is another solution to either one of the SDPs. We need to show $Z=Z^{*}.$ Since both $Z$ and $Z^{*}$ are solutions, $\langle A,H\rangle-\eta^{*}\langle\mathbf{I},H\rangle-\lambda^{*}\langle\mathbf{J},H\rangle\leq 0,$ so that $\langle D^{*},H\rangle=\langle B^{*},H\rangle=\langle S^{*},H\rangle=0.$ Therefore, by the above three points: $Z_{ii}=1$ for all inliers $i$ , and $\langle B^{*},Z\rangle=\langle S^{*},Z\rangle=0.$

Since $B^{*}_{ij}>0$ whenever $i$ and $j$ are in distinct clusters, and $Z\geq 0$ and $B^{*}\geq 0$ , the condition $\langle B^{*},Z\rangle=0$ implies that $Z_{ij}=0$ whenever $i$ and $j$ are in distinct clusters. By assumption, $\xi_{k}^{*}$ is an eigenvector of $S^{*}$ with corresponding eigenvalue zero, for $1\leq k\leq r.$ Since $\lambda_{r+1}(S)>0$ , it follows that all the other eigenvalues of $S^{*}$ are strictly positive. The condition $\langle S^{*},Z\rangle=0$ thus implies that all the other eigenvectors of $S^{*}$ are in the null space of $Z,$ so the eigenvectors of $Z$ corresponding to the positive eigenvalues of $Z$ must be in the span of $\xi^{\ast}_{1},\ldots,\xi^{\ast}_{r}.$ It follows that $Z$ is a linear combination of matrices of the form $\xi^{\ast}_{k}(\xi^{\ast}_{k^{\prime}})^{\top},$ for $k,k^{\prime}\in[r].$ It follows that $Z_{ij}=0$ if either $i$ or $j$ is an outlier vertex, or both are outlier vertices. Moreover, whenever $i$ and $j$ are in the same cluster, $Z_{ij}=Z_{ii}=1$ . In conclusion, $Z=Z^{\ast}$ . ∎

For $k\in[r]$ , denote by $C_{k}\subset[n]$ the support of the $k^{\rm th}$ cluster. Also, let $C_{0}$ denote the set of outlier vertices. For a set $T$ of vertices, let $e(i,T)\triangleq\sum_{j\in T}A_{ij}$ and $e(T^{\prime},T)=\sum_{i\in T^{\prime}}e(i,T).$ Let $k(i)$ denote the index of the cluster containing vertex $i.$ Denote the number of neighbors of $i$ in its own cluster by $s_{i}=e(i,C_{k(i)})$ and the maximum number of neighbors of $i$ in other clusters by $r_{i}=\max_{k^{\prime}\neq k(i)}e(i,C_{k^{\prime}}).$

Now, let us construct $(D^{*},B^{*},\eta^{*},\lambda^{*})$ such that the conditions of Lemma 10 hold with high probability. Notice that $d^{*}_{i}=0$ if $i$ is an outlier and $B^{*}_{C_{k}\times C_{k}}=0$ for $k\in[r]$ . In order that $(S^{*}\xi_{k}^{*})_{i}=0$ for $i\in C_{k}$ and $k\in[r]$ , we must choose:

The condition $S^{*}\xi_{k}^{*}=0$ for $k\in[r]$ also partially constrains the symmetric matrix $B^{*}.$ We should try to be economical in the choice of $B^{*}$ so that we have a chance to prove that $S^{*}\succeq 0.$

where we used the fact that for each pair of distinct $k$ and $k^{\prime}$ , each of the terms in the definition of $B^{*}_{C_{k}\times C_{k^{\prime}}}(i,j)$ is either constant in $i$ or constant in $j$ , or both, and if $k=0$ the terms are constant in $j$ and if $k^{\prime}=0$ the terms are constant in $i.$ The needed condition $d_{i}^{*}\geq 0$ involves getting a lower bound on the number of edges a vertex $i$ has to other vertices in its own cluster (we can concentrate on the smallest cluster for that purpose), while the needed condition $B\geq 0$ involves an upper bound on the number of edges between a vertex $i$ in one cluster and the vertices of a different cluster.

where $\xi_{0}$ is the indicator function for the set of outlier vertices and we used the fact that $\mathbf{J}=\mathbf{1}\mathbf{1}^{\top}$ and $\mathbf{1}=(\mathbf{1}-\xi_{0})+\xi_{0}.$ From this it is clear that if $\lambda^{*}\geq q,$ then $\lambda_{r+1}(S^{\ast})>0.$ So we will be sure to select $\lambda^{*}\geq q.$ In fact, that will be needed to ensure that $B_{ij}\geq 0$ for all $i,j.$

It remains to select $\lambda^{*}$ so that $d_{i}\geq 0$ and $B_{ij}\geq 0$ for all $i,j$ with high probability. Let $\lambda^{*}=\widetilde{\tau}\log n/n$ with $\widetilde{\tau}=b+\psi_{1}+\psi_{2}$ , where $\psi_{1}$ and $\psi_{2}$ satisfy the assumptions (28)-(31). Then, for inlier vertex $i\in C_{k}$ , in view of Lemma 1 and the definition of $I(\cdot,\cdot)$ in (27),

Turning next to $B^{\ast}_{ij}$ ’s for $(i,j)\in C_{k}\times C_{k^{\prime}},$ it suffices to consider the two following cases:

Note that $\frac{e(C_{k},C_{k^{\prime}})}{K_{k}K_{k^{\prime}}}$ will be very close to $q$ with high probability, so we can replace it by $q=\frac{b\log n}{n}$ , which is also the mean of $\frac{e(i,C_{k^{\prime}})}{K_{k^{\prime}}}$ and $\frac{e(j,C_{k})}{K_{k}}.$ Specifically, it follows from the Chernoff bound that

where $\mu=qK_{k}K_{k^{\prime}}$ and $\epsilon=\frac{2\sqrt{\log n}}{\sqrt{qK_{k}K_{k^{\prime}}}}.$ In view of Lemma 1 and the union bound,

By the assumptions $\rho_{r}I(b,b+\psi_{1})>1$ and $\rho_{r-1}I(b,b+\psi_{2})>1$ , it follows that with high probability $B^{\ast}_{C_{k}\times C_{k^{\prime}}}>0$ .

By the assumptions $\rho_{r}I(b,\widetilde{\tau})>1$ , it follows that with high probability $B^{\ast}_{C_{k}\times C_{k^{\prime}}}\geq 0$ .

In conclusion, we have constructed $(D^{*},B^{*},\eta^{*},\lambda^{*})$ such that the conditions of Lemma 10 hold with high probability. Therefore, the theorem follows by applying Lemma 10. ∎

Let $\tau=\frac{a-b}{\log(a/b)}$ for $0<a<b.$ Then $I(b,\tau)\leq(\sqrt{a}-\sqrt{b})^{2}\leq 2I(b,\tau).$

Notice that $I(a,x)+I(b,x)$ is strictly convex in $x$ . By setting the derivative to be zero, we find that it achieves its minimum value, $(\sqrt{a}-\sqrt{b})^{2},$ at $x=\sqrt{ab}.$ By definition, $I(a,\tau)=I(b,\tau)$ and $\sqrt{ab}\leq\tau.$ Thus, $(\sqrt{a}-\sqrt{b})^{2}\leq I(a,\tau)+I(b,\tau)=2I(b,\tau)$ . Moreover, $I(a,x)$ is decreasing for $x\leq a$ and $I(b,x)$ is non-negative. Therefore, $I(b,\tau)=I(a,\tau)\leq I(a,\sqrt{ab})\leq(\sqrt{a}-\sqrt{b})^{2}.$ ∎

For any $\mu>0$ and $x>0,$ $I(\mu,\mu+2x)\leq 4I(\mu,\mu+x).$

Let $f(x)=I(\mu,\mu+x)$ for $x\geq 0$ . Then $f(0)=f^{\prime}(0)=0$ and $f^{\prime\prime}(s)=\frac{1}{\mu+s}.$ Therefore,

Thus, using a change of variables $s=2t$ ,

Comparing the expressions for $f(x)$ and $f(2x)$ completes the proof. ∎

Appendix A Behavior of threshold function in (4)

Recall $\eta(\rho,a,b)$ defined in (4) which governs the sharp recovery threshold for the asymmetric binary SBM. The following lemma implies $\eta(\rho,a,b)$ is minimized at $\rho=1/2.$

For any $a>b>0$ , $\eta(\rho,a,b)$ is convex in $\rho$ over $,$ and symmetric about $\rho=1/2.$

and from this expression it is easily checked that $\eta$ is symmetric about $\rho=1/2.$

Let $\eta^{\prime},\eta^{\prime\prime}$ denote the first-order and second-order derivative of $\eta$ with respect to $\rho$ , respectively. We show that $\eta^{\prime\prime}\geq 0$ . Recall that $\gamma=\sqrt{(1-2\rho)^{2}\tau^{2}+4\rho(1-\rho)ab}$ . Hence,

Let $h(\rho)=\log\frac{(\gamma+(1-2\rho)\tau)\rho}{(\gamma-(1-2\rho)\tau)(1-\rho)}$ and then

where $(a)$ follows using the expression of $\gamma$ . Therefore,

where the last inequality follows because by letting $x=ab/\tau^{2}$ ,

Appendix B A data-driven choice of the penalization parameter in (5)

Set $\widehat{\rho}=\frac{1}{n}\sum{\mathbf{1}_{\left\{{w_{i}\leq\widehat{w}}\right\}}}$ , $\widehat{w}_{+}=\frac{1}{n}\sum w_{i}{\mathbf{1}_{\left\{{w_{i}>\widehat{w}}\right\}}}$ and $\widehat{w}_{-}=\frac{1}{n}\sum w_{i}{\mathbf{1}_{\left\{{w_{i}<\widehat{w}}\right\}}}$ , which are consistent estimates for $\rho,w_{+},w_{-}$ , respectively. From these we can readily obtain consistent estimates for $(a,b,\rho)$ whenever $\rho\neq 1/2$ . Furthermore, when $\rho=1/2$ , we claim that

Now we are ready to choose the penalty parameter $\widehat{\lambda}=\widehat{\lambda}(A)$ , so that Theorem 3 continues to hold upon replacing the deterministic $\lambda^{*}$ by $\widehat{\lambda}$ . Let

Let $\widetilde{d}_{1}=\sum_{j\neq 2}A_{1j}$ and $\widetilde{d}_{2}=\sum_{j\neq 1}A_{2j}$ . Let $\widetilde{w}_{i}=\widetilde{d}_{i}/\log n$ for $i=1,2$ . Then