A Proof Of The Block Model Threshold Conjecture

Elchanan Mossel, Joe Neeman, Allan Sly

Introduction

We consider the simplest version of the stochastic block model, namely the version with two symmetric states:

If $q=q^{\prime}$ , the stochastic block model is just an Erdős-Rényi model, but if $q\gg q^{\prime}$ then one expects that a typical graph will have two well-defined clusters.

Theoretical computer scientists’ interest in the average case analysis of the minimum-bisection problem led to intensive research on algorithms for recovering the partition (although not all of these works used the same model as us; for example, considered random regular graphs with a fixed minimum bisection size). At the same time, the model is a classical statistical model for networks with communities , and the questions of identifiability of the parameters and recovery of the clusters have been studied extensively, see e.g. . We refer the readers to for more background on the model.

2 The block models in sparse graphs

Until recently, all of the theoretical literature on the stochastic block model focused on what we call the dense case, where the average degree is of order at least $\log n$ and the graph is connected. Indeed, it is clear that connectivity is required, if we wish to label all vertices accurately. However, the case of sparse graphs with constant average degree is well motivated from the perspective of real networks, see e.g. .

Although sparse graphs are natural for modelling many large networks, the stochastic block model seems to be most difficult to analyze in the sparse setting. Despite the large body of work about this model, until recently the only result for the sparse case $q,q^{\prime}=O(\frac{1}{n})$ was that of Coja-Oghlan . Recently, Decelle et al. made some fascinating conjectures for the cluster identification problem in the sparse stochastic block model. In what follows, we will set $q=a/n$ and $q^{\prime}=b/n$ for some fixed $a,b$ . It will be useful to parameterize these by $d=(a+b)/2$ and $s=(a-b)/2$ . Note that with these parameters, $s^{2}>d$ implies that $s,d>1$ .

If $s^{2}>d$ then the clustering problem in $\mathcal{G}(n,\frac{a}{n},\frac{b}{n})$ is solvable as $n\to\infty$ , in the sense that one can a.a.s. find a bisection whose correlation with the planted bisection is bounded away from 0.

Decelle et al.’s work is based on deep but non-rigorous ideas from statistical physics. In order to identify the best bisection, they use the sum-product algorithm (also known as belief propagation). Using the cavity method, they argue that the algorithm should work, a claim that is bolstered by compelling simulation results. By contrast the best rigorous work by Coja-Oghlan showed that if $s^{2}>Cd\log d$ for a large constant $C$ , then the spectral method solves the clustering problem. (After the current article first appeared as a preprint, independent works gave simple algorithms that work when $s^{2}>Cd$ .)

What makes Conjecture 1.2 even more appealing is the fact that if it is true, it represents a threshold for the solvability of the clustering problem. Indeed it was conjectured in and proved in that if $s^{2}\leq d$ then the clustering problem in $\mathcal{G}(n,\frac{a}{n},\frac{b}{n})$ problem is not solvable as $n\to\infty$ . It was further shown in that $s^{2}=d$ represents the threshold for identifiability of the parameters $a$ and $b$ as conjectured by .

The threshold $d=s^{2}$ can be understood both in terms of spin systems and in terms of random matrices. It was first derived heuristically as a stability condition for the belief propagation algorithm . Around a typical vertex, the joint distribution of the graph labeled by the clusters is asymptotically (in the sense of local weak convergence) a Galton-Watson branching process labeled with the free Ising model. The threshold $d=s^{2}$ corresponds to the extremality or reconstruction threshold for the Ising model, the point at which information on the spin at the root can be recovered over arbitrarily long distances. In this property was used to show the impossibility of reconstructing clusters when $s^{2}\leq d$ .

In the threshold was heuristically derived by considering the spectrum of the matrix of non-backtracking walks. On an Erdős-Rényi random graph the bulk spectrum of this matrix has radius $d^{1/2}$ . When $s^{2}>d$ there is a natural construction of an approximate eigenvector of eigenvalue $s$ which thus escapes from the bulk exactly at $d=s^{2}$ . The random matrix interpretation plays a central role in our anaylsis and is discussed further in Section 2.3.

3 Notation

We write graphs as $G=(V,E)$ , where $V$ is a vertex set and $E$ is the set of edges. We write $v\sim w$ if $\{v,w\}\in E$ . If we need to speak about several graphs at the same time, we may write $V(G)$ or $E(G)$ in order to be clear that we are referring to the vertices (or edges) of the graph $G$ . If $\sigma$ is a labelling on $V$ and $U\subset V$ , then we write $\sigma_{U}$ for the restriction of $\sigma$ to $U$ .

In order not to be overwhelmed with quantifiers, we make heavy use of the asymptotic notations $o,O,\omega$ , and $\Omega$ , including in the antecedent. For example, the statement that “ $a_{n}=O(b_{n})$ implies that $c_{n}=O(d_{n})$ ” means that for every $C_{1}>0$ there exists some $C_{2}>0$ such that $a_{n}\leq C_{1}b_{n}$ for all $n$ implies that $c_{n}\leq C_{2}d_{n}$ for all $n$ . Given a collection of sequences $(a_{v,n})$ depending on some other parameter $v$ , we say that they satisfy some asymptotics uniformly in $v$ if the hidden constants or rates of convergence do not depend on $v$ . For example, “ $a_{v,n}=o(b_{n})$ uniformly in $v$ ” means that there is some sequence $c_{n}\to 0$ such that $a_{v,n}/b_{n}\leq c_{n}$ for all $v$ and all $n$ .

We write that a sequence of events holds asymptotically almost surely (or a.a.s.) if their probabilities converge to one.

Our results

Our main result is a proof of Conjecture 1.2:

Our algorithm is also computationally efficient, and can be implemented in almost linear time $O(n\log^{2}n)$ . We recently learned that Laurent Massoulié independently found a different proof of the conjecture .

We note that our proof of Theorem 2.1 actually gives slightly stronger results. First, $a$ and $b$ do not need to be fixed, but may grow slowly with $n$ :

If $a,b=n^{o(1/\log\log n)}$ and $s^{2}/d\geq\lambda>1$ for all $n$ then the clustering problem in $\mathcal{G}(n,\frac{a}{n},\frac{b}{n})$ is solvable as $n\to\infty$ , in the sense of Theorem 2.1.

Moreover, although our proof of Theorem 2.1 does not give particularly good bounds for the size of the correlation, it does show that the correlation tends to 1 as $s^{2}/d$ grows.

If $a,b=n^{o(1/\log\log n)}$ and $s^{2}/d\to\infty$ then the clustering problem in $\mathcal{G}(n,\frac{a}{n},\frac{b}{n})$ is solvable as $n\to\infty$ , in the sense that one can a.a.s. find a bisection that agrees with the planted bisection up to an error of $o(n)$ vertices.

It was conjectured in that a popular algorithm, belief propagation initialized with i.i.d. uniform messages, detects communities all the way to the threshold. However, analysis of belief propagation with random initial messages is a difficult task. Krzakala et al. argued that a novel and very efficient spectral algorithm based on a non-backtracking matrix should also detect communities all the way to the threshold. Unfortunately we were unable to follow the path suggested in and our algorithm for detection is not a spectral algorithm. Still, our analysis is based on the non-backtracking walk and we show that it can be implemented using matrix powering. The algorithm has very good theoretical running time $O(n\log^{2}n)$ but the constant in the $O$ needed for the proof that the algorithm is correct is very large, so the algorithm described is not nearly as efficient as the one in (nor have we implemented it).

A path is a sequence $u_{0},\dots,u_{k}$ of vertices such that for all $i$ , $u_{i}\neq u_{i-1}$ . (Note that we do not require vertices in a path to be connected by an edge in any given graph; thus, it might be more standard – but also rather longer – to use the term path in the complete graph instead.) We write $E(\gamma)$ for the set of $\{u_{i-1},u_{i}\}$ and $V(\gamma)$ for the set $\{u_{0},\dots,u_{k}\}$ .

A non-backtracking path is a path $u_{0},\dots,u_{k}$ such that for every $0\leq i\leq k-2$ , $u_{i}\neq u_{i+2}$ .

A self-avoiding path is a sequence of vertices $u_{0},\dots,u_{k}$ that are all distinct.

The basic intuition behind the proof is that we should be expecting a larger number of non-backtracking paths of a given length $k$ in the graph between two vertices $u$ and $v$ if they have the same label, while a smaller number of non-backtracking paths in the graph is expected if the nodes $u$ and $v$ have different labels.

Instead of working with the number of non-backtracking walks in the graph, it is more convenient to work with a rank one correction, where an edge is represented by $1-d/n$ and a non-edge by $-d/n$ . With this alteration the expected weight of each edge is .

Let $W_{e}=1_{\{e\in E(G)\}}-d/n$ . For a non-backtracking path $\gamma=u_{0},\ldots,u_{k}$ , let

Where $k$ is clear from the context, we will sometimes just write $N_{u,v}$ .

Our basic method is to show that $N_{u,v}^{(k)}$ is correlated with $\sigma_{u}\sigma_{v}$ for some $k\sim\log n$ . In order to do this, we would like to compute the expectation and variance of the $N_{u,v}$ . There is an obstacle, however, which is that on some very rare event there are many more paths in the graph than there should be; this event throws off the expectation and variance of $N_{u,v}$ . For intuition, take $k=\lceil\alpha\log n\rceil$ for some large constant $\alpha$ (in order to make our method work, we will need $\alpha$ arbitrarily large if $s^{2}$ is arbitrarily close to $d$ ). As we will show later, $N_{u,v}$ is of the order $s^{k}/n$ with high probability. However, the expectation of $N_{u,v}$ could be much larger. Indeed, the probability that an $m$ -clique containing $u$ and $v$ will appear is at least $n^{-m^{2}}=e^{-m^{2}\log n}$ . On the event of its appearance, there are at least $(m-2)^{k}=e^{(m-2)\alpha\log n}$ non-backtracking paths of length $k$ from $u$ to $v$ that stay entirely within the graph; each of these paths $\gamma$ has $X_{\gamma}\approx 1$ . If $\log(m-2)\geq 2\log s$ and $\alpha\log(m-2)>2m^{2}$ (which can be achieved by first taking $m$ large enough depending on $s$ and then taking $\alpha$ large enough depending on $m$ ), then these paths contribute an expected weight of at least

which is of a larger order than $s^{k}/n$ . For this reason, our argument for controlling $N_{u,v}$ will be divided into two parts: we will use the second moment method to control the part of $N_{u,v}$ that involves “nice” paths, and we will control the other paths by conditioning on an event that excludes cliques, along with some other problematic structures which we call tangles.

Roughly speaking, our main technical result is that $N_{u,v}$ is correlated with $\sigma_{u}\sigma_{v}$ , and that as $u$ and $v$ vary then the variables $N_{u,v}$ are essentially uncorrelated.

It follows easily from Theorem 2.8 that if $s^{2}/d\to\infty$ then we can get very accurate estimates of $\sigma_{u}\sigma_{v}$ by computing $N_{u,v}^{(k)}$ . To achieve non-trivial estimates of $\sigma_{u}\sigma_{v}$ in the case $s^{2}/d>1$ is more complicated. We will explain the procedure roughly in the next section.

2 Almost linear time algorithm

Theorem 2.8 suggests a natural way to check if two vertices are in the same cluster; this is the basis of the algorithm we develop to cluster the graph. We further show how to efficiently perform the algorithm using matrix powering.

There is an algorithm that runs in time $O(nd\log^{2}n)$ and satisfies the following guarantee: for any $\lambda>1$ there is an $\epsilon>0$ such that if $s^{2}/d\geq\lambda>1$ for all $n$ then the algorithm, given $G\sim\mathcal{G}(n,\frac{a}{n},\frac{b}{n})$ , produces a labelling $\tau$ satisfying

with probability $1-o(1)$ , where $\sigma$ is the true labelling of $G$ .

With more care in the analysis the running time could be reduced to $O(nd\log n)$ for a slightly modified algorithm.

3 Connections with random matrix theory

The preceding arguments – and also the more general random matrix theory – break down in the sparse case, where $a$ and $b$ are $O(1)$ . One reason for this is the presence of high-degree vertices: with high probability there exist vertices with degree $\Omega(\log n/\log\log n)$ , and these play havoc with the spectrum of $A$ . We emphasize that this is not merely a looseness in the analysis: naive spectral algorithms genuinely fail for sparse graphs, see e.g. for a discussion of why this happens for the stochastic block model, or for a quite different example of the connection between vertex degree and spectrum in random graphs.

Some attempts were made to modify spectral methods. A popular idea is to prune all nodes whose degree is larger than some big constant. This idea was pioneered by Feige and Ofek for a different application, and studied in our context by Coja-Oghlan , who gave a spectral algorithm that succeeds on sparse graphs but not all the way to the threshold: it requires $s^{2}>Cd\log d$ for some constant $C$ .

Krzakala et al. suggest quite a different way to “fix” the spectrum of $A$ : instead of $A$ , they consider a non-symmetric matrix that avoids the contribution of high-degree vertices by counting non-backtracking paths in the graph instead of all paths in the graph. Although simulations strongly suggested that the non-backtracking matrix had desirable spectral properties whenever $s^{2}>d$ , the proof remained elusive until very recently (and after the first appearance of this work), when Bordenave et al. gave a solution. Their result required new developments in random matrix theory, partly because the non-backtracking matrix is very sparse and partly because its entries are far from i.i.d. These methods were further developed by Bordenave , who gave a new proof of Alon’s conjecture for the second eigenvalue of random $d$ -regular graphs.

In an independent (and concurrent) work, Massoulié gave a proof of Theorem 2.1 using a spectral algorithm. He considered the matrix $M$ where $M_{uv}$ is the number of self-avoiding walks between $u$ and $v$ of length $\alpha\log n$ , for some not-too-large constant $\alpha$ . This “regularized” matrix has several advantages over the non backtracking matrix we consider. The matrix is quite dense (all degree are polynomial) and in fact is close to regular. Moreover, the matrix is symmetric which allows standard perturbation theory to apply. On the other hand, the entries of the matrix are not independent. Still, Massoulié showed how to apply the trace method and analyze the spectrum of the matrix. He proved that if $s^{2}>d$ then $M$ has a separation between its second- and third-largest eigenvalues, and that the second eigenvector is correlated with the true labelling. Hence, the spectral algorithm that rounds the second eigenvector of $M$ succeeds down to the threshold. We note that the algorithm we describe is much more efficient than the one by , which might not even be implementable in polynomial time (for example, counting self-avoiding walks is #P-complete even for fairly simple families of graphs); in any case, simply writing down the dense $n\times n$ matrix in will take time $O(n^{2})$ .

The algorithm and its running time

In this section, we will describe the algorithm and give its analysis assuming Theorem 2.8. We will begin by describing how to use the quantities $N_{u,v}^{(k)}$ to estimate the graph labelling. In Section 3.2, we will show how these quantities may be computed efficiently, thereby completing the description of our algorithm. In Section 3.3, we prove the algorithm’s correctness.

Recall that our random graph model adds a within-class edge with probability $a/n$ and a between-class edge with probability $b/n$ , where $a$ and $b$ are parameters that may grow slowly with $n$ . We set $d=(a+b)/2$ and $s=(a-b)/2$ , and assume that $s^{2}/d\geq\lambda>1$ for all $n$ .

We begin by describing a simplified version of our algorithm. This simplified version runs more slowly, but it is more intuitive and will serve to motivate the main algorithm. The basic idea is to fix a very slowly increasing sequence, say $R=R_{n}=2\lceil\log\log\log\log n\rceil$ . We write $B_{r}(v)$ for the set of vertices whose path distance to $v$ in $G$ is at most $r$ , and we write $S_{v}=B_{R}(v)\setminus B_{R-1}(v)$ . Fix a node $w^{*}$ with large degree (at least $\sqrt{\log\log n}$ , say). For every other node $v$ , consider the graph $H=H(v)$ obtained by removing $w^{*}$ and $B_{R-1}(v)$ from $G$ . Our estimate for $\sigma_{v}$ will be

The preceding algorithm has two flaws that we will correct shortly. First, the distribution of $H$ is slightly painful to work with, because after removing the node $w^{*}$ the remaining edges are no longer independent. Second, the running time of the algorithm above will be about $O(n^{2}\log n)$ , because we must compute the numbers $N_{u,w}^{(k)}$ (each of which takes time $O(n\log n)$ ) with respect to $O(n)$ different graphs $H(v)$ . This could be fixed by handling several nodes simultaneously: we could remove $\bigcup_{v}B_{R-1}(v)$ from $G$ , where the union is taken over, say, $n/\log n$ vertices $v$ . This is almost the approach that we will take, but we need to be careful that whatever we remove, the remaining graph is almost distributed according to the stochastic block model. We will ensure this by a slightly convoluted plan: instead of removing specific nodes and neighborhoods, we will remove $\delta n$ vertices from $G$ uniformly at random and look for nodes and neighborhoods that are contained in the removed part. The precise description of our algorithm follows:

Let $R=2\lceil\log\log\log\log n\rceil$ . Let $\delta^{\prime}>0$ be chosen so that $s^{2}(1-\delta^{\prime})^{2}=d(1-\delta^{\prime})$ and let $\delta=\delta^{\prime}/2$ . We will choose constants $\kappa$ and $k$ depending on $s^{\prime}$ and $d^{\prime}$ (the precise dependence will be given later). Then, we proceed as follows:

Remove at random $\lceil\sqrt{n}\rceil$ vertices $V^{\prime\prime}$ from the graph , leaving the graph $G^{\prime}=(V^{\prime},E^{\prime})$

Let $w^{\ast}$ be a node in $V^{\prime\prime}$ whose number of neighbors in $V^{\prime}$ is closest to $\lceil\sqrt{\log\log n}\rceil$ . Let $S_{\ast}$ be the set of its neighbors in $V^{\prime}$ .

For each $v\in V^{\prime}$ denote $S_{v}=B_{R}(v)\setminus B_{R-1}(v)$ .

For each $1\leq j\leq\log n$ , let $U_{j}$ be a uniformly random set of $\lceil n\delta\rceil-\lceil\sqrt{n}\rceil$ vertices of $V^{\prime}\setminus S_{\ast}$ ; set $V_{j}=V^{\prime}\setminus(S_{\ast}\cup U_{j})$ .

For each $v\in V^{\prime}$ and $1\leq j\leq\log n$ define

where $\xi_{j,v}$ are i.i.d. random variables uniform on $$ and

For each $v\in V^{\prime}$ let $J_{v}$ be the first $j$ such that $B_{R-1}(v)\cap V_{j}=\emptyset$ and $(S_{v}\cup S_{\ast})\subset V_{j}$ , and 0 if no such $j$ exists. Then set $\tau(v)=\tau_{J_{v},v}$ when $J_{v}\neq 0$ and $\tau_{J_{v},v}\neq 0$ . For all other $v\in V$ , choose $\tau(v)$ at independently at random uniformly from $\{1,-1\}$ .

We will prove Theorem 2.9 in Section 3.3 by showing that the output $\tau$ of the algorithm above is correlated with the true partition with high probability. In the following section we describe how to evaluate the $\tau$ in time $O(nd\log^{2}n)$ .

2 Efficient implementation of the algorithm

The main computational step in the algorithm above is to compute $N^{(k,j)}_{u,u^{\prime}}$ ; in this section, we will describe how to do so. First, however, note that $N^{(k,j)}_{u,u^{\prime}}$ is just $N^{(k)}_{u,u^{\prime}}$ computed on the subgraph induced by $V^{\prime}\setminus(S_{\ast}\cup V_{j})$ . In particular, it is enough to show how to compute $N^{(k)}_{u,u^{\prime}}$ efficiently.

Let $V$ be the vertex set and $A$ the adjacency matrix of the graph $G$ . We recall the definition of $N^{(k)}_{u,v}$ and introduce some related matrices

(where an empty sum is defined to be zero, so $Q^{(k,\rho)}$ is the zero matrix for $k<0$ ), and define the $4n\times 4n$ matrices

Finally, define the $4n\times n$ matrix $\mathcal{Q}_{k}$ by

For a graph on $n$ vertices and $m$ edges and for every vector $z$ the matrix $N^{(k)}z$ can be computed in time $O((m+n)k)$ .

The proof follows from the fact that (by Lemma 3.2) $N^{(k)}z$ is the first $n$ coordinates of

and that each of the matrices $\mathcal{M},\hat{\mathcal{M}}$ and $\mathcal{Q}_{0}$ is made of at most $16$ blocks, each of which is a sum of a sparse matrix with $O(n+m)$ entries and a rank $1$ matrix. Therefore, the displayed expression above can be computed with $k+1$ matrix-vector multiplications, each of which requires $O(n+m)$ time. ∎

We can now prove the running time bound in Theorem 2.9. Indeed, in iteration $j$ of the algorithm, the sum $\sum_{u\in S_{\ast}\cap V_{j},u^{\prime}\in S_{v}\cap V_{j}}N_{u,u^{\prime}}^{(k,j)}$ is the non-trivial computation that needs to be done. This sum can be read from the entries of $N^{(k)}z$ , where $N$ is computed on the graph with the removed nodes and $z$ is the indicator of the vertices in $S_{\ast}$ . By Proposition 3.3, the running time of iteration $j$ is $O((n+m)k)$ . Since there are $\log n$ iterations and we have $m=O(nd)$ and $k=O(\log n)$ we obtain that the running time of the algorithm is $O(nd\log^{2}n)$ .

We will write $N^{(k)}_{u,w,v}$ to denote the $N^{(k)}_{u,v}$ , but with the sum restricted to non-backtracking paths which move to $w$ on their first step. Then we have the recursion

By expanding the terms of the form $N^{(k-2)}_{u,w,v}$ repeatedly, we obtain by induction that

The recursion above can be written using the matrix $Q$ from (5) in the following way:

Written in terms of the matrices $\mathcal{Q}$ from (8) and $\mathcal{M},\hat{\mathcal{M}}$ defined in (6) and (7), the recursions above can be written as $\mathcal{M}\mathcal{Q}_{k}=\mathcal{Q}_{k+1}$ and $\hat{\mathcal{M}}\mathcal{Q}_{k}=\left(\begin{array}[]{cccc}N^{(k+1)}&0&0&0\\ \end{array}\right)^{T},$ as claimed. ∎

3 Correctness of the algorithm

In this section, we will prove the correctness of the algorithm assuming Theorem 2.8. We begin with some preliminary observations about the distributions of various subgraphs of $G$ . The distribution of $G^{\prime}$ is simply a stochastic block model with fewer vertices, that is $G(n-\lceil\sqrt{n}\rceil,a/n,b/n)$ . Let $G_{j}=(V_{j},E_{j})$ denote the graph obtained at iteration $j$ ; $G_{j}$ is also distributed as a stochastic block model, but we will need to say more because we will need to use $G_{j}$ conditioned on some extra properties. In particular, we need to argue that conditioned on a vertex neighborhood being removed, the distribution on the remaining graph is drawn (approximately) from the block model. The technical issue here is that the removed vertices are correlated and moreover we need to condition on some of their labels. Nevertheless, this can be handled because the neighborhood of a single vertex does not contain too many other vertices.

For a vertex $v$ in $G_{j}$ , let $U=U(v)$ denote the set $S_{v}\cup S_{\ast}$ . We will be interested in the distribution of $(G_{j},U,\sigma_{U})$ and we would like to couple it with a configuration of $(G^{\prime},U^{\prime},\sigma^{\prime}_{U^{\prime}})$ drawn from $\mathcal{G}(n-\lceil\delta n\rceil,a/n,b/n)$ and $U^{\prime}$ is some fixed set of vertices of size $|U|$ .

The proof will couple $\sigma^{\prime}$ with $\sigma_{V_{j}}$ and the edges of $G^{\prime}$ with the edges of $G_{j}$ . The coupling proceeds in the following way:

We take $\sigma^{\prime}$ and $\sigma$ to be equal on $U=U^{\prime}$ (neither one is random in either measure).

Then we try to couple all other labels so they are completely identical.

Finally, if the labels are identical, we will include exactly the same edges. This is possible since different edges are independent and the probabilities of including edges just depend on the end points.

where $n_{\pm}$ is the number of $\pm 1$ labels in $\sigma_{B_{R-1}(v)}$ . Thus

Next we note that with high probability $|S_{\ast}|=\lceil\sqrt{\log\log n}\rceil$ since the probability that there exists a vertex in $V^{\prime\prime}$ with that number of neighbors tends to one; we will condition on this event. Also, we may assume without loss of generality that $\sigma_{w^{\ast}}=+$ . We will denote

Note that $M_{\ast}$ is a sum of i.i.d. signs, each of which has expectation $\frac{s}{p}$ (since we conditioned on $\sigma_{w^{\ast}}=+$ ). Hence, a.a.s.

which means in particular that $M_{\ast}$ a.a.s. has the same sign as $s$ .

Before we proceed to the estimates that apply specifically for our algorithm, let us note a simple corollary of the first three statements of Theorem 2.8:

Take disjoint sets $U_{1},U_{2}\subset V$ that have cardinality $n^{o(1)}$ ; let $U=U_{1}\cup U_{2}$ . Under the notation and assumptions of Theorem 2.8, if $Y=\sum_{u\in U_{1},v\in U_{2}}Y_{u,v}$ then uniformly for all labellings $\sigma_{U}$ on $U$ ,

We divide the sum into three parts: the first part (containing $|U_{1}||U_{2}|$ terms) sums over $u=u^{\prime}$ and $v=v^{\prime}$ ; for this part, we apply (2). The second part (containing less than $|U_{1}|^{2}|U_{2}|+|U_{2}|^{2}|U_{1}|$ terms) sums over indices with either $u=u^{\prime}$ or $v=v^{\prime}$ ; for this part, we use the bound

and then apply (2) to each term on the right hand side. Finally, the third part ranges over distinct $u,u^{\prime},v,v^{\prime}$ (less than $|U_{1}|^{2}|V_{1}|^{2}$ terms), and we apply (3) Putting these three parts together,

We may apply the previous lemma with (4) and Chebyshev’s inequality to show that

can be used to estimate the sign of $M_{v}$ (which, recall, was defined in (9)).

For a random vertex $v$ and any $\epsilon>0$ , conditioned on $J_{v}\neq 0$ ,

Next, we control $Z_{v}-Y_{v}$ . By (4) and a union bound,

Since $|S_{v}|$ and $|S_{\ast}|$ are $n^{o(1)}$ , (4) implies that for any $t\geq 1$ , the right hand side above converges to zero. Putting it together and setting $t=\epsilon{s^{\prime}}^{R}$ (which is at least 1 for large enough $n$ ),

The purpose of this section is to show that $M_{v}$ can be used to estimate $\sigma_{v}$ . We will do this by exploiting the connection between neighborhoods in $G$ and multi-type branching processes. Since this section is the only place where we will use the theory of branching processes, we will give only a brief introduction; readers unfamiliar with this theory should consult the book by Athreya and Ney .

For notational simplicity, we will assume for now that $s>0$ . The case $s<0$ will be discussed at the end of the section. For the rest of this section, $T$ will denote a Galton-Watson branching process with $\operatorname{Poisson}(d)$ offspring distribution rooted at $\rho$ . We will assign three random labellings to the vertices of $T$ in the following way: first, divide $T$ into connected components by running bond percolation: deleting each edge independently with probability $s/d$ . Then, for each component choose a label uniformly in $\{\pm 1\}$ and assign that label to all vertices in that component. We define $\eta,\eta^{+},$ and $\eta^{-}$ respectively to be the configurations generated this way where the connected component of the root is labelled randomly, labelled $+1$ , or labelled $-1$ respectively. Let $\zeta=\eta_{\rho}$ , let $\Psi_{R}=\sum_{v\in S_{R}(\rho)}\eta_{v}$ and define $\Psi_{R}^{\pm}$ similarly.

It is well-known (and not hard to check) that the random labelling $\eta$ may also be generated in the following way: choose $\eta_{\rho}$ uniformly at random. For every child $u$ of $\eta_{\rho}$ independently, let $\eta_{u}=\eta_{\rho}$ with probability $\frac{a}{a+b}$ and otherwise let $\eta_{u}=-\eta_{\rho}$ . Then recurse this process down the tree: for every child $w$ of $u$ independently, let $\eta_{w}=\eta_{u}$ with probability $\frac{a}{a+b}$ and otherwise let $\eta_{w}=-\eta_{u}$ . The processes $\eta^{+}$ and $\eta^{-}$ may be generated similarly, except that instead of beginning with $\eta_{\rho}$ labelled randomly, we fix $\eta^{+}_{\rho}=+1$ and $\eta^{-}_{\rho}=-1$ .

Let $\xi$ be a uniform random variable on $ $that is independent of$ T $,$ \eta $, and$ \eta^{\pm} $. There exist$ \kappa>0 $and$ \epsilon>0$ such that

By symmetry, $\Psi_{R}$ is symmetric about 0 and so if $\xi$ is an independent uniform on $ $then for any$ \kappa>0$,

for small enough $\delta>0$ . By Chebyshev’s inequality, both $\Psi_{R}^{+}$ and $\Psi_{R}^{-}$ belong to $[-\kappa s^{R},\kappa s^{R}]$ with probability $1-O(\kappa^{-2})$ . Then

Finally, symmetry of $\Psi_{R}^{+}$ and $\Psi_{R}^{-}$ implies that

which completes the proof if $\kappa$ is a sufficiently large constant. ∎

For $1\leq i\leq\log n$ let $(T_{i},\eta_{i})$ be iid copies of $(T,\eta)$ above for $R=2\lceil\log\log\log\log n\rceil$ . For $v_{1},\ldots,v_{\log n}$ be uniformly chosen vertices in $V$ ,

This argument is a minor variation on a well-known argument showing the local tree-like structure of sparse graphs. We will give only a sketch, but a much more detailed argument (although for only one neighborhood) is given in .

We establish the result by coupling the two processes. By Markov’s inequality, with high probability $\sum_{i=1}^{\log n}|T_{i}|\leq\log^{2}n$ and $\sum_{i=1}^{\log n}|B_{R}(v_{i})|\leq\log^{2}n$ . Moreover, by standard arguments in sparse random graphs, $\bigcup B_{R}(v_{i})$ is a disjoint union of trees with high probability.

We reveal the branching process trees by sequentially revealing for each vertex how many children of each label it has ( $\hbox{Poisson}(a/2)$ of the same label and $\hbox{Poisson}(b/2)$ of the opposite label) down to level $R$ in a breadth-first manner.

Similarly, we can reveal the neighborhoods of the $v_{i}$ and their labels in $G$ by sequentially revealing the neighbors and labels of the currently revealed vertices. Suppose that we condition on the labels of all vertices and on the graph structure that was revealed so far, and suppose that we want to reveal the neighbors of a given vertex $u$ . With high probability, none of these revealed neighbors will belong to the already-explored set and so we will focus on $u$ ’s neighbors among the unexplored vertices. If $n^{\pm}$ are the numbers of $\pm 1$ -labelled vertices that have not yet been explored, then $u$ has $\operatorname{Bin}(n^{\sigma_{u}},a/n)$ neighbors of label $\sigma_{u}$ and $\operatorname{Bin}(n^{\sigma_{u}},b/n)$ neighbors of label $-\sigma_{u}$ . Note that $n^{\pm}$ are both in $n/2\pm n^{2/3}$ with high probability, because the original labels were biased by at most $O(n^{1/2})$ and we have revealed at most $\log^{2}n$ of them.

We couple these two processes with the usual coupling of Poisson and Binomial random variables. In each step we fail with probability $O(n^{-1/3})$ and (since there are at most $\log^{2}n$ steps) the coupling altogether fails with probability $o(1)$ . ∎

Using the coupling between graphs and trees, we will show that $\mathcal{A}_{j,v}$ is a good estimator for $\sigma_{v}$ . Later, we will show that $\mathcal{A}_{j,v}$ usually agrees with our previous estimator $\tau_{j,v}$ .

Let $v_{1},\ldots,v_{\log n}$ be a uniform sample without replacement from $V^{\prime}$ . Take the coupling in Lemma 3.8, and let $\Psi_{R,i}=\sum_{v\in S_{R}(\rho_{i})}\eta_{v}$ , where $\rho_{i}$ is the root of $T_{i}$ . Set

By Lemma 3.8, the same holds for $\mathcal{A}_{j,v_{i}}$ . If we now partition $V^{\prime}$ to sets of size $\log n$ and use the fact that $\sigma_{v_{i}}\mathcal{A}^{\prime}_{j,v_{i}}$ are $\pm 1$ , we obtain the claim of the lemma. ∎

Recall that we have been assuming $s>0$ . In the case $s<0$ , Lemma 3.9 (which is the only result from this section that we will use later) remains true. Indeed, in order to generate the $T$ and $\eta$ for the case $s<0$ , one can generate them for $|s|$ and then flip the sign of every label in an odd generation. Since $R$ is even, level $R$ of the tree is unchanged and $s^{R}=|s|^{R}$ . Thus, Lemma 3.9 remains true.

and that $J_{v}$ is the first $j$ such that $B_{R-1}(v)\cap V_{j}=\emptyset$ and $(S_{v}\cup S_{\ast})\subset V_{j}$ , and $J_{v}=0$ if no such $j$ exists.

Recall that with high probability we have that $|S_{\ast}|=\sqrt{\log\log n}$ . With high probability $|B_{R}(v)|\leq d^{2R}$ . Condition on $|B_{R}(v)|\leq d^{2R}$ . The probability that $B_{R-1}\cap V_{j}=\emptyset$ and $S_{v}\cup S_{\ast}\subset V_{j}$ is bounded below by $e^{-c(d^{2R}+|S_{\ast}|)}\geq(\log n)^{-1/2}$ . Since these are independent events given $|B_{R}(v)|$ it follows that with probability tending to one $J_{v}\neq 0$ . ∎

We now show that the indicators $\mathcal{A}_{J_{v},v}$ and $\tau_{J_{v},v}$ usually agree.

By Lemma 3.10, the probability of $J_{v}=0$ goes to zero, and therefore Lemma 3.6 implies that

The event $\mathcal{A}_{J_{v},v}\neq\tau_{J_{v},v}$ is equivalent to $\xi_{J_{v},v}$ falling outside the interval with end-points $-M_{v}/(\kappa s^{\prime R})$ and

Since with high probability $M_{\ast}$ is concentrated around $\frac{s^{\prime}}{d}|S_{\ast}|$ , Lemma 3.6 implies the latter end point converges in probability to

Therefore the probability that $-\xi_{J_{v},v}$ falls in the interval converges to as needed. ∎

We can now complete the proof of Theorem 2.9.

Combining Lemmas 3.9, 3.11 and 3.10 we have that with high probability

yielding an algorithm recovering the a constant correlation with the true partition. The running time bound was proved in Section 3.2. ∎

The proof of Theorem 2.3 (i.e., when we are far above the threshold) is rather easier, and doesn’t require the branching process tools:

Combinatorial path bounds

A crucial ingredient in the proof is obtaining bounds on the number of various types of paths (in the complete graph) in terms of how much they self-intersect, either by intersecting a previous vertex on the path or by repeating an edge of the path.

Given a path $\gamma=(v_{1},\dots,v_{k})$ , we say that an edge $(v_{i},v_{i+1})$

is new if for all $j\leq i$ , $v_{j}\neq v_{i+1}$

is old if there is some $j<i$ such that $\{v_{i},v_{i+1}\}=\{v_{j},v_{j+1}\}$ .

Otherwise, we say that $(v_{i},v_{i+1})$ is returning (in this case $v_{i+1}=v_{j}$ for $j<i$ but $\{v_{i},v_{i+1}\}$ is not one of the previous edges).

Let $k_{n}(\gamma),k_{o}(\gamma)$ and $k_{r}(\gamma)$ be the number of new, old, and returning edges respectively.

For the rest of this subsection we fix $\alpha$ and set $k=\lceil\alpha\log n\rceil$ . Note that for every new edge in a path, the number of distinct vertices in the path increases by one, as does the number of distinct edges. For a returning edge, only the number of edges increases, while for an old edge, neither increases. Therefore we easily see that:

The number of vertices visited by the path $\gamma$ is $k_{n}(\gamma)+1$ and the number of edges is $k_{n}(\gamma)+k_{r}(\gamma)$ .

Our first bound is a fairly crude one that will allow us to assume that $k_{r}$ is smaller than some constant. Note that there is no non-backtracking restriction yet.

For any constant $C$ , if $k_{r}\geq 1$ and $n$ is sufficiently large then there are at most

paths $\gamma$ of length at most $C\log n$ , with a fixed starting and ending point, and satisfying $k_{n}(\gamma)=k_{n}$ and $k_{r}(\gamma)=k_{r}$ .

The point of Lemma 4.4 is that it implies that paths with large $k_{r}$ are so rare that they do not contribute any weight. Indeed, for some $\alpha$ to be determined choose $k^{*}$ large enough (depending on $\alpha$ ) so that

It then follows from Lemma 4.4 that if $\Gamma$ is the collection of all paths of length at most $4\alpha\log n$ with $k_{r}(\gamma)\geq k^{*}$ then

Note that for any path $\gamma$ and any labelling $\sigma$ ,

(since $k_{r}(\gamma)+k_{n}(\gamma)$ is the number of edges in $\gamma$ ). Hence, we have:

Let $k^{*}=k^{*}(\alpha,d)$ be defined in (13). Then

where the sum ranges over $\gamma$ of length at most $4\alpha\log n$ and with $k_{r}(\gamma)\geq k^{*}$ .

Consider paths of fixed length $k$ ; later, we will sum over all $k\leq C\log n$ . Suppose that for all $i$ , we decide in advance whether $(v_{i},v_{i+1})$ will be new, old, or returning. There are at most $\binom{k}{k_{n}\ k_{o}\ k_{r}}$ ways to make this choice. Fix an $i$ and suppose that $v_{i}$ has already been determined. If $(v_{i},v_{i+1})$ is new then there are at most $n$ choices for $v_{i+1}$ . If $(v_{i},v_{i+1})$ is returning then there are at most $|V(\gamma)|=k_{n}+1\leq k$ choices for $v_{i+1}$ . Otherwise, $(v_{i},v_{i+1})$ is an old edge, and there are at most $k_{r}+2$ choices for $v_{i+1}$ because $k_{r}+2$ bounds the maximum degree of the final path. Hence, the total number of choices is at most

(In the case $k_{o}=0$ , we adopt the convention $(y/0)^{0}=1$ .) Now, the quantity $(y/x)^{x}$ is increasing in $x$ as long as $x\leq y/e$ . Applying this with $y=2ekk_{r}$ and the values $x=k_{o}\leq k\leq y/e$ we have

On the other hand, $ek^{2}/(nk_{r})\leq n^{-2/3}$ for sufficiently large $n$ . Hence, the total number of paths of length $k$ is at most

Summing over $k\leq C\log n$ introduces an extra factor of $C\log n$ , but this factor is cancelled out by $n^{-k_{r}/6}$ for sufficiently large $n$ . ∎

The bounds of Lemma 4.4 are not accurate when $k_{r}$ is small. Essentially, we require bounds of $n^{k_{n}-1+o(1)}$ in order to make the rest of our argument work (certainly, we can’t expect any better bounds, since every new edge but the last one has almost $n$ choices). In order to achieve this bound, we need to introduce extra structure into our paths: they need to be non-backtracking and without many tangles.

First, note that if we specify which edges are returning and we also specify the first new edge after each returning edge, then we have also determined which edges are old (because every edge after a new edge but before the next returning edge is new). Therefore, the number of ways to specify which edges are old, new, or returning is at most $k^{2k_{r}}$ . Then there are at most $k^{k_{r}}$ ways to choose the returning edges and at most $n^{k_{n}-1}$ ways to choose the new edges (since one of them must hit the final vertex, so it has no choices). So far, we have made at most $k^{3k_{r}}n^{k_{n}-1}$ choices, and these choices determine the edges traversed by $\gamma$ .

Having fixed the edges traversed by $\gamma$ , we will now count old edges. We denote by $d(v)=d(v,\gamma)$ the degree of $v$ in $\gamma$ : that is, the number of $w\in\gamma$ such that $\{w,v\}\in E(\gamma)$ . Let ${V_{\geq 3}}$ be the set of vertices with degree at least 3. Note that because $\gamma$ is non-backtracking, if an edge $(v_{i},v_{i+1})$ is old, then $v_{i+1}$ is already determined by the path up to $v_{i}$ unless $v_{i}\in{V_{\geq 3}}$ .

where $m(v)^{m_{\text{out}}(v)+m_{\text{in}}(v)}$ bounds the number of ways that we can intersperse the short arrivals and departures among all visits to $v$ , and the second inequality follows from the fact that $d(v)\leq m(v)$ . Repeating this for all $v\in{V_{\geq 3}}$ , we see that the number of ways to choose all the old edges in $d$ is at most

where the second inequality follows because every time the walk returns to its old path, it creates at most two vertices of degree higher than two (one when the walk returns, and one when it leaves again). Since $m(v)\leq k$ for every $v$ , the quantity above is bounded by $k^{2k_{r}}$ . Plugging this back into (14), we get the claimed bound. ∎

When we take second moments over various sums over paths, we will end up having to control the number of pairs of paths with certain properties. In what follows, we take two self-avoiding paths, $\gamma_{1}$ and $\gamma_{2}$ , of length $k$ . We will refine Definition 4.1 by saying that a (directed) edge $(u,v)$ of $\gamma_{2}$ is new with respect to $\gamma_{1}$ if $v\not\in V(\gamma_{1})$ . We say that $(u,v)$ is old with respect to $\gamma_{1}$ if the (undirected) edge $\{u,v\}$ appears in $\gamma_{1}$ . Otherwise, we say that that $(u,v)$ is returning with respect to $\gamma_{1}$ . We write $k_{n,\gamma_{1}}(\gamma_{2})$ , $k_{o,\gamma_{1}}(\gamma_{2})$ , and $k_{r,\gamma_{1}}(\gamma_{2})$ for the numbers of edges of these three types in $\gamma_{2}$ .

Fix vertices $u,u^{\prime},v,v^{\prime}$ (not necessarily distinct). There are at most

pairs $(\gamma_{1},\gamma_{2})$ of length- $k$ self-avoiding paths where $\gamma_{1}$ goes from $u$ to $v$ , $\gamma_{2}$ goes from $u^{\prime}$ to $v^{\prime}$ , and where $k_{n,\gamma_{1}}(\gamma_{2})=k_{n,\gamma_{1}}$ and $k_{r,\gamma_{1}}(\gamma_{2})=k_{r,\gamma_{1}}$ .

For this proof, whenever we speak of old, new, or returning edges of $\gamma_{2}$ , we mean with respect to $\gamma_{1}$ .

First, assume that $v^{\prime}$ is not an interior node of $\gamma_{1}$ . There are at most $n^{k-1}$ such choices for $\gamma_{1}$ ; fix one and consider $\gamma_{2}$ . Every sequence of old edges in $\gamma_{2}$ either occurs at the beginning of $\gamma_{2}$ , or it is preceded by a returning edge. Hence, there are at most $\binom{k}{k_{r,\gamma_{1}}}\binom{k}{k_{r,\gamma_{1}}+1}$ choices for the edge types of $\gamma_{2}$ : $\binom{k}{k_{r,\gamma_{1}}}$ choices for which edges are returning, and at most $\binom{k}{k_{r,\gamma_{1}}+1}$ choices for the end of each sequence of old edges. Each new edge has at most $n$ choices for its endpoint. In the case that $v^{\prime}\not\in\{u,v\}$ then (since it is also not an interior node of $\gamma_{1}$ ) the last edge is new, but it has no choices. Hence, there are at most $n^{k_{n,\gamma_{1}}-1_{v^{\prime}\not\in\{u,v\}}}$ choices for the new edges. Each returning edge has at most $|E(\gamma_{1})|=k$ choices for its endpoint, and every sequence of old edges has at most $2$ choices: it must follow the (self-avoiding) path $\gamma_{1}$ , but it may do so in either direction; moreover, there are at most $k_{r,\gamma_{1}}+1$ distinct sequences of old edges. Hence, the total number of choices for $\gamma_{2}$ is bounded by

pairs that satisfy the conditions of the lemma, and also the additional constraint that $v^{\prime}$ is not an interior node of $\gamma_{1}$ .

Now suppose that $v^{\prime}$ is an interior node of $\gamma_{1}$ . There are at most $kn^{k-2}$ ways to choose such a $\gamma_{1}$ . We may repeat the previous paragraph to bound the number of choices of $\gamma_{2}$ , except that this time there will be up to $n^{k_{n,\gamma_{1}}}$ choices for the new edges, because the final edge of $\gamma_{2}$ may not be new. Hence, there are at most

pairs of paths of this form, where $v^{\prime}$ is an interior node of $\gamma_{1}$ . Combined with the other case, this proves the claim. ∎

Weighted sums over self-avoiding paths

In this section, we consider the behavior of weighted sums over self-avoiding paths. In particular, we will prove (1), (2), and (3) from Theorem 2.8.

Eventually, we will need to bound (or bound) the expected weight of complicated paths. Our basic building block for these computations is the expected weight of a self-avoiding path.

Let $\zeta$ be either a self-avoiding path or a simple cycle. Let $z$ be the length of $\zeta$ and let $u,v$ be the endpoints. If $pmz=n^{o(1)}$ then uniformly with respect to $\gamma$

First, consider the case $m=1$ . For any labelling $\tau$ that is compatible with $\sigma_{u}$ and $\sigma_{v}$ ,

Since $\zeta$ is a path, if $x$ is an interior vertex of $\zeta$ then $\tau_{x}$ appears exactly twice in the product above. Since $\tau_{x}^{2}=1$ , these terms all cancel out, leaving

which proves the claim in the case $m=1$ .

Now we take the average over all assignments $\tau$ that agree with $\sigma_{u}$ and $\sigma_{v}$ :

where the sum ranges over all $2^{z-1}$ labellings $\tau$ on $\zeta$ that agree with $\sigma_{u}$ and $\sigma_{v}$ . Combining this with (15) completes the proof. ∎

2 Decomposition into segments

Although our current goal is to understand the contribution of self-avoiding paths, in order to compute the second moment in Theorem 2.8, we will need to consider the concatenation of two self-avoiding paths (which may not be self-avoiding). Therefore, we introduce the following method for decomposing a general path into its self-avoiding pieces. This decomposition will also be useful in Section 6, where we consider more complicated paths.

Consider a path $\gamma$ . We say that a collection of paths $\zeta^{(1)},\dots,\zeta^{(r)}$ is a SAW-decomposition of $\gamma$ if

each $\zeta^{(i)}$ is a self-avoiding path;

the interior vertices of each $\zeta^{(i)}$ are not contained in any other $\zeta^{(j)}$ , nor is any interior vertex of $\zeta^{(i)}$ equal to the starting or ending vertex of $\gamma$ ; and

the $\zeta^{(i)}$ cover $\gamma$ , in the sense that $E(\gamma)=\bigcup_{i}E(\zeta^{(i)})$ .

Note that since $\zeta^{(i)}$ and $\zeta^{(j)}$ share no interior vertices, every time that the path $\gamma$ begins to traverse $\zeta^{(i)}$ , it must finish traversing $\zeta^{(i)}$ . Moreover, the fact that $\zeta^{(i)}$ and $\zeta^{(j)}$ share no interior vertices implies that they are edge-disjoint, and so for each fixed $i$ , every edge in $\zeta^{(i)}$ is traversed the same number (i.e. $m_{i}$ ) of times.

There is a natural way to construct a SAW-decomposition of a path $\gamma$ . Consider a path $\gamma$ between $u$ and $v$ , and let ${V_{\geq 3}}$ be the subset of $\gamma$ ’s vertices that have degree 3 or more in $\gamma$ . Let

We call the preceding construction of $\zeta^{(1)},\dots,\zeta^{(r)}$ the canonical SAW-decomposition of $\gamma$ .

For a set of vertices $U$ , if we run the preceding construction, but with

instead of as defined in (16), then we call the resulting SAW-decomposition the $U$ -canonical SAW-decomposition of $\gamma$ .

If $\zeta^{(1)},\dots,\zeta^{(r)}$ is the canonical SAW-decomposition of $\gamma$ then $r\leq 2k_{r}(\gamma)+B(\gamma)+1$ , where $B(\gamma)$ is the number of backtracks in $\gamma$ .

If $\zeta^{(1)},\dots,\zeta^{(r)}$ is the $U$ -canonical SAW-decomposition of $\gamma$ then $r\leq 2k_{r}(\gamma)+B(\gamma)+1+|U|$ .

Every returning edge in $\gamma$ increases $r$ by at most 2, since it can create a new SAW component, and it can split an existing component into 2 pieces. Every backtrack in $\gamma$ increases $r$ by at most 1, since it can create a new SAW component. This proves the first statement; to prove the second, note that each vertex $v\in A$ creates at most one new component, since if $v\in{V_{\geq 3}}$ then it has no effect, while if $v\not\in{V_{\geq 3}}$ then it has degree at most 2 in $\gamma$ and so splitting the path that goes through $v$ introduces at most one new component. ∎

3 The weight of a SAW-decomposition

We can compute the expected weight of a SAW-decomposition by simply applying Lemma 5.1 to each component. We state the following lemma slightly more generally, so that we may also apply it to subsets of the SAW-decomposition of a path.

where $k=\sum_{i}\sum_{e\in\zeta^{(i)}}m_{e}$ .

where the last inequality follows because $1<d<s^{2}$ , and so $m_{e}>1$ implies $d<s^{2}\leq s^{m_{e}}$ . ∎

4 Proof of (1)–(3)

Now we prove the first three parts of Theorem 2.8. The claim (1) about the first moment follows from Lemma 5.1.

For the second moment, we will expand the square in $Y_{u,v}^{2}$ . Suppose $\gamma_{1}$ and $\gamma_{2}$ are a pair of self-avoiding paths of length $k$ from $u$ to $v$ . By reversing $\gamma_{2}$ and appending it to $\gamma_{1}$ , we obtain a single path ( $\gamma$ , say) from $u$ to itself which passes through $v$ and backtracks at most once (at $v$ ). We consider the set of all $\gamma$ that can be obtained in this way, and divide them into four classes:

$\Gamma_{0}$ is the collection of such paths with $k_{r}(\gamma)=0$ . These paths begin with a self-avoiding walk from $u$ to $v$ , after which they backtrack at $v$ and walk back to $u$ along exactly the same path. They have $k$ edges, $k-1$ vertices, and every edge is visited twice.

$\Gamma_{1}$ is the collection of such paths with $k_{r}=1$ . These paths consist of a simple cycle that is traversed once, with up to two “tails” that are traversed twice each.

$\Gamma_{2}$ is the collection of such paths with $2\leq k_{r}\leq k^{*}$ .

$\Gamma_{3}$ is the collection of such paths with $k_{r}>k^{*}$ .

where the second inequality follows from our choice of $k$ in Theorem 2.8. In particular, this term is of a lower order than the bound claimed in the theorem.

Next, we consider $\Gamma_{1}$ . Recall that the first $k$ steps of $\gamma\in\Gamma_{1}$ make up a simple path. Let $i$ be minimal so that the $(k+i+1)$ th step of $\gamma$ is new; let $j$ be such that the $2k-j$ th step of $\gamma$ is returning. It follows that the first $j$ edges of $\gamma$ consist of a simple path where each edge is traversed twice. The same holds for edges $k-i+1$ through $k-1$ . The rest of $\gamma$ consists of a simple cycle of length $2k-2(i+j)$ , each edge of which is traversed once. Let $\Gamma_{1}(i,j)$ denote the set of such paths. By Lemma 5.6, if $\gamma$ ’s interior does not intersect $U$ then the expected weight of $\gamma\in\Gamma_{1}(i,j)$ is bounded by

Now, $|\Gamma_{1}(i,j)|=(1+o(1))n^{2k-i-j-2}$ because $\gamma\in\Gamma_{1}(i,j)$ has $2k-i-j$ distinct vertices (including $u$ and $v$ ), and once those vertices and their order is fixed then $\gamma$ is determined. As in the argument for $\Gamma_{0}$ , the paths whose interiors intersect $U$ provide a negligible contribution, and hence

Hence, $\Gamma_{1}$ provides the main term in the claimed bound.

Summing over the $k$ choices of $k_{n}$ shows that the paths in $\Gamma_{2}$ contribute a smaller order term than the paths in $\Gamma_{1}$ .

Finally, we bound $\Gamma_{3}$ using Corollary 4.5; these terms also contribute a smaller order term.

4.2 The cross moment

$\Gamma_{0}$ are the pairs of paths that do not intersect.

$\Gamma_{1}$ are the pairs of paths that do intersect, and that satisfy $k_{r,\gamma_{1}}(\gamma_{2})\leq k^{*}$ .

$\Gamma_{2}$ are the pairs of paths that satisfy $k_{r,\gamma_{1}}(\gamma_{2})>k^{*}$ .

For $(\gamma_{1},\gamma_{2})\in\Gamma_{0}$ , the variables $X_{\gamma_{1}}$ and $X_{\gamma_{2}}$ are independent, and hence

We recall from (1) that the right hand side above is of the order $s^{2k}n^{-2}$ ; in order to prove the claim about the cross moments, we need to show that the contributions of $\Gamma_{1}$ and $\Gamma_{2}$ are of the order $s^{2k}n^{-3+o(1)}$ .

To control $\Gamma_{1}$ , we split pairs of paths according to $k_{n,\gamma_{1}}(\gamma_{2})$ . If $\Gamma_{1,k_{n,\gamma_{1}}}$ is the set of pairs of paths in $\Gamma_{1}$ with $k_{n,\gamma_{1}}(\gamma_{2})=k_{n,\gamma_{1}}$ then Lemma 4.8 implies that $|\Gamma_{1,k_{n,\gamma_{1}}}|\leq n^{k+k_{n,\gamma_{1}}-2+o(1)}$ (when applying Lemma 4.8, recall that $v^{\prime}$ is distinct from $u$ and $v$ ). By Lemma 5.6, and noting that $|E(\gamma_{1})\cup E(\gamma_{2})|\geq k+k_{n,\gamma_{1}}+1$ because $k_{r,\gamma_{1}}(\gamma_{2})\geq 1$ ,

Summing over the $k$ possible values of $k_{n,\gamma_{1}}$ adds another $n^{o(1)}$ factor, and we conclude that $\Gamma_{1}$ is a lower order term.

Finally, we control $\Gamma_{2}$ . For each pair $(\gamma_{1},\gamma_{2})\in\Gamma_{2}$ , we may create a new path $\gamma$ by joining the end of $\gamma_{1}$ to the beginning of $\gamma_{2}$ . Then $\gamma$ has length $2k+1$ and $k_{r}(\gamma)\geq k^{*}$ . Note that $|X_{\gamma}|\geq\frac{1}{n}|X_{\gamma_{1}}X_{\gamma_{2}}|$ because the new edge that we added always has $|W_{e}|\geq\frac{1}{n}$ . Hence,

where the second sum ranges over all $\gamma$ of length $2k+1$ satisfying $k_{r}(\gamma)\geq k^{*}$ . But by Corollary 4.5, the last quantity is at most $n^{-3+o(1)}$ , and so $\Gamma_{2}$ contributes a lower order term.

Weighted sums over complicated paths

For any path $\gamma$ , $\Xi\subset\Omega_{\mathcal{F}(\gamma)}$ .

Let $H$ be any fixed graph with $m$ vertices and $m+1$ edges. There are at most $n^{m}$ ways to embed $H$ into $G$ , and for each of those embeddings, the probability that all edges in $H$ appear is at most $(2d/n)^{m}$ . By a union bound, the probability that $H$ is a subgraph of $G$ is at most $n^{-1}(2d)^{m}$ .

Before proceeding to bound the weight of non-self-avoiding paths, we present one more preliminary lemma. Because we will take a second moment, we will need to handle pairs of non-self-avoiding paths. In order to do so, we need to interpret the condition $\sum_{e\in F}m_{e}\geq t$ in the definition of $\mathcal{F}(\gamma)$ for pairs of paths. In the following lemma we deal with multiple paths, so we will write $m_{e}(\gamma)$ for the number of times that the path $\gamma$ crosses the edge $e$ .

Let $\gamma_{1}$ and $\gamma_{2}$ be two paths from $u$ to $v$ of length $k$ . Let $\gamma$ be the path from $u$ to $u$ obtained by first following $\gamma_{1}$ and then following the reversal of $\gamma_{2}$ . For any $F_{1}\in\mathcal{F}(\gamma_{1})$ and $F_{2}\in\mathcal{F}(\gamma_{2})$ ,

Let $F=F_{1}\cup F_{2}$ , $F_{1}^{\prime}=(F_{1}\setminus F_{2})\cap E(\gamma_{2})$ , and $F_{2}^{\prime}=(F_{2}\setminus F_{1})\cap E(\gamma_{1})$ . Then set $H=F\setminus(F_{1}^{\prime}\cup F_{2}^{\prime})$ . Recall that $\gamma_{i}\setminus F_{i}$ has at least as many edges as vertices.

also has at least as many edges as vertices. Hence, $|H|\leq k_{r}(\gamma)-1$ .

For every $e\in F_{i}^{\prime}$ , we have $m_{e}(\gamma)\geq m_{e}(\gamma_{i})+1$ . Hence,

Let $\gamma_{1}$ and $\gamma_{2}$ be non-backtracking paths from $u$ to $v$ of length $k$ that are not self-avoiding. That is, $k_{r}(\gamma_{1}),k_{r}(\gamma_{2})\geq 1$ . Let $\gamma$ be the path obtained by first following $\gamma_{1}$ and then following $\gamma_{2}$ backwards. Let $t(\gamma_{i})$ be the number of tangles in $\gamma_{i}$ .

Recall from (13) that $k^{*}$ is a constant (depending on $s$ and $d$ ) such that paths with $k_{r}>k^{*}$ are irrelevant.

uniformly over $\gamma$ and $\sigma_{U}$ .

The term involving $X_{L}$ may be bounded by Lemma 5.6 (recalling that the combined path $\gamma$ has length $2k$ ):

Next, we turn to the first term of (18). Recall that $\Omega_{F}$ is the event that no edge in $F$ appears. Hence, if $F_{1}\in\mathcal{F}(\gamma_{1})$ , $F_{2}\in\mathcal{F}(\gamma_{2})$ , and $F=F_{1}\cup F_{2}$ then

where we applied Lemma 6.3 in the first inequality above. Hence,

Now, $|E(K)|+|E(L)|$ is the number of distinct edges traversed by $\gamma$ , which is also equal to $k_{n}(\gamma)+k_{r}(\gamma)$ . Applying this in the exponent of $n$ completes the proof. ∎

2 Proof of (4)

We will now combine Lemma 6.4 with our earlier bounds on the number of paths (Lemma 4.7) to show that the total weight of non-self-avoiding paths is negligible on the event that the graph contains no tangles.

because both sides are non-negative and, by Lemma 6.1, they agree whenever the left hand side is non-zero. Next, we expand the sum above as

Combining the last two displayed equations,

Let $\gamma=\gamma(\gamma_{1},\gamma_{2})$ be $\gamma_{1}$ concatenated with the reverse of $\gamma_{2}$ . Let $\Lambda_{1}\subset\Gamma^{\text{bad}}\times\Gamma^{\text{bad}}$ be the set of pairs $(\gamma_{1},\gamma_{2})$ such that $\gamma(\gamma_{1},\gamma_{2})$ has more than $k^{*}$ returning edges. Let $\Lambda_{2}\subset(\Gamma^{\text{bad}}\times\Gamma^{\text{bad}})\setminus\Lambda_{1}$ be the set of pairs $(\gamma_{1},\gamma_{2})$ such that $|(V(\gamma_{1})\cup V(\gamma_{2}))\cap U|\geq 2\sqrt{\log n}$ . Let $\Lambda=\Lambda_{1}\cup\Lambda_{2}$ . By Corollary 4.5,

Since $|U|=n^{o(1)}\leq n^{1/2}$ for large enough $n$ , the fraction of $\gamma_{1}\in\Gamma^{\text{bad}}$ such that $|V(\gamma_{1})\cap U|\geq\sqrt{\log n}$ is at most $n^{-c\sqrt{\log n}}$ for some constant $c>0$ . By Lemma 4.4,

We will further split this sum according to the number of tangles in $\gamma_{1}$ and $\gamma_{2}$ ; that is, we define $\Gamma^{\text{bad}}_{t}$ to be the set of $\gamma_{1}\in\Gamma^{\text{bad}}$ with $t(\gamma_{1})=t$ . We will show that for any $t_{1}$ and $t_{2}$ ,

Summing over the $k=n^{o(1)}$ possible values of $t_{1}$ and $t_{2}$ , this will imply (20) and complete the proof.

To control (21), fix $\gamma_{1}$ and consider the sum over $\gamma_{2}$ . By Lemma 4.7, there are at most $n^{k_{n}+t_{2}-1+o(1)}$ choices of $\gamma_{2}\in\Gamma^{\text{bad}}_{t_{2}}$ that satisfy $k_{n}(\gamma_{2})=k_{n}$ (denote this set by $\Gamma^{\text{bad}}_{t_{2},k_{n}}$ ). Note that the fraction of $\gamma_{2}\in\Gamma^{\text{bad}}_{t_{2},k_{n}}$ satisfying $k_{n}(\gamma)=k_{n}(\gamma_{1})+k_{n}(\gamma_{2})-m$ is at most $k^{2m}(n-2k)^{-m}$ . Indeed, $m=k_{n}(\gamma_{1})+k_{n}(\gamma_{2})-k_{n}(\gamma)$ is the number of edges that were new in $\gamma_{2}$ but not in $\gamma$ . There are at most $k^{m}$ ways to choose which edges in $\gamma_{2}$ will no longer be new and each one has at most $k^{m}$ choices for a non-new step, versus at least $(n-2k)^{m}$ choices for a new step. Set $\Gamma^{\text{bad}}_{t_{2},k_{n},m,\gamma_{1}}$ to be the paths $\gamma_{2}\in\Gamma^{\text{bad}}_{t_{2},k_{n}}$ satisfying $k_{n}(\gamma)=k_{n}(\gamma_{1})+k_{n}(\gamma_{2})-m$ . Note that if $(\gamma_{1},\gamma_{2})\not\in\Lambda$ then the total number of returning edges in $\gamma$ is at most $k^{*}=O(1)$ , and the number of vertices in $U$ intersecting $V(\gamma)$ is at most $2\sqrt{\log n}$ . Hence, Lemma 6.4 applied with $U=U\cap V(\gamma)$ implies that for any $\gamma_{1}\in\Gamma^{\text{bad}}_{t_{1}}$ and any $m$ ,

Taking the sum over $k_{n}\leq k$ and $m\leq k$ only contributes a factor of $n^{o(1)}$ ; hence,

Summing over the $n^{k_{n}+t_{1}-1+o(1)}$ possible $\gamma_{1}\in\Gamma^{\text{bad}}_{t_{1},k_{n}}$ and then over the $k$ possible values of $k_{n}$ , we see that the right hand side of (21) is bounded by $s^{2k}n^{-3+o(1)}$ , as claimed. ∎

Finally, note that we have finished the proof of Theorem 2.8. Indeed, we proved (1) at the beginning of Section 5.4, (2) in Section 5.4.1, (3) in Section 5.4.2, and we just proved (4).

Acknowledgments

The authors are grateful to Cris Moore and Lenka Zdeborová for stimulating and interesting discussions on many aspects of the block model. They also thank the Charles Bordenave and the anonymous referees for pointing out several simplifications and corrections.