Interactive Privacy via the Median Mechanism

Aaron Roth, Tim Roughgarden

Introduction

Managing a data set with sensitive but useful information, such as medical records, requires reconciling two objectives: providing utility to others, perhaps in the form of aggregate statistics; and respecting the privacy of individuals who contribute to the data set. The field of private data analysis, and in particular work on differential privacy, provides a mathematical foundation for reasoning about this utility-privacy trade-off and offers methods for non-trivial data analysis that are provably privacy-preserving in a precise sense. For a recent survey of the field, see Dwork [Dwo08].

More precisely, consider a domain $X$ and database size $n$ . A mechanism is a randomized function from the set $X^{n}$ of databases to some range. For a parameter $\alpha>0$ , a mechanism $M$ is $\alpha$ -differentially private if, for every database $D$ and fixed subset $S$ of the range of $M$ , changing a single component of $D$ changes the probability that $M$ outputs something in $S$ by at most an $e^{\alpha}$ factor. The output of a differentially private mechanism (and any analysis or privacy attack that follows) is thus essentially independent of whether or not a given individual “opts in” or “opts out” of the database.

Achieving differential privacy requires “sufficiently noisy” answers [DN03]. For example, suppose we’re interested in the result of a query $f$ — a function from databases to some range — that simply counts the fraction of database elements that satisfy some predicate $\varphi$ on $X$ . A special case of a result in Dwork et al. [DMNS06] asserts that the following mechanism is $\alpha$ -differentially private: if the underlying database is $D$ , output $f(D)+\Delta$ , where the output perturbation $\Delta$ is drawn from the Laplace distribution $\textrm{Lap}(\tfrac{1}{n\alpha})$ with density $p(y)=\tfrac{n\alpha}{2}\exp(-n\alpha|y|)$ . Among all $\alpha$ -differentially private mechanisms, this one (or rather, a discretized analog of it) maximizes user utility in a strong sense [GRS09].

What if we care about more than a single one-dimensional statistic? Suppose we’re interested in $k$ predicate queries $f_{1},\ldots,f_{k}$ , where $k$ could be large, even super-polynomial in $n$ . A natural solution is to use an independent Laplace perturbation for each query answer [DMNS06]. To maintain $\alpha$ -differential privacy, the magnitude of noise has to scale linearly with $k$ , with each perturbation drawn from $\textrm{Lap}(\tfrac{k}{n\alpha})$ . Put another way, suppose one fixes “usefulness parameters” $\epsilon,\delta$ , and insists that the mechanism is $(\epsilon,\delta)$ -useful, meaning that the outputs are within $\epsilon$ of the correct query answers with probability at least $1-\delta$ . This constrains the magnitude of the Laplace noise, and the privacy parameter $\alpha$ now suffers linearly with the number $k$ of answered queries. This dependence limits the use of this mechanism to a sublinear $k=o(n)$ number of queries.

Can we do better than independent output perturbations? For special classes of queries like predicate queries, Blum, Ligett, and Roth [BLR08] give an affirmative answer (building on techniques of Kasiviswanathan et al. [KLN+08]). Specifically, in [BLR08] the exponential mechanism of McSherry and Talwar [MT07] is used to show that, for fixed usefulness parameters $\epsilon,\delta$ , the privacy parameter $\alpha$ only has to scale logarithmically with the number of queries.More generally, linearly with the VC dimension of the set of queries, which is always at most $\log_{2}k$ . This permits simultaneous non-trivial utility and privacy guarantees even for an exponential number of queries. Moreover, this dependence on $\log k$ is necessary in every differentially private mechanism (see the full version of [BLR08]).

The mechanism in [BLR08] suffers from two drawbacks, however. First, it is non-interactive: it requires all queries $f_{1},\ldots,f_{k}$ to be given up front, and computes (noisy) outputs of all of them at once.Or rather, it computes a compact representation of these outputs in the form of a synthetic database. By contrast, independent Laplace output perturbations can obviously be implemented interactively, with the queries arriving online and each answered immediately. There is good intuition for why the non-interactive setting helps: outperforming independent output perturbations requires correlating perturbations across multiple queries, and this is clearly easier when the queries are known in advance. Indeed, prior to the present work, no interactive mechanism better than independent Laplace perturbations was known.

Second, the mechanism in [BLR08] is inefficient. Here by “efficient” we mean has running time polynomial in $n$ , $k$ , and $|X|$ ; Dwork et al. [DNR+09] prove that this is essentially the best one could hope for (under certain cryptographic assumptions). The mechanism in [BLR08] is not efficient because it requires sampling from a non-trivial probability distribution over an unstructured space of exponential size. Dwork et al. [DNR+09] recently gave an efficient (non-interactive) mechanism that is better than independent Laplace perturbations, in that the privacy parameter $\alpha$ of the mechanism scales as $2^{\sqrt{\log k}}$ with the number of queries $k$ (for fixed usefulness parameters $\epsilon,\delta$ ).

We define a new interactive differentially private mechanism for answering $k$ arbitrary predicate queries, called the median mechanism.The privacy guarantee is $(\alpha,\tau)$ -differential privacy for a negligible function $\tau$ ; see Section 2 for definitions. The basic implementation of the median mechanism interactively answers queries $f_{1},\ldots,f_{k}$ that arrive online, is $(\epsilon,\delta)$ -useful, and has privacy $\alpha$ that scales with $\log k\log|X|$ ; see Theorem 4.1 for the exact statement. These privacy and utility guarantees hold even if an adversary can adaptively choose each $f_{i}$ after seeing the mechanism’s first $i-1$ answers. This is the first interactive mechanism better than the Laplace mechanism, and its performance is close to the best possible even in the non-interactive setting.

The basic implementation of the median mechanism is not efficient, and we give an efficient implementation with a somewhat weaker utility guarantee. (The privacy guarantee is as strong as in the basic implementation.) This alternative implementation runs in time polynomial in $n$ , $k$ , and $|X|$ , and satisfies the following (Theorem 5.1): for every sequence $f_{1},\ldots,f_{k}$ of predicate queries, for all but a negligible fraction of input distributions, the efficient median mechanism is $(\epsilon,\delta)$ -useful.

This is the first efficient mechanism with a non-trivial utility guarantee and polylogarithmic privacy cost, even in the non-interactive setting.

2 The Main Ideas

The key challenge to designing an interactive mechanism that outperforms the Laplace mechanism lies in determining the appropriate correlations between different output perturbations on the fly, without knowledge of future queries. It is not obvious that anything significantly better than independent perturbations is possible in the interactive setting.

Our median mechanism and our analysis of it can be summarized, at a high level, by three facts. First, among any set of $k$ queries, we prove that there are $O(\log k\log|X|)$ “hard” queries, the answers to which completely determine the answers to all of the other queries (up to $\pm\epsilon$ ). Roughly, this holds because: (i) by a VC dimension argument, we can focus on databases over $X$ of size only $O(\log k)$ ; and (ii) every time we answer a “hard” query, the number of databases consistent with the mechanism’s answers shrinks by a constant factor, and this number cannot drop below 1 (because of the true input database). Second, we design a method to privately release an indicator vector which distinguishes between hard and easy queries online. We note that a similar private ‘indicator vector’ technique was used by Dwork et al. [DNR+09]. Essentially, the median mechanism deems a query “easy” if a majority of the databases that are consistent (up to $\pm\epsilon$ ) with the previous answers of the mechanism would answer the current query accurately. The median mechanism answers the small number of hard queries using independent Laplace perturbations. It answers an easy query (accurately) using the median query result given by databases that are consistent with previous answers. A key intuition is that if a user knows that query $i$ is easy, then it can generate the mechanism’s answer on its own. Thus answering an easy query communicates only a single new bit of information: that the query is easy. Finally, we show how to release the classification of queries as “easy” and “hard” with low privacy cost; intuitively, this is possible because (independent of the database) there can be only $O(\log k\log|X|)$ hard queries.

Our basic implementation of the median mechanism is not efficient for the same reasons as for the mechanism in [BLR08]: it requires non-trivial sampling from a set of super-polynomial size. For our efficient implementation, we pass to fractional databases, represented as fractional histograms with components indexed by $X$ . Here, we use the random walk technology of Dyer, Frieze, and Kannan [DFK91] for convex bodies to perform efficient random sampling. To explain why our utility guarantee no longer holds for every input database, recall the first fact used in the basic implementation: every answer to a hard query shrinks the number of consistent databases by a constant factor, and this number starts at $|X|^{O(\log k)}$ and cannot drop below 1. With fractional databases (where polytope volumes play the role of set sizes), the lower bound of 1 on the set of consistent (fractional) databases no longer holds. Nonetheless, we prove a lower bound on the volume of this set for almost all fractional histograms (equivalently probability distributions), which salvages the $O(\log k\log|X|)$ bound on hard queries for databases drawn from such distributions.

Preliminaries

A mechanism $M$ is $(\epsilon,\delta)$ -useful if for every sequence of queries $(f_{1},\ldots,f_{k})$ and every database $D$ , with probability at least $1-\delta$ it provides answers $a_{1},\ldots,a_{k}$ that are $\epsilon$ -accurate for $f_{1},\ldots,f_{k}$ and $D$ .

Recall that differential privacy means that changing the identity of a single element of the input database does not affect the probability of any outcome by more than a small factor. Formally, given a database $D$ , we say that a database $D^{\prime}$ of the same size is a neighbor of $D$ if it differs in only a single element: $|D\cap D^{\prime}|=|D|-1$ .

We are generally interested in the case where $\tau$ is a negligible function of some of the problem parameters, meaning one that goes to zero faster than $x^{-c}$ for every constant $c$ .

Finally, the sensitivity of a real-valued query is the largest difference between its values on neighboring databases. For example, the sensitivity of every non-trivial predicate query is precisely $1/n$ .

The Median Mechanism: Basic Implementation

We now describe the median mechanism and our basic implementation of it. As described in the Introduction, the mechanism is conceptually simple. It classifies queries as “easy” or “hard”, essentially according to whether or not a majority of the databases consistent with previous answers to hard queries would give an accurate answer to it (in which case the user already “knows the answer”). Easy queries are answered using the corresponding median value; hard queries are answered as in the Laplace mechanism.

To explain the mechanism precisely, we need to discuss a number of parameters. We take the privacy parameter $\alpha$ , the accuracy parameter $\epsilon$ , and the number $k$ of queries as input; these are hard constraints on the performance of our mechanism.We typically think of $\alpha,\epsilon$ as small constants, though our results remain meaningful for some sub-constant values of $\alpha$ and $\epsilon$ as well. We always assume that $\alpha$ is at least inverse polynomial in $k$ . Note that when $\alpha$ or $\epsilon$ is sufficiently small (at most $c/n$ for a small constant $c$ , say), simultaneously meaningful privacy and utility is clearly impossible. Our mechanism obeys these constraints with a value of $\delta$ that is inverse polynomial in $k$ and $n$ , and a value of $\tau$ that is negligible in $k$ and $n$ , provided $n$ is sufficiently large (at least polylogarithmic in $k$ and $|X|$ , see Theorem 4.1). Of course, such a result can be rephrased as a nearly exponential lower bound on the number of queries $k$ that can be successfully answered as a function of the database size $n$ .In contrast, the number of queries that the Laplace mechanism can privately and usefully answer is at most linear.

The median mechanism is shown in Figure 1, and it makes use of several additional parameters. For our analysis, we set their values to:

The denominator in (2) can be thought of as our “privacy cost” as a function of the number of queries $k$ . Needless to say, we made no effort to optimize the constants.

The value $r_{i}$ in Step 2(a) of the median mechanism is defined as

For the Laplace perturbations in Steps 2(a) and 2(d), recall that the distribution $\textrm{Lap}(\sigma)$ has the cumulative distribution function

The motivation behind the mechanism’s steps is as follows. The set $C_{i}$ is the set of size- $m$ databases consistent (up to $\pm\epsilon/50$ ) with previous answers of the mechanism to hard queries. The focus on databases with the small size $m$ is justified by a VC dimension argument, see Proposition 4.6. Steps 2(a) and 2(b) choose a random value $\hat{r}_{i}$ and a random threshold $t_{i}$ . The value $r_{i}$ in Step 2(a) is a measure of how easy the query is, with higher numbers being easier. A more obvious measure would be the fraction of databases $S$ in $C_{i-1}$ for which $|f_{i}(S)-f_{i}(D)|\leq\epsilon$ , but this is a highly sensitive statistic (unlike $r_{i}$ , see Lemma 4.9). The mechanism uses the perturbed value $\hat{r}_{i}$ rather than $r_{i}$ to privately communicate which queries are easy and which are hard. In Step 2(b), we choose the threshold $t_{i}$ at random between $3/4$ and $9/10$ . This randomly shifted threshold ensures that, for every database $D$ , there is likely to be a significant gap between $r_{i}$ and $t_{i}$ ; such gaps are useful when optimizing the privacy guarantee. Steps 2(c) and 2(d) answer easy and hard queries, respectively. Step 2(e) updates the set of databases consistent with previous answers to hard queries. We prove in Lemma 4.7 that Step 2(f) occurs with at most inverse polynomial probability.

Finally, we note that the median mechanism is defined as if the total number of queries $k$ is (approximately) known in advance. This assumption can be removed by using successively doubling “guesses” of $k$ ; this increases the privacy cost by an $O(\log k)$ factor.

Analysis of Median Mechanism

This section proves the following privacy and utility guarantees for the basic implementation of the median mechanism.

For every sequence of adaptively chosen predicate queries $f_{1},\ldots,f_{k}$ arriving online, the median mechanism is $(\epsilon,\delta)$ -useful and $(\alpha,\tau)$ -differentially private, where $\tau$ is a negligible function of $k$ and $|X|$ , and $\delta$ is an inverse polynomial function of $k$ and $n$ , provided the database size $n$ satisfies

We prove the utility and privacy guarantees in Sections 4.1 and 4.2, respectively.If desired, in Theorem 4.1 we can treat $n$ as a parameter and solve for the error $\epsilon$ . The maximum error on any query (normalized by the database size) is $O(\log k\log^{1/3}|X|/n^{1/3}\alpha^{1/3})$ ; the unnormalized error is a factor of $n$ larger.

Here we prove a utility guarantee for the median mechanism.

The median mechanism is $(\epsilon,\delta)$ -useful, where $\delta=k\exp(-\Omega(\epsilon n\alpha^{\prime}))$ .

Note that under assumption (6), $\delta$ is inverse polynomial in $k$ and $n$ .

We give the proof of Theorem 4.2 in three pieces: with high probability, every hard query is answered accurately (Lemma 4.4); every easy query is answered accurately (Lemmas 4.3 and 4.5); and the algorithm does not fail (Lemma 4.7). The next two lemmas follow from the definition of the Laplace distribution (5), our choice of $\delta$ , and trivial union bounds.

With probability at least $1-\tfrac{\delta}{2}$ , $|r_{i}-\hat{r}_{i}|\leq 1/100$ for every query $i$ .

With probability at least $1-\tfrac{\delta}{2}$ , every answer to a hard query is $(\epsilon/100)$ -accurate for $D$ .

The next lemma shows that median answers are accurate for easy queries.

If $|r_{i}-\hat{r}_{i}|\leq 1/100$ for every query $i$ , then every answer to an easy query is $\epsilon$ -accurate for $D$ .

For a query $i$ , let $G_{i-1}=\{S\in C_{i-1}:|f_{i}(D)-f_{i}(S)|\leq\epsilon\}$ denote the databases of $C_{i-1}$ on which the result of query $f_{i}$ is $\epsilon$ -accurate for $D$ . Observe that if $|G_{i-1}|\geq.51\cdot|C_{i-1}|$ , then the median value of $f_{i}$ on $C_{i-1}$ is an $\epsilon$ -accurate answer for $D$ . Thus proving the lemma reduces to showing that $\hat{r}_{i}\geq 3/4$ only if $|G_{i-1}|\geq.51\cdot|C_{i-1}|$ .

Consider a query $i$ with $|G_{i-1}|<.51\cdot|C_{i-1}|$ . Using (4), we have

Since $|r_{i}-\hat{r}_{i}|\leq 1/100$ for every query $i$ by assumption, the proof is complete. ∎

Our final lemma shows that the median mechanism does not fail and hence answers every query, with high probability; this will conclude our proof of Theorem 4.2. We need the following preliminary proposition, which instantiates the standard uniform convergence bound with the fact that the VC dimension of every set of $k$ predicate queries is at most $\log_{2}k$ [Vap96]. Recall the definition of the parameter $m$ from (1).

For every collection of $k$ predicate queries $f_{1},\ldots,f_{k}$ and every database $D$ , a database $S$ obtained by sampling points from $D$ uniformly at random will satisfy $|f_{i}(D)-f_{i}(S)|\leq\epsilon$ for all $i$ except with probability $\delta$ , provided

In particular, there exists a database $S$ of size $m$ such that for all $i\in\{1,\ldots,k\}$ , $|f_{i}(D)-f_{i}(S)|\leq\epsilon/400$ .

In other words, the results of $k$ predicate queries on an arbitrarily large database can be well approximated by those on a database with size only $O(\log k)$ .

If $|r_{i}-\hat{r}_{i}|\leq 1/100$ for every query $i$ and every answer to a hard query is $(\epsilon/100)$ -accurate for $D$ , then the median mechanism answers fewer than $20m\log|X|$ hard queries (and hence answers all queries before terminating).

The plan is to track the contraction of $C_{i}$ as hard queries are answered by the median mechanism. Initially we have $|C_{0}|\leq|X|^{m}$ . If the median mechanism answers a hard query $i$ , then the definition of the mechanism and our hypotheses yield

We then claim that the size of the set $C_{i}=\{S\in C_{i-1}:|f_{i}(S)-a_{i}|\leq\epsilon/50\}$ is at most $\tfrac{94}{100}|C_{i-1}|$ . For if not,

Iterating now shows that the number of consistent databases decreases exponentially with the number of hard queries:

On the other hand, Proposition 4.6 guarantees the existence of a database $S^{*}\in C_{0}$ for which $|f_{i}(S^{*})-f_{i}(D)|\leq\epsilon/100$ for every query $f_{i}$ . Since all answers $a_{i}$ produced by the median mechanism for hard queries $i$ are $(\epsilon/100)$ -accurate for $D$ by assumption, $|f_{i}(S^{*})-a_{i}|\leq|f_{i}(S^{*})-f_{i}(D)|+|f_{i}(D)-a_{i}|\leq\epsilon/50$ . This shows that $S^{*}\in C_{k}$ and hence $|C_{k}|\geq 1$ . Combining this with (7) gives

2 Privacy of the Median Mechanism

This section establishes the following privacy guarantee for the median mechanism.

The median mechanism is $(\alpha,\tau)$ - differentially private, where $\tau$ is a negligible function of $|X|$ and $k$ when $n$ is sufficiently large (as in (6)).

Our first lemma states that the small sensitivity of predicate queries carries over, with a $2/\epsilon$ factor loss, to the $r$ -function defined in (4).

The function $r_{i}(D)=(\sum_{S\in C}\exp(-\epsilon^{-1}|f(D)-f(S)|)/|C|$ has sensitivity $\frac{2}{\epsilon n}$ for every fixed set $C$ of databases and predicate query $f$ .

Let $D$ and $D^{\prime}$ be neighboring databases. Then

where the first inequality follows from the fact that the (predicate) query $f$ has sensitivity $1/n$ , the second from the fact that $e^{x}\leq 1+2x$ when $x\in$ , and the third from the fact that $r_{i}(D^{\prime})\leq 1$ . ∎

The next lemma identifies nice properties of “typical executions” of the median mechanism. Consider an output $({d},{a})$ of the median mechanism with a database $D$ . From $D$ and $({d},{a})$ , we can uniquely recover the values $r_{1},\ldots,r_{k}$ computed (via (4)) in Step 2(a) of the median mechanism, with $r_{i}$ depending only on the first $i-1$ components of $d$ and $a$ . We sometimes write such a value as $r_{i}(D,({d},{a}))$ , or as $r_{i}(D)$ if an output $({d},{a})$ has been fixed. Call a possible threshold $t_{i}$ good for $D$ and $({d},{a})$ if $d_{i}=0$ and $r_{i}(D,({d},{a}))\geq t_{i}+\gamma$ , where $\gamma$ is defined as in (3). Call a vector ${t}$ of possible thresholds good for $D$ and $({d},{a})$ if all but $180m\ln|X|$ of the thresholds are good for $D$ and $({d},{a})$ .

For every database $D$ , with all but negligible ( $\exp(-\Omega(\log k\log|X|/\epsilon^{2}))$ ) probability, the thresholds ${t}$ generated by the median mechanism are good for its output $({d},{a})$ .

The idea is to “charge” the probability of bad thresholds to that of answering hard queries, which are strictly limited by the median mechanism. Since the median mechanism only allows $20m\ln|X|$ of the $d_{i}$ ’s to be 1, we only need to bound the number of queries $i$ with output $d_{i}=0$ and threshold $t_{i}$ satisfying $r_{i}<t_{i}+\gamma$ , where $r_{i}$ is the value computed by the median mechanism in Step 2(a) when it answers the query $i$ .

Let $Y_{i}$ be the indicator random variable corresponding to the (larger) event that $r_{i}<t_{i}+\gamma$ . Define $Z_{i}$ to be 1 if and only if, when answering the $i$ th query, the median mechanism chooses a threshold $t_{i}$ and a Laplace perturbation $\Delta_{i}$ such that $r_{i}+\Delta_{i}<t_{i}$ (i.e., the query is classified as hard). If the median mechanism fails before reaching query $i$ , then we define $Y_{i}=Z_{i}=0$ . Set $Y=\sum_{i=1}^{k}Y_{i}$ and $Z=\sum_{i=1}^{k}Z_{i}$ . We can finish the proof by showing that $Y$ is at most $160m\ln|X|$ except with negligible probability.

Consider a query $i$ and condition on the event that $r_{i}\geq\tfrac{9}{10}$ ; this event depends only on the results of previous queries. In this case, $Y_{i}=1$ only if $t_{i}=9/10$ . But this occurs with probability $2^{-3/20\gamma}$ , which using (3) and (6) is at most $1/k$ .For simplicity, we ignore the normalizing constant in the distribution over $j$ ’s in Step 2(b), which is $\Theta(1)$ . Therefore, the expected contribution to $Y$ coming from queries $i$ with $r_{i}\geq\tfrac{9}{10}$ is at most $1$ . Since $t_{i}$ is selected independently at random for each $i$ , the Chernoff bound implies that the probability that such queries contribute more than $m\ln|X|$ to $Y$ is

Now condition on the event that $r_{i}<\tfrac{9}{10}$ . Let $T_{i}$ denote the threshold choices that would cause $Y_{i}$ to be 1, and let $s_{i}$ be the smallest such; since $r_{i}<\tfrac{9}{10}$ , $|T_{i}|\geq 2$ . For every $t_{i}\in T_{i}$ , $t_{i}>r_{i}-\gamma$ ; hence, for every $t_{i}\in T_{i}\setminus\{s_{i}\}$ , $t_{i}>r_{i}$ . Also, our distribution on the $j$ ’s in Step 2(b) ensures that $\Pr[t_{i}\in T_{i}\setminus\{s_{i}\}]\geq\tfrac{1}{2}\Pr[t_{i}\in T_{i}]$ . Since the Laplace distribution is symmetric around zero and the random choices $\Delta_{i},t_{i}$ are independent, we have

The definition of the median mechanism ensures that $Z\leq 20m\ln|X|$ with probability 1. Linearity of expectation, inequality (8), and the Chernoff bound imply that queries with $r_{i}<\tfrac{9}{10}$ contribute at most $159m\ln|X|$ to $Y$ with probability at least $1-\exp(-\Omega(\log k\log|X|/\epsilon^{2}))$ . The proof is complete. ∎

with some ${t}$ good for $({d},{a}),D$ , and where $\tau$ is the negligible function of Lemma 4.10. We complete the proof by showing that, for every neighboring database $D^{\prime}$ , possible output $({d},{a})$ , and thresholds $t$ good for $({d},{a})$ and $D$ ,

Fix a neighboring database $D^{\prime}$ , a target output $({d},{a})$ , and thresholds $t$ good for $({d},{a})$ and $D$ . The probability that the median mechanism chooses the target thresholds $t$ is independent of the underlying database, and so is the same on both sides of (9). For the rest of the proof, we condition on the event that the median mechanism uses the thresholds $t$ (both with database $D$ and database $D^{\prime}$ ).

Imagine running the median mechanism in parallel on $D,D^{\prime}$ and condition on the events $\mathcal{E}_{i-1},\mathcal{E}^{\prime}_{i-1}$ . The set $C_{i-1}$ is then the same in both runs of the mechanism, and $r_{i}(D),r_{i}(D^{\prime})$ are now fixed. Let $b_{i}$ ( $b^{\prime}_{i}$ ) be 0 if $MM(D,f)$ ( $MM(D^{\prime},f)$ ) classifies query $i$ as easy and 1 otherwise. Since $r_{i}(D^{\prime})\in[r_{i}(D)\pm\tfrac{2}{\epsilon n}]$ (Lemma 4.9) and a perturbation with distribution $\textrm{Lap}(\tfrac{2}{\alpha^{\prime}\epsilon n})$ is added to these values before comparing to the threshold $t_{i}$ (Step 2(a)),

and similarly for the events where $b_{i},b^{\prime}_{i}=1$ . Suppose that the target classification is $d_{i}=1$ (a hard query), and let $s_{i}$ and $s^{\prime}_{i}$ denote the random variables $f_{i}(D)+\textrm{Lap}(\tfrac{1}{\alpha^{\prime}n})$ and $f_{i}(D^{\prime})+\textrm{Lap}(\tfrac{1}{\alpha^{\prime}n})$ , respectively. Independence of the Laplace perturbations in Steps 2(a) and 2(d) implies that

Since the predicate query $f_{i}$ has sensitivity $1/n$ , we have

Now suppose that $d_{i}=0$ , and let $m_{i}$ denote the median value of $f_{i}$ on $C_{i-1}$ . Then $\Pr[\mathcal{E}_{i}|\mathcal{E}_{i-1}]$ is either 0 (if $m_{i}\neq a_{i}$ ) or $\Pr[b_{i}=0\,|\,\mathcal{E}_{i-1}]$ (if $m_{i}=a_{i}$ ); similarly, $\Pr[\mathcal{E}^{\prime}_{i}|\mathcal{E}^{\prime}_{i-1}]$ is either 0 or $\Pr[b^{\prime}_{i}=0\,|\,\mathcal{E}^{\prime}_{i-1}]$ . Thus the bound in (10) continues to hold (even with $e^{2\alpha^{\prime}}$ replaced by $e^{\alpha^{\prime}}$ ) when $d_{i}=0$ .

Since $\alpha^{\prime}$ is not much smaller than the privacy target $\alpha$ (recall (2)), we cannot afford to suffer the upper bound in (10) for many queries. Fortunately, for queries $i$ with good thresholds we can do much better. Consider a query $i$ such that $t_{i}$ is good for $({d},{a})$ and $D$ and condition again on $\mathcal{E}_{i-1},\mathcal{E}^{\prime}_{i-1}$ , which fixes $C_{i-1}$ and hence $r_{i}(D)$ . Goodness implies that $d_{i}=0$ , so the arguments from the previous paragraph also apply here. We can therefore assume that the median value $m_{i}$ of $f_{i}$ on $C_{i-1}$ equals $a_{i}$ and focus on bounding $\Pr[b_{i}=0\,|\,\mathcal{E}_{i-1}]$ in terms of $\Pr[b^{\prime}_{i}=0\,|\,\mathcal{E}^{\prime}_{i-1}]$ . Goodness also implies that $r_{i}(D)\geq t_{i}+\gamma$ and hence $r_{i}(D^{\prime})\geq t_{i}+\gamma-\tfrac{2}{\epsilon n}\geq t_{i}+\tfrac{\gamma}{2}$ (by Lemma 4.9). Recalling from (3) the definition of $\gamma$ , we have

and of course, $\Pr[b_{i}=0\,|\,\mathcal{E}_{i-1}]\leq 1$ .

Applying (10) to the bad queries — at most $180m\ln|X|$ of them, since $t$ is good for $({d},{a})$ and $D$ — and (11) to the rest, we can derive

which completes the proof of both the inequality (9) and the theorem. $\blacksquare$

The Median Mechanism: Efficient Implementation

The basic implementation of the median mechanism runs in time $|X|^{\Theta(\log k\log(1/\epsilon)/\epsilon^{2})}$ . This section provides an efficient implementation, running in time polynomial in $n$ , $k$ , and $|X|$ , although with a weaker usefulness guarantee.

Assume that the database size $n$ satisfies (6). For every sequence of adaptively chosen predicate queries $f_{1},\ldots,f_{k}$ arriving online, the efficient implementation of the median Mechanism is $(\alpha,\tau)$ -differentially private for a negligible function $\tau$ . Moreover, for every fixed set $f_{1},\ldots,f_{k}$ of queries, it is $(\epsilon,\delta)$ -useful for all but a negligible fraction of fractional databases (equivalently, probability distributions).

We note however that even for the small fraction of fractional histograms for which the efficient median mechanism may not satisfy our usefulness guarantee, it does not output incorrect answers: it merely halts after having answered a sufficiently large number of queries using the Laplace mechanism. Therefore, even for this small fraction of databases, the efficient median mechanism is an improvement over the Laplace mechanism: in the worst case, it simply answers every query using the Laplace mechanism before halting, and in the best case, it is able to answer many more queries.

We give a high-level overview of the proof of Theorem 5.1 which we then make formal. First, why isn’t the median mechanism a computationally efficient mechanism? Because $C_{0}$ has super-polynomial size $|X|^{m}$ , and computing $r_{i}$ in Step 2(a), the median value in Step 2(c), and the set $C_{i}$ in Step 2(e) could require time proportional to $|C_{0}|$ . An obvious idea is to randomly sample elements of $C_{i-1}$ to approximately compute $r_{i}$ and the median value of $f_{i}$ on $C_{i-1}$ ; while it is easy to control the resulting sampling error and preserve the utility and privacy guarantees of Section 4, it is not clear how to sample from $C_{i-1}$ efficiently.

We now give a formal analysis of the efficient implementation.

We redefine the sets $C_{i}$ to represent databases that can contain points fractionally, as opposed to the finite set of small discrete databases. Equivalently, we can view the sets $C_{i}$ as containing probability distributions over the set of points $X$ .

We generalize our query functions $f_{i}$ to fractional histograms in the natural way:

The update operation after a hard query $i$ is answered is the same as in the basic implementation:

Note that each updating operation after a hard query merely intersects $C_{i-1}$ with the pair of halfspaces:

and so $C_{i}$ is a convex polytope for each $i$ .

Since $C_{i}$ is given as the intersection of a set of explicit halfspaces, we have a simple membership oracle to determine whether a given point ${F}\in C_{i}$ : we simply check that ${F}$ lies on the appropriate side of each of the halfspaces. This takes time poly $(|X|,m)$ , since the number of halfspaces defining $C_{i}$ is linear in the number of answers to hard queries given before time $i$ , which is never more than $20m\ln|X|$ . Moreover, for each $i$ we have $C_{i}\subseteq C_{0}\subset mB_{1}^{|X|}\subset mB_{2}^{|X|}\subset|X|B_{2}^{|X|}.$ Finally, we can safely assume that $B_{2}^{X}\subseteq C_{i}$ by simply considering the convex set $C^{\prime}_{i}=C_{i}+B_{2}^{X}$ instead. This will not affect our results.

Therefore, we can implement the median mechanism in time poly $(|X|,k)$ by using sets $C_{i}$ as defined in this section, and sampling from them using the grid walk of [DFK91]. Estimation error in computing $r_{i}$ and the median value of $f_{i}$ on $C_{i-1}$ by random sampling rather than brute force is easily controlled via the Chernoff bound and can be incorporated into the proofs of Lemmas 4.3 and 4.5 in the obvious way. It remains to prove a continuous version of Lemma 4.7 to show that the efficient implementation of the median mechanism is $(\epsilon,\delta)$ -useful on all but a negligibly small fraction of fractional histograms $F$ .

2 Usefulness for Almost All Distributions

We now prove an analogue of Lemma 4.7 to establish a usefulness guarantee for the efficient version of the median mechanism.

With respect to any set of $k$ queries $f_{1},\ldots,f_{k}$ and for any ${F}^{*}\in C_{0}$ , define

as the set of points that agree up to an additive $\epsilon$ factor with ${F}^{*}$ on every query $f_{i}$ .

For any ${F}^{*}$ , $\textrm{Good}_{\epsilon}({F}^{*})$ is a convex polytope contained inside $C_{0}$ . We will prove that the efficient version of the median mechanism is $(\epsilon,\delta)$ -useful for a database $D$ if

We first prove that (12) holds for almost every fractional histogram. For this, we need a preliminary lemma.

Let $\mathcal{L}$ denote the set of integer points inside $C_{0}$ . Then with respect to an arbitrary set of $k$ queries,

Every rational valued point ${F}\in C_{0}$ corresponds to some (large) database $D\subset X$ by scaling ${F}$ to an integer-valued histogram. Irrational points can be arbitrarily approximated by such a finite database. By Proposition 4.6, for every set of $k$ predicates $f_{1},\ldots,f_{k}$ , there is a database $F^{*}\subset X$ with $|F^{*}|=m$ such that for each $i$ , $|f_{i}(F^{*})-f_{i}(F)|\leq\epsilon/400$ . Recalling that the histograms corresponding to databases of size at most $m$ are exactly the integer points in $C_{0}$ , the proof is complete. ∎

All but an $|X|^{-m}$ fraction of fractional histograms $F$ satisfy

Consider a randomly selected fractional histogram $F^{*}\in C_{0}$ . For any $F\in\mathcal{B}$ we have:

Since $|\mathcal{B}|\leq|\mathcal{L}|\leq|X|^{m}$ , by a union bound we can conclude that except with probability $\frac{1}{|X|^{m}}$ , $F^{*}\not\in\textrm{Good}_{\epsilon/400}(F)$ for any $F\in\mathcal{B}$ . However, by Lemma 5.2, $F^{*}\in\textrm{Good}_{\epsilon/400}({F^{\prime}})$ for some $F^{\prime}\in\mathcal{L}$ . Therefore, except with probability $1/|X|^{m}$ , $F^{\prime}\in\mathcal{L}\setminus\mathcal{B}$ . Thus, since $\textrm{Good}_{\epsilon/400}({F^{\prime}})\subseteq\textrm{Good}_{\epsilon/200}({F^{*}})$ , except with negligible probability, we have:

We are now ready to prove the analogue of Lemma 4.7 for the efficient implementation of the median mechanism.

For every set of $k$ queries $f_{1},\ldots,f_{k}$ , for all but an $O(|X|^{-m})$ fraction of fractional histograms $F$ , the efficient implementation of the median mechanism guarantees that: The mechanism answers fewer than $40m\log|X|$ hard queries, except with probability $k\exp(-\Omega(\epsilon n\alpha^{\prime}))$ ,

We assume that all answers to hard queries are $\epsilon/100$ accurate, and that $|r_{i}-\hat{r}_{i}|\leq\frac{1}{100}$ for every $i$ . By Lemmas 4.3 and 4.4 — the former adapted to accommodate approximating $r_{i}$ via random sampling — we are in this case except with probability $k\exp(-\Omega(\epsilon n\alpha^{\prime}))$ .

We analyze how the volume of $C_{i}$ contracts with the number of hard queries answered. Suppose the mechanism answers a hard query at time $i$ . Then:

Since all answers to hard queries are $\epsilon/100$ accurate, it must be that $\textrm{Good}_{\epsilon/100}(D)\in C_{k}$ . Therefore, for an input database $D$ that satisfies (12) — and this is all but an $O(|X|^{-m})$ fraction of them, by Lemma 5.3 — we have

Combining inequalities (13) and (14) yields

Lemmas 4.4, 4.5, and 5.4 give the following utility guarantee.

For every set $f_{1},\ldots,f_{k}$ of queries, for all but a negligible fraction of fractional histograms $F$ , the efficient implementation of the median mechanism is $(\epsilon,\delta)$ -useful with $\delta=k\exp(-\Omega(\epsilon n\alpha^{\prime}))$ .

3 Usefulness for Finite Databases

Fractional histograms correspond to probability distributions over $X$ . Lemma 5.3 shows that most probability distributions are ‘good’ for the efficient implementation of the Median Mechanism; in fact, more is true. We next show that finite databases sampled from randomly selected probability distributions also have good volume properties. Together, these lemmas show that the efficient implementation of the median mechanism will be able to answer nearly exponentially many queries with high probability, in the setting in which the private database $D$ is drawn from some ‘typical’ population distribution. DatabaseSample( $|D|$ ):

Select a fractional point $F\in C_{0}$ uniformly at random.

Sample and return a database $D$ of size $|D|$ by drawing each $x\in D$ independently at random from the probability distribution over $X$ induced by $F$ (i.e. sample $x_{i}\in X$ with probability proportional to $F_{i}$ ).

For $|D|$ as in (6) (as required for the Median Mechanism), a database sampled by DatabaseSample( $|D|$ ) satisfies (12) except with probability at most $O(|X|^{-m})$ .

By lemma 5.3, except with probability $|X|^{-m}$ , the fractional histogram $F$ selected in step $1$ satisfies

By lemma 4.6, when we sample a database $D$ of size $|D|\geq O((\log|X|\log^{3}k\log 1/\epsilon)/\epsilon^{3})$ from the probability distribution induced by $F$ , except with probability $\delta=O(k|X|^{-\log^{3}k/\epsilon})$ , $\textrm{Good}_{\epsilon/200}(F)\subset\textrm{Good}_{\epsilon/100}(D)$ , which gives us condition (12). ∎

We would like an analogue of lemma 5.3 that holds for all but a diminishing fraction of finite databases (which correspond to lattice points within $C_{0}$ ) rather than fractional points in $C_{0}$ , but it is not clear how uniformly randomly sampled lattice points distribute themselves with respect to the volume of $C_{0}$ . If $n>>|X|$ , then the lattice will be fine enough to approximate the volume of $C_{0}$ , and lemma 5.3 will continue to hold. We now show that small uniformly sampled databases will also be good for the efficient version of the median mechanism. Here, small means $n=o(\sqrt{|X|})$ , which allows for databases which are still polynomial in the size of $X$ . A tighter analysis is possible, but we opt instead to give a simple argument.

For every $n$ such that $n$ satisfies (6) and $n=o(\sqrt{|X|})$ , all but an $O(n^{2}/|X|)$ fraction of databases $D$ of size $|D|=n$ satisfy condition (12).

Finally, let $B$ be the event that database $D$ fails to satisfy (12). We have:

where the last equality follows from lemma 5.6, which states that $\Pr_{N}[B]$ is negligibly small. ∎

We observe that we can substitute either of the above lemmas for lemma 5.3 in the proof of lemma 5.4 to obtain versions of Thoerem 5.5:

For every set $f_{1},\ldots,f_{k}$ of queries, for all but a negligible fraction of databases sampled by DatabaseSample, the efficient implementation of the median mechanism is $(\epsilon,\delta)$ -useful with $\delta=k\exp(-\Omega(\epsilon n\alpha^{\prime}))$ .

For every set $f_{1},\ldots,f_{k}$ of queries, for all but an $n^{2}/|X|$ fraction of uniformly randomly sampled databases of size $n$ , the efficient implementation of the median mechanism is $(\epsilon,\delta)$ -useful with $\delta=k\exp(-\Omega(\epsilon n\alpha^{\prime}))$ .

Conclusion

We have shown that in the setting of predicate queries, interactivity does not pose an information theoretic barrier to differentially private data release. In particular, our dependence on the number of queries $k$ nearly matches the optimal dependence of $\log k$ achieved in the offline setting by [BLR08]. We remark that our dependence on other parameters is not necessarily optimal: in particular, [DNR+09] achieves a better (and optimal) dependence on $\epsilon$ . We have also shown how to implement our mechanism in time poly $(|X|,k)$ , although at the cost of sacrificing worst-case utility guarantees. The question of an interactive mechanism with poly $(|X|,k)$ runtime and worst-case utility guarantees remains an interesting open question. More generally, although the lower bounds of [DNR+09] seem to preclude mechanisms with run-time poly $(\log|X|)$ from answering a superlinear number of generic predicate queries, the question of achieving this runtime for specific query classes of interest (offline or online) remains largely open. Recently a representation-dependent impossibility result for the class of conjunctions was obtained by Ullman and Vadhan [UV10]: either extending this to a representation-independent impossibility result, or circumventing it by giving an efficient mechanism with a novel output representation would be very interesting.

Acknowledgments

The first author wishes to thank a number of people for useful discussions, including Avrim Blum, Moritz Hardt, Katrina Ligett, Frank McSherry, and Adam Smith. He would particularly like to thank Moritz Hardt for suggesting trying to prove usefulness guarantees for a continuous version of the BLR mechanism, and Avrim Blum for suggesting the distribution from which we select the threshold in the median mechanism.