Density Estimation in Infinite Dimensional Exponential Families

Bharath Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Aapo Hyvärinen, Revant Kumar

Introduction

In this paper, we consider an infinite dimensional generalization (Canu and Smola, 2005; Fukumizu, 2009) of (1),

where the function space $\mathcal{F}$ is defined as

being the cumulant generating function, and $(\mathcal{H},\langle\cdot,\cdot\rangle_{\mathcal{H}})$ a reproducing kernel Hilbert space (RKHS) (Aronszajn, 1950) with $k$ as its reproducing kernel. While various generalizations are possible for different choices of $\mathcal{F}$ (e.g., an Orlicz space as in Pistone and Sempi, 1995), the connection of $\mathcal{P}$ to the natural exponential family in (1) is particularly enlightening when $\mathcal{H}$ is an RKHS. This is due to the reproducing property of the kernel, $f(x)=\langle f,k(x,\cdot)\rangle_{\mathcal{H}}$ , through which $k(x,\cdot)$ takes the role of the sufficient statistic. In fact, it can be shown (see Section 3 and Example 1 for more details) that every $\mathscr{P}_{\text{fin}}$ is generated by $\mathcal{P}$ induced by a finite dimensional RKHS $\mathcal{H}$ , and therefore the family $\mathcal{P}$ with $\mathcal{H}$ being an infinite dimensional RKHS is a natural infinite dimensional generalization of $\mathscr{P}_{\text{fin}}$ . Furthermore, this generalization is particularly interesting as in contrast to $\mathscr{P}_{\text{fin}}$ , it can be shown that $\mathcal{P}$ is a rich class of densities (depending on the choice of $k$ and therefore $\mathcal{H}$ ) that can approximate a broad class of probability densities arbitrarily well (see Propositions 1, 13 and Corollary 2). This generalization is not only of theoretical interest, but also has implications for statistical and machine learning applications. For example, in Bayesian non-parametric density estimation, the densities in $\mathcal{P}$ are chosen as prior distributions on a collection of probability densities (e.g., see van der Vaart and van Zanten, 2008). $\mathcal{P}$ has also found applications in nonparametric hypothesis testing (Gretton et al., 2012; Fukumizu et al., 2008) and dimensionality reduction (Fukumizu et al., 2004, 2009) through the mean and covariance operators, which are obtained as the first and second Fréchet derivatives of $A(f)$ (see Fukumizu, 2009, Section 1.2.3). Recently, the infinite dimensional exponential family, $\mathcal{P}$ has been used to develop a gradient-free adaptive MCMC algorithm based on Hamiltonian Monte Carlo (Strathmann et al., 2015) and also has been used in the context of learning the structure of graphical models (Sun et al., 2015).

Motivated by the richness of the infinite dimensional generalization and its statistical applications, it is of interest to model densities by $\mathcal{P}$ , and therefore the goal of this paper is to estimate unknown densities by elements in $\mathcal{P}$ when $\mathcal{H}$ is an infinite dimensional RKHS. Formally, given i.i.d. random samples $(X_{a})^{n}_{a=1}$ drawn from an unknown density $p_{0}$ , the goal is to estimate $p_{0}$ through $\mathcal{P}$ . Throughout the paper, we refer to case of $p_{0}\in\mathcal{P}$ as well-specified, in contrast to the misspecified case where $p_{0}\notin\mathcal{P}$ . The setting is useful because $\mathcal{P}$ is a rich class of densities that can approximate a broad class of probability densities arbitrarily well, hence it may be widely used in place of non-parametric density estimation methods (e.g., kernel density estimation (KDE)). In fact, through numerical simulations, we show in Section 6 that estimating $p_{0}$ through $\mathcal{P}$ performs better than KDE, and that the advantage of the proposed estimator grows with increasing dimensionality.

A related work was carried out by Barron and Sheu (1991)—also see references therein—where the goal is to estimate a density, $p_{0}$ by approximating its logarithm as an expansion in terms of basis functions, such as polynomials, splines or trigonometric series. Similar to Fukumizu (2009), Barron and Sheu proposed the ML estimator $p_{\hat{f}_{m}}$ , where

and $\mathcal{F}_{m}$ is the linear space of dimension $m$ spanned by the chosen basis functions. Under the assumption that $\log p_{0}$ has square-integrable derivatives up to order $r$ , they showed that $KL(p_{0}\|p_{\hat{f}_{m}})=O_{p_{0}}(n^{-2r/(2r+1)})$ with $m=n^{1/(2r+1)}$ for each of the approximating families, where $KL(p\|q)=\int p(x)\log(p(x)/q(x))\,dx$ is the KL divergence between $p$ and $q$ . Similar work was carried out by Gu and Qiu (1993), who assumed that $\log p_{0}$ lies in an RKHS, and proposed an estimator based on penalized MLE, with consistency and rates established in Jensen-Shannon divergence. Though these results are theoretically interesting, these estimators are obtained via a procedure similar to that in Fukumizu (2009), and therefore suffers from the practical drawbacks discussed above.

The discussion so far shows that the MLE approach to learning $p_{0}\in\mathcal{P}$ results in estimators that are of limited practical interest. To alleviate this, one can treat the problem of estimating $p_{0}\in\mathcal{P}$ in a completely non-parametric fashion by using KDE, which is well-studied (Tsybakov, 2009, Chapter 1) and easy to implement. This approach ignores the structure of $\mathcal{P}$ , however, and is known to perform poorly for moderate to large $d$ (Wasserman, 2006, Section 6.5) (see also Section 6 of this paper).

To understand the advantages associated with the score matching method, let us consider the problem of density estimation where the data generating distribution (say $p_{0}$ ) belongs to $\mathscr{P}_{\text{fin}}$ in (1). In other words, given random samples $(X_{a})^{n}_{a=1}$ drawn i.i.d. from $p_{0}:=p_{\theta_{0}}$ , the goal is to estimate $\theta_{0}$ as $\hat{\theta}_{n}$ , and use $p_{\hat{\theta}_{n}}$ as an estimator of $p_{0}$ . While the MLE approach is well-studied and enjoys nice statistical properties in asymptopia (i.e., asymptotically unbiased, efficient, and normally distributed), the computation of $\hat{\theta}_{n}$ can be intractable in many situations as discussed above. In particular, this is the case for $p_{\theta}(x)=\frac{r_{\theta}(x)}{A(\theta)}$ where $r_{\theta}\geq 0$ for all $\theta\in\Theta$ , $A(\theta)=\int_{\Omega}r_{\theta}(x)\,dx$ , and the functional form of $r$ is known (as a function of $\theta$ and $x$ ); yet we do not know how to easily compute $A$ , which is often analytically intractable. In this setting (which is exactly the setting of this paper), assuming $p_{\theta}$ to be differentiable (w.r.t. $x$ ), and $\int_{\Omega}p_{0}(x)\|\nabla\log p_{\theta}(x)\|^{2}_{2}\,dx<\infty,\,\forall\,\theta\in\Theta$ , $J(p_{0}\|p_{\theta})=:J(\theta)$ in (3) reduces to

through integration by parts (see Hyvärinen, 2005, Theorem 1), under appropriate regularity conditions on $p_{0}$ and $p_{\theta}$ for all $\theta\in\Theta$ . Here $\partial^{2}_{i}\log p_{\theta}(x):=\frac{\partial^{2}}{\partial x^{2}_{i}}\log p_{\theta}(x)$ . The main advantage of the objective in (3) (and also (4)) is that when it is applied to the situation discussed above where $p_{\theta}(x)=\frac{r_{\theta}(x)}{A(\theta)}$ , $J(\theta)$ is independent of $A(\theta)$ , and an estimate of $\theta_{0}$ can be obtained by simply minimizing the empirical counterpart of $J(\theta)$ , given by

Since $J_{n}(\theta)$ is also independent of $A(\theta)$ , $\hat{\theta}_{n}=\arg\min_{\theta\in\Theta}J_{n}(\theta)$ may be easily computable, unlike the MLE. We would like to highlight that while the score matching approach may have computational advantages over MLE, it only estimates $p_{\theta}$ up to the scaling factor $A(\theta)$ , and therefore requires the approximation or computation of $A(\theta)$ through numerical integration to estimate $p_{\theta}$ . Note that this issue (of computing $A(\theta)$ through numerical integration) exists even with MLE, but not with KDE. In score matching, however, numerical integration is needed only once, while MLE would typically require a functional form of the log-partition function which is approximated through numerical integration at every step of an iterative optimization algorithm (for example, see (2)), thus leading to major computational savings. An important application that does not require the computation of $A(\theta)$ is in finding modes of the distribution, which has recently become very popular in image processing (Comaniciu and Meer, 2002), and has already been investigated in the score matching framework (Sasaki et al., 2014). Similarly, in sampling methods such as sequential Monte Carlo (Doucet et al., 2001), it is often the case that the evaluation of unnormalized densities is sufficient to calculate required importance weights.

2 Contributions

(i) We present an estimate of $p_{0}\in\mathcal{P}$ in the well-specified case through the minimization of Fisher divergence, in Section 4. First, we show that estimating $p_{0}:=p_{f_{0}}$ using the score matching method reduces to estimating $f_{0}$ by solving a simple finite-dimensional linear system (Theorems 4 and 5). Hyvärinen (2007) obtained a similar result for $\mathscr{P}_{\text{fin}}$ where the estimator is obtained by solving a linear system, which in the case of Gaussian family matches the MLE (Hyvärinen, 2005). The estimator obtained in the infinite dimensional case is not a simple extension of its finite-dimensional counterpart, however, as the former requires an appropriate regularizer (we use $\|\cdot\|^{2}_{\mathcal{H}}$ ) to make the problem well-posed. We would like to highlight that to the best of our knowledge, the proposed estimator is the first practically computable estimator of $p_{0}$ with consistency guarantees (see below). (ii) In contrast to Hyvärinen (2007) where no guarantees on consistency or convergence rates are provided for the density estimator in $\mathscr{P}_{\text{fin}}$ , we establish in Theorem 6 the consistency and rates of convergence for the proposed estimator of $f_{0}$ , and use these to prove consistency and rates of convergence for the corresponding plug-in estimator of $p_{0}$ (Theorems 7 and B.2), even when $\mathcal{H}$ is infinite dimensional. Furthermore, while the estimator of $f_{0}$ (and therefore $p_{0}$ ) is obtained by minimizing the Fisher divergence, the resultant density estimator is also shown to be consistent in KL divergence (and therefore in Hellinger and total-variation distances) and we provide convergence rates in all these distances.

Formally, we show that the proposed estimator $\hat{f}_{n}$ is converges as

if $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ , where $\mathcal{R}(A)$ denotes the range or image of an operator $A$ , $\alpha=\min\{\frac{1}{4},\frac{\beta}{2\beta+2}\}$ , and $C:=\sum^{d}_{i=1}\int_{\Omega}\partial_{i}k(x,\cdot)\otimes\partial_{i}k(x,\cdot)\,p_{0}(x)\,dx$ is a Hilbert-Schmidt operator on $\mathcal{H}$ (see Theorem 4) with $k$ being the reproducing kernel and $\otimes$ denoting the tensor product. When $\mathcal{H}$ is a finite-dimensional RKHS, we show that the estimator enjoys parametric rates of convergence, i.e.,

Note that the convergence rates are obtained under a non-classical smoothness assumption on $f_{0}$ , namely that it lies in the image of certain fractional power of $C$ , which reduces to a more classical assumption if we choose $k$ to be a Matérn kernel (see Section 2 for its definition), as it induces a Sobolev space. In Section 4.2, we discuss in detail the smoothness assumption on $f_{0}$ for the Gaussian (Example 2) and Matérn (Example 3) kernels. Another interesting point to observe is that unlike in the classical function estimation methods (e.g., kernel density estimation and regression), the rates presented above for the proposed estimator tend to saturate for $\beta>1$ ( $\beta>\frac{1}{2}$ w.r.t. $J$ ), with the best rate attained at $\beta=1$ ( $\beta=\frac{1}{2}$ w.r.t. $J$ ), which means the smoothness of $f_{0}$ is not fully captured by the estimator. Such a saturation behavior is well-studied in the inverse problem literature (Engl et al., 1996) where it has been attributed to the choice of regularizer. In Section 4.3, we discuss alternative regularization strategies using ideas from Bauer et al. (2007), which covers non-parametric least squares regression: we show that for appropriately chosen regularizers, the above mentioned rates hold for any $\beta>0$ , and do not saturate for the aforementioned ranges of $\beta$ (see Theorem 9). (iii) In Section 5, we study the problem of density estimation in the misspecified setting, i.e., $p_{0}\notin\mathcal{P}$ , which is not addressed in Hyvärinen (2007) and Fukumizu (2009). Using a more sophisticated analysis than in the well-specified case, we show in Theorem 12 that $J(p_{0}\|p_{\hat{f}_{n}})\rightarrow\inf_{p\in\mathcal{P}}J(p_{0}\|p)$ as $n\rightarrow\infty$ . Under an appropriate smoothness assumption on $\log\frac{p_{0}}{q_{0}}$ (see the statement of Theorem 12 for details), we show that $J(p_{0}\|p_{\hat{f}_{n}})\rightarrow 0$ as $n\rightarrow\infty$ along with a rate for this convergence, even though $p_{0}\notin\mathcal{P}$ . However, unlike in the well-specified case, where the consistency is obtained not only in $J$ but also in other distances, we obtain convergence only in $J$ for the misspecified case. Note that while Barron and Sheu (1991) considered the estimation of $p_{0}$ in the misspecified setting, the results are restricted to the approximating families consisting of polynomials, splines, or trigonometric series. Our results are more general, as they hold for abstract RKHSs. (iv) In Section 6, we present preliminary numerical results comparing the proposed estimator with KDE in estimating a Gaussian and mixture of Gaussians, with the goal of empirically evaluating performance as $d$ gets large for a fixed sample size. In these two estimation problems, we show that the proposed estimator outperforms KDE, and the advantage grows as $d$ increases. Inspired by this preliminary empirical investigation, our proposed estimator (or computationally efficient approximations) has been used by Strathmann et al. (2015) in a gradient-free adaptive MCMC sampler, and by Sun et al. (2015) for graphical model structure learning. These applications demonstrate the practicality and performance of the proposed estimator.

Finally, we would like to make clear that our principal goal is not to construct density estimators that improve uniformly upon KDE, but to provide a novel flexible modeling technique for approximating an unknown density by a rich parametric family of densities, with the parameter being infinite dimensional, in contrast to the classical approach of finite dimensional approximation.

Various notations and definitions that are used throughout the paper are collected in Section 2. The proofs of the results are provided in Section 8, along with some supplementary results in an appendix.

Definitions & Notation

where $i$ denotes the imaginary unit $\sqrt{-1}$ .

where $\Gamma$ is the Gamma function, and $\mathfrak{K}_{v}$ is the modified Bessel function of the third kind of order $v$ ( $v$ controls the smoothness of $k$ ).

Approximation of Densities by 𝒫𝒫\mathcal{P}

In this section, we first show that every finite dimensional exponential family, $\mathscr{P}_{\text{fin}}$ is generated by the family $\mathcal{P}$ induced by a finite dimensional RKHS, which naturally leads to the infinite dimensional generalization of $\mathscr{P}_{\text{fin}}$ when $\mathcal{H}$ is an infinite dimensional RKHS. Next, we investigate the approximation properties of $\mathcal{P}$ in Proposition 1 and Corollary 2 when $\mathcal{H}$ is an infinite dimensional RKHS.

The following are some popular examples of probability distributions that belong to $\mathscr{P}_{\emph{fin}}$ . Here we show the corresponding RKHSs $(\mathcal{H},k)$ that generate these distributions. In some of these examples, we choose $q_{0}(x)=1$ and ignore the fact that $q_{0}$ is a probability distribution as assumed in the definition of $\mathcal{P}$ .

Beta: $\Omega=(0,1)$ , $k(x,y)=\log x\log y+\log(1-x)\log(1-y)$ .

Binomial: $\Omega=\{0,\ldots,m\}$ , $k(x,y)=xy$ , $q_{0}(x)=2^{-m}{m\choose c}$ .

While Example 1 shows that all popular probability distributions are contained in $\mathcal{P}$ for an appropriate choice of finite-dimensional $\mathcal{H}$ , it is of interest to understand the richness of $\mathcal{P}$ (i.e., what class of distributions can be approximated arbitrarily well by $\mathcal{P}$ ?) when $\mathcal{H}$ is an infinite dimensional RKHS. This is addressed by the following result, which is proved in Section 8.1.

Then $\mathcal{P}$ is dense in $\mathcal{P}_{0}$ w.r.t. Kullback-Leibler divergence, total variation ( $L^{1}$ norm) and Hellinger distances. In addition, if $q_{0}\in L^{1}(\Omega)\cap L^{r}(\Omega)$ for some $1<r\leq\infty$ , then $\mathcal{P}$ is also dense in $\mathcal{P}_{0}$ w.r.t. $L^{r}$ norm.

Suppose $k(x,\cdot)\in C_{0}(\Omega),\,\forall\,x\in\Omega$ and (5) holds. Then $\mathcal{P}$ is dense in $\mathcal{P}_{c}$ w.r.t. KL divergence, TV and Hellinger distances. Moreover, if $q_{0}\in L^{1}(\Omega)\cap L^{r}(\Omega)$ for some $1<r\leq\infty$ , then $\mathcal{P}$ is also dense in $\mathcal{P}_{c}$ w.r.t. $L^{r}$ norm.

By choosing $\Omega$ to be compact and $q_{0}$ to be a uniform distribution on $\Omega$ , Corollary 2 reduces to an easily interpretable result that any continuous density $p_{0}$ on $\Omega$ can be approximated arbitrarily well by densities in $\mathcal{P}$ in KL, Hellinger and $L^{r}$ ( $1\leq r\leq\infty$ ) distances.

Similar to the results so far, an approximation result for $\mathcal{P}$ can also be obtained w.r.t. Fisher divergence (see Proposition 13). Since this result is heavily based on the notions and results developed in Section 5, we defer its presentation until that section. Briefly, this result states that if $\mathcal{H}$ is sufficiently rich (i.e., dense in an appropriate class of functions), then any $p\in C^{1}(\Omega)$ with $J(p\|q_{0})<\infty$ can be approximated arbitrarily well by elements in $\mathcal{P}$ w.r.t. Fisher divergence, where $q_{0}\in C^{1}(\Omega)$ .

Density Estimation in 𝒫𝒫\mathcal{P}: Well-specified Case

In this section, we present our score matching estimator for an unknown density $p_{0}:=p_{f_{0}}\in\mathcal{P}$ (well-specified case) from i.i.d. random samples $(X_{a})^{n}_{a=1}$ drawn from it. This involves choosing the minimizer of the (empirical) Fisher divergence between $p_{0}$ and $p_{f}\in\mathcal{P}$ as the estimator, $\hat{f}$ which we show in Theorem 5 to be obtained by solving a simple finite-dimensional linear system. In contrast, we would like to remind the reader that the MLE is infeasible in practice due to the difficulty in handling $A(f)$ . The consistency and convergence rates of $\hat{f}\in\mathcal{F}$ and the plug-in estimator $p_{\hat{f}}$ are provided in Section 4.1 (see Theorems 6 and 7). Before we proceed, we list the assumptions on $p_{0}$ , $q_{0}$ and $\mathcal{H}$ that we need in our analysis.

$p_{0}$ is continuously extendible to $\overline{\Omega}$ . $k$ is twice continuously differentiable on $\Omega\times\Omega$ with continuous extension of $\partial^{\alpha,\alpha}k$ to $\overline{\Omega}\times\overline{\Omega}$ for $|\alpha|\leq 2$ .

$\partial_{i}\partial_{i+d}k(x,x)p_{0}(x)=0$ for $x\in\partial\Omega$ and $\sqrt{\partial_{i}\partial_{i+d}k(x,x)}p_{0}(x)=o(\|x\|^{1-d}_{2})$ as $x\in\Omega$ , $\|x\|_{2}\rightarrow\infty$ for all $i\in[d]$ .

( $\varepsilon$ -Integrability) For some $\varepsilon\geq 1$ and $\forall\,i\in[d]$ , $\partial_{i}\partial_{i+d}k(x,x),\sqrt{\partial^{2}_{i}\partial^{2}_{i+d}k(x,x)}$ and $\sqrt{\partial_{i}\partial_{i+d}k(x,x)}\partial_{i}\log q_{0}(x)\in L^{\varepsilon}(\Omega,p_{0}),$ where $q_{0}\in C^{1}(\Omega).$

Under these assumptions, the following result—proved in Section 8.3—shows that the problem of estimating $p_{0}$ through the minimization of Fisher divergence reduces to the problem of estimating $f_{0}$ through a weighted least squares minimization in $\mathcal{H}$ (see parts (i) and (ii)). This motivates the minimization of the regularized empirical weighted least squares (see part (iv)) to obtain an estimator $f_{\lambda,n}$ of $f_{0}$ , which is then used to construct the plug-in estimate $p_{f_{\lambda,n}}$ of $p_{0}$ .

Suppose (A)–(D) hold with $\varepsilon=1$ . Then $J(p_{0}\|p_{f})<\infty$ for all $f\in\mathcal{F}$ . In addition, the following hold. (i) For all $f\in\mathcal{F}$ ,

where $C:\mathcal{H}\rightarrow\mathcal{H}$ , $C:=\int_{\Omega}p_{0}(x)\sum^{d}_{i=1}\partial_{i}k(x,\cdot)\otimes\partial_{i}k(x,\cdot)\,dx$ is a trace-class positive operator with

and $f_{0}$ satisfies $Cf_{0}=-\xi$ . (iii) For any $\lambda>0$ , a unique minimizer $f_{\lambda}$ of $J_{\lambda}(f):=J(f)+\frac{\lambda}{2}\|f\|^{2}_{\mathcal{H}}$ over $\mathcal{H}$ exists and is given by

(iv) (Estimator of $f_{0}$ ) Given samples $(X_{a})^{n}_{a=1}$ drawn i.i.d. from $p_{0}$ , for any $\lambda>0$ , the unique minimizer $f_{\lambda,n}$ of $\hat{J}_{\lambda}(f):=\hat{J}(f)+\frac{\lambda}{2}\|f\|^{2}_{\mathcal{H}}$ over $\mathcal{H}$ exists and is given by

where $\hat{J}(f):=\frac{1}{2}\langle f,\hat{C}f\rangle_{\mathcal{H}}+\langle f,\hat{\xi}\rangle_{\mathcal{H}}+J(p_{0}\|q_{0})$ , $\hat{C}:=\frac{1}{n}\sum^{n}_{a=1}\sum^{d}_{i=1}\partial_{i}k(X_{a},\cdot)\otimes\partial_{i}k(X_{a},\cdot)$ and

An advantage of the alternate formulation of $J(f)$ in Theorem 4(ii) over (6) is that it provides a simple way to obtain an empirical estimate of $J(f)$ —by replacing $C$ and $\xi$ by their empirical estimators, $\hat{C}$ and $\hat{\xi}$ respectively—from finite samples drawn i.i.d. from $p_{0}$ , which is then used to obtain an estimator of $f_{0}$ . Note that the empirical estimate of $J(f)$ , i.e., $\hat{J}(f)$ depends only on $\hat{C}$ and $\hat{\xi}$ which in turn depend on the known quantities, $k$ and $q_{0}$ , and therefore $f_{\lambda,n}$ in Theorem 4(iv) should in principle be computable. In practice, however, it is not easy to compute the expression for $f_{\lambda,n}=-(\hat{C}+\lambda I)^{-1}\hat{\xi}$ as it involves solving an infinite dimensional linear system. In Theorem 5 (proved in Section 8.4), we provide an alternative expression for $f_{\lambda,n}$ as a solution of a simple finite-dimensional linear system (see (7) and (8)), using the general representer theorem (see Theorem A.2). It is interesting to note that while the solution to $J(f)$ in Theorem 4(ii) is obtained by solving a non-linear system, $Cf_{0}=-\xi$ (the system is non-linear as $C$ depends on $p_{0}$ which in turn depends on $f_{0}$ ), its estimator $f_{\lambda,n}$ proposed in Theorem 4, is obtained by solving a simple linear system. In addition, we would like to highlight the fact that the proposed estimator, $f_{\lambda,n}$ is precisely the Tikhonov regularized solution (which is well-studied in the theory of linear inverse problems) to the ill-posed linear system $\hat{C}f=-\hat{\xi}$ . We further discuss the choice of regularizer in Section 4.3 using ideas from the inverse problem literature.

An important remark we would like to make about Theorem 4 is that though $J(f)$ in (6) is valid only for $f\in\mathcal{F}$ , as it is obtained from $J(p_{0}\|p_{f})$ where $p_{0},p_{f}\in\mathcal{P}$ , the expression $\langle f-f_{0},C(f-f_{0})\rangle_{\mathcal{H}}$ is valid for any $f\in\mathcal{H}$ , as it is finite under the assumption that (D) holds with $\varepsilon=1$ . Therefore, in Theorem 4(iii, iv), $f_{\lambda}$ and $f_{\lambda,n}$ are obtained by minimizing $J_{\lambda}$ and $\hat{J}_{\lambda}$ over $\mathcal{H}$ instead of over $\mathcal{F}$ , as the latter does not yield a nice expression (unlike $f_{\lambda}$ and $f_{\lambda,n}$ , respectively). However, there is no guarantee that $f_{\lambda,n}\in\mathcal{F}$ , and so the density estimator $p_{f_{\lambda,n}}$ may not be valid. While this is not an issue when studying the convergence of $\|f_{\lambda,n}-f_{0}\|_{\mathcal{H}}$ (see Theorem 6), the convergence of $p_{f_{\lambda,n}}$ to $p_{0}$ (in various distances) needs to be handled slightly differently depending on whether the kernel is bounded or not (see Theorems 7 and B.2). Note that when the kernel is bounded, we obtain $\mathcal{F}=\mathcal{H}$ , which implies $p_{f_{\lambda,n}}$ is valid.

Let $f_{\lambda,n}=\arg\inf_{f\in\mathcal{H}}\hat{J}_{\lambda}(f)$ , where $\hat{J}_{\lambda}(f)$ is defined in Theorem 4(iv) and $\lambda>0$ . Then

where $\hat{\xi}$ is defined in Theorem 4(iv) and $\bm{\beta}=(\beta_{(a-1)d+i})_{a,i}$ is obtained by solving

with $(\bm{G})_{(a-1)d+i,(b-1)d+j}=\partial_{i}\partial_{j+d}k(X_{a},X_{b})\,\,\,\text{and}$

We would like to highlight that though $f_{\lambda,n}$ requires solving a simple linear system in (8), it can still be computationally intensive when $d$ and $n$ are large as $\bm{G}$ is a $nd\times nd$ matrix. This is still a better scenario than that of MLE, however, since computationally efficient methods exist to solve large linear systems such as (8), whereas MLE can be intractable due to the difficulty in handling the log-partition function (though it can be approximated). On the other hand, MLE is statistically well-understood, with consistency and convergence rates established in general for the problem of density estimation (van de Geer, 2000) and in particular for the problem at hand (Fukumizu, 2009). In order to ensure that $f_{\lambda,n}$ and $p_{f_{\lambda,n}}$ are statistically useful, in the following section, we investigate their consistency and convergence rates under some smoothness conditions on $f_{0}$ .

Suppose (A)–(D) with $\varepsilon=2$ hold. (i) If $f_{0}\in\overline{\mathcal{R}(C)}$ , then $\left\|f_{\lambda,n}-f_{0}\right\|_{\mathcal{H}}\stackrel{{\scriptstyle p_{0}}}{{\rightarrow}}0\,\,\text{as}\,\,\lambda\to 0,\,\lambda\sqrt{n}\to\infty\,\,\text{and}\,\,n\to\infty.$ (ii) If $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ , then for $\lambda=n^{-\max\left\{\frac{1}{4},\frac{1}{2(\beta+1)}\right\}}$ ,

(iii) If $\|C^{-1}\|<\infty$ , then for $\lambda=n^{-\frac{1}{2}}$ , $\|f_{\lambda,n}-f_{0}\|_{\mathcal{H}}=O_{p_{0}}(n^{-1/2})$ as $n\rightarrow\infty$ .

(i) While Theorem 6 (proved in Section 8.5) provides an asymptotic behavior for $\|f_{\lambda,n}-f_{0}\|_{\mathcal{H}}$ under conditions that depend on $p_{0}$ (and are therefore not easy to check in practice), a non-asymptotic bound on $\|f_{\lambda,n}-f_{0}\|_{\mathcal{H}}$ that holds for all $n\geq 1$ can be obtained under stronger assumptions through an application of Bernstein’s inequality in separable Hilbert spaces. For the sake of simplicity, we provided asymptotic results which are obtained through an application of Chebyshev’s inequality. (ii) The proof of Theorem 6(i) involves decomposing $\|f_{\lambda,n}-f_{0}\|_{\mathcal{H}}$ into an estimation error part, $\mathcal{E}(\lambda,n):=\|f_{\lambda,n}-f_{\lambda}\|_{\mathcal{H}}$ , and an approximation error part, $\mathcal{A}_{0}(\lambda):=\|f_{\lambda}-f_{0}\|_{\mathcal{H}}$ , where $f_{\lambda}=(C+\lambda I)^{-1}Cf_{0}$ . While $\mathcal{E}(\lambda,n)\rightarrow 0$ as $\lambda\rightarrow 0$ , $\lambda\sqrt{n}\rightarrow\infty$ and $n\rightarrow\infty$ without any assumptions on $f_{0}$ (see the proof in Section 8.5 for details), it is not reasonable to expect $\mathcal{A}_{0}(\lambda)\rightarrow 0$ as $\lambda\rightarrow 0$ without assuming $f_{0}\in\overline{\mathcal{R}(C)}$ . This is because, if $f_{0}$ lies in the null space of $C$ , then $f_{\lambda}$ is zero irrespective of $\lambda$ and therefore cannot approximate $f_{0}$ . (iii) The condition $f_{0}\in\overline{\mathcal{R}(C)}$ is difficult to check in practice as it depends on $p_{0}$ (which in turn depends on $f_{0}$ ). However, since the null space of $C$ is just constant functions if the kernel is bounded and $\emph{supp}(q_{0})=\Omega$ (see Lemma 14 in Section 8.6 for details), assuming $1\notin\mathcal{H}$ yields that $\overline{\mathcal{R}(C)}=\mathcal{H}$ and therefore consistency can be attained under conditions that are easy to impose in practice. As mentioned in Remark 3(iii), the condition $1\notin\mathcal{H}$ ensures identifiability and a sufficient condition for it to hold is $k\in C_{0}(\Omega\times\Omega)$ , which is satisfied by Gaussian, Matérn and inverse multiquadric kernels. (iv) It is well known that convergence rates are possible only if the quantity of interest (here $f_{0}$ ) satisfies some additional conditions. In function estimation, this additional condition is classically imposed by assuming $f_{0}$ to be sufficiently smooth, e.g., $f_{0}$ lies in a Sobolev space of certain smoothness. By contrast, the smoothness condition in Theorem 6(ii) is imposed in an indirect manner by assuming $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ —so that the results hold for abstract RKHSs and not just Sobolev spaces—which then provides a rate, with the best rate being $n^{-1/4}$ that is attained when $\beta\geq 1$ . While such a condition has already been used in various works (Caponnetto and Vito, 2007; Smale and Zhou, 2007; Fukumizu et al., 2013) in the context of non-parametric least squares regression, we explore it in more detail in Proposition 8, and Examples 2 and 3. Note that this condition is common in the inverse problem theory (see Engl, Hanke, and Neubauer, 1996), and it naturally arises here through the connection of $f_{\lambda,n}$ being a Tikhonov regularized solution to the ill-posed linear system $\hat{C}f=-\hat{\xi}$ . An interesting observation about the rate is that it does not improve with increasing $\beta$ (for $\beta>1$ ), in contrast to the classical results in function estimation (e.g., kernel density estimation and kernel regression) where the rate improves with increasing smoothness. This issue is discussed in detail in Section 4.3. (v) Since $\|C^{-1}\|<\infty$ only if $\mathcal{H}$ is finite-dimensional, we recover the parametric rate of $n^{-1/2}$ in a finite-dimensional situation with an automatic choice for $\lambda$ as $n^{-1/2}$ .

While Theorem 6 provides statistical guarantees for parameter convergence, the question of primary interest is the convergence of $p_{f_{\lambda,n}}$ to $p_{0}$ . This is guaranteed by the following result, which is proved in Section 8.6.

Suppose (A)–(D) with $\varepsilon=2$ hold and $\|k\|_{\infty}:=\sup_{x\in\Omega}k(x,x)<\infty$ . Assume $\emph{supp}(q_{0})=\Omega$ . Then the following hold: (i) For any $1<r\leq\infty$ with $q_{0}\in L^{1}(\Omega)\cap L^{r}(\Omega)$ ,

In addition, if $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ , then for $\lambda=n^{-\max\left\{\frac{1}{4},\frac{1}{2(\beta+1)}\right\}}$ ,

as $n\rightarrow\infty$ where $\theta_{n}:=n^{-\min\left\{\frac{1}{4},\frac{\beta}{2(\beta+1)}\right\}}$ . (ii) $J(p_{0}\|p_{f_{\lambda,n}})\rightarrow 0\,\,\text{as}\,\,\lambda n\rightarrow\infty,\,\lambda\rightarrow 0\,\,\text{and}\,\,n\rightarrow\infty.$ In addition, if $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta\geq 0$ , then for $\lambda=n^{-\max\left\{\frac{1}{3},\frac{1}{2(\beta+1)}\right\}}$ ,

(iii) If $\|C^{-1}\|<\infty$ , then $\theta_{n}=n^{-\frac{1}{2}}$ and $J(p_{0}\|p_{f_{\lambda,n}})=O_{p_{0}}(n^{-1})$ with $\lambda=n^{-\frac{1}{2}}$ .

While Theorem 7 addresses the case of bounded kernels, the case of unbounded kernels requires a technical modification. The reason for this modification, as alluded to in the discussion following Theorem 4, is due to the fact that $f_{\lambda,n}$ may not be in $\mathcal{F}$ when $k$ is unbounded, and therefore the corresponding density estimator, $p_{f_{\lambda,n}}$ may not be well-defined. In order to keep the main ideas intact, we discuss the unbounded case in detail in Section B.2 in Appendix B.

2 Range Space Assumption

While Theorems 6 and 7 are satisfactory from the point of view of consistency, we believe the presented rates are possibly not minimax optimal since these rates are valid for any RKHS that satisfies the conditions (A)–(D) and does not capture the smoothness of $k$ (and therefore the corresponding $\mathcal{H}$ ). In other words, the rates presented in Theorems 6 and 7 should depend on the decay rate of the eigenvalues of $C$ which in turn effectively captures the smoothness of $\mathcal{H}$ . However, we are not able to obtain such a result—see the remark following the proof of Theorem 6 for a discussion. While these rates do not reflect the intrinsic smoothness of $\mathcal{H}$ , they are obtained under the smoothness assumption, i.e., range space condition that $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ . This condition is quite different from the classical smoothness conditions that appear in non-parametric function estimation. While the range space assumption has been made in various earlier works (e.g., Caponnetto and Vito (2007); Smale and Zhou (2007); Fukumizu et al. (2013) in the context of non-parametric least square regression), in the following, we investigate the implicit smoothness assumptions that it makes on $f_{0}$ in our context. To this end, first it is easy to show (see the proof of Proposition B.3 in Section B.3) that

Then $f_{0}\in\mathcal{R}(C)$ implies $f_{0}\in\mathcal{G}\subset\mathcal{H}$ .

In the following, we apply the above result in two examples involving Gaussian and Matérn kernels to get insights into the range space assumption.

Let $\psi(x)=e^{-\sigma\|x\|^{2}}$ with $\mathcal{H}_{\sigma}$ as its corresponding RKHS (see Section 2 for its definition). By Proposition 8, it is easy to verify that $f_{0}\in\mathcal{R}(C)$ implies $f_{0}\in\mathcal{H}_{\alpha}\subset\mathcal{H}_{\sigma}$ for $\frac{\sigma}{2}<\alpha\leq\sigma$ . Since $\mathcal{H}_{\beta}\subset\mathcal{H}_{\gamma}$ for $\beta<\gamma$ (i.e., Gaussian RKHSs are nested), $f_{0}\in\mathcal{R}(C)$ ensures that $f_{0}$ lies in $\mathcal{H}_{\frac{\sigma}{2}+\epsilon}$ for arbitrary small $\epsilon>0$ .

While Example 3 provides some understanding about the minimax optimality of $f_{\lambda,n}$ under additional assumptions on $f_{0}$ , the problem is not completely resolved. In the following section, however, we show that the rate in Theorem 6 is not optimal for $\beta>1$ , and that improved rates can be obtained by choosing the regularizer appropriately.

3 Choice of Regularizer

We understand from the characterization of $\mathcal{R}(C^{\beta})$ in (9) that larger $\beta$ values yield smoother functions in $\mathcal{H}$ . However, the smoothness of $f_{0}\in\mathcal{R}(C^{\beta})$ for $\beta>1$ is not captured in the rates in Theorem 6(ii), where the rate saturates at $\beta=1$ providing the best possible rate of $n^{-1/4}$ (irrespective of the size of $\beta$ ). This is unsatisfactory on the part of the estimator, as it does not effectively capture the smoothness of $f_{0}$ , i.e., the estimator is not adaptive to the smoothness of $f_{0}$ . We remind the reader that the estimator $f_{\lambda,n}$ is obtained by minimizing the regularized empirical Fisher divergence (see Theorem 4(iv)) yielding $f_{\lambda,n}=-(\hat{C}+\lambda I)^{-1}\hat{\xi}$ , which can be seen as a heuristic to solve the (non-linear) inverse problem $Cf_{0}=-\xi$ (see Theorem 4(ii)) from finite samples, by replacing $C$ and $\xi$ with their empirical counterparts. This heuristic, which ensures that the finite sample inverse problem is well-posed, is popular in inverse problem literature under the name of Tikhonov regularization (Engl et al., 1996, Chapter 5). Note that Tikhonov regularization helps to make the ill-posed inverse problem a well-posed one by approximating $\alpha^{-1}$ by $(\alpha+\lambda)^{-1}$ , $\lambda>0$ , where $\alpha^{-1}$ appears as the inverse of the eigenvalues of $C$ while computing $C^{-1}$ . In other words, if $\hat{C}$ is invertible, then an estimate of $f_{0}$ can be obtained as $\hat{f}_{n}=-\hat{C}^{-1}\hat{\xi}$ , i.e., $\hat{f}_{n}=-\sum_{i\in I}\frac{\langle\hat{\xi},\hat{\phi}_{i}\rangle_{\mathcal{H}}}{\hat{\alpha}_{i}}\hat{\phi}_{i},$ where $(\hat{\alpha}_{i})_{i\in I}$ and $(\hat{\phi}_{i})_{i\in I}$ are the eigenvalues and eigenvectors of $\hat{C}$ respectively. However, $\hat{C}$ being a rank $n$ operator defined on $\mathcal{H}$ (which can be infinite dimensional) is not invertible and therefore the regularized estimator is constructed as $f_{\lambda,n}=-g_{\lambda}(\hat{C})\hat{\xi}$ where $g_{\lambda}(\hat{C})$ is defined through functional calculus (see Engl, Hanke, and Neubauer, 1996, Section 2.3) as

The constant $\eta_{0}$ is called the qualification of $g_{\lambda}$ which is what determines the point of saturation of $g_{\lambda}$ . We show in Theorem 9 that if $g_{\lambda}$ has a finite qualification, then the resultant estimator cannot fully exploit the smoothness of $f_{0}$ and therefore the rate of convergence will suffer for $\beta>\eta_{0}$ . Given $g_{\lambda}$ that satisfies (E), we construct our estimator of $f_{0}$ as

Note that the above estimator can be obtained by using the data dependent regularizer, $\frac{1}{2}\langle f,((g_{\lambda}(\hat{C}))^{-1}-\hat{C})f\rangle_{\mathcal{H}}$ in the minimization of $\hat{J}(f)$ defined in Theorem 4(iv), i.e.,

However, unlike $f_{\lambda,n}$ for which a simple form is available in Theorem 5 by solving a linear system, we are not able to obtain such a nice expression for $f_{g,\lambda,n}$ . The following result (proved in Section 8.8) presents an analog of Theorems 6 and 7 for the new estimators, $f_{g,\lambda,n}$ and $p_{f_{g,\lambda,n}}$ .

Suppose (A)–(E) hold with $\varepsilon=2$ . (i) If $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ , then for any $\lambda\geq n^{-1/2}$ ,

where $\theta_{n}:=n^{-\min\left\{\frac{\beta}{2(\beta+1)},\frac{\eta_{0}}{2(\eta_{0}+1)}\right\}}$ with $\lambda=n^{-\max\left\{\frac{1}{2(\beta+1)},\frac{1}{2(\eta_{0}+1)}\right\}}$ . In addition, if $\|k\|_{\infty}<\infty$ , then for any $1<r\leq\infty$ with $q_{0}\in L^{1}(\Omega)\cap L^{r}(\Omega)$ ,

(ii) If $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta\geq 0$ , then for any $\lambda\geq n^{-1/2}$ ,

with $\lambda=n^{-\frac{1}{\min\{2\beta+2,2\eta_{0}+1\}}}$ . (iii) If $\|C^{-1}\|<\infty$ , then for any $\lambda\geq n^{-1/2}$ ,

with $\theta_{n}=n^{-\frac{1}{2}}$ and $\lambda=n^{-\frac{1}{\min\{2,2\eta_{0}\}}}$ .

Theorem 9 shows that if $g_{\lambda}$ has infinite qualification, then smoothness of $f_{0}$ is fully captured in the rates and as $\beta\rightarrow\infty$ , we attain $O_{p_{0}}(n^{-1/2})$ rate for $\|f_{g,\lambda,n}-f_{0}\|_{\mathcal{H}}$ in contrast to $n^{-1/4}$ (similar improved rates are also obtained for $p_{f_{g,\lambda,n}}$ in various distances) in Theorem 6. In the following example, we present two choices of $g_{\lambda}$ that improve on Tikhonov regularization. We refer the reader to Rosasco et al. (2005, Section 3.1) for more examples of $g_{\lambda}$ .

(i) Tikhonov regularization involves $g_{\lambda}(\alpha)=(\alpha+\lambda)^{-1}$ for which it is easy to verify that $\eta_{0}=1$ and therefore the rates saturate at $\beta=1$ , leading to the results in Theorems 6 and 7. (ii) Showalter’s method and spectral cut-off use

respectively for which it is easy to verify that $\eta_{0}=+\infty$ (see Engl, Hanke, and Neubauer, 1996, Examples 4.7 & 4.8 for details) and therefore improved rates are obtained for $\beta>1$ in Theorem 9 compared to that of Tikhonov regularization.

Density Estimation in 𝒫𝒫\mathcal{P}: Misspecified Case

In this section, we analyze the misspecified case where $p_{0}\notin\mathcal{P}$ , which is a more reasonable case than the well-specified one, as in practice it is not easy to check whether $p_{0}\in\mathcal{P}$ . To this end, we consider the same estimator $p_{f_{\lambda,n}}$ as considered in the well-specified case where $f_{\lambda,n}$ is obtained from Theorem 5. The following result shows that $J(p_{0}\|p_{f_{\lambda,n}})\rightarrow\inf_{p\in\mathcal{P}}J(p_{0}\|p)$ as $\lambda\rightarrow 0$ , $\lambda n\rightarrow\infty$ and $n\rightarrow\infty$ under the assumption that there exists $f^{*}\in\mathcal{F}$ such that $J(p_{0}\|p_{f^{*}})=\inf_{p\in\mathcal{P}}J(p_{0}\|p)$ . We present the result for bounded kernels although it can be easily extended to unbounded kernels as in Theorem B.2. Also, the presented result for Tikhonov regularization extends easily to $p_{f_{g,\lambda,n}}$ using the ideas in the proof of Theorem 9. Note that unlike in the well-specified case where convergence in other distances can be shown even though the estimator is constructed from $J$ , it is difficult to show such a result in the misspecified case.

Let $p_{0},\,q_{0}\in C^{1}(\Omega)$ be probability densities such that $J(p_{0}\|q_{0})<\infty$ where $\Omega$ satisfies (A). Assume that (B), (C) and (D) with $\varepsilon=2$ hold. Suppose $\|k\|_{\infty}<\infty$ , $\emph{supp}(q_{0})=\Omega$ and there exists $f^{\ast}\in\mathcal{F}$ such that

Then for an estimator $p_{f_{\lambda,n}}$ constructed from random samples $(X_{a})^{n}_{a=1}$ drawn i.i.d. from $p_{0}$ , where $f_{\lambda,n}$ is defined in (7)—also see Theorem 4(iv)—with $\lambda>0$ , we have

In addition, if $f^{*}\in\mathcal{R}(C^{\beta})$ for some $\beta\geq 0$ , then

with $\lambda=n^{-\max\left\{\frac{1}{3},\frac{1}{2(\beta+1)}\right\}}$ . If $\|C^{-1}\|<\infty$ , then for $\lambda=n^{-\frac{1}{2}}$ ,

While the above result is useful and interesting, the assumption about the existence of $f^{*}$ is quite restrictive. This is because if $p_{0}$ (which is not in $\mathcal{P}$ ) belongs to a family $\mathcal{Q}$ where $\mathcal{P}$ is dense in $\mathcal{Q}$ w.r.t. $J$ , then there is no $f\in\mathcal{H}$ that attains the infimum, i.e., $f^{*}$ does not exist and therefore the proof technique employed in Theorem 10 will fail. In the following, we present a result (Theorem 12) that does not require the existence of $f^{\ast}$ but attains the same result as in Theorem 10, but requiring a more complicated proof. Before we present Theorem 12, we need to introduce some notation.

To this end, let us return to the objective function under consideration,

where $f_{\star}=\log\frac{p_{0}}{q_{0}}$ and $p_{0}\notin\mathcal{P}$ . Define

This is a reasonable class of functions to consider as under the condition $J(p_{0}\|q_{0})<\infty$ , it is clear that $f_{\star}\in\mathcal{W}_{2}(\Omega,p_{0})$ . Endowed with a semi-norm,

$\mathcal{W}_{2}(\Omega,p_{0})$ is a vector space of functions, from which a normed space can be constructed as follows. Let us define $f,f^{\prime}\in\mathcal{W}_{2}(\Omega,p_{0})$ to be equivalent, i.e., $f\sim f^{\prime}$ , if $\|f-f^{\prime}\|_{\mathcal{W}_{2}}=0$ . In other words, $f\sim f^{\prime}$ if and only if $f$ and $f^{\prime}$ differ by a constant $p_{0}$ -almost everywhere. Now define the quotient space $\mathcal{W}^{\sim}_{2}(\Omega,p_{0}):=\left\{[f]_{\sim}:f\in\mathcal{W}_{2}(\Omega,p_{0})\right\}$ where $[f]_{\sim}:=\{f^{\prime}\in\mathcal{W}_{2}(\Omega,p_{0}):f\sim f^{\prime}\}$ denotes the equivalence class of $f$ . Defining $\|[f]_{\sim}\|_{\mathcal{W}^{\sim}_{2}}:=\|f\|_{\mathcal{W}_{2}}$ , it is easy to verify that $\|\cdot\|_{\mathcal{W}^{\sim}_{2}}$ defines a norm on $\mathcal{W}^{\sim}_{2}(p_{0})$ . In addition, endowing the following bilinear form on $\mathcal{W}^{\sim}_{2}(\Omega,p_{0})$

makes it a pre-Hilbert space. Let $W_{2}(\Omega,p_{0})$ be the Hilbert space obtained by completion of $\mathcal{W}^{\sim}_{2}(\Omega,p_{0})$ . As shown in Proposition 11 below, under some assumptions, a continuous mapping $I_{k}:\mathcal{H}\to W_{2}(\Omega,p_{0}),f\mapsto[f]_{\sim}$ can be defined, which is injective modulo constant functions. Since addition of a constant does not contribute to $p_{f}$ , the space $W_{2}(\Omega,p_{0})$ can be regarded as a parameter space extended from $\mathcal{H}$ . In addition to $I_{k}$ , Proposition 11 (proved in Section 8.10) describes the adjoint of $I_{k}$ and relevant self-adjoint operators, which will be useful in analyzing $p_{f_{\lambda,n}}$ in Theorem 12.

In addition, $I_{k}$ and $S_{k}$ are Hilbert-Schmidt and therefore compact. Also, $E_{k}:=S_{k}I_{k}$ and $T_{k}:=I_{k}S_{k}$ are compact, positive and self-adjoint operators on $\mathcal{H}$ and $W_{2}(\Omega,p_{0})$ respectively where

and the restriction of $T_{k}$ to $\mathcal{W}^{\sim}_{2}(\Omega,p_{0})$ is given by

Note that for $[h]_{\sim}\in\mathcal{W}^{\sim}_{2}(\Omega,p_{0})$ , the derivatives $\partial_{i}h$ do not depend on the choice of a representative element almost surely w.r.t. $p_{0}$ , and thus the above integrals are well defined. Having constructed $W_{2}(\Omega,p_{0})$ , it is clear that $J(p_{0}\|p_{f})=\frac{1}{2}\|[f_{\star}]_{\sim}-I_{k}f\|^{2}_{W_{2}}$ , which means estimating $p_{0}$ is equivalent to estimating $f_{\star}\in W_{2}(\Omega,p_{0})$ by $f\in\mathcal{F}$ . With all these preparations, we are now ready to present a result (proved in Section 8.11) on consistency and convergence rate for $p_{f_{\lambda,n}}$ without assuming the existence of $f^{\ast}$ .

Let $p_{0},\,q_{0}\in C^{1}(\Omega)$ be probability densities such that $J(p_{0}\|q_{0})<\infty$ . Assume that (A)–(D) hold with $\varepsilon=2$ and $\chi:=d\sup_{x\in\Omega,i\in[d]}\partial_{i}\partial_{i+d}k(x,x)<\infty$ . Then the following hold. (i) As $\lambda\rightarrow 0,\,\lambda n\rightarrow\infty\,\,\text{and}\,\,n\rightarrow\infty,$ $J(p_{0}\|p_{f_{\lambda,n}})\rightarrow\inf_{p\in\mathcal{P}}J(p_{0}\|p).$ (ii) Define $f_{\star}:=\log\frac{p_{0}}{q_{0}}$ . If $[f_{\star}]_{\sim}\in\overline{\mathcal{R}(T_{k})}$ , then

In addition, if $[f_{\star}]_{\sim}\in\mathcal{R}(T^{\beta}_{k})$ for some $\beta>0$ , then for $\lambda=n^{-\max\left\{\frac{1}{3},\frac{1}{2\beta+1}\right\}}$

. (iii) If $\|E^{-1}_{k}\|<\infty$ and $\|T^{-1}_{k}\|<\infty$ , then $J(p_{0}\|p_{f_{\lambda,n}})=O_{p_{0}}\left(n^{-1}\right)$ with $\lambda=n^{-\frac{1}{2}}$ .

Based on the observation (i) in the above remark that $\inf_{p\in\mathcal{P}}J(p_{0}\|p)=0$ if $I_{k}(\mathcal{H})$ is dense in $W_{2}(\Omega,p_{0})$ w.r.t. $\|\cdot\|_{W_{2}}$ , it is possible to obtain an approximation result for $\mathcal{P}$ (similar to those discussed in Section 3) w.r.t. Fisher divergence as shown below, whose proof is provided in Section 8.12.

Numerical Simulations

We have proposed an estimator of $p_{0}$ that is obtained by minimizing the regularized empirical Fisher divergence and presented its consistency along with convergence rates. As discussed in Section 1, however one can simply ignore the structure of $\mathcal{P}$ and estimate $p_{0}$ in a completely non-parametric fashion, for example using the kernel density estimator (KDE). In fact, consistency and convergence rates of KDE are also well-studied (Tsybakov, 2009, Chapter 1) and the kernel density estimator is very simple to compute—requiring only $O(n)$ computations—compared to the proposed estimator, which is obtained by solving a linear system of size $nd\times nd$ . This raises questions about the applicability of the proposed estimator in practice, though it is very well known that KDE performs poorly for moderate to large $d$ (Wasserman, 2006, Section 6.5). In this section, we numerically demonstrate that the proposed score matching estimator performs significantly better than the KDE, and in particular, that the advantage with the proposed estimator grows as $d$ gets large. Note further that the maximum likelihood approach of Barron and Sheu (1991) and Fukumizu (2009) does not yield estimators that are practically feasible, and therefore to the best of our knowledge, the proposed estimator is the only viable estimator for estimating densities through $\mathcal{P}$ .

In the following, we consider two simple scenarios of estimating a multivariate normal and mixture of normals using the proposed estimator and demonstrate the superior performance of the proposed estimator over KDE. Inspired by this preliminary empirical investigation, recently, the proposed estimator has been explored in two concrete applications of gradient-free adaptive MCMC sampler (Strathmann et al., 2015) and graphical model structure learning (Sun et al., 2015) where the superiority of working with the infinite dimensional family is demonstrated. We would like to again highlight that the goal of this work is not to construct density estimators that improve upon KDE but to provide a novel modeling technique of approximating an unknown density by a rich parametric family of densities with the parameter being infinite dimensional in contrast to the classical approach of finite dimensional approximation.

through the score matching approach and KDE, and compare their estimation accuracies. Here $\phi_{d}(x;\mu,\Sigma)$ is the p.d.f. of $N(\mu,\Sigma I_{d})$ . By choosing the kernel, $k(x,y)=\exp(-\frac{\|x-y\|^{2}_{2}}{2\sigma^{2}})+r(x^{T}y+c)^{2},$ which is a Gaussian plus polynomial of degree 2, it is easy to verify that Gaussian distributions lie in $\mathcal{P}$ , and therefore the first problem considers the well-specified case while the second problem deals with the misspecified case. In our simulations, we chose $r=0.1$ , $c=0.5$ , $\alpha=4$ and $\beta=-4$ . The base measure of the exponential family is $N(0,10^{2}I_{d})$ . The bandwidth parameter $\sigma$ is chosen by cross-validation (CV) of the objective function $\hat{J}_{\lambda}$ (see Theorem 4(iv)) within the parameter set $\{0.1,0.2,0.4,0.6,0.8,1,1.2,1.4,1.6\}\times\sigma_{*}$ , where $\sigma_{*}$ is the median of pairwise distances of data, and the regularization parameter $\lambda$ is set as $\lambda=0.1\times n^{-1/3}$ with sample size $n$ . For KDE, the Gaussian kernel is used for the smoothing kernel, and the bandwidth parameter is chosen by CV from $\{0.02,0.04,0.06,0.08,0.1,0.2,0.4,0.6,0.8,1.0\}\times\sigma_{*}$ ; where for both the methods, 5-fold CV is applied.

Since it is difficult to accurately estimate the normalization constant in the proposed method, we use two methods to evaluate the accuracy of estimation. One is the objective function for the score matching method,

and the other is correlation of the estimator with the true density function,

where $R$ is a probability distribution. For $R$ , we use the empirical distribution based on 10000 random samples drawn i.i.d. from $p_{0}(x)$ .

Summary & Discussion

We have considered an infinite dimensional generalization, $\mathcal{P}$ , of the finite-dimensional exponential family, where the densities are indexed by functions in a reproducing kernel Hilbert space (RKHS), $\mathcal{H}$ . We showed that $\mathcal{P}$ is a rich object that can approximate a large class of probability densities arbitrarily well in Kullback-Leibler divergence, and addressed the main question of estimating an unknown density, $p_{0}$ from finite samples drawn i.i.d. from it, in well-specified ( $p_{0}\in\mathcal{P}$ ) and misspecified ( $p_{0}\notin\mathcal{P}$ ) settings. We proposed a density estimator based on minimizing the regularized version of the empirical Fisher divergence, which results in solving a simple finite-dimensional linear system. Our estimator provides a computationally efficient alternative to maximum likelihood based estimators, which suffer from the computational intractability of the log-partition function. The proposed estimator is also shown to empirically outperform the classical kernel density estimator, with advantage increasing as the dimension of the space increases. In addition to these computational and empirical results, we have established the consistency and convergence rates under certain smoothness assumptions (e.g., $\log p_{0}\in\mathcal{R}(C^{\beta})$ ) for both well-specified and misspecified scenarios.

Proofs

We provide proofs of the results presented in Sections 3–5.

Sriperumbudur et al. (2011, Proposition 5) showed that $\mathcal{H}$ is dense in $C_{0}(\Omega)$ w.r.t. uniform norm if and only if $k$ satisfies (5). Therefore, the denseness in $L^{1}$ , KL and Hellinger distances follow trivially from Lemma A.1. For $L^{r}$ norm ( $r>1$ ), the denseness follows by using the bound $\|p_{f}-p_{g}\|_{L^{r}(\Omega)}\leq 2e^{2\|f-g\|_{\infty}}e^{2\|f\|_{\infty}}\|f-g\|_{\infty}\|q_{0}\|_{L^{r}(\Omega)}$ obtained from Lemma A.1(i) with $f\in C_{0}(\Omega)$ and $g\in\mathcal{H}$ . $\blacksquare$

2 Proof of Corollary 2

For any $p\in\mathcal{P}_{c}$ , define $p_{\delta}:=\frac{p+\delta q_{0}}{1+\delta}$ . Note that $p_{\delta}(x)>0$ for all $x\in\Omega$ and $\|p-p_{\delta}\|_{L^{r}(\Omega)}=\frac{\delta\|p-q_{0}\|_{L^{r}(\Omega)}}{1+\delta}$ , implying that $\lim_{\delta\rightarrow 0}\|p-p_{\delta}\|_{L^{r}(\Omega)}=0$ for any $1\leq r\leq\infty$ . This means, for any $\epsilon>0$ , $\exists\delta_{\epsilon}>0$ such that for any $0<\theta<\delta_{\epsilon}$ , we have $\|p-p_{\theta}\|_{L^{r}(\Omega)}\leq\epsilon$ , where $p_{\theta}(x)>0$ for all $x\in\Omega$ .

We now prove the denseness in KL divergence by noting that

which implies $\lim_{\delta\rightarrow 0}KL(p\|p_{\delta})=0$ . This implies, for any $\epsilon>0$ , $\exists\delta_{\epsilon}>0$ such that for any $0<\theta<\delta_{\epsilon}$ , $KL(p\|p_{\theta})\leq\epsilon$ . Arguing as above, we have $p_{\theta}\in\mathcal{P}_{0}$ , i.e., there exists $f\in C_{0}(\Omega)$ such that $p_{\theta}=\frac{e^{f}q_{0}}{\int e^{f}q_{0}\,dx}$ . Since $\mathcal{H}$ is dense in $C_{0}(\Omega)$ , for any $f\in C_{0}(\Omega)$ and any $\epsilon>0$ , there exists $g\in\mathcal{H}$ such that $\|f-g\|_{\infty}\leq\epsilon$ . For $p_{g}\in\mathcal{P}$ , since $\int p\,\log\frac{p_{\theta}}{p_{g}}\,dx\leq\left\|\log\frac{p_{\theta}}{p_{g}}\right\|_{\infty}\leq 2\|f-g\|_{\infty}\leq 2\epsilon,$ we have

3 Proof of Theorem 4

(i) By the reproducing property of $\mathcal{H}$ , since $\partial_{i}f(x)=\left\langle f,\partial_{i}k(x,\cdot)\right\rangle_{\mathcal{H}}$ for all $i\in[d]$ , it is easy to verify that

where in the second line, we used $\langle a,b\rangle^{2}_{H}=\langle a,b\rangle_{H}\langle a,b\rangle_{H}=\langle a,(b\otimes b)a\rangle_{H}$ for $a,b\in H$ with $H$ being a Hilbert space and

Observe that for all $x\in\Omega$ , $C_{x}$ is a Hilbert-Schmidt operator as $\|C_{x}\|_{HS}\leq\sum^{d}_{i=1}\left\|\partial_{i}k(x,\cdot)\right\|^{2}_{\mathcal{H}}$ $=\sum^{d}_{i=1}\partial_{i}\partial_{i+d}k(x,x)<\infty$ and $(f-f_{0})\otimes(f-f_{0})$ is also Hilbert-Schmidt as $\|(f-f_{0})\otimes(f-f_{0})\|_{HS}=\|f-f_{0}\|^{2}_{\mathcal{H}}<\infty$ . Therefore, (11) is equivalent to

Since the first condition in (D) implies $\int_{\Omega}\|C_{x}\|_{HS}p_{0}(x)\,dx<\infty$ , $C_{x}$ is $p_{0}$ -integrable in the Bochner sense (see Diestel and Uhl, 1977, Definition 1 and Theorem 2), and therefore it follows from Diestel and Uhl (1977, Theorem 6) that

where $C:=\int_{\Omega}C_{x}\;p_{0}(x)\,dx$ is the Bochner integral of $C_{x}$ , thereby yielding (6).

which means $C$ is trace-class and therefore compact. Here, we used monotone convergence theorem in $(*)$ and Parseval’s identity in $(**)$ . Note that $C$ is positive since $\langle f,Cf\rangle_{\mathcal{H}}=\int_{\Omega}p_{0}(x)\left\|\nabla f\right\|^{2}_{2}\,dx\geq 0,\,\forall\,f\in\mathcal{H}.$ (ii) From (6), we have $J(f)=\frac{1}{2}\langle f,Cf\rangle_{\mathcal{H}}-\langle f,Cf_{0}\rangle_{\mathcal{H}}+\frac{1}{2}\langle f_{0},Cf_{0}\rangle_{\mathcal{H}}$ . Using $\partial_{i}f_{0}(x)=\partial_{i}\log p_{0}(x)-\partial_{i}\log q_{0}(x)$ for all $i\in[d]$ , we obtain that for any $f\in\mathcal{H}$ ,

where $(b)$ follows from integration by parts under (C) and the equality in $(c)$ is valid as $\xi_{x}$ is Bochner $p_{0}$ -integrable under (D) with $\varepsilon=1$ . Therefore $Cf_{0}=-\xi$ . For the third term, $\langle f_{0},Cf_{0}\rangle_{\mathcal{H}}=\int_{\Omega}p_{0}(x)\sum^{d}_{i=1}\left(\partial_{i}f_{0}(x)\right)^{2}\,dx$ and the result follows. (iii) Define $c_{0}:=J(p_{0}\|q_{0})$ . For any $\lambda>0$ , it is easy to verify that

Clearly, $J_{\lambda}(f)$ is minimized if and only if $(C+\lambda I)^{1/2}f=-(C+\lambda I)^{-1/2}\xi$ and therefore $f_{\lambda}=-(C+\lambda I)^{-1}\xi$ is the unique minimizer of $J_{\lambda}(f)$ . (iv) Since $(iv)$ is similar to $(iii)$ with $C$ replaced by $\hat{C}$ and $\xi$ replaced by $\hat{\xi}$ , we obtain $f_{\lambda,n}=(\hat{C}+\lambda I)^{-1}\hat{\xi}$ . $\blacksquare$

4 Proof of Theorem 5

We prove the result based on the general representer theorem (Theorem A.2). From Theorem 4(iv), we have

where $V(\theta_{1},\ldots,\theta_{nd},\theta_{nd+1}):=\frac{1}{2n}\sum^{n}_{a=1}\sum^{d}_{i=1}\theta^{2}_{(a-1)d+i}+\theta_{nd+1}$ , $\phi_{(a-1)d+i}:=\partial_{i}k(X_{a},\cdot),\,a\in[n],\,i\in[d]$ and $\phi_{nd+1}:=\hat{\xi}$ . Therefore, it follows from Theorem A.2 that

with $\bm{K}=\begin{pmatrix}\bm{G}\,&\,\bm{h}\\ \bm{h}^{T}\,&\,\|\hat{\xi}\|^{2}_{\mathcal{H}}\end{pmatrix}.$ Since $\nabla V\begin{pmatrix}\bm{z}\\ t\end{pmatrix}=\begin{pmatrix}\frac{1}{n}\bm{z}\\ 1\end{pmatrix}$ , (15) reduces to $\lambda\delta+1=0$ and $\lambda\bm{\beta}+\frac{1}{n}\bm{G\beta}+\frac{\delta}{n}\bm{h}=0$ yielding $\delta=-\frac{1}{\lambda}$ and $(\frac{1}{n}\bm{G}+\lambda I)\bm{\beta}=\frac{1}{n\lambda}\bm{h}$ . $\blacksquare$

Instead of using the general representer theorem (Theorem A.2), it is possible to see that the standard representer theorem (Kimeldorf and Wahba, 1971; Schölkopf et al., 2001) gives a similar, but slightly different linear system, and the solutions are the same if $\bm{K}$ is non-singular. The general representer theorem yields that $\bm{\beta}$ and $\delta$ are solution to $\bm{F}\begin{pmatrix}\bm{\beta}\\ \delta\end{pmatrix}=\begin{pmatrix}\bm{0}\\ 1\end{pmatrix}$ , where $\bm{F}=\begin{pmatrix}\frac{1}{n}\bm{G}+\lambda I\,&\,\frac{1}{n}\bm{h}\\ \bm{0}^{T}\,&\,\lambda\end{pmatrix}$ . On the other hand, by using the standard representer theorem, it is easy to show that $f_{\lambda,n}$ has the form in (14) with $\delta$ and $\bm{\beta}$ being solution to $\bm{K}\bm{F}\begin{pmatrix}\bm{\beta}\\ \delta\end{pmatrix}=\bm{K}\begin{pmatrix}\bm{0}\\ 1\end{pmatrix}$ . Clearly, both the solutions match if $\bm{K}$ is invertible while the latter has many solutions if $\bm{K}$ is not invertible.

5 Proof of Theorem 6

where we used $\lambda f_{\lambda}=C(f_{0}-f_{\lambda})$ in ( $\ast$ ). Define $S_{1}:=\|(\hat{C}+\lambda I)^{-1}(C-\hat{C})(f_{\lambda}-f_{0})\|_{\mathcal{H}}$ , $S_{2}:=\|(\hat{C}+\lambda I)^{-1}(\hat{\xi}-\xi)\|_{\mathcal{H}}$ and $S_{3}:=\|(\hat{C}+\lambda I)^{-1}(C-\hat{C})f_{0}\|_{\mathcal{H}}$ so that

where $\mathcal{A}_{0}(\lambda):=\|f_{\lambda}-f_{0}\|_{\mathcal{H}}$ . We now bound $S_{1}$ , $S_{2}$ and $S_{3}$ using Proposition A.4. Note that $C=\int_{\Omega}C_{x}\,p_{0}(x)\,dx$ where $C_{x}$ is defined in (12) is a positive, self-adjoint, trace-class operator and (D) (with $\varepsilon=2$ ) implies that

Using the bounds in $S_{1}$ , $S_{2}$ and $S_{3}$ in (16), we obtain

(i) By Proposition A.3(i), we have that $\mathcal{A}_{0}(\lambda)\rightarrow 0$ as $\lambda\rightarrow 0$ if $f_{0}\in\overline{\mathcal{R}(C)}$ . Therefore, it follows from (20) that $\|f_{\lambda,n}-f_{0}\|_{\mathcal{H}}\rightarrow 0$ as $\lambda\rightarrow 0$ , $\lambda\sqrt{n}\rightarrow\infty$ and $n\rightarrow\infty$ . (ii) If $f_{0}\in\mathcal{R}(C^{\beta})$ for $\beta>0$ , it follows from Proposition A.3(ii) that

and therefore the result follows by choosing $\lambda=n^{-\max\left\{\frac{1}{4},\frac{1}{2(\beta+1)}\right\}}$ . (iii) Note that

It follows from Proposition A.4(v) that $\|C(\hat{C}+\lambda I)^{-1}\|\lesssim 1$ for $n\geq\frac{c}{\lambda^{2}}$ where $c$ is a sufficiently large constant that depends on $\sum^{d}_{i=1}\int_{\Omega}(\partial_{i}\partial_{i+d}k(x,x))^{2}p_{0}(x)\,dx$ but not on $n$ and $\lambda$ . Using the bounds on $\|(C-\hat{C})(f_{\lambda}-f_{0})\|_{\mathcal{H}}$ , $\|\hat{\xi}-\xi\|_{\mathcal{H}}$ and $\|(C-\hat{C})f_{0}\|_{\mathcal{H}}$ from part (i) and the bound on $\|C(f_{\lambda}-f_{0})\|_{\mathcal{H}}$ from Proposition A.3(ii), we therefore obtain

as $n\rightarrow\infty$ and the result follows. $\blacksquare$

Under slightly strong assumptions on the kernel, the bound on $S_{1}$ in (17) can be improved to obtain $S_{1}=O_{p_{0}}(n^{-1/2})$ while the one on $S_{3}$ in (19) can be refined to obtain $S_{3}=O_{p_{0}}\left(\sqrt{\frac{\mathcal{N}(\lambda)}{\lambda n}}\right)$ where $\mathcal{N}(\lambda):=\emph{Tr}((C+\lambda I)^{-1}C)$ is the intrinsic dimension of $\mathcal{H}$ . Using the fact that $\mathcal{N}(\lambda)\leq\frac{1}{\lambda}$ , it is easy to verify that the latter is an improved bound than the one in (19). In addition $S_{3}$ dominates $S_{1}$ . However, if $S_{2}$ in (18) is not improved, then $S_{2}$ dominates $S_{3}$ , thereby resulting in a bound that does not capture the smoothness of $k$ (or the corresponding $\mathcal{H}$ ). Unfortunately, even with a refined analysis (not reported here), we are not able to improve the bound on $S_{2}$ wherein the difficulty lies with handling $\xi$ .

6 Proof of Theorem 7

Before we prove the result, we present a lemma.

Since $\sup_{x\in\Omega}k(x,x)<\infty$ , it implies that, for every $f\in\mathcal{H}$ , $\int_{\Omega}e^{f(x)}q_{0}(x)\,dx<\infty$ and hence $\mathcal{F}=\mathcal{H}$ . Also, under the assumptions on $k$ and $q_{0}$ , it is easy to verify that $\text{supp}(p_{0})=\Omega$ , which implies

In the following, we obtain a bound on $J(p_{0}\|p_{f_{\lambda,n}})=\frac{1}{2}\|\sqrt{C}(f_{\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}$ . While one can trivially use the bound $\|\sqrt{C}(f_{\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}\leq\|\sqrt{C}\|^{2}\|f_{\lambda,n}-f_{0}\|^{2}_{\mathcal{H}}$ to obtain a rate on $J(p_{0}\|p_{f_{\lambda,n}})$ through the result in Theorem 6(ii), a better rate can be obtained by carefully bounding $\|\sqrt{C}(f_{\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}$ as shown below. Consider

7 Proof of Proposition 8

Observation 1: By (Wendland, 2005, Theorem 10.12), we have

where $f^{\wedge}$ is defined in $L^{2}$ sense. Since

Observation 3: For $g\in\mathcal{H}$ , we have for all $j\in[d]$ ,

Observation 4: For any $g\in\mathcal{G}$ , we have

which implies $g\in\mathcal{H}$ , i.e., $\mathcal{G}\subset\mathcal{H}$ .

We now use these observations to prove the result. Since $f_{0}\in\mathcal{R}(C)$ , there exists $g\in\mathcal{H}$ such that $f_{0}=Cg$ , which means

where we have invoked generalized Young’s inequality (Folland, 1999, Proposition 8.9) in ( $\ast$ ), Hausdorff-Young inequality (Folland, 1999, p. 253) in ( $\ast\ast$ ), and observation 3 combined with $(iv)$ in ( $\ddagger$ ). This shows that $f_{0}\in\mathcal{R}(C)\Rightarrow f_{0}\in\mathcal{G}$ , i.e., $\mathcal{R}(C)\subset\mathcal{G}$ . $\blacksquare$

8 Proof of Theorem 9

To prove Theorem 9, we need the following lemma (De Vito et al., 2012, Lemma 5), which is due to Andreas Maurer.

Proof of Theorem 9. (i) The proof follows the ideas in the proof of Theorem 10 in Bauer et al. (2007), which is a more general result dealing with the smoothness condition, $f_{0}\in\mathcal{R}(\Theta(C))$ where $\Theta$ is operator monotone. Recall that $\Theta$ is operator monotone on $[0,b]$ if for any pair of self-adjoint operators $U$ , $V$ with spectra in $[0,b]$ such that $U\leq V$ , we have $\Theta(U)\leq\Theta(V)$ , where “ $\leq$ ” is the partial ordering for self-adjoint operators on some Hilbert space $H$ , which means for any $f\in H$ , $\langle f,Uf\rangle_{H}\leq\langle f,Vf\rangle_{H}$ . In our case, we adapt the proof for $\Theta(C)=C^{\beta}$ . Define $r_{\lambda}(\alpha):=g_{\lambda}(\alpha)\alpha-1$ . Since $f_{0}\in\mathcal{R}(C^{\beta})$ , there exists $h\in\mathcal{H}$ such that $f_{0}=C^{\beta}h$ , which yields

We now bound $(A)$ – $(D)$ . Since $(A)\leq\|g_{\lambda}(\hat{C})\|\|\hat{\xi}-\xi\|_{\mathcal{H}}$ , we have $(A)=O_{p_{0}}\left(\frac{1}{\lambda\sqrt{n}}\right)$ where we used $(b)$ in (E) and the bound on $\|\hat{\xi}-\xi\|_{\mathcal{H}}$ from the proof of Theorem 6(i). Similarly, $(B)\leq\|g_{\lambda}(\hat{C})\|\|(\hat{C}-C)f_{0}\|_{\mathcal{H}}$ implies $(B)=O_{p_{0}}\left(\frac{1}{\lambda\sqrt{n}}\right)$ where $(b)$ in (E) and Proposition A.4(i) are invoked. Also, $(d)$ in (E) implies that

We now consider two cases: $\beta\leq 1$ : Since $\alpha\mapsto\alpha^{\theta}$ is operator monotone on $[0,\chi]$ for $0\leq\theta\leq 1$ , by Theorem 1 in Bauer et al. (2007), there exists a constant $c_{\theta}$ such that $\|\hat{C}^{\theta}-C^{\theta}\|\leq c_{\theta}\|\hat{C}-C\|^{\theta}\leq c_{\theta}\|\hat{C}-C\|^{\theta}_{HS}$ . We now obtain a bound on $\|\hat{C}-C\|_{HS}$ . To this end, consider

which by Chebyshev’s inequality implies that

and therefore $(D)=O_{p_{0}}(n^{-\beta/2})$ . Since $\lambda\geq n^{-1/2}$ , we have $(D)=O_{p_{0}}(\lambda^{\beta})$ . $\beta>1$ : Since $\alpha\mapsto\alpha^{\theta}$ is Lipschitz on $[0,\chi]$ for $\theta\geq 1$ , by Lemma 15, $\|C^{\beta}-\hat{C}^{\beta}\|\leq\|C^{\beta}-\hat{C}^{\beta}\|_{HS}\leq\beta\chi^{\beta-1}\|C-\hat{C}\|_{HS}$ and therefore $(C)=O_{p_{0}}(n^{-1/2})$ .

Collecting all the above bounds, we obtain

and the result follows. The proofs of the claims involving $L^{r}$ , $h$ and $KL$ follow exactly the same ideas as in the proof of Theorem 7 by using the above bound on $\|f_{g,\lambda,n}-f_{0}\|_{\mathcal{H}}$ in Lemma A.1. (ii) We now bound $J(p_{0}\|p_{f_{g,\lambda,n}})=\|\sqrt{C}(f_{g,\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}$ as follows. Note that

We bound $\|(I^{\prime})\|_{\mathcal{H}}$ as

where we used the fact that $\alpha\mapsto\sqrt{\alpha}$ is operator monotone along with $\lambda\geq n^{-1/2}$ . Using (26), $\|(II^{\prime})\|_{\mathcal{H}}$ can be bounded as

with $\|\hat{C}f_{0}+\hat{\xi}\|$ and $\|C^{\beta}-\hat{C}^{\beta}\|$ bounded as in part (i) above. Here $(a\vee b):=\max\{a,b\}$ . Combining $\|(I^{\prime})\|_{\mathcal{H}}$ and $\|(II^{\prime})\|_{\mathcal{H}}$ , we obtain the required result. (iii) The proof follows the ideas in the proof of Theorems 6 and 7. Consider $f_{g,\lambda,n}-f_{0}=-g_{\lambda}(\hat{C})(\hat{C}f_{0}+\hat{\xi})+r_{\lambda}(\hat{C})f_{0}$ so that

Therefore $\|f_{g,\lambda,n}-f_{0}\|_{\mathcal{H}}=O_{p_{0}}(n^{-1/2})+O\left(\lambda^{\min\{1,\eta_{0}\}}\right)$ where we used the fact that $\lambda\geq n^{-1/2}$ and the result follows. $\blacksquare$

9 Proof of Theorem 10

Before we analyze $J(p_{0}\|p_{f_{\lambda,n}})$ , we need a small calculation for notational convenience. For any probability densities $p,q\in C^{1}$ , it is clear that $\sqrt{2J(p\|q)}=\left\|\left\|\nabla\log p-\nabla\log q\right\|_{2}\right\|_{L^{2}(p)}$ . We generalize this by defining

Clearly, if $\mu=p$ , then $J(p\|q\|\mu)$ matches with $J(p\|q)$ . Therefore, for probability densities $p,q,r\in C^{1}$ ,

where $\mathcal{A}^{*}(\lambda)=\|\sqrt{C}(f_{\lambda}-f^{*})\|_{\mathcal{H}}$ . The result simply follows from the proof of Theorem 7, where we showed that $\|\sqrt{C}(f_{\lambda,n}-f_{\lambda})\|_{\mathcal{H}}=O_{p_{0}}\left(\frac{1}{\sqrt{\lambda n}}\right)$ and $\mathcal{A}^{*}(\lambda)=O(\lambda^{\min\{1,\beta+\frac{1}{2}\}})$ if $f^{\ast}\in\mathcal{R}(C^{\beta})$ for $\beta\geq 0$ as $\lambda\rightarrow 0$ , $n\rightarrow\infty$ . When $\|C^{-1}\|<\infty$ , we bound $\|\sqrt{C}(f_{\lambda,n}-f^{*})\|_{\mathcal{H}}$ in (8.9) as $\|\sqrt{C}\|\|f_{\lambda,n}-f^{*}\|_{\mathcal{H}}$ where $\|f_{\lambda,n}-f^{*}\|_{\mathcal{H}}$ is in turn bounded as in (21). $\blacksquare$

10 Proof of Proposition 11

which means $f\in\mathcal{W}_{2}(\Omega,p_{0})$ and therefore $[f]_{\sim}\in W_{2}(\Omega,p_{0})$ . Since $\|I_{k}f\|_{W_{2}}=\|[f]_{\sim}\|_{\mathcal{W}^{\sim}_{2}}=\|f\|_{\mathcal{W}_{2}}\leq c\|f\|_{\mathcal{H}}<\infty$ where $c$ is some constant, it is clear that $I_{k}$ is a continuous map from $\mathcal{H}$ to $W_{2}(\Omega,p_{0})$ . The adjoint $S_{k}:W_{2}(\Omega,p_{0})\rightarrow\mathcal{H}$ of $I_{k}:\mathcal{H}\rightarrow W_{2}(\Omega,p_{0})$ is defined by the relation $\langle S_{k}f,g\rangle_{\mathcal{H}}=\langle f,I_{k}g\rangle_{W_{2}},\,\,f\in W_{2}(\Omega,p_{0}),\,g\in\mathcal{H}$ . If $f:=[h]_{\sim}\in\mathcal{W}^{\sim}_{2}(\Omega,p_{0})$ , then

For $y\in\Omega$ and $g=k(\cdot,y)$ , this yields

We now show that $I_{k}$ is Hilbert-Schmidt. Since $\mathcal{H}$ is separable, let $(e_{l})_{l\geq 1}$ be an ONB of $\mathcal{H}$ . Then we have

which proves that $I_{k}$ is Hilbert-Schmidt (hence compact) and therefore $S_{k}$ is also Hilbert-Schmidt and compact. The other assertions about $S_{k}I_{k}$ and $I_{k}S_{k}$ are straightforward. $\blacksquare$

11 Proof of Theorem 12

By slight abuse of notation, $f_{\star}$ is used to denote $[f_{\star}]_{\sim}$ in the proof for simplicity. For $f\in\mathcal{F}$ , we have

Since $k$ satisfies (C) it is easy to verify that $\langle S_{k}f_{\star},f\rangle_{\mathcal{H}}=\langle f,-\xi\rangle_{\mathcal{H}},\,\forall\,f\in\mathcal{H}$ (see proof of Theorem 4(ii)). This implies $S_{k}f_{\star}=-\xi$ and

where $\xi$ is defined in Theorem 4(ii), and $E_{k}$ is precisely the operator $C$ defined in Theorem 4(ii). Following the proof of Theorem 4(ii), for $\lambda>0$ , it is easy to show that the unique minimizer of the regularized objective, $J(p_{0}\|p_{f})+\frac{\lambda}{2}\|f\|^{2}_{\mathcal{H}}$ exists and is given by

We would like to reiterate that (29) and (30) also match with their counterparts in Theorem 4 and therefore as in Theorem 4(iv), an estimator of $f_{\star}$ is given by $f_{\lambda,n}=-(\hat{E}_{k}+\lambda I)^{-1}\hat{\xi}$ . In other words, this is the same as in Theorem 4(iv) since $\hat{E}_{k}=\hat{C}$ , and can be solved by a simple linear system provided in Theorem 5. Here $\hat{E}_{k}$ is the empirical estimator of $E_{k}$ . Now consider

where $\mathcal{B}(\lambda):=\|I_{k}f_{\lambda}-f_{\star}\|_{W_{2}}$ . The proof now proceeds using the following decomposition, equivalent to the one used in the proof of Theorem 6(i), i.e.,

where we used (30) in $(\dagger)$ . $\hat{S}_{k}f_{\star}$ is well-defined as it is the empirical version of the restriction of $S_{k}$ to $\mathcal{W}^{\sim}_{2}(p_{0})$ . Since $S_{k}f_{\star}-E_{k}f_{\lambda}=S_{k}(f_{\star}-I_{k}f_{\lambda})$ and $\hat{S}_{k}f_{\star}-\hat{E}_{k}f_{\lambda}=\hat{S}_{k}(f_{\star}-I_{k}f_{\lambda})$ , we have

for $n\geq\frac{c}{\lambda^{2}}$ where $c$ is a sufficiently large constant that does not depend on $n$ and $\lambda$ . Following the proof of Proposition A.4(i), we have

wherein the first term is zero as $S_{k}f_{\star}+\xi=0$ and since

the integral in the second term is finite because of (D) and $f^{*}\in W_{2}(\Omega,p_{0})$ . Therefore, an application of Chebyshev’s inequality yields

We now show that $\|(\hat{S}_{k}-S_{k})(f_{\star}-I_{k}f_{\lambda})\|_{\mathcal{H}}=O_{p_{0}}(\mathcal{B}(\lambda)n^{-1/2})$ . To this end, define $g:=f_{\star}-I_{k}f_{\lambda}$ and consider

which therefore yields the claim through an application of Chebyshev’s inequality. Using this along with (33) and (34) in (32), and using the resulting bound in (31) yields

(i) We bound $\mathcal{B}(\lambda)$ as follows. First note that

and so for any $h\in\mathcal{H}$ , we have

From $(T_{k}+\lambda I)^{-1}T_{k}=I_{k}(E_{k}+\lambda I)^{-1}S_{k}$ and $S_{k}I_{k}h=E_{k}h$ , we have

where the inequality follows from Proposition A.3(ii). Using (37) and (38) in (36), we obtain $\mathcal{B}(\lambda)\leq\|f_{\star}-I_{k}h\|_{W_{2}}+\|h\|_{\mathcal{H}}\sqrt{\lambda}$ , using which in (35) yields

Since the above inequality holds for any $h\in\mathcal{H}$ , we therefore have

Since $J(p_{0}\|p_{f_{\lambda,n}})\geq\inf_{p\in\mathcal{P}}J(p_{0}\|p)$ , we have that $J(p_{0}\|p_{f_{\lambda,n}})\rightarrow\inf_{p\in\mathcal{P}}J(p_{0}\|p)$ as $\lambda\rightarrow 0$ , $\lambda n\rightarrow\infty$ and $n\rightarrow\infty$ . (ii) Recall $\mathcal{B}(\lambda)$ from (i). From Proposition A.3(i) it follows that $\mathcal{B}(\lambda)\rightarrow 0$ as $\lambda\rightarrow 0$ if $f_{\star}\in\overline{\mathcal{R}(T_{k})}$ . Therefore, (35) reduces to $\sqrt{2\,J(p_{0}\|p_{f_{\lambda,n}})}\leq O_{p_{0}}\left(\frac{1}{\sqrt{\lambda n}}\right)+\mathcal{B}(\lambda)$ and the consistency result follows. If $f_{\star}\in\mathcal{R}(T^{\beta}_{k})$ for some $\beta>0$ , then the rates follow from Proposition A.3 by noting that $\mathcal{B}(\lambda)\leq\max\{1,\|T_{k}\|^{\beta-1}\}\lambda^{\min\{1,\beta\}}\|T^{-\beta}_{k}f_{\star}\|_{W_{2}}$ and choosing $\lambda=n^{-\max\left\{\frac{1}{3},\frac{1}{2\beta+1}\right\}}$ . (iii) This simply follows from an analysis similar to the one used in the proof of Theorem 6(iii). $\blacksquare$

12 Proof of Proposition 13

For any $p\in\mathcal{P}_{\text{FD}}$ , define $f:=\log\frac{p}{q_{0}}$ , which implies that $[f]_{\sim}\in W_{2}(p)$ . Since $I_{k}(\mathcal{H})$ is dense in $W_{2}(p)$ , we have for any $\epsilon>0$ , there exists $g\in\mathcal{H}$ such that $\|[f]_{\sim}-I_{k}g\|_{W_{2}}\leq\sqrt{2\epsilon}$ . For a given $g\in\mathcal{H}$ , pick $p_{g}\in\mathcal{P}$ . Therefore,

Appendix A Appendix: Technical Results

In this appendix, we present some technical results that are used in the proofs.

In the following result, claims (iii) and (iv) are quoted from Lemma 3.1 of van der Vaart and van Zanten (2008).

$\|p_{f}-p_{g}\|_{L^{r}(\Omega)}\leq 2e^{2\|f-g\|_{\infty}}e^{2\min\{\|f\|_{\infty},\|g\|_{\infty}\}}\|f-g\|_{\infty}\|q_{0}\|_{L^{r}(\Omega)}$ for any $1\leq r\leq\infty$ ;

$\|p_{f}-p_{g}\|_{L^{1}(\Omega)}\leq 2e^{\|f-g\|_{\infty}}\|f-g\|_{\infty}$ ;

$KL(p_{f}\|p_{g})\leq c\,\|f-g\|^{2}_{\infty}e^{\|f-g\|_{\infty}}\left(1+\|f-g\|_{\infty}\right)$ where $c$ is a universal constant;

$h(p_{f},p_{g})\leq e^{\|f-g\|_{\infty}/2}\|f-g\|_{\infty}.$

(i) Define $B(f):=\int e^{f}q_{0}\,dx$ . Consider

Since $\|e^{f}q_{0}\|_{L^{r}(\Omega)}\leq e^{\|f\|_{\infty}}\|q_{0}\|_{L^{r}(\Omega)}$ and $B(f)\geq e^{-\|f\|_{\infty}}$ , from (A.2) we obtain

where we used $\max\{a,b\}\leq\min\{a,b\}+|a-b|$ for $a,b\geq 0$ in the last line above. (ii) This simply follows from (A.2) by using $r=1$ .

A.2 General Representer Theorem

The following is the general representer theorem for abstract Hilbert spaces.

and thus $A^{*}\bm{\alpha}=\sum^{m}_{i=1}\alpha_{i}\phi_{i}$ . Therefore $AA^{*}\bm{\alpha}=\sum^{m}_{j=1}\alpha_{j}A\phi_{j}=\sum^{m}_{j=1}\alpha_{j}(\langle\phi_{j},\phi_{i}\rangle_{H})^{m}_{i=1}$ , and so for every $i\in[m]$ , $(AA^{*}\bm{\alpha})_{i}=\sum^{m}_{j=1}\langle\phi_{j},\phi_{i}\rangle_{H}\alpha_{j}$ and hence $AA^{*}=\bm{K}$ .

The following result is quite well-known in the linear inverse problem theory (Engl et al., 1996).

Let $C$ be a bounded, self-adjoint compact operator on a separable Hilbert space $H$ . For $\lambda>0$ and $f\in H$ , define $f_{\lambda}:=(C+\lambda I)^{-1}Cf$ and $\mathcal{A}_{\theta}(\lambda):=\|C^{\theta}(f_{\lambda}-f)\|_{H}$ for $\theta\geq 0$ . Then the following hold.

For any $\theta>0$ , $\mathcal{A}_{\theta}(\lambda)\rightarrow 0$ as $\lambda\rightarrow 0$ and if $f\in\overline{\mathcal{R}(C)}$ , then $\mathcal{A}_{0}(\lambda)\rightarrow 0$ as $\lambda\rightarrow 0$ .

If $f\in\mathcal{R}(C^{\beta})$ for $\beta\geq 0$ and $\beta+\theta>0$ , then

by the dominated convergence theorem. For any $\theta>0$ , we have

Let $f=f_{R}+f_{N}$ where $f_{R}\in\overline{\mathcal{R}(C^{\theta})}$ , $f_{N}\in\overline{\mathcal{R}(C^{\theta})}^{\perp}$ if $0<\theta\leq 1$ and $f_{R}\in\overline{\mathcal{R}(C)}$ , $f_{N}\in\overline{\mathcal{R}(C)}^{\perp}$ if $\theta\geq 1$ . Then

as $\lambda\rightarrow 0$ . (ii) If $f\in\mathcal{R}(C^{\beta})$ , then there exists $g\in H$ such that $f=C^{\beta}g$ . This yields

On the other hand, for $\beta+\theta\geq 1$ , we have

Using the above in (A.3) yields the result.

A.4 Bound on the Norm of Certain Operators and Functions

The following result is used in many places throughout the paper. We would like to highlight that special cases of this result are known, e.g., see the proof of Theorem 4 in Caponnetto and Vito (2007) where concentration inequalites are obtained for the quantities in Proposition A.4 using Bernstein’s inequality. Here, we provide asymptotic statements using Chebyshev’s inequality.

$\|R^{\alpha}(R+\lambda I)^{-\theta}\|\leq\lambda^{\alpha-\theta}$ .

$\|\hat{R}^{\alpha}(\hat{R}+\lambda I)^{-\theta}\|\leq\lambda^{\alpha-\theta}$ .

where the last inequality follows from (iii). Using (A.5) in (A.4), we obtain

The result therefore follows by an application of Chebyshev’s inequality. (v) We use the idea in Step 2.1 of the proof of Theorem 4 in Caponnetto and Vito (2007), where $R^{\alpha}(\hat{R}+\lambda I)^{-1}$ is written equivalently as follows: Note that $\hat{R}+\lambda I=(\hat{R}-R)+(R+\lambda I)$ , which implies

and so $R^{\alpha}(\hat{R}+\lambda I)^{-1}=R^{\alpha}(R+\lambda I)^{-1}\left(I-(R-\hat{R})(R+\lambda I)^{-1}\right)^{-1}$ . Using the Von Neumann series representation, we have

A.5 Interpolation Space

In this section, we briefly recall the definition of interpolation spaces of the real method. To this end, let $E_{0}$ and $E_{1}$ be two arbitrary Banach spaces that are continuously embedded in some topological (Hausdorff) vector space $\mathcal{E}$ . Then, for $x\in E_{0}+E_{1}:=\{x_{0}+x_{1}:x_{0}\in E_{0},\,x_{1}\in E_{1}\}$ and $t>0$ , the $K$ -functional of the real interpolation method (see Bennett and Sharpley, 1988, Definition 1.1, p. 293) is defined by

Suppose $E$ and $F$ are two Banach spaces that satisfy $F\hookrightarrow E$ (i.e., $F\subset E$ and the inclusion operator $\text{id}:F\rightarrow E$ is continuous), then the $K$ -functional reduces to

The $K$ -functional can be used to define interpolation norms, for $0<\theta<1$ , $1\leq s\leq\infty$ and $x\in E_{0}+E_{1}$ , as

Moreover, the corresponding interpolation spaces (Bennett and Sharpley, 1988, Definition 1.7, p. 299) are defined as

Appendix B Appendix: Miscellaneous Results

In this appendix, we present the proofs of some claims that we made in Sections 1, 4 and 5.

The following result provides a relationship between Fisher and Kullback-Leibler divergences.

as $\|x\|_{2}\rightarrow\infty$ for all $i\in[d]$ where $\alpha=1-d$ . Then

Under the conditions mentioned on $p_{t}$ and $q_{t}$ , it can be shown that

See Theorem 1 in Lyu (2009) for a proof. The above identity is a simple generalization of de Bruijn’s identity that relates the Fisher information to the derivative of the Shannon entropy (see Cover and Thomas, 1991, Theorem 16.6.2). Integrating w.r.t. $t$ on both sides of (B.2), we obtain $KL(p_{t}\|q_{t})\Big{|}^{\infty}_{t=0}=-\int^{\infty}_{0}J(p_{t}\|q_{t})\,dt$ which yields the equality in (B.1) as $KL(p_{t}\|q_{t})\rightarrow 0$ as $t\rightarrow\infty$ and $KL(p_{t}\|q_{t})\rightarrow KL(p\|q)$ as $t\rightarrow 0$ .

To handle the case of unbounded $k$ , in the following, we assume that there exists a positive constant $M$ such that $\|f_{0}\|_{\mathcal{H}}\leq M$ , so that an estimator of $f_{0}$ can be constructed as

where $\hat{J}_{\lambda}$ is defined in Theorem 4(iv). This modification yields a valid estimator $p_{\breve{f}_{\lambda,n}}$ as long as $k$ satisfies $\int_{\Omega}e^{M\sqrt{k(x,x)}}q_{0}(x)\,dx<\infty,$ since this implies $\breve{f}_{\lambda,n}\in\mathcal{F}$ . The construction of $\breve{f}_{\lambda,n}$ requires the knowledge of $M$ , however, which we assume is known a priori. Using the representer theorem in RKHS, it can be shown (see Section B.2.1) that

where $\breve{\delta}$ and $\breve{\bm{\beta}}$ are obtained by solving the following quadratically constrained quadratic program (QCQP),

with $\Delta:=(\bm{h},\|\hat{\xi}\|^{2}_{\mathcal{H}})$ , $\Theta:=(\bm{\beta},\delta)$ and $\bm{K}$ , $\bm{H}$ being defined in the proof of Theorem 5 and the remark following it. The following result investigates the consistency and convergence rates for $p_{\breve{f}_{\lambda,n}}$ .

Let $M\geq\|f_{0}\|_{\mathcal{H}}$ be a fixed constant, and $\breve{f}_{n,\lambda}$ be a clipped estimator given by (B.3). Suppose (A)–(D) with $\varepsilon=2$ hold. Let $\emph{supp}(q_{0})=\Omega$ and $\int_{\Omega}e^{M\sqrt{k(x,x)}}q_{0}(x)\,dx<\infty$ . Define $\eta(x)=\sqrt{k(x,x)}e^{M\sqrt{k(x,x)}}$ . Then, as $\lambda\sqrt{n}\rightarrow\infty,\,\lambda\rightarrow 0\,\,\text{and}\,\,n\rightarrow\infty$ ,

$\|p_{\breve{f}_{\lambda,n}}-p_{0}\|_{L^{1}(\Omega)}\rightarrow 0$ , $KL(p_{0}\|p_{\breve{f}_{\lambda,n}})\rightarrow 0$ if $\eta\in L^{1}(\Omega,q_{0})$ ;

for $1<r\leq\infty$ , $\|p_{\breve{f}_{\lambda,n}}-p_{0}\|_{L^{r}(\Omega)}\rightarrow 0$ if $\eta q_{0}\in L^{1}(\Omega)\cap L^{r}(\Omega)$ and $e^{M\sqrt{k(\cdot,\cdot)}}q_{0}\in L^{r}(\Omega)$ ;

$h(p_{\breve{f}_{\lambda,n}},p_{0})\rightarrow 0$ if $\sqrt{k(\cdot,\cdot)}\eta\in L^{1}(\Omega,q_{0})$ ;

$J(p_{0}\|p_{\breve{f}_{\lambda,n}})\rightarrow 0$ .

In addition, if $f_{0}\in\mathcal{R}(C^{\beta})$ for some $\beta>0$ , then $\|p_{\breve{f}_{\lambda,n}}-p_{0}\|_{L^{r}(\Omega)}=O_{p_{0}}(\theta_{n}),\,h(p_{0},p_{\breve{f}_{\lambda,n}})=O_{p_{0}}(\theta_{n}),\,KL(p_{0}\|p_{\breve{f}_{\lambda,n}})=O_{p_{0}}(\theta_{n})\,\,\text{and}\,\,J(p_{0}\|p_{\breve{f}_{\lambda,n}})=O_{p_{0}}(\theta^{2}_{n})$ where $\theta_{n}:=n^{-\min\left\{\frac{1}{4},\frac{\beta}{2(\beta+1)}\right\}}$ with $\lambda=n^{-\max\left\{\frac{1}{4},\frac{1}{2(\beta+1)}\right\}}$ assuming the respective conditions in (i)-(iii) above hold.

For any $x\in\Omega$ , since $|f_{0}(x)|\leq\|f_{0}\|_{\mathcal{H}}\sqrt{k(x,x)}\leq M\sqrt{k(x,x)}$ and $|\breve{f}_{\lambda,n}(x)|\leq M\sqrt{k(x,x)}$ , we have

where we used the fact that $|e^{x}-e^{y}|\leq e^{a}|x-y|$ for $x,y\in[-a,a]$ and $\eta(x):=\sqrt{k(x,x)}e^{M\sqrt{k(x,x)}}$ . In the following, we obtain bounds for $\bigl{\|}p_{\breve{f}_{\lambda,n}}-p_{0}\bigr{\|}_{L^{r}(\Omega)}$ for any $1\leq r\leq\infty$ , $h(p_{\breve{f}_{\lambda,n}},p_{0})$ and $KL(p_{0}\|p_{\breve{f}_{\lambda,n}})$ in terms of $\|\breve{f}_{\lambda,n}-f_{0}\bigr{\|}_{\mathcal{H}}$ . Define $B(f):=\int_{\Omega}e^{f}q_{0}\,dx$ . Since $k$ satisfies $\int_{\Omega}e^{M\sqrt{k(x,x)}}q_{0}(x)\,dx<\infty$ , then it is clear that $\breve{f}_{\lambda,n}\in\mathcal{F}$ as $B(\breve{f}_{\lambda,n})<\infty$ since

Similarly, it is easy to verify that $B(f_{0})<\infty$ . (i) Recalling (A.1), we have

Using (B.4), we bound $|B(\breve{f}_{\lambda,n})-B(f_{0})|$ as

Also for any $f\in\mathcal{H}$ with $\|f\|_{\mathcal{H}}\leq M$ , we have $B(f)\geq\int_{\Omega}e^{-M\sqrt{k(x,x)}}q_{0}(x)\,dx=:\theta$ , where $\theta>0$ . Again using (B.4), we have

and $\|e^{f_{0}}q_{0}\|_{L^{r}(\Omega)}\leq\|e^{M\sqrt{k(x,x)}}q_{0}\|_{L^{r}(\Omega)}$ . Therefore,

where the above inequality is obtained by carrying out and simplifying the decomposition as in (A.1). Using (B.4), we therefore have

(iv) As $f_{0},\,\breve{f}_{\lambda,n}\in\mathcal{F}$ , by Theorem 4, we obtain $J(p_{0}\|p_{\breve{f}_{\lambda,n}})=\frac{1}{2}\|\sqrt{C}(\breve{f}_{\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}\leq\frac{1}{2}\|\sqrt{C}\|^{2}\|\breve{f}_{\lambda,n}-f_{0}\|^{2}_{\mathcal{H}}.$ Note that we have bounded the various distances between $p_{\breve{f}_{\lambda,n}}$ and $p_{0}$ in terms of $\|\breve{f}_{\lambda,n}-f_{0}\|_{\mathcal{H}}$ . Since $\breve{f}_{\lambda,n}=f_{\lambda,n}$ with probability converging to 1, the assertions on consistency are proved by Theorem 6(i) in combination with Lemma 14—as we did not explicitly assume $f_{0}\in\overline{\mathcal{R}(C)}$ —and the rates follow from Theorem 6(iii).

The following observations can be made while comparing the scenarios of using bounded vs. unbounded kernels in the problem of estimating $p_{0}$ through Theorems 7 and B.2. First, the consistency results in $L^{r}$ , Hellinger and KL distances are the same but for additional integrability conditions on $k$ and $q_{0}$ . The additional integrability conditions are not too difficult to hold in practice as they involve $k$ and $q_{0}$ which can be chosen appropriately. However, the unbounded situation in Theorem B.2 requires the knowledge of $M$ which is usually not known. On the other hand, the consistency result in $J$ in Theorem B.2 is slightly weaker than in Theorem 7. This may be an artifact of our analysis as we are not able to adapt the bounding technique used in the proof of Theorem 7 to bound $J(p_{0}\|p_{\breve{f}_{\lambda,n}})=\frac{1}{2}\|\sqrt{C}(\breve{f}_{\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}$ as it critically depends on the boundedness of $k$ . Therefore, we used a trivial bound of $J(p_{0}\|p_{\breve{f}_{\lambda,n}})=\frac{1}{2}\|\sqrt{C}(\breve{f}_{\lambda,n}-f_{0})\|^{2}_{\mathcal{H}}\leq\frac{1}{2}\|\sqrt{C}\|^{2}\|\breve{f}_{\lambda,n}-f_{0}\|^{2}_{\mathcal{H}}$ , which yields the result through Theorem 6(i). Due to the same reason, we also obtain a slower rate of convergence in $J$ . Second, the rate of convergence in KL is slower than in Theorem B.2, which again may be an artifact of our analysis. The convergence rate for KL in Theorem 7 is based on the application of Theorem 6(ii) in Lemma A.1, where the bound on KL in Lemma A.1 critically uses the boundedness to upper bound KL in terms of squared Hellinger distance.

Any $f\in\mathcal{H}$ can be decomposed as $f=f_{\|}+f_{\perp}$ where

which is a closed subset of $\mathcal{H}$ and $f_{\perp}\in\mathcal{H}^{\perp}_{\|}:=\left\{g\in\mathcal{H}:\langle g,h\rangle_{\mathcal{H}}=0,\,\forall\,h\in\mathcal{H}_{\|}\right\}$ so that $\mathcal{H}=\mathcal{H}_{\|}\oplus\mathcal{H}^{\perp}_{\|}$ . Since the objective function in (B.3) matches with the one in Theorem 5, using the above decomposition in (B.3), it is easy to verify that $\hat{J}$ depends only on $f_{\|}\in\mathcal{H}_{\|}$ so that (B.3) reduces to

and $\breve{f}_{\lambda,n}=\breve{f}^{\|}_{\lambda,n}+\breve{f}^{\perp}_{\lambda,n}$ . Since $f_{\|}$ is of the form in (14), using it in (B.5), it is easy to show that $\hat{J}_{\lambda}(f_{\|})+\frac{\lambda}{2}\|f_{\|}\|^{2}_{\mathcal{H}}=\frac{1}{2}\Theta^{T}\bm{H}\Theta+\Theta^{T}\Delta$ . Similarly, it can be shown that $\|f_{\|}\|^{2}_{\mathcal{H}}=\Theta^{T}\bm{K}\Theta$ . Since $f_{\perp}$ appears in (B.5) only through $\|f_{\perp}\|^{2}_{\mathcal{H}}$ , (B.5) reduces to

where $\breve{f}^{\|}_{\lambda,n}$ is constructed as in (14) using $\Theta_{\|}$ and $\breve{f}^{\perp}_{\lambda,n}$ is such that $\|\breve{f}^{\perp}_{\lambda.n}\|^{2}_{\mathcal{H}}=c_{\perp}$ . The necessary and sufficient conditions for the optimality of $(\Theta_{\|},c_{\perp})$ is given by the following Karush-Kuhn-Tucker conditions,

Combining the dual feasibility and stationary conditions, we have $\eta=\tau-\frac{\lambda}{2}\geq 0$ , i.e., $\tau\geq\frac{\lambda}{2}$ . Using this in the complementary slackness involving $\tau$ and $c_{\perp}$ , it follows that $c_{\perp}=0$ . Since $\|\breve{f}^{\perp}_{\lambda,n}\|^{2}=c_{\perp}$ , we have $\breve{f}^{\perp}_{\lambda,n}=0$ , i.e., $\breve{f}_{\lambda,n}$ is completely determined by $\breve{f}^{\|}_{\lambda,n}$ . Therefore $\breve{f}^{\|}_{\lambda,n}$ is of the form in (14) and (B.6) reduces to a quadratically constrained quadratic program.

where $\mathcal{R}(C^{0}):=\mathcal{H}$ , and the spaces $\mathcal{R}(C^{\beta})$ and $\left[\mathcal{R}(C^{\lfloor\beta\rfloor}),\mathcal{R}(C^{\lceil\beta\rceil})\right]_{\beta-\lfloor\beta\rfloor,2}$ have equivalent norms.

To prove Proposition B.3, we need the following result which we quote from Steinwart and Scovel (2012, Lemma 6.3) (also see Tartar, 2007, Lemma 23.1) that interpolates $L^{2}$ -spaces whose underlying measures are absolutely continuous with respect to a measure $\nu$ .

Let $\nu$ be a measure on a measurable space $\Theta$ and $w_{0}:\Theta\rightarrow[0,\infty)$ and $w_{1}:\Theta\rightarrow[0,\infty)$ be measurable functions. For $0<\beta<1$ , define $w_{\beta}:=w^{1-\beta}_{0}w^{\beta}_{1}$ . Then we have

and the norms on these two spaces are equivalent. Moreover, this result still holds for weights $w_{0}:\Theta\rightarrow(0,\infty)$ and $w_{1}:\Theta\rightarrow[0,\infty]$ , if one uses the convention $0\cdot\infty:=0$ in the definition of the weighted spaces.

Proof of Proposition B.3. The proof is based on the ideas used in the proof of Theorem 4.6 in Steinwart and Scovel (2012). Recall that by the Hilbert-Schmidt theorem, $C$ has the following representation,

where $\theta_{i}:=\phi_{i}$ if $i\in I$ and $\theta_{i}:=\psi_{i}$ if $i\in J$ with $a_{i}:=\langle f,\theta_{i}\rangle_{\mathcal{H}}$ . Let $\beta>0$ . By definition, $g\in\mathcal{R}(C^{\beta})$ is equivalent to $\exists h\in\mathcal{H}$ such that $g=C^{\beta}h$ , i.e.,

where $\alpha:=(\alpha_{i})_{i\in I}$ . Let us equip this space with the bilinear form

It is easy to verify that $(\alpha^{\beta}_{i}\phi_{i})_{i\in I}$ is an ONB of $\mathcal{R}(C^{\beta})$ . Also since $\mathcal{R}(C^{\beta_{1}})\subset\mathcal{R}(C^{\beta_{2}})$ for $0<\beta_{2}<\beta_{1}<\infty$ and $\text{id}:\mathcal{R}(C^{\beta_{1}})\rightarrow\mathcal{R}(C^{\beta_{2}})$ is continuous, i.e., for any $g\in\mathcal{R}(C^{\beta_{1}})$ ,

and so $\mathcal{R}(C^{\beta_{1}})\hookrightarrow\mathcal{R}(C^{\beta_{2}})$ . Similarly, we can show that $\mathcal{R}(C)\hookrightarrow\mathcal{H}$ . In the following, we first prove the result for $0<\beta<1$ and then for $\beta>1$ .

(a) $0<\beta<1$ : For any $f\in\mathcal{H}$ and $g\in\mathcal{R}(C)$ , we have

where we define $c_{i}:=0$ for $i\in J$ . For $t>0$ , we find

From this we immediately obtain the equivalence

where $0<\beta<1$ . Applying the second part of Lemma B.4 to the counting measure on $I\cup J$ yields

from which we obtain the following equivalence

For $d=1$ , this implies $(\partial_{j}f)p=0$ a.e. and so $\|f\|_{W_{2}}=0$ .

Acknowledgments

Theorem A.2 and the proof are due to an anonymous reviewer. We thank Dougal Sutherland for a careful reading of the paper which helped to fix minor errors. A part of the work was carried out while BKS was a Research Fellow in the Statistical Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge. BKS thanks Richard Nickl for many valuable comments and insightful discussions. KF is supported in part by MEXT Grant-in-Aid for Scientific Research on Innovative Areas 25120012.