Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Arthur Jacot, Franck Gabriel, Clément Hongler

Introduction

Artificial neural networks (ANNs) have achieved impressive results in numerous areas of machine learning. While it has long been known that ANNs can approximate any function with sufficiently many hidden neurons (11; 14), it is not known what the optimization of ANNs converges to. Indeed the loss surface of neural networks optimization problems is highly non-convex: it has a high number of saddle points which may slow down the convergence (5). A number of results (3; 17; 18) suggest that for wide enough networks, there are very few “bad” local minima, i.e. local minima with much higher cost than the global minimum. More recently, the investigation of the geometry of the loss landscape at initialization has been the subject of a precise study (12). The analysis of the dynamics of training in the large-width limit for shallow networks has seen recent progress as well (15). To the best of the authors knowledge, the dynamics of deep networks has however remained an open problem until the present paper: see the contributions section below.

A particularly mysterious feature of ANNs is their good generalization properties in spite of their usual over-parametrization (20). It seems paradoxical that a reasonably large neural network can fit random labels, while still obtaining good test accuracy when trained on real data (23). It can be noted that in this case, kernel methods have the same properties (1).

In the infinite-width limit, ANNs have a Gaussian distribution described by a kernel (16; 4; 7; 13; 6). These kernels are used in Bayesian inference or Support Vector Machines, yielding results comparable to ANNs trained with gradient descent (2; 13). We will see that in the same limit, the behavior of ANNs during training is described by a related kernel, which we call the neural tangent network (NTK).

We study the network function $f_{\theta}$ of an ANN, which maps an input vector to an output vector, where $\theta$ is the vector of the parameters of the ANN. In the limit as the widths of the hidden layers tend to infinity, the network function at initialization, $f_{\theta}$ converges to a Gaussian distribution (16; 4; 7; 13; 6).

In this paper, we investigate fully connected networks in this infinite-width limit, and describe the dynamics of the network function $f_{\theta}$ during training:

During gradient descent, we show that the dynamics of $f_{\theta}$ follows that of the so-called kernel gradient descent in function space with respect to a limiting kernel, which only depends on the depth of the network, the choice of nonlinearity and the initialization variance.

The convergence properties of ANNs during training can then be related to the positive-definiteness of the infinite-width limit NTK. In the case when the dataset is supported on a sphere, we prove this positive-definiteness using recent results on dual activation functions (4). The values of the network function $f_{\theta}$ outside the training set is described by the NTK, which is crucial to understand how ANN generalize.

For a least-squares regression loss, the network function $f_{\theta}$ follows a linear differential equation in the infinite-width limit, and the eigenfunctions of the Jacobian are the kernel principal components of the input data. This shows a direct connection to kernel methods and motivates the use of early stopping to reduce overfitting in the training of ANNs.

Finally we investigate these theoretical results numerically for an artificial dataset (of points on the unit circle) and for the MNIST dataset. In particular we observe that the behavior of wide ANNs is close to the theoretical limit.

Neural networks

In this paper, we assume that the input distribution $p^{in}$ is the empirical distribution on a finite dataset $x_{1},...,x_{N}$ , i.e the sum of Dirac measures $\frac{1}{N}\sum_{i=0}^{N}\delta_{x_{i}}$ .

where the nonlinearity $\sigma$ is applied entrywise. The scalar $\beta>0$ is a parameter which allows us to tune the influence of the bias on the training.

Kernel gradient

The kernel $K$ is positive definite with respect to $||\cdot||_{p^{in}}$ if $||f||_{p^{in}}>0\implies||f||_{K}>0$ .

The kernel gradient $\nabla_{K}C|_{f_{0}}\in\mathcal{F}$ is defined as $\Phi_{K}\left(\partial_{f}^{in}C|_{f_{0}}\right)$ . In contrast to $\partial_{f}^{in}C$ which is only defined on the dataset, the kernel gradient generalizes to values $x$ outside the dataset thanks to the kernel $K$ :

A time-dependent function $f(t)$ follows the kernel gradient descent with respect to $K$ if it satisfies the differential equation

During kernel gradient descent, the cost $C(f(t))$ evolves as

Convergence to a critical point of $C$ is hence guaranteed if the kernel $K$ is positive definite with respect to $||\cdot||_{p^{in}}$ : the cost is then strictly decreasing except at points such that $||d|_{f(t)}||_{p^{in}}=0$ . If the cost is convex and bounded from below, the function $f(t)$ therefore converges to a global minimum as $t\to\infty$ .

As a starting point to understand the convergence of ANN gradient descent to kernel gradient descent in the infinite-width limit, we introduce a simple model, inspired by the approach of (19).

A kernel $K$ can be approximated by a choice of $P$ random functions $f^{(p)}$ sampled independently from any distribution on $\mathcal{F}$ whose (non-centered) covariance is given by the kernel $K$ :

The partial derivatives of the parametrization are given by

Optimizing the cost $C\circ F^{lin}$ through gradient descent, the parameters follow the ODE:

As a result the function $f_{\theta(t)}^{lin}$ evolves according to

Neural tangent kernel

For ANNs trained using gradient descent on the composition $C\circ F^{(L)}$ , the situation is very similar to that studied in the Section 3.1. During training, the network function $f_{\theta}$ evolves along the (negative) kernel gradient

with respect to the neural tangent kernel (NTK)

However, in contrast to $F^{lin}$ , the realization function $F^{(L)}$ of ANNs is not linear. As a consequence, the derivatives $\partial_{\theta_{p}}F^{(L)}(\theta)$ and the neural tangent kernel depend on the parameters $\theta$ . The NTK is therefore random at initialization and varies during training, which makes the analysis of the convergence of $f_{\theta}$ more delicate.

In the next subsections, we show that, in the infinite-width limit, the NTK becomes deterministic at initialization and stays constant during training. Since $f_{\theta}$ at initialization is Gaussian in the limit, the asymptotic behavior of $f_{\theta}$ during training can be explicited in the function space $\mathcal{F}$ .

As observed in (16; 4; 7; 13; 6), the output functions $f_{\theta,i}$ for $i=1,...,n_{L}$ tend to iid Gaussian processes in the infinite-width limit (a proof in our setup is given in the appendix):

For a network of depth $L$ at initialization, with a Lipschitz nonlinearity $\sigma$ , and in the limit as $n_{1},...,n_{L-1}\to\infty$ , the output functions $f_{\theta,k}$ , for $k=1,...,n_{L}$ , tend (in law) to iid centered Gaussian processes of covariance $\Sigma^{(L)}$ , where $\Sigma^{(L)}$ is defined recursively by:

taking the expectation with respect to a centered Gaussian process $f$ of covariance $\Sigma^{(L)}$ .

Strictly speaking, the existence of a suitable Gaussian measure with covariance $\Sigma^{(L)}$ is not needed: we only deal with the values of $f$ at $x,x^{\prime}$ (the joint measure on $f(x),f(x^{\prime})$ is simply a Gaussian vector in 2D). For the same reasons, in the proof of Proposition 1 and Theorem 1, we will freely speak of Gaussian processes without discussing their existence.

The first key result of our paper (proven in the appendix) is the following: in the same limit, the Neural Tangent Kernel (NTK) converges in probability to an explicit deterministic limit.

For a network of depth $L$ at initialization, with a Lipschitz nonlinearity $\sigma$ , and in the limit as the layers width $n_{1},...,n_{L-1}\to\infty$ , the NTK $\Theta^{(L)}$ converges in probability to a deterministic limiting kernel:

taking the expectation with respect to a centered Gaussian process $f$ of covariance $\Sigma^{(L)}$ , and where $\dot{\sigma}$ denotes the derivative of $\sigma$ .

By Rademacher’s theorem, $\dot{\sigma}$ is defined everywhere, except perhaps on a set of zero Lebesgue measure.

Note that the limiting $\Theta^{(L)}_{\infty}$ only depends on the choice of $\sigma$ , the depth of the network and the variance of the parameters at initialization (which is equal to $1$ in our setting).

2 Training

Our second key result is that the NTK stays asymptotically constant during training. This applies for a slightly more general definition of training: the parameters are updated according to a training direction $d_{t}\in\mathcal{F}$ :

In the case of gradient descent, $d_{t}=-d|_{f_{\theta(t)}}$ (see Section 3), but the direction may depend on another network, as is the case for e.g. Generative Adversarial Networks (10). We only assume that the integral $\int_{0}^{T}\|d_{t}\|_{p^{in}}dt$ stays stochastically bounded as the width tends to infinity, which is verified for e.g. least-squares regression, see Section 5.

Assume that $\sigma$ is a Lipschitz, twice differentiable nonlinearity function, with bounded second derivative. For any $T$ such that the integral $\int_{0}^{T}\|d_{t}\|_{p^{in}}dt$ stays stochastically bounded, as $n_{1},...,n_{L-1}\to\infty$ , we have, uniformly for $t\in[0,T]$ ,

As a consequence, in this limit, the dynamics of $f_{\theta}$ is described by the differential equation

As the proof of the theorem (in the appendix) shows, the variation during training of the individual activations in the hidden layers shrinks as their width grows. However their collective variation is significant, which allows the parameters of the lower layers to learn: in the formula of the limiting NTK $\Theta^{(L+1)}_{\infty}(x,x^{\prime})$ in Theorem 1, the second summand $\Sigma^{(L+1)}$ represents the learning due to the last layer, while the first summand represents the learning performed by the lower layers.

As discussed in Section 3, the convergence of kernel gradient descent to a critical point of the cost $C$ is guaranteed for positive definite kernels. The limiting NTK is positive definite if the span of the derivatives $\partial_{\theta_{p}}F^{(L)}$ , $p=1,...,P$ becomes dense in $\mathcal{F}$ w.r.t. the $p^{in}$ -norm as the width grows to infinity. It seems natural to postulate that the span of the preactivations of the last layer (which themselves appear in $\partial_{\theta_{p}}F^{(L)}$ , corresponding to the connection weights of the last layer) becomes dense in $\mathcal{F}$ , for a large family of measures $p^{in}$ and nonlinearities (see e.g. (11; 14) for classical theorems about ANNs and approximation). In the case when the dataset is supported on a sphere, the positive-definiteness of the limiting NTK can be shown using Gaussian integration techniques and existing positive-definiteness criteria, as given by the following proposition, proven in Appendix A.4:

Least-squares regression

Given a goal function $f^{*}$ and input distribution $p^{in}$ , the least-squares regression cost is

Theorems 1 and 2 apply to an ANN trained on such a cost. Indeed the norm of the training direction $\|d(f)\|_{p^{in}}=\|f^{*}-f\|_{p^{in}}$ is strictly decreasing during training, bounding the integral. We are therefore interested in the behavior of a function $f_{t}$ during kernel gradient descent with a kernel $K$ (we are of course especially interested in the case $K=\Theta^{(L)}_{\infty}\otimes Id_{n_{L}}$ ):

The solution of this differential equation can be expressed in terms of the map $\Pi:f\mapsto\Phi_{K}\left(\left<f,\cdot\right>_{p^{in}}\right)$ :

where $e^{-t\Pi}=\sum_{k=0}^{\infty}\frac{(-t)^{k}}{k!}\Pi^{k}$ is the exponential of $-t\Pi$ . If $\Pi$ can be diagonalized by eigenfunctions $f^{(i)}$ with eigenvalues $\lambda_{i}$ , the exponential $e^{-t\Pi}$ has the same eigenfunctions with eigenvalues $e^{-t\lambda_{i}}$ .

For a finite dataset $x_{1},...,x_{N}$ of size $N$ , the map $\Pi$ takes the form

The map $\Pi$ has at most $Nn_{L}$ positive eigenfunctions, and they are the kernel principal components $f^{(1)},...,f^{(Nn_{L})}$ of the data with respect to to the kernel $K$ (21; 22). The corresponding eigenvalues $\lambda_{i}$ is the variance captured by the component.

Decomposing the difference $(f^{*}-f_{0})=\Delta^{0}_{f}+\Delta^{1}_{f}+...+\Delta^{Nn_{L}}_{f}$ along the eigenspaces of $\Pi$ , the trajectory of the function $f_{t}$ reads

where $\Delta^{0}_{f}$ is in the kernel (null-space) of $\Pi$ and $\Delta^{i}_{f}\propto f^{(i)}$ .

The above decomposition can be seen as a motivation for the use of early stopping. The convergence is indeed faster along the eigenspaces corresponding to larger eigenvalues $\lambda_{i}$ . Early stopping hence focuses the convergence on the most relevant kernel principal components, while avoiding to fit the ones in eigenspaces with lower eigenvalues (such directions are typically the ‘noisier’ ones: for instance, in the case of the RBF kernel, lower eigenvalues correspond to high frequency functions).

with the $Nn_{l}$ -vectors $\kappa_{x,k}$ , $y^{*}$ and $y_{0}$ given by

The first term, the mean, has an important statistical interpretation: it is the maximum-a-posteriori (MAP) estimate given a Gaussian prior on functions $f_{k}\sim\mathcal{N}(0,\Theta^{(L)}_{\infty})$ and the conditions $f_{k}(x_{i})=f^{*}_{k}(x_{i})$ . Equivalently, it is equal to the kernel ridge regression (22) as the regularization goes to zero ( $\lambda\to 0$ ). The second term is a centered Gaussian whose variance vanishes on the points of the dataset.

Numerical experiments

In the following numerical experiments, fully connected ANNs of various widths are compared to the theoretical infinite-width limit. We choose the size of the hidden layers to all be equal to the same value $n:=n_{1}=...=n_{L-1}$ and we take the ReLU nonlinearity $\sigma(x)=\max(0,x)$ .

In the first two experiments, we consider the case $n_{0}=2$ . Moreover, the input elements are taken on the unit circle. This can be motivated by the structure of high-dimensional data, where the centered data points often have roughly the same norm The classical example is for data following a Gaussian distribution $\mathcal{N}(0,Id_{n_{0}})$ : as the dimension $n_{0}$ grows, all data points have approximately the same norm $\sqrt{n_{0}}$ ..

In all experiments, we took $n_{L}=1$ (note that by our results, a network with $n_{L}$ outputs behaves asymptotically like $n_{L}$ networks with scalar outputs trained independently). Finally, the value of the parameter $\beta$ is chosen as $0.1$ , see Remark 1.

The first experiment illustrates the convergence of the NTK $\Theta^{(L)}$ of a network of depth $L=4$ for two different widths $n=500,10000$ . The function $\Theta^{(4)}(x_{0},x)$ is plotted for a fixed $x_{0}=(1,0)$ and $x=(cos(\gamma),sin(\gamma))$ on the unit circle in Figure 2. To observe the distribution of the NTK, $10$ independent initializations are performed for both widths. The kernels are plotted at initialization $t=0$ and then after $200$ steps of gradient descent with learning rate $1.0$ (i.e. at $t=200$ ). We approximate the function $f^{*}(x)=x_{1}x_{2}$ with a least-squares cost on random $\mathcal{N}(0,1)$ inputs.

For the wider network, the NTK shows less variance and is smoother. It is interesting to note that the expectation of the NTK is very close for both networks widths. After $200$ steps of training, we observe that the NTK tends to “inflate”. As expected, this effect is much less apparent for the wider network ( $n=10000$ ) where the NTK stays almost fixed, than for the smaller network ( $n=500$ ).

2 Kernel regression

For a regression cost, the infinite-width limit network function $f_{\theta(t)}$ has a Gaussian distribution for all times $t$ and in particular at convergence $t\to\infty$ (see Section 5). We compared the theoretical Gaussian distribution at $t\to\infty$ to the distribution of the network function $f_{\theta(T)}$ of a finite-width network for a large time $T=1000$ . For two different widths $n=50,1000$ and for $10$ random initializations each, a network is trained on a least-squares cost on $4$ points of the unit circle for $1000$ steps with learning rate $1.0$ and then plotted in Figure 2.

We also approximated the kernels $\Theta_{\infty}^{(4)}$ and $\Sigma^{(4)}$ using a large-width network ( $n=10000$ ) and used them to calculate and plot the 10th, 50th and 90-th percentiles of the $t\to\infty$ limiting Gaussian distribution.

The distributions of the network functions are very similar for both widths: their mean and variance appear to be close to those of the limiting distribution $t\to\infty$ . Even for relatively small widths ( $n=50$ ), the NTK gives a good indication of the distribution of $f_{\theta(t)}$ as $t\to\infty$ .

3 Convergence along a principal component

We now illustrate our result on the MNIST dataset of handwritten digits made up of grayscale images of dimension $28\times 28$ , yielding a dimension of $n_{0}=784$ .

We computed the first 3 principal components of a batch of $N=512$ digits with respect to the NTK of a high-width network $n=10000$ (giving an approximation of the limiting kernel) using a power iteration method. The respective eigenvalues are $\lambda_{1}=0.0457$ , $\lambda_{2}=0.00108$ and $\lambda_{3}=0.00078$ . The kernel PCA is non-centered, the first component is therefore almost equal to the constant function, which explains the large gap between the first and second eigenvaluesIt can be observed numerically, that if we choose $\beta=1.0$ instead of our recommended $0.1$ , the gap between the first and the second principal component is about ten times bigger, which makes training more difficult.. The next two components are much more interesting as can be seen in Figure 3(a), where the batch is plotted with $x$ and $y$ coordinates corresponding to the 2nd and 3rd components.

We have seen in Section 5 how the convergence of kernel gradient descent follows the kernel principal components. If the difference at initialization $f_{0}-f^{*}$ is equal (or proportional) to one of the principal components $f^{(i)}$ , then the function will converge along a straight line (in the function space) to $f^{*}$ at an exponential rate $e^{-\lambda_{i}t}$ .

We tested whether ANNs of various widths $n=100,1000,10000$ behave in a similar manner. We set the goal of the regression cost to $f^{*}=f_{\theta(0)}+0.5f^{(2)}$ and let the network converge. At each time step $t$ , we decomposed the difference $f_{\theta(t)}-f^{*}$ into a component $g_{t}$ proportional to $f^{(2)}$ and another one $h_{t}$ orthogonal to $f^{(2)}$ . In the infinite-width limit, the first component decays exponentially fast $||g_{t}||_{p^{in}}=0.5e^{-\lambda_{2}t}$ while the second is null ( $h_{t}=0$ ), as the function converges along a straight line.

As expected, we see in Figure 3(b) that the wider the network, the less it deviates from the straight line (for each width $n$ we performed two independent trials). As the width grows, the trajectory along the 2nd principal component (shown in Figure 3(c)) converges to the theoretical limit shown in blue.

A surprising observation is that smaller networks appear to converge faster than wider ones. This may be explained by the inflation of the NTK observed in our first experiment. Indeed, multiplying the NTK by a factor $a$ is equivalent to multiplying the learning rate by the same factor. However, note that since the NTK of large-width network is more stable during training, larger learning rates can in principle be taken. One must hence be careful when comparing the convergence speed in terms of the number of steps (rather than in terms of the time $t$ ): both the inflation effect and the learning rate must be taken into account.

Conclusion

This paper introduces a new tool to study ANNs, the Neural Tangent Kernel (NTK), which describes the local dynamics of an ANN during gradient descent. This leads to a new connection between ANN training and kernel methods: in the infinite-width limit, an ANN can be described in the function space directly by the limit of the NTK, an explicit constant kernel $\Theta^{(L)}_{\infty}$ , which only depends on its depth, nonlinearity and parameter initialization variance. More precisely, in this limit, ANN gradient descent is shown to be equivalent to a kernel gradient descent with respect to $\Theta^{(L)}_{\infty}$ . The limit of the NTK is hence a powerful tool to understand the generalization properties of ANNs, and it allows one to study the influence of the depth and nonlinearity on the learning abilities of the network. The analysis of training using NTK allows one to relate convergence of ANN training with the positive-definiteness of the limiting NTK and leads to a characterization of the directions favored by early stopping methods.

Acknowledgements

The authors thank K. Kytölä for many interesting discussions. The second author was supported by the ERC CG CRITICAL. The last author acknowledges support from the ERC SG Constamis, the NCCR SwissMAP, the Blavatnik Family Foundation and the Latsis Foundation.

References

Appendix A Appendix

This appendix is dedicated to proving the key results of this paper, namely Proposition 1 and Theorems 1 and 2, which describe the asymptotics of neural networks at initialization and during training.

We study the limit of the NTK as $n_{1},...,n_{L-1}\to\infty$ sequentially, i.e. we first take $n_{1}\to\infty$ , then $n_{2}\to\infty$ , etc. This leads to much simpler proofs, but our results could in principle be strengthened to the more general setting when $\min(n_{1},...,n_{L-1})\to\infty$ .

A natural choice of convergence to study the NTK is with respect to the operator norm on kernels:

where the expectation is taken over two independent $x,x^{\prime}\sim p^{in}$ . This norm depends on the input distribution $p^{in}$ . In our setting, $p^{in}$ is taken to be the empirical measure of a finite dataset of distinct samples $x_{1},...,x_{N}$ . As a result, the operator norm of $K$ is equal to the leading eigenvalue of the $Nn_{L}\times Nn_{L}$ Gram matrix $\left(K_{kk^{\prime}}(x_{i},x_{j})\right)_{k,k^{\prime}<n_{L},i,j<N}$ . In our setting, convergence in operator norm is hence equivalent to pointwise convergence of $K$ on the dataset.

It has already been observed that the output functions $f_{\theta,i}$ for $i=1,...,n_{L}$ tend to iid Gaussian processes in the infinite-width limit.

For a network of depth $L$ at initialization, with a Lipschitz nonlinearity $\sigma$ , and in the limit as $n_{1},...,n_{L-1}\to\infty$ sequentially, the output functions $f_{\theta,k}$ , for $k=1,...,n_{L}$ , tend (in law) to iid centered Gaussian processes of covariance $\Sigma^{(L)}$ , where $\Sigma^{(L)}$ is defined recursively by:

taking the expectation with respect to a centered Gaussian process $f$ of covariance $\Sigma^{(L)}$ .

We prove the result by induction. When $L=1$ , there are no hidden layers and $f_{\theta}$ is a random affine function of the form:

All output functions $f_{\theta,k}$ are hence independent and have covariance $\Sigma^{(1)}$ as needed.

conditioned on the values of $\alpha^{(L)}$ are iid centered Gaussians with covariance

By the law of large numbers, as $n_{L}\to\infty$ , this covariance tends in probability to the expectation

In particular the covariance is deterministic and hence independent of $\alpha^{(L)}$ . As a consequence, the conditioned and unconditioned distributions of $f_{\theta,i}$ are equal in the limit: they are iid centered Gaussian of covariance $\Sigma^{(L+1)}$ . ∎

In the infinite-width limit, the neural tangent kernel, which is random at initialization, converges in probability to a deterministic limit.

For a network of depth $L$ at initialization, with a Lipschitz nonlinearity $\sigma$ , and in the limit as the layers width $n_{1},...,n_{L-1}\to\infty$ sequentially, the NTK $\Theta^{(L)}$ converges in probability to a deterministic limiting kernel:

taking the expectation with respect to a centered Gaussian process $f$ of covariance $\Sigma^{(L)}$ , and where $\dot{\sigma}$ denotes the derivative of $\sigma$ .

The proof is again by induction. When $L=1$ , there is no hidden layer and therefore no limit to be taken. The neural tangent kernel is a sum over the entries of $W^{(0)}$ and those of $b^{(0)}$ :

For the first sum let us observe that by the chain rule:

By the law of large numbers, as $n_{L}\to\infty$ , this tends to its expectation which is equal to

It is then easy to see that the second part of the neural tangent kernel, the sum over $W^{(L)}$ and $b^{(L)}$ converges to $\Sigma^{(L+1)}\delta_{kk^{\prime}}$ as $n_{1},...,n_{L}\to\infty$ . ∎

A.2 Asymptotics during Training

Given a training direction $t\mapsto d_{t}\in\mathcal{F}$ , a neural network is trained in the following manner: the parameters $\theta_{p}$ are initialized as iid $\mathcal{N}(0,1)$ and follow the differential equation:

In this context, in the infinite-width limit, the NTK stays constant during training:

As a consequence, in this limit, the dynamics of $f_{\theta}$ is described by the differential equation

As in the previous theorem, the proof is by induction on the depth of the network. When $L=1$ , the neural tangent kernel does not depend on the parameters, it is therefore constant during training.

To apply the induction hypothesis, we now need to bound $\lVert\frac{1}{\sqrt{n_{L}}}W^{(L)}(t)\rVert_{op}$ . For this, we use the following lemma, which is proven in Appendix A.3 below:

From this lemma, to bound $\lVert\frac{1}{\sqrt{n_{L}}}W^{(L)}(t)\rVert_{op}$ , it is hence enough to bound $\lVert\frac{1}{\sqrt{n_{L}}}W^{(L)}(0)\rVert_{op}$ . From the law of large numbers, we obtain that the norm of each of the $n_{L+1}$ rows of $W^{(L)}(0)$ is bounded, and hence that $\lVert\frac{1}{\sqrt{n_{L}}}W^{(L)}(0)\rVert_{op}$ is bounded (keep in mind that $n_{L+1}$ is fixed, while $n_{1},\ldots,n_{L}$ grow).

From the above considerations, we can apply the induction hypothesis to the smaller network, yielding, in the limit as $n_{1},\ldots,n_{L}\to\infty$ (sequentially), that the dynamics is governed by the constant kernel $\Theta^{(L)}_{\infty}$ :

At the same time, the parameters of the last layer evolve according to

Now, observing that the operator norm of $\Phi_{\Theta_{\infty}^{(L)}}$ is equal to $||\Theta_{\infty}^{(L)}||_{op}$ , defined in the introduction of Appendix A, and using the Cauchy-Schwarz inequality, we get

where the sup norm $\lVert\cdot\rVert_{\infty}$ is defined by $\left\lVert f\right\lVert_{\infty}=\sup_{x}|f(x)|.$

To bound both quantities simultaneously, study the derivative of the quantity

where, in the first inequality, we have used that $|\dot{\sigma}|\leq c$ and, in the second inequality, that the sum $\lVert W^{(L)}_{i}(t)\rVert_{2}+||\alpha^{(L)}_{i}(t)||_{p^{in}}$ is bounded by $A(t)$ . Applying Grönwall’s Lemma, we now get

We can now use these bounds to control the variation of the NTK and to prove the theorem. To understand how the NTK evolves, we study the evolution of the derivatives with respect to the parameters. The derivatives with respect to the bias parameters of the top layer $\partial_{b^{(L)}_{j}}f_{\theta,j^{\prime}}$ are always equal to $\delta_{jj^{\prime}}$ . The derivatives with respect to the connection weights of the top layer are given by

Finally let us study the derivatives with respect to the parameters of the lower layers

Their contribution to the NTK $\Theta^{(L+1)}_{jj^{\prime}}(x,x^{\prime})$ is

By the induction hypothesis, the NTK of the smaller network $\Theta^{(L)}$ tends to $\Theta^{(L)}_{\infty}\delta_{ii^{\prime}}$ as $n_{1},...,n_{L-1}\to\infty$ . The contribution therefore becomes

A.3 A Priori Control during Training

The goal of this section is to prove Lemma 1, which is a key ingredient in the proof of Theorem 2. Let us first recall it:

At all times, the evolution of the preactivations and weights is given by:

where the layer-wise training directions $d^{(1)},\ldots,d^{(L)}$ are defined recursively by

Set $w^{(k)}(t):=\left\|\frac{1}{\sqrt{n_{k}}}W^{(k)}(t)\right\|_{op}$ and $a^{(k)}\left(t\right):=\left\|\frac{1}{\sqrt{n_{k}}}\alpha^{\left(k\right)}\left(t\right)\right\|_{p^{in}}$ . The identities of the previous step yield the following recursive bounds:

where $c$ is the Lipschitz constant of $\sigma$ . These bounds lead to

For the subnetworks NTKs we have the recursive bounds

This allows one to bound the derivative of $A\left(t\right)$ as follows:

The polynomial control we obtained on the derivative of $A\left(t\right)$ now allows one to use (a nonlinear form of, see e.g. ) Grönwall’s Lemma: we obtain that $A\left(t\right)$ stays uniformly bounded on $\left[0,\tau\right]$ for some $\tau=\tau\left(n_{1},\ldots,n_{L}\right)>0$ , and that $\tau\to T$ as $\min\left(n_{1},\ldots,n_{L}\right)\to\infty$ , owing to the $\frac{1}{\sqrt{\min\left\{1,\ldots,n_{L}\right\}}}$ in front of the polynomial. Since $A\left(t\right)$ is bounded, the differential bound on $A\left(t\right)$ gives that the derivative $\partial_{t}A\left(t\right)$ converges uniformly to on $\left[0,\tau\right]$ for any $\tau<T$ , and hence $A\left(t\right)\to A\left(0\right)$ . This concludes the proof of the lemma.

This subsection is devoted to the proof of Proposition 2, which we now recall:

A key ingredient for the proof of Proposition 2 is the following Lemma, which comes from .

If the expansion of $\mu$ in Hermite polynomials $\left(h_{i}\right)_{i\geq 0}$ is given by $\mu=\sum_{i=0}^{\infty}a_{i}h_{i}$ , we have

The other key ingredient for proving Proposition 2 is the following theorem, which is a slight reformulation of Theorem 1(b) in , which itself is a generalization of a classical result of Schönberg:

is positive-definite for any $n_{0}\geq 1$ if and only if the coefficients $b_{n}$ are strictly positive for infinitely many even and infinitely many odd integers $n$ .

With Lemma 2 and Theorem 3 above, we are now ready to prove Proposition 2.

We first decompose the limiting NTK $\Theta^{\left(L\right)}$ recursively, relate its positive-definiteness to that of the activation kernels, then show that the positive-definiteness of the activation kernels at level $2$ implies that of the higher levels, and finally show the positive-definiteness at level $2$ using Lemma 2 and Theorem 3:

Observe that for any $L\geq 1$ , using the notation of Theorem 1, we have

Note that the kernel $\dot{\Sigma}^{\left(L\right)}\Theta^{\left(L\right)}$ is positive semi-definite, being the product of two positive semi-definite kernels. Hence, if we show that $\Sigma^{\left(L+1\right)}$ is positive-definite, this implies that $\Theta^{\left(L+1\right)}$ is positive-definite.

By definition, with the notation of Proposition 1 we have

Hence the left-hand side only vanishes if $\sum c_{i}\sigma\left(f\left(x_{i}\right)\right)$ is almost surely zero. If $\Sigma^{\left(L\right)}$ is positive-definite, the Gaussian $\left(f\left(x_{i}\right)\right)_{i=1,\ldots d}$ is non-degenerate, so this only occurs when $c_{1}=\cdots=c_{d}=0$ since $\sigma$ is assumed to be non-constant. This shows that the positive-definiteness of $\Sigma^{\left(L+1\right)}$ is implied by that of $\Sigma^{\left(L\right)}$ . By induction, if $\Sigma^{\left(2\right)}$ is positive-definite, we obtain that all $\Sigma^{\left(L\right)}$ with $L\geq 2$ are positive-definite as well. By the first step this hence implies that $\Theta^{\left(L\right)}$ is positive-definite as well.

Writing the expansion of $\mu$ in Hermite polynomials $\left(h_{i}\right)_{i\geq 0}$

we obtain that $\hat{\mu}$ is given by the power series

Since $\sigma$ is non-polynomial, so is $\mu$ , and as a result, there is an infinite number of nonzero $a_{i}$ ’s in the above sum.

where the $a_{i}$ ’s are the coefficients of the Hermite expansion of $\mu$ . Now, observe that by the previous step, the power series expansion of $\nu$ contains both an infinite number of nonzero even terms and an infinite number of nonzero odd terms. This enables one to apply Theorem 3 to obtain that $\Sigma^{\left(2\right)}$ is indeed positive-definite, thereby concluding the proof.