Understanding and Mitigating the Tradeoff Between Robustness and Accuracy

Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, Percy Liang

Introduction

Adversarial training methods (Goodfellow et al., 2015; Madry et al., 2017) attempt to improve the robustness of neural networks against adversarial examples (Szegedy et al., 2014) by augmenting the training set (on-the-fly) with perturbed examples that preserve the label but that fool the current model. While such methods decrease the robust error, the error on worst-case perturbed inputs, they have been observed to cause an undesirable increase in the standard error, the error on unperturbed inputs (Madry et al., 2018; Zhang et al., 2019; Tsipras et al., 2019).

In this work, we provide a different explanation for the tradeoff between standard and robust error that takes generalization from finite data into account. We first consider a linear model where the true linear function has zero standard and robust error. Adversarial training augments the original training set with extra data, consisting of samples $(x_{\text{ext}},y)$ where the perturbations $x_{\text{ext}}$ are consistent, meaning that the conditional distribution stays constant $P_{\mathsf{y}}(\cdot\mid x_{\text{ext}})=P_{\mathsf{y}}(\cdot\mid x)$ . We show that even in this simple setting, the augmented estimator, i.e. the minimum norm interpolant of the augmented data (standard + extra data), could have a larger standard error than that of the standard estimator, which is the minimum norm interpolant of the standard data alone. We found this surprising given that adding consistent perturbations enforces the predictor to satisfy invariances that the true model exhibits. One might think adding this information would only restrict the hypothesis class and thus enable better generalization, not worse.

We show that this tradeoff stems from overparameterization. If the restricted hypothesis class (by enforcing invariances) is still overparameterized, the inductive bias of the estimation procedure (e.g., the norm being minimized) plays a key role in determining the generalization of a model.

Figure 2 shows an illustrative example of this phenomenon with cubic smoothing splines. The predictor obtained via standard training (dashed blue) is a line that captures the global structure and obtains low error. Training on augmented data with locally consistent perturbations of the training data (crosses) restricts the hypothesis class by encouraging the predictor to fit the local structure of the high density points. Within this set, the cubic splines predictor (solid orange) minimizes the second derivative on the augmented data, compromising the global structure and performing badly on the tails (Figure 2(b)). More generally, as we characterize in Section 3, the tradeoff stems from the inductive bias of the minimum norm interpolant, which minimizes a fixed norm independent of the data, while the standard error depends on the geometry of the covariates.

Recent works (Carmon et al., 2019; Najafi et al., 2019; Uesato et al., 2019) introduced robust self-training (RST), a robust variant of self-training that overcomes the sample complexity barrier of learning a model with low robust error by leveraging extra unlabeled data. In this paper, our theoretical understanding of the tradeoff between standard and robust error in linear regression motivates RST as a method to improve robust error without sacrificing standard error. In Section 4.2, we prove that RST eliminates the tradeoff for linear regression—RST does not increase standard error compared to the standard estimator while simultaneously achieving the best possible robust error, matching the standard error (see Figure 2(c) for the effect of RST on the spline problem). Intuitively, RST regularizes the predictions of the robust estimator towards that of the standard estimator on the unlabeled data thereby eliminating the tradeoff.

As previous works only focus on the empirical evaluation of the gains in robustness via RST, we systematically evaluate the effect of RST on both the standard and robust error on Cifar-10 when using unlabeled data from Tiny Images as sourced in Carmon et al. (2019). We expand upon empirical results in two ways. First, we study the effect of the labeled training set sizes and and find that the RST improves both robust and standard error over vanilla adversarial training across all sample sizes. RST offers maximum gains at smaller sample sizes where vanilla adversarial training increases the standard error the most. Second, we consider an additional family of perturbations over random and adversarial rotation/translations and find that RST offers gains in both robust and standard error.

Setup

for consistent perturbations $T(x)$ that satisfy

In this work, we focus on interpolating estimators in highly overparameterized models, motivated by modern machine learning models that achieve near zero training loss (on both standard and extra data). Interpolating estimators for linear regression have been studied in many recent works such as (Ma et al., 2018; Belkin et al., 2018; Hastie et al., 2019; Liang & Rakhlin, 2018; Bartlett et al., 2019). We present our results for interpolating estimators with minimum Euclidean norm, but our analysis directly applies to more general Mahalanobis norms via suitable reparameterization (see Appendix A).

Analysis in the linear regression setting

In this section, we compare the standard errors of the standard estimator and the augmented estimator in noiseless linear regression. We begin with a simple toy example that describes the intuition behind our results (Section 3.1) and provide a more complete characterization in Section 3.2. This section focuses only on the standard error of both estimators; we revisit the robust error together with the standard error in Section 4.

Suppose we augment with an extra data point $X_{\text{ext}}==e_{1}+e_{2}$ which lies in $\text{Null}(X_{\text{std}})$ (black dashed line in Figure 3). The augmented estimator $\hat{\theta}_{\text{aug}}$ still fits the standard data $X_{\text{std}}$ and thus $(\hat{\theta}_{\text{aug}})_{3}=\theta^{\star}_{3}=(\hat{\theta}_{\text{std}})_{3}$ . Due to fitting the extra data $X_{\text{ext}}$ , $\hat{\theta}_{\text{aug}}$ (orange vector in Figure 3) must also satisfy an additional constraint $X_{\text{ext}}\hat{\theta}_{\text{aug}}=X_{\text{ext}}\theta^{\star}$ . The crucial observation is that additional constraints along one direction ( $e_{1}+e_{2}$ in this case) could actually increase parameter error along other directions. For example, let’s consider the direction $e_{2}$ in Figure 3. Note that fitting $X_{\text{ext}}$ makes $\hat{\theta}_{\text{aug}}$ have a large component along $e_{2}$ . Now if $\theta^{\star}_{2}$ is small (precisely, $\theta^{\star}_{2}<\theta^{\star}_{1}/3$ ), $\hat{\theta}_{\text{aug}}$ has a larger parameter error along $e_{2}$ than $\hat{\theta}_{\text{std}}$ , which was simply zero (Figure 3 (a)). Conversely, if the true component $\theta^{\star}_{2}$ is large enough (precisely, $\theta^{\star}_{2}>\theta^{\star}_{1}/3$ ), the parameter error of $\hat{\theta}_{\text{aug}}$ along $e_{2}$ is smaller than that of $\hat{\theta}_{\text{std}}$ .

The contribution of different components of the parameter error to the standard error is scaled by the population covariance $\Sigma$ (see Equation 4). For simplicity, let $\Sigma=\operatorname*{diag}([\lambda_{1},\lambda_{2},\lambda_{3}])$ . In our example, the parameter error along $e_{3}$ is zero since both estimators interpolate the standard training point $X_{\text{std}}=e_{1}=3$ . Then, the ratio between $\lambda_{1}$ and $\lambda_{2}$ determines which component of the parameter error contributes more to the standard error.

Putting the two effects together, we see that when $\theta^{\star}_{2}$ is small as in Fig 3(a), $\hat{\theta}_{\text{aug}}$ has larger parameter error than $\hat{\theta}_{\text{std}}$ in the direction $e_{2}$ . If $\lambda_{2}\gg\lambda_{1}$ , error in $e_{2}$ is weighted much more heavily in the standard error and consequently $\hat{\theta}_{\text{aug}}$ would have a larger standard error. Precisely, we have

We present a formal characterization of this tradeoff in general in the next section.

2 General characterizations

In this section, we precisely characterize when the augmented estimator $\hat{\theta}_{\text{aug}}$ that fits extra training data points $X_{\text{ext}}$ in addition to the standard points $X_{\text{std}}$ has higher standard error than the standard estimator $\hat{\theta}_{\text{std}}$ that only fits $X_{\text{std}}$ . In particular, this enables us to understand when there is a “tradeoff” where the augmented estimator $\hat{\theta}_{\text{aug}}$ has lower robust error than $\hat{\theta}_{\text{std}}$ by virtue of fitting perturbations, but has higher standard error. In Section 3.1, we illustrated how the parameter error of $\hat{\theta}_{\text{aug}}$ could be larger than $\hat{\theta}_{\text{std}}$ in some directions, and if these directions are weighted heavily in the population covariance $\Sigma$ , the standard error of $\hat{\theta}_{\text{aug}}$ would be larger.

Formally, let us define the parameter errors $\Delta_{\text{std}}\stackrel{{\scriptstyle\rm def}}{{=}}\hat{\theta}_{\text{std}}-\theta^{\star}$ and $\Delta_{\text{aug}}\stackrel{{\scriptstyle\rm def}}{{=}}\hat{\theta}_{\text{aug}}-\theta^{\star}$ . Recall that the standard errors are

where $\Sigma$ is the population covariance of the underlying inputs drawn from $P_{\mathsf{x}}$ .

To characterize the effect of the inductive bias of minimum norm interpolation on the standard errors, we define the following projection operators: $\Pi_{\text{std}}^{\perp}$ , the projection matrix onto $\text{Null}(X_{\text{std}})$ and $\Pi_{\text{aug}}^{\perp}$ , the projection matrix onto $\text{Null}([X_{\text{ext}};X_{\text{std}}])$ (see formal definition in Appendix B). Since $\hat{\theta}_{\text{aug}}$ and $\hat{\theta}_{\text{std}}$ are minimum norm interpolants, $\Pi_{\text{std}}^{\perp}\hat{\theta}_{\text{std}}=0$ and $\Pi_{\text{aug}}^{\perp}\hat{\theta}_{\text{aug}}=0$ . Further, in noiseless linear regression, $\hat{\theta}_{\text{std}}$ and $\hat{\theta}_{\text{aug}}$ have no error in the span of $X_{\text{std}}$ and $[X_{\text{std}};X_{\text{ext}}]$ respectively. Hence,

Our main result relies on the key observation that for any vector $u$ , $\Pi_{\text{std}}^{\perp}u$ can be decomposed into a sum of two orthogonal components $v$ and $w$ such that $\Pi_{\text{std}}^{\perp}u=v+w$ with $w=\Pi_{\text{aug}}^{\perp}u$ and $v=\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}u$ . This is because $\text{Null}([X_{\text{std}};X_{\text{ext}}])\subseteq\text{Null}(X_{\text{std}})$ and thus $\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp}=\Pi_{\text{aug}}^{\perp}$ . Now setting $u=\theta^{\star}$ and using the error expressions in Equation 6 and Equation 7 gives a precise characterization of the difference in the standard errors of $\hat{\theta}_{\text{std}}$ and $\hat{\theta}_{\text{aug}}$ .

The difference in the standard errors of the standard estimator $\hat{\theta}_{\text{std}}$ and augmented estimator $\hat{\theta}_{\text{aug}}$ can be written as follows.

where $v=\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}\theta^{\star}$ and $w=\Pi_{\text{aug}}^{\perp}\theta^{\star}$ .

The proof of Theorem 1 is in Appendix B.3. The increase in standard error of the augmented estimator can be understood in terms of the vectors $w$ and $v$ defined in Theorem 1. The first term $v^{\top}\Sigma v$ is always positive, and corresponds to the decrease in the standard error of the augmented estimator $\hat{\theta}_{\text{aug}}$ by virtue of fitting extra training points in some directions. However, the second term $2w^{\top}\Sigma v$ can be negative and intuitively measures the cost of a possible increase in the parameter error along other directions (similar to the increase along $e_{2}$ in the simple setting of Figure 3(a)). When the cost outweighs the benefit, the standard error of $\hat{\theta}_{\text{aug}}$ is larger. Note that both the cost and benefit is determined by $\Sigma$ which governs how the parameter error affects the standard error.

We can use the above expression (Theorem 1) for the difference in standard errors of $\hat{\theta}_{\text{aug}}$ and $\hat{\theta}_{\text{std}}$ to characterize different “safe” conditions under which augmentation with extra data does not increase the standard error. See Appendix B.7 for a proof.

The following conditions are sufficient for $L_{\text{std}}(\hat{\theta}_{\text{aug}})\leq L_{\text{std}}(\hat{\theta}_{\text{std}})$ , i.e. the standard error does not increase when fitting augmented data.

The population covariance $\Sigma$ is identity.

The augmented data $[X_{\text{std}};X_{\text{ext}}]$ spans the entire space, or equivalently $\Pi_{\text{aug}}^{\perp}=0$ .

We would like to draw special attention to the first condition. When $\Sigma=I$ , notice that the norm that governs the standard error (Equation 6) matches the norm that is minimized by the interpolants (Equation 5). Intuitively, the estimators have the “right” inductive bias; under this condition, the augmented estimator $\hat{\theta}_{\text{aug}}$ does not have higher standard error. In other words, the observed increase in the standard error of $\hat{\theta}_{\text{aug}}$ can be attributed to the “wrong” inductive bias. In Section 4, we will use this understanding to propose a method of robust training which does not increase standard error over standard training.

Finally, we relate the magnitude of increase in standard error of the augmented estimator to the complexity of the true model.

For a given $X_{\text{std}},X_{\text{ext}},\Sigma$ ,

for some scalar $\gamma>0$ that depends on $X_{\text{std}},X_{\text{ext}},\Sigma$ .

Robust self-training

We now use insights from Section 3 to construct estimators with low robust error without increasing the standard error. While Section 3 characterized the effect of adding extra data $X_{\text{ext}}$ in general, in this section we consider robust training which augments the dataset with extra data $X_{\text{ext}}$ that are consistent perturbations of the standard training data $X_{\text{std}}$ .

Since the standard estimator has small standard error, a natural strategy to mitigate the tradeoff is to regularize the augmented estimator to be closer to the standard estimator. The choice of distance between the estimators we regularize is very important. Recall from Section 3.1 that the population covariance $\Sigma$ determines how the parameter error affects the standard error. This suggests using a regularizer that incorporates information about $\Sigma$ .

We first revisit the recently proposed robust self-training (RST) (Carmon et al., 2019; Najafi et al., 2019; Uesato et al., 2019) that incorporates additional unlabeled data via pseudo-labels from a standard estimator. Previous work only focused on the effectiveness of RST in improving the robust error. In Section 4.2, we prove that in linear regression, RST eliminates the tradeoff between standard and robust error (Theorem 2). The proof hinges on the connection between RST and the idea of regularizing towards the standard estimator discussed above. In particular, we show that the RST objective can be rewritten as minimizing a suitable $\Sigma$ -induced distance to the standard estimator.

We first describe the general two-step robust self-training (RST) procedure (Carmon et al., 2019; Uesato et al., 2019) for a parameteric model $f_{\theta}$ :

It is convenient to summarize the robust self-training estimator $\hat{\theta}_{\text{rst}}$ as the minimizer of a weighted combination of four separate losses as follows. We define the losses on the labeled dataset $\{(x_{i},y_{i})\}_{i=1}^{n}$ as

for fixed scalars $\alpha,\beta,\gamma,\lambda\geq 0$ .

2 Robust self-training for linear regression

We now return to the noiseless linear regression as described in Section 2 and specialize the general RST estimator described in Equation (10) to this setting. We prove that RST eliminates the decrease in standard error in this setting while achieving low robust error by showing that RST appropriately regularizes the augmented estimator towards the standard estimator.

Figure 5 shows the four losses of RST in this special case of linear regression.

Obtaining this specialized estimator from the general RST estimator in Equation (10) involves the following steps. First, for convenience of analysis, we assume access to the population covariance $\Sigma$ via infinite unlabeled data and thus replace the finite sample losses on the unlabeled data $\hat{L}_{\text{std-unlab}}(\theta),\hat{L}_{\text{rob-unlab}}(\theta)$ by their population losses $L_{\text{std-unlab}}(\theta),L_{\text{rob-unlab}}(\theta)$ . Second, the general RST objective minimizes some weighted combination of four losses. When specializing to the case of noiseless linear regression, since $\hat{L}_{\text{std, lab}}(\theta^{\star})=0$ , rather than minimizing $\alpha\hat{L}_{\text{std-lab}}(\theta^{\star})$ , we set the coefficients on the losses such that the estimator satisfies a hard constraint $\hat{L}_{\text{std-lab}}(\theta^{\star})=0$ . This constraint which enforces interpolation on the labeled dataset $y_{i}=x_{i}^{\top}\theta~{}\forall i=1,\ldots n$ allows us to rewrite the robust loss (Equation 9) on the labeled examples equivalently as a self-consistency loss defined independent of labels.

Since $\theta^{\star}$ is invariant on perturbations $T(x)$ by definition, we have $\hat{L}_{\text{rob-lab}}(\theta^{\star})=0$ and thus we introduce a constraint $\hat{L}_{\text{rob-lab}}(\theta)=0$ in the estimator.

For the losses on the unlabeled data, since the pseudo-labels are not perfect, we minimize $L_{\text{std-unlab}}$ in the objective instead of enforcing a hard constraint on $L_{\text{std-unlab}}$ . However, similarly to the robust loss on labeled data, we can reformulate the robust loss on unlabeled samples $L_{\text{rob-unlab}}$ as a self-consistency loss that does not use pseudo-labels. By definition, $L_{\text{rob-unlab}}(\theta^{\star})=0$ and thus we enforce $L_{\text{rob-unlab}}(\theta)=0$ in the specialized estimator.

We now study the standard and robust error of the linear regression RST estimator defined above in Equation (4.2).

Assume the noiseless linear model $y=x^{\top}\theta^{\star}$ . Let $\theta_{\textup{int-std}}$ be an arbitrary interpolant of the standard data, i.e. $X_{\text{std}}\theta_{\textup{int-std}}=y_{\text{std}}$ . Then

Simultaneously, $L_{\text{rob}}(\hat{\theta}_{\text{rst}})=L_{\text{std}}(\hat{\theta}_{\text{rst}})$ .

The crux of the proof is that the optimization objective of RST is an inductive bias that regularizes the estimator to be close to the standard estimator, weighing directions by their contribution to the standard error via $\Sigma$ . To see this, we rewrite

By incorporating an appropriate $\Sigma$ -induced regularizer while satisfying constraints on the robust losses, RST ensures that the standard error of the estimator never exceeds the standard error of $\hat{\theta}_{\text{std}}$ . The robust error of any estimator is lower bounded by its standard error, and this gap can be arbitrarily large for the standard estimator. However, the robust error of the RST estimator matches the lower bound of its standard error which in turn is bounded by the standard error of the standard estimator and hence is small. To provide some graphical intuition for the result, see Figure 2 that visualizes the RST estimator on the cubic splines interpolation problem that exemplifies the increase in standard error upon augmentation. RST captures the global structure and obtains low standard error by matching $\hat{\theta}_{\text{std}}$ (straight line) on unlabeled inputs. Simultaneously, RST enforces invariance on local transformations on both labeled and unlabeled inputs, and obtains low robust error by capturing the local structure across the domain.

The constraint on the standard loss on labeled data simply corresponds to interpolation on the standard labeled data. The constraints on the robust self-consistency losses involve a maximization over a set of transformations. In the case of linear regression, such constraints can be equivalently represented by a set of at most $d$ linear constraints, where $d$ is the dimension of the covariates. Further, with this finite set of constraints, we only require access to the covariance $\Sigma$ in order to constrain the population robust loss. Appendix D gives a practical iterative algorithm that computes the RST estimator for linear regression reminiscent of adversarial training in the semi-supervised setting.

3 Empirical evaluation of RST

Both RST+AT and RST+TRADES have lower robust and standard error than their supervised counterparts AT and TRADES across all perturbation types. This mirrors the theoretical analysis of RST in linear regression (Theorem 2) where the RST estimator has small robust error while provably not sacrificing standard error, and never obtaining larger standard error than the standard estimator.

Recall that our work motivates studying the tradeoff between robust and standard error while taking generalization from finite data into account. We showed that the gap in the standard error of a standard estimator and that of a robust estimator is large for small training set sizes and decreases as the labeled dataset is larger (Figure 1). We now study the effect of RST as we vary the training set size in Figure 6. We find that RST+AT has lower standard error than standard training across all sample sizes for small $\epsilon$ , while simultaneously achieving lower robust error than AT (see Appendix E.2.1). In the small data regime where vanilla adversarial training hurts the standard error the most, we find that RST+AT gives about 3x more absolute improvement than in the large data regime. We note that this set of experiments are complementary to the experiments in (Schmidt et al., 2018) which study the effect of the training set size only on robust error.

We also test the effect of RST on perturbations where robust training slightly improves standard error rather than hurting it. Since RST regularizes towards the standard estimator, one might suspect that the improvements from robust training disappear with RST. In particular, we consider spatial transformations $T(x)$ that consist of simultaneous rotations and translations. We use two common forms of robust training for spatial perturbations, where we approximately maximize over $T(x)$ with either adversarial (worst-of-10) or random augmentations (Yang et al., 2019; Engstrom et al., 2019). Table 1 (right) presents the results. In the regime where vanilla robust training does not hurt standard error, RST in fact further improves the standard error by almost 1% and the robust error by 2-3% over the standard and robust estimators for both forms of robust training. Thus in settings where vanilla robust training improves standard error, RST seems to further amplify the gains while in settings where vanilla robust training hurts standard error, RST mitigates the harmful effect.

The RST estimator minimizes both a robust loss and a standard loss on the unlabeled data with pseudo-labels (bottom row, Figure 5). Both of these losses are necessary to simultaneously the standard and robust error over vanilla supervised robust training. Standard self-training, which only uses standard loss on unlabeled data, has very high robust error ( $\approx 100\%$ ). Similarly, Robust Consistency Training, an extension of Virtual Adversarial Training (Miyato et al., 2018) that only minimizes a robust self-consistency loss on unlabeled data, marginally improves the robust error but actually hurts standard error (Table 1).

Related Work

In concurrent and independent work, Min et al. (2020) also study the effect of dataset size on the tradeoff. They prove that in a “strong adversary” regime, there is a tradeoff even with infinite data, as the perturbations are large enough to change the ground truth target. They also identify a “weak adversary” regime (smaller perturbations) where the gap in standard error between robust and standard estimators first increases and then decreases, with no tradeoff in the infinite data limit. Similar to our work, this provides an example of a tradeoff due to generalization from finite data. However, their experimental validation of the tradeoff trends is restricted to simulated settings and they do not study how to mitigate the tradeoff.

To the best of our knowledge, ours is the first work that theoretically studies how to mitigate the tradeoff between standard and robust error. While robust self-training (RST) was proposed in recent works (Carmon et al., 2019; Najafi et al., 2019; Uesato et al., 2019) as a way to improve robust error, we prove that RST eliminates the tradeoff between standard and robust error in noiseless linear regression and systematically study the effect on RST on the tradeoff with several different perturbations and adversarial training algorithms on Cifar-10.

Interpolated Adversarial Training (IAT) (Lamb et al., 2019) and Neural Architecture Search (NAS) (Cubuk et al., 2017) were proposed to mitigate the tradeoff bbetween standard and robust error empirically. IAS considers a different training algorithm based on Mixup, NAS (Cubuk et al., 2017) uses RL to search for more robust architectures. In Table 1, we also report the standard and robust errors of these methods. RST, IAT and NAS are incomparable as they find different tradeoffs between standard and robust error. Recently, Xie et al. (2020) showed that adversarial training with appropriate batch normalization (AdvProp) with small perturbations can actually improve standard error. However, since they only aim to improve and evaluate the standard error, it is unclear if the robust error improves. We believe that since RST provides a complementary statistical perspective on the tradeoff, it can be combined with methods like IAT, NAS or AdvProp to see further gains in standard and robust errors. We leave this to future work.

Conclusion

We study the commonly observed increase in standard error upon adversarial training due to generalization from finite data in a well-specified setting with consistent perturbations. Surprisingly, we show that methods that augment the training data with consistent perturbations, such as adversarial training, can increase the standard error even in the simple setting of noiseless linear regression where the true linear function has zero standard and robust error. Our analysis reveals that the mismatch between the inductive bias of models and the underlying distribution of the inputs causes the standard error to increase even when the augmented data is perfectly labeled. This insight motivates a method that provably eliminates the tradeoff in linear regression by incorporating an appropriate regularizer that utilizes the distribution of the inputs. While not immediately apparent, we show that this is a special case of the recently proposed robust self-training (RST) procedure that uses additional unlabeled data to estimate the distribution of the inputs. Previous works view RST as a method to improve the robust error by increasing the sample size. Our work provides some theoretical justification for why RST improves both the standard and robust error, thereby mitigating the tradeoff between accuracy and robustness. How to best utilize unlabeled data, and whether sufficient unlabeled data can completely eliminate the tradeoff remain open questions.

We are grateful to Tengyu Ma, Yair Carmon, Ananya Kumar, Pang Wei Koh, Fereshte Khani, Shiori Sagawa and Karan Goel for valuable discussions and comments. This work was funded by an Open Philanthropy Project Award and NSF Frontier Award as part of the Center for Trustworthy Machine Learning (CTML). AR was supported by Google Fellowship and Open Philanthropy AI Fellowship. SMX was supported by an NDSEG Fellowship. FY was supported by the Institute for Theoretical Studies ETH Zurich and the Dr. Max Rossler and the Walter Haefner Foundation. FY and JCD were supported by the Office of Naval Research Young Investigator Awards.

References

Appendix A Transformations to handle arbitrary matrix norms

Consider a more general minimum norm estimator of the following form. Given inputs $X$ and corresponding targets $y$ as training data, we study the interpolation estimator,

Appendix B Standard error of minimum norm interpolants

The projection operators $\Pi_{\text{std}}^{\perp}$ and $\Pi_{\text{aug}}^{\perp}$ are formally defined as follows.

B.2 Invariant transformations may have arbitrary nullspace components

B.3 Proof of Theorem 1

by decomposition of $\Pi_{\text{std}}^{\perp}\theta^{\star}=v+w$ where $v=\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}\theta^{\star}$ and $w=\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp}\theta^{\star}$ . Note that the error difference does scale with $\|\theta^{\star}\|^{2}$ , although the sign of the difference does not.

B.4 Proof of Corollary 1

Corollary 1 presents three sufficient conditions under which the standard error of the augmented estimator $L_{\text{std}}(\hat{\theta}_{\text{aug}})$ is never larger than the standard error of the standard estimator $L_{\text{std}}(\hat{\theta}_{\text{std}})$ .

When the population covariance $\Sigma=I$ , from Theorem 1, we see that

since $v=\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}\theta^{\star}$ and $w=\Pi_{\text{aug}}^{\perp}\theta^{\star}$ are orthogonal.

When $\Pi_{\text{aug}}^{\perp}=0$ , the vector $w$ in Theorem 1 is , and hence we get

We prove the eigenvector condition in Section B.7 which studies the effect of augmenting with a single extra point in general.

B.5 Proof of Proposition 1

The proof of Proposition 1 is based on the following two lemmas that are also useful for characterization purposes in Corollary 2.

If a PSD matrix $\Sigma$ has non-equal eigenvalues, one can find two unit vectors $w,v$ for which the following holds

Hence, there exists a combination of original and augmentation dataset $X_{\text{std}},X_{\text{ext}}$ such that condition (19) holds for two directions $v\in\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}})$ and $w\in\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp})=\text{Col}(\Pi_{\text{aug}}^{\perp})$ .

Note that neither $w$ nor $v$ can be eigenvectors of $\Sigma$ in order for both conditions in equation (19) to hold. Given a population covariance, fixed original and augmentation data for which condition (19) holds, we can now explicitly construct $\theta^{\star}$ for which augmentation increases standard error.

where $\beta_{i}$ are constants that depend on $X_{\text{std}},X_{\text{ext}},\Sigma$ .

Proposition 1 follows directly from the second statement of Lemma 2 by minimizing the bound (20) with respect to $c_{1}$ which is a free parameter to be chosen during construction of $\theta^{\star}$ (see proof of Lemma (2). The minimum is attained for $c_{1}=2\sqrt{(\beta_{1}+1)(\beta_{2}c^{2})}$ . We hence conclude that $\theta^{\star}$ needs to be sufficiently more complex than a good standard solution, i.e. $\|\theta^{\star}\|^{2}_{2}-\|\hat{\theta}_{\text{std}}\|^{2}_{2}>\gamma c$ where $\gamma>0$ is a constant that depends on the $X_{\text{std}},X_{\text{ext}}$ .

B.6 Proof of technical lemmas

In this section we prove the technical lemmas that are used to prove Theorem 1.

Any vector $\Pi_{\text{std}}^{\perp}\theta\in\text{Null}(\Sigma_{\text{std}})$ can be decomposed into orthogonal components $\Pi_{\text{std}}^{\perp}\theta=\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp}\theta+\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}\theta$ . Using the minimum-norm property, we can then always decompose the (rotated) augmented estimator $\hat{\theta}_{\text{aug}}\in\text{Col}(\Pi_{\text{aug}}^{\perp})=\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp})$ and true parameter $\theta^{\star}$ by

where we define “ext” as the set of basis vectors which span $\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}})$ and respectively “rest” for $\text{Null}(\Sigma_{\text{aug}})$ . Requiring the standard error increase to be some constant $c>0$ can be rewritten using identity (16) as follows

The left hand side of equation (21) is always positive, hence it is necessary for this equality to hold with any $c>0$ , that there exists at least one pair $i,j$ such that $w_{j}^{\top}\Sigma v_{i}\neq 0$ and one direction of the iff statement is proved.

For the other direction, we show that if there exist $v\in\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}})$ and $w\in\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp})$ for which condition (19) holds (wlog we assume that the $w^{\top}\Sigma v<0$ ) we can construct a $\theta^{\star}$ for which the inequality (8) in Theorem 1 holds as follows:

It is then necessary by our assumption that $\xi_{j}\zeta_{i}w_{j}^{\top}\Sigma v_{i}>0$ for at least some $i,j$ . We can then set $\zeta_{i}>0$ such that $\|\hat{\theta}_{\text{aug}}-\hat{\theta}_{\text{std}}\|^{2}=\|\zeta\|^{2}=c_{1}>0$ , i.e. that the augmented estimator is not equal to the standard estimator (else obviously there can be no difference in error and equality (21) cannot be satisfied for any desired error increase $c>0$ ).

The choice of $\xi$ minimizing $\|\theta^{\star}-\hat{\theta}_{\text{aug}}\|^{2}=\sum_{j}\xi_{j}^{2}$ that also satisfies equation (21) is an appropriately scaled vector in the direction of $x=W^{\top}\Sigma V\zeta$ where we define $W:=[w_{1},\dots,w_{|\text{rest}|}]$ and $V:=[v_{1},\dots,v_{|\text{ext}|}]$ . Defining $c_{0}=\zeta^{\top}V^{\top}\Sigma V\zeta$ for convenience and then setting

which is well-defined since $x\neq 0$ , yields a $\theta^{\star}$ such that augmentation increases standard error. It is thus necessary for $L_{\text{std}}(\hat{\theta}_{\text{aug}})-L_{\text{std}}(\hat{\theta}_{\text{std}})=c$ that

By assuming existence of $i,j$ such that $\xi_{j}\zeta_{i}w_{j}^{\top}\Sigma v_{i}\neq 0$ , we are guaranteed that $\lambda^{2}_{\max}(W^{\top}\Sigma V)>0$ .

Note due to construction we have $\|\theta^{\star}\|_{2}^{2}=\|\hat{\theta}_{\text{std}}\|_{2}^{2}+\sum_{i}\zeta_{i}^{2}+\sum_{j}\xi_{j}^{2}$ and plugging in the choice of $\xi_{j}$ in equation (22) we have

Setting $\beta_{1}=\left[1+\frac{\lambda_{\min}^{2}(V^{\top}\Sigma V)}{4\lambda^{2}_{\max}(W^{\top}\Sigma V)}\right]$ , $\beta_{2}=\frac{1}{4\lambda^{2}_{\max}(W^{\top}\Sigma V)}$ yields the result.

B.6.2 Proof of Lemma 1

Let $\lambda_{1},\dots,\lambda_{m}$ be the $m$ non-zero eigenvalues of $\Sigma$ and $u_{i}$ be the corresponding eigenvectors. Then choose $v$ to be any combination of the eigenvectors $v=U\beta$ where $U=[u_{1},\dots,u_{m}]$ where at least $\beta_{i},\beta_{j}\neq 0$ for $\lambda_{i}\neq\lambda_{j}$ . We next construct $w=U\alpha$ by choosing $\alpha$ as follows such that the inequality in (19) holds:

and $\alpha_{k}=0$ for $k\neq i,j$ . Then we have that $\alpha^{\top}\beta=0$ and hence $w^{\top}v=0$ . Simultaneously

which concludes the proof of the first statement.

We now prove the second statement by constructing $\Sigma_{\text{std}}=X_{\text{std}}^{\top}X_{\text{std}},\Sigma_{\text{ext}}=X_{\text{ext}}^{\top}X_{\text{ext}}$ using $w,v$ . We can then obtain $X_{\text{std}},X_{\text{ext}}$ using any standard decomposition method to obtain $X_{\text{std}},X_{\text{ext}}$ . We construct $\Sigma_{\text{std}},\Sigma_{\text{ext}}$ using $w,v$ . Without loss of generality, we can make them simultaneously diagonalizable. We construct a set of eigenvectors that is the same for both matrices paired with different eigenvalues. Let the shared eigenvectors include $w,v$ . Then if we set the corresponding eigenvalues $\lambda_{w}(\Sigma_{\text{ext}})=0,\lambda_{v}(\Sigma_{\text{ext}})>0$ and $\lambda_{w}(\Sigma_{\text{std}})=0,\lambda_{v}(\Sigma_{\text{std}})=0$ , then $\lambda_{w}(\Sigma_{\text{aug}})=0$ such that $w\in\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}}^{\perp})$ and $v\in\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}})$ . This shows the second statement. With this, we can design a $\theta^{\star}$ for which augmentation increases standard error as in Lemma 2.

B.7 Characterization Corollary 2

A simpler case to analyze is when we only augment with one extra data point. The following corollary characterizes which single augmentation directions lead to higher prediction error for the augmented estimator.

The following characterizations hold for augmentation directions that do not cause the standard error of the augmented estimator to be higher than the original estimator.

(in terms of ratios of inner products) For a given $\theta^{\star}$ , data augmentation does not increase the standard error of the augmented estimator for a single augmentation direction $x_{\text{ext}}$ if

(in terms of eigenvectors) Data augmentation does not increase standard error for any $\theta^{\star}$ if $\Pi_{\text{std}}^{\perp}x_{\text{ext}}$ is an eigenvector of $\Sigma$ . However if one augments in the direction of a mixture of eigenvectors of $\Sigma$ with different eigenvalues, there exists $\theta^{\star}$ such that augmentation increases standard error.

(depending on well-conditioning of $\Sigma$ ) If $\frac{\lambda_{\max}(\Sigma)}{\lambda_{\min}(\Sigma)}\leq 2$ and $\Pi_{\text{std}}^{\perp}\theta^{\star}$ is an eigenvector of $\Sigma$ , then no augmentations $x_{\text{ext}}$ increase standard error.

The form in Equation (23) compares ratios of inner products of $\Pi_{\text{std}}^{\perp}x_{\text{ext}}$ and $\Pi_{\text{std}}^{\perp}\theta^{\star}$ in two spaces: the one in the numerator is weighted by $\Sigma$ whereas the denominator is the standard inner product. Thus, if $\Sigma$ scales and rotates rather inhomogeneously, then augmenting with $x_{\text{ext}}$ may hurt standard error. Here again, if $\Sigma=\gamma I$ for $\gamma>0$ , then the condition must hold.

Note that for a single augmentation point $X_{\text{ext}}=x_{\text{ext}}^{\top}$ , the orthogonal decomposition of $\Pi_{\text{std}}^{\perp}\theta^{\star}$ into $\text{Col}(\Pi_{\text{aug}}^{\perp})$ and $\text{Col}(\Pi_{\text{std}}^{\perp}\Pi_{\text{aug}})$ is defined by $v=\frac{{\Pi_{\text{std}}^{\perp}x_{\text{ext}}}^{\top}\theta^{\star}}{\|{\Pi_{\text{std}}^{\perp}x_{\text{ext}}}\|^{2}}{\Pi_{\text{std}}^{\perp}x_{\text{ext}}}$ and $w=\Pi_{\text{std}}^{\perp}\theta^{\star}-v$ respectively. Plugging back into into identity (16) then yields the following condition for safe augmentations:

Rearranging the terms yields inequality (23).

Safe augmentation directions for specific choices of $\theta^{\star}$ and $\Sigma$ are illustrated in Figure 3.

B.7.2 Proof of Corollary 2 (b)

Assume that $\Pi_{\text{std}}^{\perp}x_{\text{ext}}$ is an eigevector of $\Sigma$ with eigenvalue $\lambda>0$ . We have

for any $\theta^{\star}$ . Hence by Corollary 2 (a), the standard error doesn’t increase by augmenting with eigenvectors of $\Sigma$ for any $\theta^{\star}$ .

When the single augmentation direction $v$ is not an eigenvector of $\Sigma$ , by Lemma 1 one can find $w$ such that $w^{\top}\Sigma v\neq 0$ . The proof in Lemma 1 gives an explicit construction for $w$ such that condition (19) holds and the result then follows directly by Lemma 2.

B.7.3 Proof of Corollary 2 (c)

Suppose $\Sigma\Pi_{\text{std}}^{\perp}\theta^{\star}=\lambda\Pi_{\text{std}}^{\perp}\theta^{\star}$ for some $\lambda_{\min}(\Sigma)\leq\lambda\leq\lambda_{\max}(\Sigma)$ . Then starting with the expression (23),

by applying $\frac{\lambda_{\max}(\Sigma)}{\lambda_{\min}(\Sigma)}\leq 2$ . Thus when $\Pi_{\text{std}}^{\perp}\theta^{\star}$ is an eigenvector of $\Sigma$ , there are no augmentations $x_{\text{ext}}$ that increase the standard error.

Appendix C Details for spline staircase

We describe the data distribution, augmentations, and model details for the spline experiment in Figure 1 and toy scenario in Figure 2. Finally, we show that we can construct a simplified family of spline problems where the ratio between standard errors of the augmented and standard estimators increases unboundedly as the number of stairs.

We describe the data distribution in terms of the one-dimensional input $t$ , and by the one-to-one correspondence with spline basis features $x=X(t)$ , this also defines the distribution of spline features $x\in\mathcal{X}$ . Let $w\in\Delta_{s}$ define a distribution over $\mathcal{T}_{\text{line}}$ where $\Delta_{s}$ is the probability simplex of dimension $s$ . We define the data distribution with the following generative process for one sample $t$ . First, sample a point $i$ from $\mathcal{T}_{\text{line}}$ according to the categorical distribution described by $w$ , such that $i\sim\text{Categorical}(w)$ . Second, sample $t$ by perturbing $i$ with probability $\delta$ such that

The sampled $t$ is in $\mathcal{T}_{\text{line}}$ with probability $1-\delta$ and $\mathcal{T}_{\text{line}}^{c}$ with probability $\delta$ , where we choose $\delta$ to be small.

C.2 Spline model

Our hypothesis class is the family of cubic B-splines as defined in (Friedman et al., 2001). Cubic B-splines are piecewise cubic functions, where the endpoints of each cubic function are called the knots. In our example, we fix the knots to be $[0,\epsilon,1,\dots,s-1,s-1+\epsilon]$ , which places a knot on every point in $\mathcal{T}$ . This ensures that the function class contains an interpolating function on all $t\in\mathcal{T}$ , i.e. for some $\theta^{\star}$ ,

for the standard estimator and the corresponding augmented problem to obtain the augmented estimator.

C.3 Evaluating Corollary 2 (a) for splines

In the spline staircase, the local perturbations can be thought of as fitting high frequency noise in the function space, where fitting them causes a global change in the function.

Suppose the original training set consists of two points, $X_{\text{std}}=[X_{M}(0),X_{M}(1)]^{\top}$ . We study the effect of augmenting point $x_{\text{ext}}$ in terms of $q_{i}$ above. First, we find that the first two eigenvectors corresponding to linear functions satisfy $\Pi_{\text{std}}^{\perp}q_{1}=\Pi_{\text{std}}^{\perp}q_{2}=0$ . Intuitively, this is because the standard estimator is linear. For ease of visualization, we consider the 2D space in $\text{Null}(\Sigma)$ spanned by $\Pi_{\text{std}}^{\perp}q_{3}$ (global direction, low frequency) and $\Pi_{\text{std}}^{\perp}q_{2s}$ (local direction, high frequency). The matrix $\Pi_{\text{lg}}=[\Pi_{\text{std}}^{\perp}q_{3},~{}\Pi_{\text{std}}^{\perp}q_{2s}]^{\top}$ projects onto this space. Note that the same results hold when projecting onto all $\Pi_{\text{std}}^{\perp}q_{i}$ in $\text{Null}(\Sigma)$ .

In terms of the simple 3-D example in Section 3.1, the global direction corresponds to the costly direction with large eigenvalue, as changes in global structure heavily affect the standard error. Figure 8 plots the projections $\Pi_{\text{lg}}\theta^{\star}$ and $\Pi_{\text{lg}}X_{\text{ext}}$ for different $X_{\text{ext}}$ . When $\theta^{\star}$ has high frequency variations and is complex, $\Pi_{\text{lg}}\theta^{\star}=(\theta^{\star}-\hat{\theta}_{\text{std}})$ is aligned with the local dimension. For $x_{\text{ext}}$ immediately local to training points, the projection $\Pi_{\text{lg}}x_{\text{ext}}$ (orange vector in Figure 8) has both local and global components. Augmenting these local perturbations introduces error in the global component. For other $x_{\text{ext}}$ farther from training points, $\Pi_{\text{lg}}x_{\text{ext}}$ (blue vector in Figure 8) is almost entirely global and perpendicular to $\theta^{\star}-\hat{\theta}_{\text{std}}$ , leaving bias unchanged. Thus, augmenting data close to original data cause estimators to fit local components at the cost of the costly global component which changes overall structure of the predictor like in Figure 2(middle). The choice of inductive bias in the $M$ –norm being minimized results in eigenvectors of $\Sigma$ that correspond to local and global components, dictating this tradeoff.

C.4 Data augmentation can be quite painful for splines

We construct a family of spline problems such that as the number the augmented estimator has much higher error than the standard estimator. We assume that our predictors are from the full family of cubic splines.

We define a modified domain with continuous intervals $\mathcal{T}=\cup_{t=0}^{s-1}[t,t+\epsilon]$ . Considering only $s$ which is a multiple of 2, we sample the original data set as described in Section C.1 with the following probability mass $w$ :

for $\gamma\in[0,1)$ . We define a probability distribution $P_{\mathcal{T}}$ on $\mathcal{T}$ for a random variable $T$ by setting $T=Z+S(Z)$ where $Z\sim\text{Categorical}(w)$ and the $Z$ -dependent perturbation $S(z)$ is defined as

We obtain the training dataset $X_{\text{std}}=\{X(t_{1}),\dots,X(t_{n})\}$ by sampling $t_{i}\sim P_{\mathcal{T}}$ .

Consider a modified augmented estimator for the splines problem, where for each point $t_{i}$ we augment with the entire interval $[\lfloor t_{i}\rfloor,\lfloor t_{i}\rfloor+\epsilon]$ with $\epsilon\in[0,1/2)$ and the estimator is enforced to output $f_{\hat{\theta}}(x)=y_{i}=\lfloor t_{i}\rfloor$ for all $x$ in the interval $[\lfloor t_{i}\rfloor,\lfloor t_{i}\rfloor+\epsilon]$ . Additionally, suppose that the ratio $s/n=O(1)$ between the number of stairs $s$ and the number of samples $n$ is constant.

In this simplified setting, we can show that the standard error of the augmented estimator grows while the standard error of the standard estimator decays to 0.

Let the setting be defined as above. Then with the choice of $\delta=\frac{\log(s^{7})-\log(s^{7}-1)}{s}$ and $\gamma=c/s$ for a constant $c\in[0,1)$ , the ratio between standard errors is lower bounded as

which goes to infinity as $s\rightarrow\infty$ . Furthermore, $R(\hat{\theta}_{\text{std}})\rightarrow 0$ as $s\rightarrow\infty$ .

We first lower bound the standard error of the augmented estimator. Define $E_{1}$ as the event that only the lower half of the stairs is sampled, i.e. $\{t:t<s/2\}$ , which occurs with probability $(1-\gamma)^{n}$ . Let $t^{\star}=\max_{i}\lfloor t_{i}\rfloor$ be the largest “stair” value seen in the training set. Note that the min-norm augmented estimator will extrapolate with zero derivative for $t\geq\max_{i}\lfloor t_{i}\rfloor$ . This is because on the interval $[t^{\star},t^{\star}+\epsilon]$ , the augmented estimator is forced to have zero derivative, and the solution minimizing the second derivative of the prediction continues with zero derivative for all $t\geq t^{\star}$ . In the event $E_{1}$ , $t^{\star}\leq s/2-1$ , where $t^{*}=s/2-1$ achieves the lowest error in this event. As a result, on the points in the second half of the staircase, i.e. $t=\{t\in\mathcal{T}:t>\frac{s}{2}-1\}$ , the augmented estimator incurs large error:

Therefore the standard error of the augmented estimator is bounded by

where in the first line, we note that the error on each interval is the same and the probability of each interval is $(1-\delta)\frac{\gamma}{s/2}+\epsilon\frac{\delta}{\epsilon}\cdot\frac{\gamma}{s/2}=\frac{\gamma}{s/2}$ .

Next we upper bound the standard error of the standard estimator. Define $E_{2}$ to be the event where all points are sampled from $\mathcal{T}_{\text{line}}$ , which occurs with probability $(1-\delta)^{n}$ . In this case, the standard estimator is linear and fits the points on $\mathcal{T}_{\text{line}}$ with zero error, while incurring error for all points not in $\mathcal{T}_{\text{line}}$ . Note that the probability density of sampling a point not in $\mathcal{T}_{\text{line}}$ is either $\frac{\delta}{\epsilon}\cdot\frac{1-\gamma}{s/2}$ or $\frac{\delta}{\epsilon}\cdot\frac{\gamma}{s/2}$ , which we upper bound as $\frac{\delta}{\epsilon}\cdot\frac{1}{s/2}$ .

Therefore for event $E_{2}$ , the standard error is bounded as

since $\log(s^{7})-\log(s^{7}-1)\leq 1$ for $s\geq 2$ . For the complementary event $E_{2}^{c}$ , note that cubic spline predictors can grow only as $O(t^{3})$ , with error at most $O(t^{6})$ . Therefore the standard error for case $E_{2}^{c}$ is bounded as

Thus overall, $R(\hat{\theta}_{\text{std}})=O(1/s)$ and combining the bounds yields the result. ∎

Appendix D Robust Self-Training

We define the linear robust self-training estimator from Equation (4.2) and expand all the terms.

Notice that for unlabeled components of the estimator, we assume access to the data distribution $P_{\mathsf{x}}$ and thus optimize the population quantities.

As we show in the next subsection, we can rewrite the robust self-training estimator into the following reduced form, more directly connecting to the general analysis of adding extra data $X_{\text{ext}}$ in min-norm linear regression.

for the appropriate choice of $X_{\text{ext}}$ , as shown in Section D.1. Here, we can interpret $X_{\text{ext}}$ as the difference between the perturbed inputs and original inputs. These are perturbations which we want the model to be invariant to, and hence output zero.

We give an algorithm for constructing $X_{\text{ext}}$ which enforces the population robustness constraints. Suppose we are given $\Sigma$ , the population covariance of $P_{\mathsf{x}}$ . In robust self-training, we enforce that the model is consistent over perturbations of the labeled data $X_{\text{std}}$ and (infinite) unlabeled data. To do this, we add linear constraints of the form $x_{\text{adv}}^{\top}\theta-x^{\top}\theta=0$ , where $x_{\text{adv}}\in T(x)$ for all $x$ . We can view these linear constraints as augmenting the dataset with input-target pairs $(x_{\text{ext}},0)$ where $x_{\text{ext}}=x_{\text{adv}}-x$ . By assumption, $x_{\text{ext}}^{\top}\theta^{\star}=0$ so these augmentations fit into our data augmentation framework.

However, when we enforce these constraints over the entire population $P_{\mathsf{x}}$ or when there are an infinite number of transformations in $T(x)$ , a naive implementation requires augmenting with infinitely many points. Noting that the space of augmentations $x_{\text{ext}}$ satisfying $x_{\text{ext}}^{\top}\theta^{\star}=0$ is a linear subspace, we can instead summarize the augmentations with a basis that spans the transformations. Let the space of perturbations be $\mathcal{T}=\cup_{x\in\text{supp}(P_{\mathsf{x}}),x_{\text{adv}}\in T(x)}x_{\text{adv}}-x$ . Note that this space of perturbations also contains perturbations of the original data $X_{\text{std}}$ if $X_{\text{std}}$ is in the support of $P_{\mathsf{x}}$ . If $X_{\text{std}}$ is not in the support of $P_{\mathsf{x}}$ , the behavior of the estimator on these points do not affect standard or robust error. Assuming that we can efficiently optimize over $\mathcal{T}$ , we construct the basis by an iterative procedure reminiscent of adversarial training.

Set $t=0$ . Initialize $\theta^{t}=\theta_{\textup{int-std}}$ and $(X_{\text{ext}})_{0}$ as an empty matrix.

At iteration $t$ , solve for $x_{\text{ext}}^{t}=\operatorname*{arg\,max}_{x_{\text{ext}}\in\mathcal{T}}(x_{\text{ext}}^{\top}\theta^{t})^{2}$ . If the objective is unbounded, choose any $x_{\text{ext}}^{t}$ such that $x_{\text{ext}}^{\top}\theta^{t}\neq 0$ .

If ${\theta^{t}}^{\top}x_{\text{ext}}^{t}=0$ , stop and return $(X_{\text{ext}})_{t}$ .

Otherwise, add $x_{\text{ext}}^{t}$ as a row in $(X_{\text{ext}})_{t}$ . Increment $t$ and let $\theta^{t}$ solve (31) with $X_{\text{ext}}=(X_{\text{ext}})_{t}$ .

In each iteration, we search for a perturbation that the current $\theta^{t}$ is not invariant to. If we can find such a perturbation, we add it to the constraint set in $(X_{\text{ext}})_{t}$ . We stop when we cannot find such a perturbation, implying that the rows of $(X_{\text{ext}})_{t}$ and $X_{\text{std}}$ span $\mathcal{T}$ . The final RST estimator solves (31) using $X_{\text{ext}}$ returned from this procedure.

This procedure terminates within $O(d)$ iterations. To see this, note that $\theta^{t}$ is orthogonal to all rows of $(X_{\text{ext}})_{t}$ . Any vector in the span of $(X_{\text{ext}})_{t}$ is orthogonal to $\theta^{t}$ . Thus, if ${\theta^{t}}^{\top}x_{\text{ext}}^{t}\neq 0$ , then $x_{\text{ext}}^{t}$ must not be in the span of $(X_{\text{ext}})_{t}$ . At most $d-\text{rank}(X_{\text{std}})$ such new directions can be added until $(X_{\text{ext}})_{t}$ is full rank. When $(X_{\text{ext}})_{t}$ is full rank, ${\theta^{t}}^{\top}x_{\text{ext}}^{t}=0$ must hold and the algorithm terminates.

D.2 Proof of Theorem 2

In this section, we prove Theorem 2, which we reproduce here. See 2

We work with the RST estimator in the form from Equation (31). We note that our result applies generally to any extra data $X_{\text{ext}},y_{\text{ext}}$ . We define $\Sigma_{\text{std}}=X_{\text{std}}^{\top}X_{\text{std}}$ . Let $\{u_{i}\}$ be an orthonormal basis of the kernel $\text{Null}(\Sigma_{\text{std}}+X_{\text{ext}}^{\top}X_{\text{ext}})$ and $\{v_{i}\}$ be an orthonormal basis for $\text{Null}(\Sigma_{\text{std}})\setminus\mathop{\rm span}(\{u_{i}\})$ . Let $U$ and $V$ be the linear operators defined by $Uw=\sum_{i}u_{i}w_{i}$ and $Vw=\sum_{i}v_{i}w_{i}$ , respectively, noting that $U^{\top}V=0$ . Defining $\Pi_{\text{std}}^{\perp}:=(I-\Sigma_{\text{std}}^{\dagger}\Sigma_{\text{std}})$ to be the projection onto the null space of $X_{\text{std}}$ , we see that there are unique vectors $\rho,\alpha$ such that

Using the representations (32) we may provide an alternative formulation for the augmented estimator (D), using this to prove the theorem. Indeed, writing $\theta_{\textup{int-std}}-\hat{\theta}_{\text{rst}}=U(w-\rho)+V(z-\lambda)$ , we immediately have that the estimator has the form (32c), with the choice

The optimality conditions for this quadratic imply that

Now, recall that the standard error of a vector $\theta$ is $R(\theta)=(\theta-\theta^{\star})^{\top}\Sigma(\theta-\theta^{\star})=\left\|{\theta-\theta^{\star}}\right\|^{2}_{\Sigma}$ , using Mahalanobis norm notation. In particular, a few quadratic expansions yield

where step $(i)$ used that $(U(w-\rho))^{\top}\Sigma V=(V(\lambda-z))^{\top}\Sigma V$ from the optimality conditions (33).

Finally, we consider the rightmost term in equality (34). Again using the optimality conditions (33), we have

by Cauchy-Schwarz. Revisiting equality (34), we obtain

where we used that $x_{\text{adv}}^{\top}\theta^{\star}=x^{\top}\theta^{\star}$ by assumption. Since $L_{\text{rob}}(\hat{\theta}_{\text{rst}})\geq L_{\text{std}}(\hat{\theta}_{\text{rst}})$ , $\hat{\theta}_{\text{rst}}$ has perfect consistency, achieving the lowest possible robust error (matching the standard error). ∎

D.3 Different instantiations of the general RST procedure

In the first variant, RST + PG-AT, we use multiclass logistic loss (cross-entropy) as the standard loss. The robust loss is the maximum cross-entropy loss between any perturbed input (within the set of tranformations $T(\cdot)$ ) and the label (pseudo-label in the case of unlabeled data). We set the weights such that the estimator can be written as follows.

D.3.2 TRADES

Appendix E Experimental Details

For spline simulations in Figure 2 and Figure 1, we implement the optimization of the standard and robust objectives using the basis described in (Friedman et al., 2001). The penalty matrix $M$ computes second-order finite differences of the parameters $\theta$ . We solve the min-norm objective directly using CVXPY (Diamond & Boyd, 2016). Each point in Figure 1(a) represents the average standard error over 25 trials of randomly sampled training datasets between $22$ and $1000$ samples. Shaded regions represent 1 standard deviation.

E.2 RST experiments

We compare the standard error of the augmented estimator with an estimator trained using RST. We apply RST to adversarial training algorithms in Cifar-10 using 500k unlabeled examples sourced from Tiny Images, as in (Carmon et al., 2019).

We use Wide ResNet 40-2 models (Zagoruyko & Komodakis, 2016) while varying the number of samples in Cifar-10. We sub-sample CIFAR-10 by factors of $\{1,2,5,8,10,20,40\}$ in Figure 1(a) and $\{1,2,5,8,10\}$ in Figure 1(b). We report results averaged from 2 trials for each sub-sample factor. All models are trained for 200 epochs with respect to the size of the labeled training dataset and all achieve almost 100% standard and robust training accuracy.

We evaluate the robustness of models to the strong PGD-attack with $40$ steps and $5$ restarts. In Figure 1(b), we used a simple heuristic to set the regularization strength on unlabeled data $\lambda$ in Equation (D.3.1) to be $\lambda=\min(0.9,p)$ where $p\in$ is the fraction of the original Cifar-10 dataset sampled. We set $\beta=0.5$ . Intuitively, we give more weight to the unlabeled data when the original dataset is larger, meaning that the standard estimator produces more accurate pseudo-labels.

Figure 9 shows that the robust accuracy of the RST model improves about 5-15% percentage points above the robust model (trained using PGD adversarial training) for all subsamples, including the full dataset (Tables 2,3).

We use a smaller model due to computational constraints enforced by adversarial training. Since the model is small, we could only fit adversarially augmented examples with small $\epsilon=2/255$ , while existing baselines use $\epsilon=8/255$ . Note that even for $\epsilon=2/255$ , adversarial data augmentation leads to an increase in standard error. We show that RST can fix this. While ensuring models are robust is an important goal in itself, in this work, we view adversarial training through the lens of covariate-shifted data augmentation and study how to use augmented data without increasing standard error. We show that RST preserves the other benefits of some kinds of data augmentation like increased robustness to adversarial examples.

E.2.3 Adversarial and random rotation/translations

In Table 1 (right), we use RST for adversarial and random rotation/translations, denoting these transformations as $x_{\text{adv}}$ in Equation (D.3.1). The attack model is a grid of rotations of up to 30 degrees and translations of up to $\sim 10\%$ of the image size. The grid consists of 31 linearly spaced rotations and 5 linearly spaced translations in both dimensions. The Worst-of-10 model samples 10 uniformly random transformations of each input and augment with the one where the model performs the worst (causes an incorrect prediction, if it exists). The Random model samples 1 random transformation as the augmented input. All models (besides cited models) use the WRN-40-2 architecture and are trained for 200 epochs. We use the same hyperparameters $\lambda,\beta$ as in E.2.2 for Equation (D.3.1).

Appendix F Comparison to standard self-training algorithms

The main objective of RST is to allow to perform robust training without sacrificing standard accuracy. This is done by regularizing an augmented estimator to provide labels close to a standard estimator on the unlabeled data. This is closely related to but different two broad kinds of semi-supervised learning.

Self-training (pseudo-labeling): Classical self-training does not deal with data augmentation or robustness. We view RST as a a generalization of self-training in the context of data augmentations. Here the pseudolabels are generated by a standard non-augmented estimator that is not trained on the labeled augmented points. In contrast, standard self-training would just use all labeled data to generate pseudo-labels. However, since some augmentations cause a drop in standard accuracy, and hence this would generate worse pseudo-labels than RST.

Robust consistency training: Another popular semi-supervised learning strategy is based on enforcing consistency in a model’s predictions across various perturbations of the unlabeled data (Miyato et al., 2018; Xie et al., 2019; Sajjadi et al., 2016; Laine & Aila, 2017)). RST is similar in spirit, but has an additional crucial component. We generate pseudo-labels first by performing standard training, and rather than enforcing simply consistency across perturbations, RST enforces that the unlabeled data and perturbations are matched with the pseudo-labels generated.

We begin with a 3-dimensional construction and then increase the number of dimensions. Let the domain of possible values be $\mathcal{X}=\{\mathbf{x_{1}},\mathbf{x_{2}},\mathbf{x_{3}}\}$ where

Define the data distribution through the generative process for the random feature vector $\mathbf{x}$

where $0<\delta<1$ and $\epsilon>0$ . Define the optimal linear predictor $\theta^{\star}=\mathbf{1}$ to be the all-ones vector, such that in all cases, $\mathbf{x}^{\top}\theta^{\star}=2+\delta$ . We define the consistent perturbations as

The augmented estimator will add all possible consistent perturbations of the training set as extra data $X_{\text{ext}}$ . For example, if $\mathbf{x_{1}}$ is in the training set, then the augmented estimator will add $\mathbf{x_{2}}$ as extra data since $\mathbf{x_{2}}\in T(\mathbf{x_{1}})$ . The standard error is measured by mean squared error.

We give some intuition for how augmentation can hurt standard error in this 3-dimensional example. Define $E_{1}$ to be the event that we draw $n$ samples with value $\mathbf{x_{1}}$ . Given $E_{1}$ , the standard and augmented estimators are

G.2 Construction for general d𝑑d

We construct the example by sampling $\mathbf{x}$ in 3 dimensions and then repeating the vector $d$ times. In particular, the samples are realizations of the random vector $[\mathbf{x};\mathbf{x};\mathbf{x};\dots;\mathbf{x}]$ which have dimension $3d$ and every block of 3 coordinates have the same values. Under this setup, we can show that there is a family of problems such that the difference between standard errors of the augmented and standard estimators grows to infinity as $d,n\rightarrow\infty$ .

Let the setting be defined as above, where the dimension $d$ and number of samples $n$ are such that $n/d\rightarrow\gamma$ approaches a constant. Let $p=1/d^{2}$ , $\epsilon=1/d^{3}$ , and $\delta$ be a constant. Then the ratio between standard errors of the augmented and standard estimators grows as

We define an event where the augmented estimator has high error relative to the standard estimator and bound the ratio between the standard errors of the standard and augmented estimators given this event. Define $E_{1}$ as the event that we have $n$ samples where all samples are $[\mathbf{x_{1}};\mathbf{x_{1}};\dots;\mathbf{x_{1}}]$ . The standard and augmented estimators are the corresponding repeated versions

The event $E_{1}$ occurs with probability $(1-p)^{n}+(p-\epsilon)^{n}$ . It is straightforward to verify that the respective standard errors are

and that the ratio between standard errors is

The ratio between standard errors is bounded by

as $n,d\rightarrow\infty$ , where we used Bernoulli’s inequality in the second to last step. ∎