Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yunjuan Wang

dblp:31/560 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 55% Learning theory · 16% Motion planning and robot control · 9%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
3.652025
Backdoor Attacks in Token Selection of Attention Mechanism · ICML 2025
Stability and Generalization of Adversarial Training for Shallow Neural Networks with Smooth Activation · NeurIPS 2024
Benign Overfitting in Adversarial Training of Neural Networks · ICML 2024
Machine learning › Learning theory
generalization bounds
2.032024
Stability and Generalization of Adversarial Training for Shallow Neural Networks with Smooth Activation · NeurIPS 2024
On the Stability and Generalization of Meta-Learning · NeurIPS 2024
Robust Learning for Data Poisoning Attacks · ICML 2021
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
1.522024
Stability and Generalization of Adversarial Training for Shallow Neural Networks with Smooth Activation · NeurIPS 2024
Benign Overfitting in Adversarial Training of Neural Networks · ICML 2024
Robotics › Motion planning and robot control
stability analysis
1.522024
Stability and Generalization of Adversarial Training for Shallow Neural Networks with Smooth Activation · NeurIPS 2024
On the Stability and Generalization of Meta-Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.322024
Benign Overfitting in Adversarial Training of Neural Networks · ICML 2024
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Trustworthy machine learning › robustness › data poisoning
backdoor attack
0.912025
Backdoor Attacks in Token Selection of Attention Mechanism · ICML 2025
Security and privacy of machine learning › adversarial attack
backdoor attack
0.912025
Backdoor Attacks in Token Selection of Attention Mechanism · ICML 2025
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarially robust learning
0.812024
Adversarially Robust Hypothesis Transfer Learning · ICML 2024
Machine learning › Learning theory › overfitting
benign overfitting
0.812024
Benign Overfitting in Adversarial Training of Neural Networks · ICML 2024
Machine learning › Transfer learning and domain adaptation › parameter-based transfer learning
hypothesis transfer learning
0.812024
Adversarially Robust Hypothesis Transfer Learning · ICML 2024
Machine learning › Transfer learning and domain adaptation › instance weighting
importance weighting
0.712023
Leveraging Importance Weights in Subset Selection · ICLR 2023
Machine learning › Efficient and distributed learning
subset selection
0.712023
Leveraging Importance Weights in Subset Selection · ICLR 2023
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.612022
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Deep learning architectures and training › training dynamics
lazy training regime
0.612022
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Deep learning architectures and training
overparameterized neural network
0.612022
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Trustworthy machine learning › robustness
poisoning attack defense
0.512021
Robust Learning for Data Poisoning Attacks · ICML 2021

Methods — techniques the papers use, named apart from their topics

gradient descent · 2.5token selection analysis · 1.7regularized empirical risk minimization · 1.5proximal stochastic adversarial training · 0.8moreau envelope smoothing · 0.8generalization bounds · 0.8early stopping · 0.8algorithmic stability · 0.8adversarial training · 0.8importance weighting · 0.7
YearPublicationVenuePosition
2025 Backdoor Attacks in Token Selection of Attention Mechanism
abstract
Despite the remarkable success of large foundation models across a range of tasks, they remain susceptible to security threats such as backdoor attacks. By injecting poisoned data containing specific triggers during training, adversaries can manipulate model predictions in a targeted manner. While prior work has focused on empirically designing and evaluating such attacks, a rigorous theoretical understanding of when and why they succeed is lacking. In this work, we analyze backdoor attacks that exploit the token selection process within attention mechanisms--a core component of transformer-based architectures. We show that single-head self-attention transformers trained via gradient descent can interpolate poisoned training data. Moreover, we prove that when the backdoor triggers are sufficiently strong but not overly dominant, attackers can successfully manipulate model predictions. Our analysis characterizes how adversaries manipulate token selection to alter outputs and identifies the theoretical conditions under which these attacks succeed. We validate our findings through experiments on synthetic datasets.
Yunjuan Wang, Raman Arora
ICML1
2025 When Does Curriculum Learning Help? A Theoretical Perspective
abstract
Curriculum learning has emerged as an effective strategy to enhance the training efficiency and generalization of machine learning models. However, its theoretical underpinnings remain relatively underexplored. In this work, we develop a theoretical framework for curriculum learning based on biased regularized empirical risk minimization (RERM), identifying conditions under which curriculum learning provably improves generalization. We introduce a sufficient condition that characterizes a "good" curriculum and analyze a multi-task curriculum framework, where solving a sequence of convex tasks can facilitate better generalization. We also demonstrate how these theoretical insights translate to practical benefits when using stochastic gradient descent (SGD) as an optimization method. Beyond convex settings, we explore the utility of curriculum learning for non-convex tasks. Empirical evaluations on synthetic datasets and MNIST validate our theoretical findings and highlight the practical efficacy of curriculum-based training.
Raman Arora, Yunjuan Wang, Kaibo Zhang
NeurIPS2
2024 Adversarially Robust Hypothesis Transfer Learning
abstract
In this work, we explore Hypothesis Transfer Learning (HTL) under adversarial attacks. In this setting, a learner has access to a training dataset of size $n$ from an underlying distribution $\mathcal{D}$ and a set of auxiliary hypotheses. These auxiliary hypotheses, which can be viewed as prior information originating either from expert knowledge or as pre-trained foundation models, are employed as an initialization for the learning process. Our goal is to develop an adversarially robust model for $\mathcal{D}$. We begin by examining an adversarial variant of the regularized empirical risk minimization learning rule that we term A-RERM. Assuming a non-negative smooth loss function with a strongly convex regularizer, we establish a bound on the robust generalization error of the hypothesis returned by A-RERM in terms of the robust empirical loss and the quality of the initialization. If the initialization is good, i.e., there exists a weighted combination of auxiliary hypotheses with a small robust population loss, the bound exhibits a fast rate of $\mathcal{O}(1/n)$. Otherwise, we get the standard rate of $\mathcal{O}(1/\sqrt{n})$. Additionally, we provide a bound on the robust excess risk which is similar in nature, albeit with a slightly worse rate. We also consider solving the problem using a practical variant, namely proximal stochastic adversarial training, and present a bound that depends on the initialization. This bound has the same dependence on the sample size as the ARERM bound, except for an additional term that depends on the size of the adversarial perturbation.
Yunjuan Wang, Raman Arora
ICML1
2024 Benign Overfitting in Adversarial Training of Neural Networks
abstract
Benign overfitting is the phenomenon wherein none of the predictors in the hypothesis class can achieve perfect accuracy (i.e., non-realizable or noisy setting), but a model that interpolates the training data still achieves good generalization. A series of recent works aim to understand this phenomenon for regression and classification tasks using linear predictors as well as two-layer neural networks. In this paper, we study such a benign overfitting phenomenon in an adversarial setting. We show that under a distributional assumption, interpolating neural networks found using adversarial training generalize well despite inference-time attacks. Specifically, we provide convergence and generalization guarantees for adversarial training of two-layer networks (with smooth as well as non-smooth activation functions) showing that under moderate $\ell_2$ norm perturbation budget, the trained model has near-zero robust training loss and near-optimal robust generalization error. We support our theoretical findings with an empirical study on synthetic and real-world data.
Yunjuan Wang, Kaibo Zhang, Raman Arora
ICML1
2024 On the Stability and Generalization of Meta-Learning
abstract
We focus on developing a theoretical understanding of meta-learning. Given multiple tasks drawn i.i.d. from some (unknown) task distribution, the goal is to find a good pre-trained model that can be adapted to a new, previously unseen, task with little computational and statistical overhead. We introduce a novel notion of stability for meta-learning algorithms, namely *uniform meta-stability*. We instantiate two uniformly meta-stable learning algorithms based on regularized empirical risk minimization and gradient descent and give explicit generalization bounds for convex learning problems with smooth losses and for weakly convex learning problems with non-smooth losses. Finally, we extend our results to stochastic and adversarially robust variants of our meta-learning algorithm.
Yunjuan Wang, Raman Arora
NeurIPS1
2024 Stability and Generalization of Adversarial Training for Shallow Neural Networks with Smooth Activation
abstract
Adversarial training has emerged as a popular approach for training models that are robust to inference-time adversarial attacks. However, our theoretical understanding of why and when it works remains limited. Prior work has offered generalization analysis of adversarial training, but they are either restricted to the Neural Tangent Kernel (NTK) regime or they make restrictive assumptions about data such as (noisy) linear separability or robust realizability. In this work, we study the stability and generalization of adversarial training for two-layer networks **without any data distribution assumptions** and **beyond the NTK regime**. Our findings suggest that for networks with *any given initialization* and *sufficiently large width*, the generalization bound can be effectively controlled via early stopping. We further improve the generalization bound by leveraging smoothing using Moreau’s envelope.
Kaibo Zhang, Yunjuan Wang, Raman Arora
NeurIPS2
2023 Leveraging Importance Weights in Subset Selection
Gui Citovsky, Giulia DeSalvo, Sanjiv Kumar, Srikumar Ramalingam, Afshin Rostamizadeh, Yunjuan Wang
ICLR6
2022 Adversarial Robustness is at Odds with Lazy Training
abstract
Recent works show that adversarial examples exist for random neural networks [Daniely and Schacham, 2020] and that these examples can be found using a single step of gradient ascent [Bubeck et al., 2021]. In this work, we extend this line of work to ``lazy training'' of neural networks -- a dominant model in deep learning theory in which neural networks are provably efficiently learnable. We show that over-parametrized neural networks that are guaranteed to generalize well and enjoy strong computational guarantees remain vulnerable to attacks generated using a single step of gradient ascent.
Yunjuan Wang, Enayat Ullah, Poorya Mianjy, Raman Arora
NeurIPS1
2021 Robust Learning for Data Poisoning Attacks
abstract
We investigate the robustness of stochastic approximation approaches against data poisoning attacks. We focus on two-layer neural networks with ReLU activation and show that under a specific notion of separability in the RKHS induced by the infinite-width network, training (finite-width) networks with stochastic gradient descent is robust against data poisoning attacks. Interestingly, we find that in addition to a lower bound on the width of the network, which is standard in the literature, we also require a distribution-dependent upper bound on the width for robust generalization. We provide extensive empirical evaluations that support and validate our theoretical results.
Yunjuan Wang, Poorya Mianjy, Raman Arora
ICML1