EDBT 2026 Demo / reviewers in the wild / expert
Yuyang Deng
dblp:261/9253
· DBLP profile ↗
17ranked-venue papers
10as first author
16since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 10 first-author · 14 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MSLA-XLS-R: Enhancing hierarchical SSL representations for audio deepfake detection
Haoyang Meng, Yunqi Tang, Yuyang Deng |
Comput. Speech Lang. | 4 |
| 2026 | CoRe-DoS: Inference-time denial-of-service attack against retrieval-augmented generation
Haocheng Sun, Mingfeng Li, Yuyang Deng |
Comput. Networks | 4 |
| 2025 | Stochastic Compositional Minimax Optimization with Provable Convergence GuaranteesabstractStochastic compositional minimax problems are prevalent in machine learning, yet there exist only limited established findings on the convergence of this class of problems. In this paper, we propose a formal definition of the stochastic compositional minimax problem, which involves optimizing a minimax loss with a compositional structure either in primal, dual, or both primal and dual variables. We introduce a simple yet effective algorithm, stochastically Corrected stOchastic gradient Descent Ascent (CODA), which is a primal-dual type algorithm with compositional correction steps, and establish its convergence rate in the aforementioned three settings. We also propose a variance reduced variant, CODA+, which achieves the best-known rate on nonconvex-strongly-concave and nonconvex-concave compositional minimax problems. This work initiates the theoretical study of the stochastic compositional minimax problem in various settings and may inform modern machine learning scenarios such as domain adaptation or robust model-agnostic meta-learning. Yuyang Deng, Fuli Qiao, Mehrdad Mahdavi |
AISTATS | 1 |
| 2025 | Mixed-Sample SGD: an End-to-end Analysis of Supervised Transfer LearningabstractTheoretical works on supervised transfer learning (STL)---where the learner has access to labeled samples from both source and target distributions---have for the most part focused on statistical aspects of the problem, while efficient optimization has received less attention.
We consider the problem of designing an SGD procedure
for STL that alternates sampling between source and target data, while maintaining statistical transfer guarantees without prior knowledge of the quality of the source data.
A main algorithmic difficulty is in understanding how to design such an adaptive sub-sampling mechanism at each SGD step, to automatically gain from the source when it is informative, or bias towards the target and avoid negative transfer when the source is less informative.
We show that, such a mixed-sample SGD procedure is feasible for general prediction tasks with convex losses, rooted in tracking an abstract sequence of constrained convex programs that serve to maintain the desired transfer guarantees.
We instantiate these results in the concrete setting of linear regression with square loss, and show that the procedure converges, with $1/\sqrt{T}$ rate, to a solution whose statistical performance on the target is adaptive to the a priori unknown quality of the source. Experiments with synthetic and real datasets support the theory. Yuyang Deng, Samory Kpotufe |
NeurIPS | 1 |
| 2025 | Dynamic algorithmic awareness based on FAT evaluation: Heuristic intervention and multidimensional predictionabstractAbstract As the widespread use of algorithms and artificial intelligence (AI) technologies, understanding the interaction process of human–algorithm interaction becomes increasingly crucial. From the human perspective, algorithmic awareness is recognized as a significant factor influencing how users evaluate algorithms and engage with them. In this study, a formative study identified four dimensions of algorithmic awareness: conceptions awareness (AC), data awareness (AD), functions awareness (AF), and risks awareness (AR). Subsequently, we implemented a heuristic intervention and collected data on users' algorithmic awareness and FAT (fairness, accountability, and transparency) evaluation in both pre‐test and post‐test stages (N = 622). We verified the dynamics of algorithmic awareness and FAT evaluation through fuzzy clustering and identified three patterns of FAT evaluation changes: “Stable high rating pattern,” “Variable medium rating pattern,” and “Unstable low rating pattern.” Using the clustering results and FAT evaluation scores, we trained classification models to predict different dimensions of algorithmic awareness by applying different machine learning techniques, namely Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), Linear Discriminant Analysis (LDA), and XGBoost (XGB). Comparatively, experimental results show that the SVM algorithm accomplishes the task of predicting the four dimensions of algorithmic awareness with better results and interpretability. Its F1 scores are 0.6377, 0.6780, 0.6747, and 0.75. These findings hold great potential for informing human‐centered algorithmic practices and HCI design. Jing Liu 0071, Dan Wu 0003, Guoye Sun, Yuyang Deng |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2024 | On the Generalization Ability of Unsupervised PretrainingabstractRecent advances in unsupervised learning have shown that unsupervised pre-training, followed by fine-tuning, can improve model generalization. However, a rigorous understanding of how the representation function learned on an unlabeled dataset affects the generalization of the fine-tuned model is lacking. Existing theoretical research does not adequately account for the heterogeneity of the distribution and tasks in pre-training and fine-tuning stage. To bridge this gap, this paper introduces a novel theoretical framework that illuminates the critical factor influencing the transferability of knowledge acquired during unsupervised pre-training to the subsequent fine-tuning phase, ultimately affecting the generalization capabilities of the fine-tuned model on downstream tasks. We apply our theoretical framework to analyze generalization bound of two distinct scenarios: Context Encoder pre-training with deep neural networks and Masked Autoencoder pre-training with deep transformers, followed by fine-tuning on a binary classification task. Finally, inspired by our findings, we propose a novel regularization method during pre-training to further enhances the generalization of fine-tuned model. Overall, our results contribute to a better understanding of unsupervised pre-training and fine-tuning paradigm, and can shed light on the design of more effective pre-training algorithms. Yuyang Deng, Junyuan Hong, Mehrdad Mahdavi |
AISTATS | 1 |
| 2024 | Collaborative Learning with Different Labeling FunctionsabstractWe study a variant of Collaborative PAC Learning, in which we aim to learn an accurate classifier for each of the $n$ data distributions, while minimizing the number of samples drawn from them in total. Unlike in the usual collaborative learning setup, it is not assumed that there exists a single classifier that is simultaneously accurate for all distributions. We show that, when the data distributions satisfy a weaker realizability assumption, which appeared in (Crammer & Mansour, 2012) in the context of multi-task learning, sample-efficient learning is still feasible. We give a learning algorithm based on Empirical Risk Minimization (ERM) on a natural augmentation of the hypothesis class, and the analysis relies on an upper bound on the VC dimension of this augmented class. In terms of the computational efficiency, we show that ERM on the augmented hypothesis class is $\mathsf{NP}$-hard, which gives evidence against the existence of computationally efficient learners in general. On the positive side, for two special cases, we give learners that are both sample- and computationally-efficient. Yuyang Deng, Mingda Qiao |
ICML | 1 |
| 2024 | Generating and encouraging: An effective framework for solving class imbalance in multimodal emotion recognition conversation
Qianer Li, Peijie Huang, Yuhong Xu, Yuyang Deng, Shangjian Yin |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Ta-Adapter: Enhancing few-shot CLIP with task-aware encoders
Yifan Zhang 0009, Yuyang Deng, Jianfeng Lin 0004, Binqiang Huang, Wenhao Yu 0001 |
Pattern Recognit. | 3 |
| 2023 | Early ChatGPT User Portrait through the Lens of DataabstractSince its launch, ChatGPT has achieved remarkable success as a versatile conversational AI platform, drawing millions of users worldwide and garnering widespread recognition across academic, industrial, and general communities. This paper aims to point a portrait of early GPT users and understand how they evolved. Specific questions include their topics of interest and their potential careers; and how this changes over time. We conduct a detailed analysis of real-world ChatGPT datasets with multi-turn conversations between users and ChatGPT. Through a multi-pronged approach, we quantify conversation dynamics by examining the number of turns, then gauge sentiment to understand user sentiment variations, and finally employ Latent Dirichlet Allocation (LDA) to discern overarching topics within the conversation. By understanding shifts in user demographics and interests, we aim to shed light on the changing nature of human-AI interaction and anticipate future trends in user engagement with language models. Yuyang Deng, Ni Zhao |
IEEE Big Data | 1 |
| 2023 | Mixture Weight Estimation and Model Prediction in Multi-source Multi-target Domain AdaptationabstractWe consider a problem of learning a model from multiple sources with the goal to perform
well on a new target distribution. Such problem arises in
learning with data collected from multiple sources (e.g. crowdsourcing) or
learning in distributed systems, where the data can be highly heterogeneous. The
goal of learner is to mix these data sources in a target-distribution aware way and
simultaneously minimize the empirical risk on the mixed source. The literature has made some tangible advancements in establishing
theory of learning on mixture domain. However, there are still two unsolved problems. Firstly, how to estimate the optimal mixture of sources, given a target domain; Secondly, when there are numerous target domains, we have to solve empirical risk minimization for each target on possibly unique mixed source data , which is computationally expensive. In this paper we address both problems efficiently and with guarantees.
We cast the first problem, mixture weight estimation as convex-nonconcave compositional minimax, and propose an efficient stochastic
algorithm with provable stationarity guarantees.
Next, for the second problem, we identify that for certain regime,
solving ERM for each target domain individually can be avoided, and instead parameters for a target optimal
model can be viewed as a non-linear function on
a space of the mixture coefficients.
To this end, we show that in offline setting, a GD-trained overparameterized neural network can provably learn such function.
Finally, we also consider an online setting and propose an label efficient online algorithm, which predicts parameters for new models given arbitrary sequence of mixing coefficients, while enjoying optimal regret. Yuyang Deng, Ilja Kuzborskij, Mehrdad Mahdavi |
NeurIPS | 1 |
| 2023 | Distributed Personalized Empirical Risk MinimizationabstractThis paper advocates a new paradigm Personalized Empirical Risk Minimization (PERM) to facilitate learning from heterogeneous data sources without imposing stringent constraints on computational resources shared by participating devices. In PERM, we aim at learning a distinct model for each client by personalizing the aggregation of local empirical losses by effectively estimating the statistical discrepancy among data distributions, which entails optimal statistical accuracy for all local distributions and overcomes the data heterogeneity issue. To learn personalized models at scale, we propose a distributed algorithm that replaces the standard model averaging with model shuffling to simultaneously optimize
PERM objectives for all devices. This also allows to learn distinct model architectures (e.g., neural networks with different number of parameters) for different clients, thus confining to underlying memory and compute resources of individual clients. We rigorously analyze the convergence of proposed algorithm and conduct experiments that corroborates the effectiveness of proposed paradigm. Yuyang Deng, Mohammad Mahdi Kamani, Pouria Mahdavinia, Mehrdad Mahdavi |
NeurIPS | 1 |
| 2023 | Understanding Deep Gradient Leakage via Inversion Influence FunctionsabstractDeep Gradient Leakage (DGL) is a highly effective attack that recovers private training images from gradient vectors.
This attack casts significant privacy challenges on distributed learning from clients with sensitive data, where clients are required to share gradients.
Defending against such attacks requires but lacks an understanding of when and how privacy leakage happens, mostly because of the black-box nature of deep networks.
In this paper, we propose a novel Inversion Influence Function (I$^2$F) that establishes a closed-form connection between the recovered images and the private gradients by implicitly solving the DGL problem.
Compared to directly solving DGL, I$^2$F is scalable for analyzing deep networks, requiring only oracle access to gradients and Jacobian-vector products.
We empirically demonstrate that I$^2$F effectively approximated the DGL generally on different model architectures, datasets, modalities, attack implementations, and perturbation-based defenses.
With this novel tool, we provide insights into effective gradient perturbation directions, the unfairness of privacy protection, and privacy-preferred model initialization.
Our codes are provided in https://github.com/illidanlab/inversion-influence-function. Haobo Zhang 0002, Junyuan Hong, Yuyang Deng, Mehrdad Mahdavi |
NeurIPS | 3 |
| 2022 | Local SGD Optimizes Overparameterized Neural Networks in Polynomial TimeabstractIn this paper we prove that Local (S)GD (or FedAvg) can optimize deep neural networks with Rectified Linear Unit (ReLU) activation function in polynomial time. Despite the established convergence theory of Local SGD on optimizing general smooth functions in communication-efficient distributed optimization, its convergence on non-smooth ReLU networks still eludes full theoretical understanding. The key property used in many Local SGD analysis on smooth function is gradient Lipschitzness, so that the gradient on local models will not drift far away from that on averaged model. However, this decent property does not hold in networks with non-smooth ReLU activation function. We show that, even though ReLU network does not admit gradient Lipschitzness property, the difference between gradients on local models and average model will not change too much, under the dynamics of Local SGD. We validate our theoretical results via extensive experiments. This work is the first to show the convergence of Local SGD on non-smooth functions, and will shed lights on the optimization theory of federated training of deep neural networks. Yuyang Deng, Mohammad Mahdi Kamani, Mehrdad Mahdavi |
AISTATS | 1 |
| 2022 | Tight Analysis of Extra-gradient and Optimistic Gradient Methods For Nonconvex Minimax ProblemsabstractDespite the established convergence theory of Optimistic Gradient Descent Ascent (OGDA) and Extragradient (EG) methods for the convex-concave minimax problems, little is known about the theoretical guarantees of these methods in nonconvex settings. To bridge this gap, for the first time, this paper establishes the convergence of OGDA and EG methods under the nonconvex-strongly-concave (NC-SC) and nonconvex-concave (NC-C) settings by providing a unified analysis through the lens of single-call extra-gradient methods. We further establish lower bounds on the convergence of GDA/OGDA/EG, shedding light on the tightness of our analysis. We also conduct experiments supporting our theoretical results. We believe our results will advance the theoretical understanding of OGDA and EG methods for solving complicated nonconvex minimax real-world problems, e.g., Generative Adversarial Networks (GANs) or robust neural networks training. Pouria Mahdavinia, Yuyang Deng, Haochuan Li, Mehrdad Mahdavi |
NeurIPS | 2 |
| 2021 | Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication EfficiencyabstractLocal SGD is a promising approach to overcome the communication overhead in distributed learning by reducing the synchronization frequency among worker nodes. Despite the recent theoretical advances of local SGD in empirical risk minimization, the efficiency of its counterpart in minimax optimization remains unexplored. Motivated by large scale minimax learning problems, such as adversarial robust learning and GANs, we propose local Stochastic Gradient Descent Ascent (local SGDA), where the primal and dual variables can be trained locally and averaged periodically to significantly reduce the number of communications. We show that local SGDA can provably optimize distributed minimax problems in both homogeneous and heterogeneous data with reduced number of communications and establish convergence rates under strongly-convex-strongly-concave and nonconvex-strongly-concave settings. In addition, we propose a novel variant, dubbed as local SGDA+, to solve nonconvex-nonconcave problems. We also give corroborating empirical evidence on different distributed minimax problems. Yuyang Deng, Mehrdad Mahdavi |
AISTATS | 1 |
| 2020 | Distributionally Robust Federated AveragingabstractIn this paper, we study communication efficient distributed algorithms for distributionally robust federated learning via periodic averaging with adaptive sampling. In contrast to standard empirical risk minimization, due to the minimax structure of the underlying optimization problem, a key difficulty arises from the fact that the global parameter that controls the mixture of local losses can only be updated infrequently on the global stage. To compensate for this, we propose a Distributionally Robust Federated Averaging (DRFA) algorithm that employs a novel snapshotting scheme to approximate the accumulation of history gradients of the mixing parameter. We analyze the convergence rate of DRFA in both convex-linear and nonconvex-linear settings. We also generalize the proposed idea to objectives with regularization on the mixture parameter and propose a proximal variant, dubbed as DRFA-Prox, with provable convergence rates. We also analyze an alternative optimization method for regularized case in strongly-convex-strongly-concave and non-convex (under PL condition)-strongly-concave settings. To the best of our knowledge, this paper is the first to solve distributionally robust federated learning with reduced communication, and to analyze the efficiency of local descent methods on distributed minimax problems. We give corroborating experimental evidence for our theoretical results in federated learning settings. Yuyang Deng, Mohammad Mahdi Kamani, Mehrdad Mahdavi |
NeurIPS | 1 |