Mengyang Li 0001

dblp:166/3118-1 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-8958-3163ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Difficulty-Aware Learning Curve Extrapolation
Mengyang Li 0001, Pinlong Zhao
AAAI1
2026 What Do LLMs Learn First? Asymmetric Learning Dynamics of Input Complexity and Output Ambiguity in Preference Alignment
abstract
Direct Preference Optimization (DPO) has become a standard approach for aligning large language models with human preferences, yet existing methods treat all preference pairs uniformly during training.We identify two distinct sources of learning difficulty: Input Complexity (IC), capturing prompt understanding challenges, and Output Ambiguity (OA), measuring preference discrimination difficulty.Through systematic analysis, we demonstrate that these dimensions induce asymmetric learning dynamics, with IC-related competencies developing rapidly in early training while OA-related competencies emerge more gradually.Building on this observation, we propose DECOPO, a training framework that maintains separate, adaptive pacing schedules for each dimension.Experiments on UltraFeedback show that DE-COPO achieves 42.3% length-controlled win rate on AlpacaEval 2.0 and 7.66 on MT-Bench, outperforming curriculum baselines by 2.1% and 0.21 points respectively, while matching full-data baseline performance with only 75% of training samples.
Mengyang Li 0001, Pinlong Zhao
ACL (1)1
2026 Unequal Vulnerability: The Differential Impact of Label Flipping Attacks Across Classes
abstract
Label flipping attacks stand as a potent and practical threat to the integrity of machine learning models. While extensive research has focused on designing sophisticated attack and defense mechanisms, the underlying factors that govern a model's susceptibility remain underexplored. This paper reveals a critical phenomenon: the impact of label flipping attacks is highly differential across classes, strongly correlated with the intrinsic confusability between the source and target classes. We provide a rigorous theoretical analysis, demonstrating that a lower standardized separation between classes fundamentally leads to greater vulnerability. Grounded in this insight, we propose Confusability-Aware Contrastive Learning (CACL), a targeted defense that maximizes the feature-space separation for the most vulnerable class pairs. Extensive experiments validate the strong link between class separability and vulnerability, and show that CACL significantly mitigates the attack's impact while providing superior protection for the most susceptible classes. Our code is available at https://github.com/Pinlong-Zhao/Unequal-Vulnerability.
Pinlong Zhao, Mengyang Li 0001, Pengfei Jiao, Huijun Tang, Ou Wu 0001
WWW2
2025 Toward learnable and interpretable data Shapley valuation for deep learning
Mengyang Li 0001, Weiyao Zhu, Ou Wu 0001
Knowl. Based Syst.1
2025 Delving Into the Training Dynamics for Image Classification
abstract
In recent years, there has been an increase in exploring and applying the training dynamics (TD) of deep neural networks (DNNs). Current studies typically rely on quite limited TD quantities and apply their sequences to understand or aid training. This study investigates how to create more effective TD representations, and then apply them to improve the training process of real learning tasks. Specifically, first, an epoch-wise vector comprising 142-dimensional TD quantities, such as loss, is extracted for each sample. Second, a new learning strategy with both self-supervised and supervised learning is designed to learn the deep TD representation of each sample on 200 typical image classification tasks. Third, two novel methods for noisy label detection and imbalance learning, respectively, are presented based on deep TD representations. Our study reveals that neighborhoods and logits are the most important TD quantities, unlike the traditional research that focuses on loss and margin. Moreover, our method based on deep TD representations achieves better performance and demonstrates that high-level TD quantities can facilitate understanding model training, leading to improvements in practical learning tasks, such as noisy label detection and imbalance learning. All the codes are available at https://github.com/limengyang1992/TD_Exploring.
Mengyang Li 0001, Xiaoling Zhou, Ou Wu 0001
IEEE Trans. Image Process.1
2024 Revisiting the Effective Number Theory for Imbalanced Learning
abstract
Imbalanced learning is a traditional yet hot research subarea in machine learning. There are a huge number of imbalanced learning methods proposed in previous literature. This study focuses on one of the most popular imbalanced learning strategies, namely, sample reweighting. The key issue is how to calculate the weights of samples in training. While most studies have relied on intuitive theoretical or heuristic inspirations, few studies have attempted to establish a comprehensive theoretical path for weight calculation. A recent study utilizes the effective number theory for random covering to construct a theoretical weighting framework. In this study, we conduct a deep analysis to theoretically reveal the defects in the existing effective number-based weighting theory. An enhanced effective number theory is established in which data scatter and covering offset among different categories are involved. Subsequently, a new weight calculation manner is proposed based on our new theory, yielding a new loss, namely, NENum loss. In this loss, weights are sample-wise instead of category-wise used in the existing effective number-based weighting. Furthermore, another novel loss that combines weighting and logit perturbation is designed inspired the limitations of the NENum loss. Meta learning is employed to optimize the concrete calculation based on sample-wise training dynamics. We conduct extensive experiments on benchmark imbalanced and standard data corpora. Results validate the reasonableness of our enhanced theory and the effectiveness of the proposed methodology.
Ou Wu 0001, Mengyang Li 0001
IEEE Trans. Knowl. Data Eng.2
2024 Investigating the Sample Weighting Mechanism Using an Interpretable Weighting Framework
abstract
Training deep learning models with unequal sample weights has been shown to enhance model performance in various typical learning scenarios, particularly for imbalanced and noisy-label learning scenarios. A deep understanding of the weighting mechanism facilitates the application of existing weighting strategies and illuminates the design of new weighting strategies for real learning tasks. Scholars have focused on exploring existing weighting methods. However, their studies mainly establish how the weights of samples influence the model training. Little headway is made on the weighting mechanism, i.e., which and how the characteristics of a sample influence its weight. In this study, we adopt a data-driven approach to investigate the weighting mechanism by utilizing an interpretable weighting framework. First, a wide range of sample characteristics is extracted from the classifier network during training. Second, the extracted characteristics are fed into a new neural regression tree (NRT), which is a tree model implemented by a neural network, and its output is the weight of the input sample. Third, the NRT is trained using meta-learning within the whole training process. Once the NRT is learned, the weighting mechanism, including the importance of weighting characteristics, prior modes, and specific weighting rules, can be obtained. We conduct extensive experiments on benchmark noisy and imbalanced data corpora. A package of weighting mechanisms is derived from the learned NRT. Furthermore, our proposed interpretable weighting framework exhibits superior performance in comparison to existing weighting strategies.
Xiaoling Zhou, Ou Wu 0001, Mengyang Li 0001
IEEE Trans. Knowl. Data Eng.3
2024 Class-Level Logit Perturbation
abstract
Features, logits, and labels are the three primary data when a sample passes through a deep neural network (DNN). Feature perturbation and label perturbation receive increasing attention in recent years. They have been proven to be useful in various deep learning approaches. For example, (adversarial) feature perturbation can improve the robustness or even generalization capability of learned models. However, limited studies have explicitly explored for the perturbation of logit vectors. This work discusses several existing methods related to class-level logit perturbation. A unified viewpoint between regular/irregular data augmentation and loss variations incurred by logit perturbation is established. A theoretical analysis is provided to illuminate why class-level logit perturbation is useful. Accordingly, new methodologies are proposed to explicitly learn to perturb logits for both the single-label and multilabel classification tasks. Meta-learning is also leveraged to determine the regular or irregular augmentation for each class. Extensive experiments on benchmark image classification datasets and their long-tail versions indicated the competitive performance of our learning method. As it only perturbs on logit, it can be used as a plug-in to fuse with any existing classification algorithms. All the codes are available at https://github.com/limengyang1992/lpl.
Mengyang Li 0001, Fengguang Su, Ou Wu 0001, Ji Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Investigating annotation noise for named entity recognition
Yingchun Ye, Mengyang Li 0001, Ji Zhang 0001, Ou Wu 0001
Neural Comput. Appl.3
2022 Logit Perturbation
abstract
Features, logits, and labels are the three primary data when a sample passes through a deep neural network. Feature perturbation and label perturbation receive increasing attention in recent years. They have been proven to be useful in various deep learning approaches. For example, (adversarial) feature perturbation can improve the robustness or even generalization capability of learned models. However, limited studies have explicitly explored for the perturbation of logit vectors. This work discusses several existing methods related to logit perturbation. Based on a unified viewpoint between positive/negative data augmentation and loss variations incurred by logit perturbation, a new method is proposed to explicitly learn to perturb logits. A comparative analysis is conducted for the perturbations used in our and existing methods. Extensive experiments on benchmark image classification data sets and their long-tail versions indicated the competitive performance of our learning method. In addition, existing methods can be further improved by utilizing our method.
Mengyang Li 0001, Fengguang Su, Ou Wu 0001, Ji Zhang 0001
AAAI1
2022 Two-Level LSTM for Sentiment Analysis With Lexicon Embedding and Polar Flipping
abstract
Sentiment analysis is a key component in various text mining applications. Numerous sentiment classification techniques, including conventional and deep-learning-based methods, have been proposed in the literature. In most existing methods, a high-quality training set is assumed to be given. Nevertheless, constructing a high-quality training set that consists of highly accurate labels is challenging in real applications. This difficulty stems from the fact that text samples usually contain complex sentiment representations, and their annotation is subjective. We address this challenge in this study by leveraging a new labeling strategy and utilizing a two-level long short-term memory network to construct a sentiment classifier. Lexical cues are useful for sentiment analysis, and they have been utilized in conventional studies. For example, polar and negation words play important roles in sentiment analysis. A new encoding strategy, that is, ρ -hot encoding, is proposed to alleviate the drawbacks of one-hot encoding and, thus, effectively incorporate useful lexical cues. Moreover, the sentimental polarity of a word may change in different sentences due to label noise or context. A flipping model is proposed to model the polar flipping of words in a sentence. We compile three Chinese datasets on the basis of our label strategy and proposed methodology. Experiments demonstrate that the proposed method outperforms state-of-the-art algorithms on both benchmark English data and our compiled Chinese data.
Ou Wu 0001, Tao Yang 0033, Mengyang Li 0001
IEEE Trans. Cybern.3