Ying-Peng Tang

dblp:234/7906 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-1529-9714ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Efficient and distributed learning · 64% Learning paradigms · 13% Learning theory · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Energy systems and smart grids · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
active learning
3.252026
Active Learning for Multiple Target Models · IEEE Trans. Pattern Anal. Mach. Intell. 2026
One-shot Active Learning Based on Lewis Weight Sampling for Multiple Deep Models · ICLR 2024
Active Learning for Multiple Target Models · NeurIPS 2022
Machine learning › Efficient and distributed learning
federated learning
1.722025
Class-wise Balancing Data Replay for Federated Class-Incremental Learning · NeurIPS 2025
Efficient Heterogeneity-Aware Federated Active Data Selection · ICML 2025
Machine learning › Learning paradigms › incremental learning
data replay
0.912025
Class-wise Balancing Data Replay for Federated Class-Incremental Learning · NeurIPS 2025
Machine learning › Efficient and distributed learning › federated learning › label-efficient federated learning
federated active learning
0.912025
Efficient Heterogeneity-Aware Federated Active Data Selection · ICML 2025
Machine learning › Efficient and distributed learning › federated learning › federated continual learning
federated class-incremental learning
0.912025
Class-wise Balancing Data Replay for Federated Class-Incremental Learning · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
0.812024
One-shot Active Learning Based on Lewis Weight Sampling for Multiple Deep Models · ICLR 2024
Machine learning › Efficient and distributed learning › active learning
label complexity
0.612022
Active Learning for Multiple Target Models · NeurIPS 2022
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
human-in-the-loop learning
0.512021
Dual Active Learning for Both Model and Data Selection · IJCAI 2021
Machine learning › Learning paradigms › curriculum learning
self-paced learning
0.412019
Self-Paced Active Learning: Query the Right Thing at the Right Time · AAAI 2019
Machine learning › Learning paradigms
class imbalance
0.312025
Class-wise Balancing Data Replay for Federated Class-Incremental Learning · NeurIPS 2025
Distributed systems › distributed machine learning
distributed training
0.312025
QFEVAL: Quantum Federated Ensembled Variational Adaptive Learning for Dynamic Security Assessment in Cyber-Physical Systems · IEEE J. Sel. Areas Commun. 2025
Distributed systems › distributed machine learning
federated learning
0.312025
QFEVAL: Quantum Federated Ensembled Variational Adaptive Learning for Dynamic Security Assessment in Cyber-Physical Systems · IEEE J. Sel. Areas Commun. 2025
Machine learning › Optimization for machine learning › hyperparameter optimization
combined algorithm selection and hyperparameter optimization
0.112021
Dual Active Learning for Both Model and Data Selection · IJCAI 2021
Machine learning › Optimization for machine learning
hyperparameter optimization
0.112021
Dual Active Learning for Both Model and Data Selection · IJCAI 2021

Methods — techniques the papers use, named apart from their topics

variational quantum circuit · 1.7quantum machine learning · 1.7federated learning · 1.7disagreement-based sampling · 1.6agnostic active learning · 1.0temperature scaling · 0.9leverage score sampling · 0.9contrastive learning · 0.9class-wise balancing · 0.9active learning · 0.9FedSVD · 0.9lewis weight sampling · 0.8l_p regression · 0.8
YearPublicationVenuePosition
2026 Active Learning for Multiple Target Models
abstract
We present a novel setting of active learning (AL) where multiple target models are simultaneously learned. This setting arises in real-world applications where machine learning systems require training multiple models on the same labeled dataset to accommodate diverse devices with varying computational resources. However, traditional AL methods are often limited by their model dependence and non-transferability. In this paper, we address the question of whether an effective AL method can be designed for multiple target models. We analyze the query complexity of active and passive learning in this setting and demonstrate the potential for AL to achieve improved query complexity. Based on this insight, we further propose an agnostic AL sampling strategy which selects examples located in the joint disagreement regions of different target models. Experimental evaluations on classification and regression benchmarks validate the effectiveness of our approach over traditional AL methods.
Sheng-Jun Huang, Yi Li 0002, Ying-Peng Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Efficient Heterogeneity-Aware Federated Active Data Selection
abstract
Federated Active Learning (FAL) aims to learn an effective global model, while minimizing label queries. Owing to privacy requirements, it is challenging to design effective active data selection schemes due to the lack of cross-client query information. In this paper, we bridge this important gap by proposing the Federated Active data selection by LEverage score sampling (FALE) method. It is designed for regression tasks in the presence of non-i.i.d. client data to enable the server to select data globally in a privacy-preserving manner. Based on FedSVD, FALE aims to estimate the utility of unlabeled data and perform data selection via leverage score sampling. Besides, a secure model learning framework is designed for federated regression tasks to exploit supervision. FALE can operate without requiring an initial labeled set and select the instances in a single pass, significantly reducing communication overhead. Theoretical analyze establishes the query complexity for FALE to achieve constant factor approximation and relative error approximation. Extensive experiments on 11 benchmark datasets demonstrate significant improvements of FALE over existing state-of-the-art methods.
Ying-Peng Tang, Chao Ren 0006, Xiaoli Tang 0001, Sheng-Jun Huang, Han Yu 0001
ICML1
2025 Class-wise Balancing Data Replay for Federated Class-Incremental Learning
abstract
Federated Class Incremental Learning (FCIL) aims to collaboratively process continuously increasing incoming tasks across multiple clients. Among various approaches, data replay has become a promising solution, which can alleviate forgetting by reintroducing representative samples from previous tasks. However, their performance is typically limited by class imbalance, both within the replay buffer due to limited global awareness and between replayed and newly arrived classes. To address this issue, we propose a class-wise balancing data replay method for FCIL (FedCBDR), which employs a global coordination mechanism for class-level memory construction and reweights the learning objective to alleviate the aforementioned imbalances. Specifically, FedCBDR has two key components: 1) the global-perspective data replay module reconstructs global representations of prior task knowledge in a privacy-preserving manner, which then guides a class-aware and importance-sensitive sampling strategy to achieve balanced replay; 2) Subsequently, to handle class imbalance across tasks, the task-aware temperature scaling module adaptively adjusts the temperature of logits at both class and instance levels based on task dynamics, which reduces the model’s overconfidence in majority classes while enhancing its sensitivity to minority classes. Experimental results verified that FedCBDR achieves balanced class-wise sampling under heterogeneous data distributions and improves generalization under task imbalance between earlier and recent tasks, yielding a 2%-15% Top-1 accuracy improvement over six state-of-the-art methods.
Zhuang Qi, Ying-Peng Tang, Lei Meng 0001, Han Yu 0001, Xiaoxiao Li 0001, Xiangxu Meng
NeurIPS2
2025 QFEVAL: Quantum Federated Ensembled Variational Adaptive Learning for Dynamic Security Assessment in Cyber-Physical Systems
abstract
In the era of smart cyber-physical grid, dynamic insecurity risk has become a significant concern due to the increasing integration of renewable energy sources and the inherent uncertainties in smart grid. Dynamic security assessment (DSA) has been adopted to hedge against such risks by estimating the stability of large-scale smart grids. Existing DSA approaches often involve complex high dimensional models which incur high communication and computational costs, hindering their practical adoption. In this paper, we address these limitations with the Quantum Federated Ensembled Variational Adaptive Learning (QFEVAL) approach for smart grid DSA. QFEVAL is designed to combine quantum machine learning and federated learning to handle the differential-algebraic equations that describe smart grid stability, providing an efficient way to deal with high-dimensional data and uncertainties. QFEVAL enables the training of the hybrid quantum-classical neural networks on distributed DSA datasets located at different nodes in smart grids, without requiring large numbers of parameters to be transmitted. QFEVAL accurately predicts the stability of the smart grid under various conditions, enabling the implementation of preventive stability control measures. Through extensive experiments, we demonstrate that QFEVAL achieves comparable performance to 9 state-of-the-art DSA approaches with more than 2 orders of magnitude fewer model parameter transmissions. QFEVAL paves the way for reliable, secure, and continuous electricity supply, offering a robust solution to the challenges of DSA in smart grids.
Chao Ren 0006, Ying-Peng Tang, Yulan Gao, Xian Sun 0001, Kun Fu 0001, Mikael Skoglund, Zhao Yang Dong, Han Yu 0001, Anran Li 0001, Ming Xiao 0001
IEEE J. Sel. Areas Commun.2
2024 One-shot Active Learning Based on Lewis Weight Sampling for Multiple Deep Models
abstract
Active learning (AL) for multiple target models aims to reduce labeled data querying while effectively training multiple models concurrently. Existing AL algorithms often rely on iterative model training, which can be computationally expensive, particularly for deep models. In this paper, we propose a one-shot AL method to address this challenge, which performs all label queries without repeated model training. Specifically, we extract different representations of the same dataset using distinct network backbones, and actively learn the linear prediction layer on each representation via an $\ell_p$-regression formulation. The regression problems are solved approximately by sampling and reweighting the unlabeled instances based on their maximum Lewis weights across the representations. An upper bound on the number of samples needed is provided with a rigorous analysis for $p\in [1, +\infty)$. Experimental results on 11 benchmarks show that our one-shot approach achieves competitive performances with the state-of-the-art AL methods for multiple target models.
Sheng-Jun Huang, Yi Li 0002, Ying-Peng Tang
ICLR4
2023 MUS-CDB: Mixed Uncertainty Sampling With Class Distribution Balancing for Active Annotation in Aerial Object Detection
abstract
Recent aerial object detection models rely on a large amount of labeled training data, which requires unaffordable manual labeling costs in large aerial scenes with dense objects. Active learning effectively reduces the data labeling cost by selectively querying the informative and representative unlabelled samples. However, existing active learning methods are mainly with class-balanced settings and image-based querying for generic object detection tasks, which are less applicable to aerial object detection scenarios due to the long-tailed class distribution and dense small objects in aerial scenes. In this paper, we propose a novel active learning method for cost-effective aerial object detection. Specifically, both object-level and image-level informativeness are considered in the object selection to refrain from redundant and myopic querying. Besides, an easy-to-use class-balancing criterion is incorporated to favor the minority objects to alleviate the long-tailed class distribution problem in model training. We further devise a training loss to mine the latent knowledge in the unlabeled image regions. Extensive experiments are conducted on the DOTA-v1.0 and DOTA-v2.0 benchmarks to validate the effectiveness of the proposed method. For the ReDet, KLD, and SASM detectors on the DOTA-v2.0 dataset, the results show that our proposed MUS-CDB method can save nearly 75% of the labeling cost while achieving comparable performance to other active learning methods in terms of mAP. Code is publicly online.
Dong Liang 0008, Jing-Wei Zhang, Ying-Peng Tang, Sheng-Jun Huang
IEEE Trans. Geosci. Remote. Sens.3
2023 QBox: Partial Transfer Learning With Active Querying for Object Detection
abstract
Object detection requires plentiful data annotated with bounding boxes for model training. However, in many applications, it is difficult or even impossible to acquire a large set of labeled examples for the target task due to the privacy concern or lack of reliable annotators. On the other hand, due to the high-quality image search engines, such as Flickr and Google, it is relatively easy to obtain resource-rich unlabeled datasets, whose categories are a superset of those of target data. In this article, to improve the target model with cost-effective supervision from source data, we propose a partial transfer learning approach QBox to actively query labels for bounding boxes of source images. Specifically, we design two criteria, i.e., informativeness and transferability, to measure the potential utility of a bounding box for improving the target model. Based on these criteria, QBox actively queries the labels of the most useful boxes from the source domain and, thus, requires fewer training examples to save the labeling cost. Furthermore, the proposed query strategy allows annotators to simply labeling a specific region, instead of the whole image, and, thus, significantly reduces the labeling difficulty. Extensive experiments are performed on various partial transfer benchmarks and a real COVID-19 detection task. The results validate that QBox improves the detection accuracy with lower labeling cost compared to state-of-the-art query strategies for object detection.
Ying-Peng Tang, Xiu-Shen Wei, Borui Zhao, Sheng-Jun Huang
IEEE Trans. Neural Networks Learn. Syst.1
2022 Active Learning for Multiple Target Models
abstract
We describe and explore a novel setting of active learning (AL), where there are multiple target models to be learned simultaneously. In many real applications, the machine learning system is required to be deployed on diverse devices with varying computational resources (e.g., workstation, mobile phone, edge devices, etc.), which leads to the demand of training multiple target models on the same labeled dataset. However, it is generally believed that AL is model-dependent and untransferable, i.e., the data queried by one model may be less effective for training another model. This phenomenon naturally raises a question "Does there exist an AL method that is effective for multiple target models?" In this paper, we answer this question by theoretically analyzing the label complexity of active and passive learning under the setting with multiple target models, and conclude that AL does have potential to achieve better label complexity under this novel setting. Based on this insight, we further propose an agnostic AL sampling strategy to select the examples located in the joint disagreement regions of different target models. The experimental results on the OCR benchmarks show that the proposed method can significantly surpass the traditional active and passive learning methods under this challenging setting.
Ying-Peng Tang, Sheng-Jun Huang
NeurIPS1
2021 Dual Active Learning for Both Model and Data Selection
abstract
To learn an effective model with less training examples, existing active learning methods typically assume that there is a given target model, and try to fit it by selecting the most informative examples. However, it is less likely to determine the best target model in prior, and thus may get suboptimal performance even if the data is perfectly selected. To tackle with this practical challenge, this paper proposes a novel framework of dual active learning (DUAL) to simultaneously perform model search and data selection. Specifically, an effective method with truncated importance sampling is proposed for Combined Algorithm Selection and Hyperparameter optimization (CASH), which mitigates the model evaluation bias on the labeled data. Further, we propose an active query strategy to label the most valuable examples. The strategy on one hand favors discriminative data to help CASH search the best model, and on the other hand prefers informative examples to accelerate the convergence of winner models. Extensive experiments are conducted on 12 openML datasets. The results demonstrate the proposed method can effectively learn a superior model with less labeled examples.
Ying-Peng Tang, Sheng-Jun Huang
IJCAI1
2019 Self-Paced Active Learning: Query the Right Thing at the Right Time
abstract
Active learning queries labels from the oracle for the most valuable instances to reduce the labeling cost. In many active learning studies, informative and representative instances are preferred because they are expected to have higher potential value for improving the model. Recently, the results in self-paced learning show that training the model with easy examples first and then gradually with harder examples can improve the performance. While informative and representative instances could be easy or hard, querying valuable but hard examples at early stage may lead to waste of labeling cost. In this paper, we propose a self-paced active learning approach to simultaneously consider the potential value and easiness of an instance, and try to train the model with least cost by querying the right thing at the right time. Experimental results show that the proposed approach is superior to state-of-the-art batch mode active learning methods.
Ying-Peng Tang, Sheng-Jun Huang
AAAI1