Shao-Yuan Li

dblp:79/1523 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-0610-8568ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail Recognition
abstract
Long-tail recognition remains challenging for pre-trained foundation models like CLIP, which often suffer from performance degradation under imbalanced data. This stems not only from the overfitting/underfitting issues during fine-tuning but, more fundamentally, from the inherent bias inherited from the long-tail distribution of their massive pre-training datasets. To address this, we propose HGLTR (Hierarchy-Guided Long-Tail Recognition), a novel framework that calibrates pre-trained models by injecting objective class hierarchy knowledge. We argue that the semantic proximity defined by a hierarchy provides a robust, data-independent prior to counteract model bias. Our method is specifically designed for vision-language models' dual-modality architecture. At the feature level, we align image embeddings with a hierarchy-guided text similarity structure. At the classifier level, we employ a distillation loss to regularize predictions using soft labels derived from the hierarchy. This dual-level injection effectively transfers knowledge from head to tail classes. Experiments on ImageNet-LT, Places-LT, and iNaturalist 2018 demonstrate that HGLTR achieves state-of-the-art performance, particularly in tail-classes accuracy, highlighting the importance of leveraging structural priors to calibrate foundation models for real-world data.
Jinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan, Zijian Tao, Songcan Chen, Kangkan Wang
AAAI2
2026 GSH3D: Efficient 3D Gaussian human generation from 2D image collections
Kangkan Wang, Miao Zhao, Shao-Yuan Li
Knowl. Based Syst.3
2026 Filter, Obstruct, and Dilute: Defending Against Backdoor Attacks on Semi-Supervised Learning
abstract
Recent studies have demonstrated that semi-supervised learning (SSL) is highly vulnerable to backdoor attacks, where adversaries can manipulate up to 90% of model predictions through just a tiny fraction of poisoned training data. Despite the widespread adoption of SSL in safety-critical applications, effective defenses against such attacks remain limited. In this paper, we present a comprehensive defense framework designed to protect SSL against sophisticated backdoor attacks. Our work begins with a systematic analysis of backdoor mechanisms in SSL from two critical perspectives: (1) how attackers establish persistent correlations between triggers and target classes; (2) how triggers are introduced and resist removal at the data level. Our investigation reveals that, unlike supervised learning, SSL backdoor attacks (1) uniquely exploit pseudo-labeling mechanisms to establish stronger trigger-target correlations, and (2) demonstrate remarkable resilience at the data level, with triggers potentially appearing in any frequency band (low, medium, or high). Based on these insights, we introduce Backdoor Invalidator (BI), a defense framework that integrates three novel techniques: complementary learning, trigger mix-up, and dual domain filtering, which collectively obstruct, dilute, and filter the influence of backdoor attacks in both feature learning and data processing. Through extensive evaluation against state-of-the-art attacks, BI significantly reduces the average attack success rate while maintaining comparable accuracy on clean data. We also provide theoretical guarantees for BI’s generalization capability and demonstrate its practical deployability as a plug-in component. The code of this work is available at https://github.com/wxr99/Backdoor Invalidator4SSL.
Xinrui Wang 0003, Wenhai Wan, Xiang Li 0119, Chuanxing Geng, Shao-Yuan Li, Songcan Chen
IEEE Trans. Inf. Forensics Secur.5
2025 MLC-NC: Long-Tailed Multi-Label Image Classification Through the Lens of Neural Collapse
abstract
Long-tailed (LT) data distribution is common in multi-label image classification (MLC) and can significantly impact the performance of classification models. One reason is the challenge of learning unbiased instance representations (i.e. features) for imbalanced datasets. Additionally, the co-occurrence of head/tail classes within the same instance, along with complex label dependencies, introduces further challenges. In this work, we delve into this problem through the lens of neural collapse (NC). NC refers to a phenomenon where the last-layer features and classifier of a deep neural network model exhibit a simplex Equiangular Tight Frame (ETF) structure during its terminal training phase. This structure creates an optimal linearly separable state. However, this phenomenon typically occurs in balanced datasets but rarely applies to the typical imbalanced problem. To induce NC properties under Long-tailed multi-label classification (LT-MLC) conditions, we propose an approach named MLC-NC, which aims to learn high-quality data representations and improve the model’s generalization ability. Specifically, MLC-NC accounts for the fact that different labels correspond to different feature parts located in images. MLC-NC extracts class-wise features from each instance through a cross-attention mechanism. To guide the features toward the ETF structure, we introduce visual-semantic feature alignment with a fixed ETF structured label embedding, which helps to learn evenly distributed class centers. To reduce within-class feature variation, we introduce collapse calibration within a lower-dimensional feature space. To mitigate classification bias, we concatenate features and feed them into a binarized fixed ETF classifier. As an orthogonal approach to existing methods, MLC-NC can be seamlessly integrated into various frameworks. Extensive experiments on widely-used benchmarks demonstrate the effectiveness of our method.
Zijian Tao, Shao-Yuan Li, Wenhai Wan, Jinpeng Zheng, Jia-Yao Chen, Sheng-Jun Huang, Songcan Chen
AAAI2
2025 Cut out and Replay: A Simple yet Versatile Strategy for Multi-Label Online Continual Learning
abstract
Multi-Label Online Continual Learning (MOCL) requires models to learn continuously from endless multi-label data streams, facing complex challenges including persistent catastrophic forgetting, potential missing labels, and uncontrollable imbalanced class distributions. While existing MOCL methods attempt to address these challenges through various techniques, \textit{they all overlook label-specific region identifying and feature learning} - a fundamental solution rooted in multi-label learning but challenging to achieve in the online setting with incremental and partial supervision. To this end, we first leverage the inherent structural information of input data to evaluate and verify the innate localization capability of different pre-trained models. Then, we propose CUTER (CUT-out-and-Experience-Replay), a simple yet versatile strategy that provides fine-grained supervision signals by further identifying, strengthening and cutting out label-specific regions for efficient experience replay. It not only enables models to simultaneously address catastrophic forgetting, missing labels, and class imbalance challenges, but also serves as an orthogonal solution that seamlessly integrates with existing approaches. Extensive experiments on multiple multi-label image benchmarks demonstrate the superiority of our proposed method. The code is available at \href{https://github.com/wxr99/Cut-Replay}{https://github.com/wxr99/Cut-Replay}
Shao-Yuan Li, Songcan Chen
ICML2
2025 DM-POSA: Enhancing Open-World Test-Time Adaptation with Dual-Mode Matching and Prompt-Based Open Set Adaptation
abstract
The need to generalize the pre-trained deep learning models to unknown test-time data distributions has spurred research into test-time adaptation (TTA). Existing studies have mainly focused on closed-set TTA with only covariate shifts, while largely overlooking open-set TTA that involves semantic shifts, i.e., unknown open-set classes. However, addressing adaptation to unknown classes is crucial for open-world safety-critical applications such as autonomous driving. In this paper, we emphasize that accurate identification of the open-set samples is rather challenging in TTA. The entanglement of semantic shift and covariate shift mutually confuse the network’s discriminative capability. This co-interference further exacerbates considering the single-pass data nature and low latency requirements. With this under standing, we propose Dual-mode Matching and Prompt-based Open Set Adaptation (DM-POSA) for open-set TTA to enhance discriminative feature learning and unknown classes distinguishment with minimal time cost. DM-POSA identifies open-set samples via dual-mode matching strategies, including model-parameter-based and feature space-based matching. It also optimizes the model with a random pairing discrepancy loss, enhancing the distributional difference between open-set and closed-set samples, thus improving the model’s ability to recognize unknown categories. Extensive experiments show the superiority of DM-POSA over state-of-the-art baselines on both closed-set class adaptation and open-set class detection.
Shao-Yuan Li, Chuanxing Geng, Sheng-Jun Huang, Songcan Chen
IJCAI2
2025 Robust domain adaptation with noisy and shifted label distribution
Shao-Yuan Li, Shi-Ji Zhao, Zheng-Tao Cao, Sheng-Jun Huang, Songcan Chen
Frontiers Comput. Sci.1
2025 Robust contrastive knowledge distillation for long-tailed noisy class labels
Shao-Yuan Li, Jinpeng Zheng, Mingguang Zhang, Shaofang Li, Kangkan Wang
Knowl. Based Syst.1
2025 Continual learning in the presence of repetition
abstract
Continual learning (CL) provides a framework for training models in ever-evolving environments. Although re-occurrence of previously seen objects or tasks is common in real-world problems, the concept of repetition in the data stream is not often considered in standard benchmarks for CL. Unlike with the rehearsal mechanism in buffer-based strategies, where sample repetition is controlled by the strategy, repetition in the data stream naturally stems from the environment. This report provides a summary of the CLVision challenge at CVPR 2023, which focused on the topic of repetition in class-incremental learning. The report initially outlines the challenge objective and then describes three solutions proposed by finalist teams that aim to effectively exploit the repetition in the stream to learn continually. The experimental results from the challenge highlight the effectiveness of ensemble-based solutions that employ multiple versions of similar modules, each trained on different but overlapping subsets of classes. This report underscores the transformative potential of taking a different perspective in CL by employing repetition in the data stream to foster innovative strategy design. • An overview of the continual learning challenge of the CLVision workshop at CVPR 2023. • Novel benchmarks focussing on the topic of repetition in continual learning. • Description and discussion of the strategies submitted by the winning teams. • The results highlight the remarkable effectiveness of ensemble-based solutions.
Hamed Hemati, Lorenzo Pellegrini, Xiaotian Duan, Fangfang Xia, Marc Masana, Benedikt Tscheschner, Eduardo E. Veas, Shao-Yuan Li, Sheng-Jun Huang, Vincenzo Lomonaco, Gido M. van de Ven
Neural Networks11
2025 Handling Noisy Annotation for Remote Sensing Semantic Segmentation via Boundary-Aware Knowledge Distillation
abstract
In recent years, image segmentation has made significant progress, but acquiring annotated data is still a considerable challenge, especially in remote sensing imagery (RSI). The complex structure and inter-category confusion of RSI increase the time-consuming and cost of pixel-level annotation, and noisy annotations inevitably appear. This paper proposes a boundary-aware knowledge distillation method (BAKD) to handle noisy annotations by evaluating their uncertainty. BAKD consists of two core strategies: Predictive Confidence Evaluation (PCE) and Boundary-annotated Reliability Evaluation (BRE). The predictive confidence jointly decided by the teacher and student networks reflects the annotation’s uncertainty. The boundary-annotated reliability directly measures the annotation’s uncertainty based on the distance from the annotation to the semantic boundary. Leveraging these two types of uncertainty information, BAKD assigns each sample a comprehensive boundary-aware weight to identify samples with potential noisy annotations. This alleviates the impact of noisy annotation on the model’s training and improves its generalization performance. Experimental results show that BAKD achieves competitive semantic segmentation performance on the Potsdam and Vaihingen benchmarks compared with the state-of-the-art KD methods. In addition, BAKD can be easily integrated into semantic segmentation methods based on KD, extending their applicability in handling noisy annotations. Codes are available at https://github.com/sunyueue/BAKD.git.
Dong Liang 0008, Shao-Yuan Li, Songcan Chen, Sheng-Jun Huang
IEEE Trans. Geosci. Remote. Sens.3
2024 Unlocking the Power of Open Set: A New Perspective for Open-Set Noisy Label Learning
abstract
Learning from noisy data has attracted much attention, where most methods focus on closed-set label noise. However, a more common scenario in the real world is the presence of both open-set and closed-set noise. Existing methods typically identify and handle these two types of label noise separately by designing a specific strategy for each type. However, in many real-world scenarios, it would be challenging to identify open-set examples, especially when the dataset has been severely corrupted. Unlike the previous works, we explore how models behave when faced with open-set examples, and find that a part of open-set examples gradually get integrated into certain known classes, which is beneficial for the separation among known classes. Motivated by the phenomenon, we propose a novel two-step contrastive learning method CECL (Class Expansion Contrastive Learning) which aims to deal with both types of label noise by exploiting the useful information of open-set examples. Specifically, we incorporate some open-set examples into closed-set classes to enhance performance while treating others as delimiters to improve representative ability. Extensive experiments on synthetic and real-world datasets with diverse label noise demonstrate the effectiveness of CECL.
Wenhai Wan, Xinrui Wang 0003, Ming-Kun Xie, Shao-Yuan Li, Sheng-Jun Huang, Songcan Chen
AAAI4
2024 NanoAdapt: Mitigating Negative Transfer in Test Time Adaptation with Extremely Small Batch Sizes
Shao-Yuan Li, Sheng-Jun Huang
IJCAI2
2024 Forgetting, Ignorance or Myopia: Revisiting Key Challenges in Online Continual Learning
abstract
Online continual learning (OCL) requires the models to learn from constant, endless streams of data. While significant efforts have been made in this field, most were focused on mitigating the \textit{catastrophic forgetting} issue to achieve better classification ability, at the cost of a much heavier training workload. They overlooked that in real-world scenarios, e.g., in high-speed data stream environments, data do not pause to accommodate slow models. In this paper, we emphasize that \textit{model throughput}-- defined as the maximum number of training samples that a model can process within a unit of time -- is equally important. It directly limits how much data a model can utilize and presents a challenging dilemma for current methods. With this understanding, we revisit key challenges in OCL from both empirical and theoretical perspectives, highlighting two critical issues beyond the well-documented catastrophic forgetting: (\romannumeral1) Model's ignorance: the single-pass nature of OCL challenges models to learn effective features within constrained training time and storage capacity, leading to a trade-off between effective learning and model throughput; (\romannumeral2) Model's myopia: the local learning nature of OCL on the current task leads the model to adopt overly simplified, task-specific features and \textit{excessively sparse classifier}, resulting in the gap between the optimal solution for the current task and the global objective. To tackle these issues, we propose the Non-sparse Classifier Evolution framework (NsCE) to facilitate effective global discriminative feature learning with minimal time cost. NsCE integrates non-sparse maximum separation regularization and targeted experience replay techniques with the help of pre-trained models, enabling rapid acquisition of new globally discriminative features. Extensive experiments demonstrate the substantial improvements of our framework in performance, throughput and real-world practicality.
Xinrui Wang 0003, Chuanxing Geng, Wenhai Wan, Shao-Yuan Li, Songcan Chen
NeurIPS4
2024 UNM: A Universal Approach for Noisy Multi-Label Learning
abstract
Multi-label image classification relies on a large-scale, well-maintained dataset, which may easily be mislabeled due to various subjective reasons. Existing methods for coping with noise usually focus on improving the model robustness in the case of single-label noise. However, compared with noisy single-label learning, noisy multi-label learning is more practical and challenging. To reduce the negative impact of noisy multi-annotations, we propose a universal approach for noisy multi-label learning (UNM). In UNM, we propose the label-wise embedding network which investigates the semantic alignment between label embeddings and their corresponding output features to learn robust feature representations. Meanwhile, mining the co-occurrence of multi-labels is also added to regularize the noisy network predictions. We cyclically change the fitting status of our label-wise embedding network to distinguish the noisy samples and generate pseudo labels for them. As a result, UNM provides an effective way to exploit the label-wise features and semantic label embeddings in noisy scenarios. To verify the generalizability of our method, we also test our method on Partial Multi-label Learning (PML) and Multi-label Learning with Missing Labels (MLML). Extensive experiments on benchmark datasets including Microsoft COCO, Pascal VOC, and Visual Genome explicitly validate the proposed method.
Jia-Yao Chen, Shao-Yuan Li, Sheng-Jun Huang, Songcan Chen, Lei Wang 0226, Ming-Kun Xie
IEEE Trans. Knowl. Data Eng.2
2023 Learning from crowds with sparse and imbalanced annotations
Ye Shi 0004, Shao-Yuan Li, Sheng-Jun Huang
Mach. Learn.2
2020 Incremental Multi-Label Learning with Active Queries
Sheng-Jun Huang, Guo-Xiang Li, Wen-Yu Huang, Shao-Yuan Li
J. Comput. Sci. Technol.4
2019 Multi-Label Learning from Crowds
abstract
We consider multi-label crowdsourcing learning in two scenarios. In the first scenario, we aim at inferring instances' groundtruth given the crowds' annotations. We propose two approaches NAM/RAM (Neighborhood/Relevance Aware Multi-label crowdsourcing) modeling the crowds' expertise and label correlations from different perspectives. Extended from single-label crowdsourcing methods, NAM models the crowds' expertise on individual labels, but based on the idea that for rational workers, their annotations for instances similar in the feature space should also be similar, NAM utilizes information from the feature space and incorporates the local influence of neighborhoods' annotations. Noting that the crowds tend to act in an effort-saving manner while labeling multiple labels, i.e., rather than carefully annotating every proper label, they would prefer scanning and tagging a few most relevant labels, RAM models the crowds' expertise as their ability to distinguish the relevance between label pairs. In the second scenario, we care about cost-efficient crowdsourcing where the labeling and learning process are conducted in tandem. We extend NAM/RAM to the active paradigm and propose instance, label, and worker selection criteria such that the labeling cost is significantly saved compared to passive learning without labeling control. The proposals' effectiveness are validated on simulated and real data.
Shao-Yuan Li, Yuan Jiang 0001, Nitesh V. Chawla, Zhi-Hua Zhou
IEEE Trans. Knowl. Data Eng.1
2018 Multi-label Crowdsourcing Learning with Incomplete Annotations
Shao-Yuan Li, Yuan Jiang 0001
PRICAI (1)1
2017 BayDNN: Friend Recommendation with Bayesian Personalized Ranking Deep Neural Network
abstract
Friendship is the cornerstone to build a social network. In online social networks, statistics show that the leading reason for user to create a new friendship is due to recommendation. Thus the accuracy of recommendation matters. In this paper, we propose a Bayesian Personalized Ranking Deep Neural Network (BayDNN) model for friend recommendation in social networks. With BayDNN, we achieve significant improvement on two public datasets: Epinions and Slashdot. For example, on Epinions dataset, BayDNN significantly outperforms the state-of-the-art algorithms, with a 5% improvement on NDCG over the best baseline.
Daizong Ding, Mi Zhang 0001, Shao-Yuan Li, Jie Tang 0001, Xiaotie Chen, Zhi-Hua Zhou
CIKM3
2017 Obtaining High-Quality Label by Distinguishing between Easy and Hard Items in Crowdsourcing
abstract
Crowdsourcing systems make it possible to hire voluntary workers to label large-scale data by offering them small monetary payments. Usually, the taskmaster requires to collect high-quality labels, while the quality of labels obtained from the crowd may not satisfy this requirement. In this paper, we study the problem of obtaining high-quality labels from the crowd and present an approach of learning the difficulty of items in crowdsourcing, in which we construct a small training set of items with estimated difficulty and then learn a model to predict the difficulty of future items. With the predicted difficulty, we can distinguish between easy and hard items to obtain high-quality labels. For easy items, the quality of their labels inferred from the crowd could be high enough to satisfy the requirement; while for hard items, the crowd could not provide high-quality labels, it is better to choose a more knowledgable crowd or employ specialized workers to label them. The experimental results demonstrate that the proposed approach by learning to distinguish between easy and hard items can significantly improve the label quality.
Wei Wang 0028, Xiang-Yu Guo, Shao-Yuan Li, Yuan Jiang 0001, Zhi-Hua Zhou
IJCAI3
2014 Partial Multi-View Clustering
abstract
Real data are often with multiple modalities or comingfrom multiple channels, while multi-view clusteringprovides a natural formulation for generating clustersfrom such data. Previous studies assumed that each exampleappears in all views, or at least there is one viewcontaining all examples. In real tasks, however, it is oftenthe case that every view suffers from the missing ofsome data and therefore results in many partial examples,i.e., examples with some views missing. In this paper,we present possibly the first study on partial multiviewclustering. Our proposed approach, PVC, worksby establishing a latent subspace where the instancescorresponding to the same example in different viewsare close to each other, and similar instances (belongingto different examples) in the same view should bewell grouped. Experiments on two-view data demonstratethe advantages of our proposed approach.
Shao-Yuan Li, Yuan Jiang 0001, Zhi-Hua Zhou
AAAI1
2004 Stability analysis and design of T-S fuzzy control system with simplified linear rule consequent
abstract
Stability and design issues of simple T-S fuzzy control system with simplified linear rule consequent (TSS) are investigated. A systematic approach to find a common matrix P for TSS fuzzy system is presented, where system matrix Ai is decomposed into proportional part Ai and the remainder delta Ai. Hence an iterative approach to find a common matrix P for pairwise commutative Ai's can be used. The stability of the global system is guaranteed if delta Ai satisfies certain conditions. Qualitative instructions for TSS control system design are summarized. A physical example is given to illustrate the issues discussed throughout the paper.
Shao-Yuan Li
IEEE Trans. Syst. Man Cybern. Part B2
2003 Soft sensor modeling for slab temperature estimation
abstract
This paper investigated a soft sensor modeling approach based on radial basis function networks (RBFN). A fuzzy c-means (FCM) clustering algorithm is used to classify training vectors into several clusters, each cluster is trained by a radial basis function network, and membership values are used for combining several network outputs to obtain the final result. In the online stage, membership values are computed using an adaptive fuzzy clustering algorithm for the new vector. The proposed approach has been applied to the slab temperature estimation in a practical walking beam reheating furnace. Simulation results show that the approach is effective.
Xi-Huai Wang, Shao-Yuan Li
FUZZ-IEEE2