Yang Lu 0009

dblp:16/6317-9 · DBLP profile ↗
← Back
99ranked-venue papers
10as first author
83since 2021 · last 2026
0000-0002-3497-9611ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 4 first-author · 53 since 2021Artificial intelligence and machine learning · 45 · 7 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Security and privacy · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
abstract
Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing methods achieve promising results by leveraging CLIP’s vision-language alignment, they rarely investigate the performance boundaries of CLIP for dense prediction from an interpretability mechanisms perspective. In this work, we systematically investigate CLIP's internal mechanisms and identify a critical phenomenon: analogous to human distraction, CLIP diverts significant attention resources from target regions to irrelevant tokens. Our analysis reveals that these tokens arise from dimension-specific over-activation; filtering them enhances CLIP's dense prediction performance. Consequently, we propose Refocusing CLIP (RF-CLIP), a training-free approach that emulates human distraction-refocusing behavior to redirect attention from distraction tokens back to target regions, thereby refining CLIP's multimodal alignment granularity. Our method achieves SOTA performance on eight benchmarks while maintaining high inference efficiency.
Jiahao Li 0003, Yang Lu 0009, Yachao Zhang 0001, Fangyong Wang, Yuan Xie 0006, Yanyun Qu
AAAI2
2026 Joint Implicit and Explicit Language Learning for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition (PAR) has received increasing attention due to its wide application in video surveillance and pedestrian analysis. Some text-enhanced methods tackle this task by converting attributes into language descriptions to facilitate interactive learning between attributes and visual images. However, these generic languages fail to uniquely describe different pedestrian images, missing individual characteristics. In this paper, we propose a Joint Implicit and Explicit Language Guidance Enhancement Learning (JGEL) method, which converts each pedestrian image into a language description with dual language learning to effectively learn enhanced attribute information. Specifically, we first propose an Implicit Language Guidance Learning (ILGL) stream. It projects visual image features into the text embedding space to generate pseudo-word tokens, implicitly modeling image attributes and providing personalized descriptions. Moreover, we propose an Explicit Attribute Enhancement Learning (EAEL) stream to guide the generated pseudo-word tokens obtained by ILGL explicitly aligned with pedestrian attributes, which can effectively align the pseudo-word tokens with the attribute concepts in the text embedding space. Extensive experiments show that JGEL has significant advantages in improving the performance of PAR and the challenging zero-shot PAR task.
Yang Lu 0009, Yan Yan 0001, Hanzi Wang
AAAI3
2026 Break the Tie: Learning Cluster-Customized Category Relationships for Categorical Data Clustering
abstract
Categorical attributes with qualitative values are ubiquitous in cluster analysis of real datasets. Unlike the Euclidean distance of numerical attributes, the categorical attributes lack well-defined relationships of their possible values (also called categories interchangeably), which hampers the exploration of compact categorical data clusters. Although most attempts are made for developing appropriate distance metrics, they typically assume a fixed topological relationship between categories when learning distance metrics, which limits their adaptability to varying cluster structures and often leads to suboptimal clustering performance. This paper, therefore, breaks the intrinsic relationship tie of attribute categories and learns customized distance metrics suitable for flexibly and accurately revealing various cluster distributions. As a result, the fitting ability of the clustering algorithm is significantly enhanced, benefiting from the learnable category relationships. Moreover, the learned category relationships are proved to be Euclidean distance metric-compatible, enabling a seamless extension to mixed datasets that include both numerical and categorical attributes. Comparative experiments on 12 real benchmark datasets with significance tests show the superior clustering accuracy of the proposed method with an average ranking of 1.25, which is significantly higher than the 5.21 ranking of the best-performing methods. Code and extended version with detailed proofs are provided online.
Mingjie Zhao 0003, Zhanpei Huang, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Weifeng Su, Yiu-Ming Cheung
AAAI3
2026 HyReaL: Clustering Attributed Graph via Hyper-complex Space Representation Learning
Yang Lu 0009, Mengke Li 0001, Cuie Yang, Yiqun Zhang 0006, Yiu-Ming Cheung
DASFAA (2)2
2026 Bridging land and sea: A latent diffusion framework for high-resolution ocean floor mapping
Duo Shuai, Qiang Deng, Yang Lu 0009, Qixian Zhong, Zhonglei Wang
Expert Syst. Appl.4
2026 Global-Local Disturbance Decoupling for Federated Facial Expression Recognition
abstract
Most existing facial expression recognition (FER) methods are designed for centralized model training on largescale data. Unfortunately, accessing massive facial expression data can be difficult due to privacy concerns in practice. In this paper, we study an important but little-explored task, federated FER, which allows us to train an FER model with decentralized expression data. To this end, we propose a novel global-local disturbance decoupling (GLDD) method for federated FER. Specifically, for local disturbance decoupling on each client, we develop a dual-branch feature decoupling network consisting of a backbone network, an expression branch, and a disturbance branch, to perform local FER. In the disturbance branch, we design an entropy-guided feature encoding module to extract priorbased disturbance features. This greatly facilitates the extraction of client-specific disturbance features. For global disturbance decoupling on the server, we introduce orthogonal decoupling, which is a global-level feature disentanglement technique that enforces mutual orthogonality between the global expression feature class centers and the global disturbance feature centers, thereby eliminating cross-client disturbances across clients and yields decoupled global expression feature class centers. These centers are then used to retrain the global model, substantially enhancing disturbance invariance and classification performance. By jointly performing disturbance decoupling at local and global levels, our method effectively addresses the unique challenges of heterogeneous expression data and heterogeneous disturbances in federated FER. Experimental results on two real-world facial expression databases show that, on the federated FER task, our method significantly outperforms several state-of-the-art federated learning methods and FER methods. The code will be released soon.
Hu Ding 0005, Yan Yan 0001, Yang Lu 0009, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Affect. Comput.3
2025 MaskViM: Domain Generalized Semantic Segmentation with State Space Models
abstract
Domain Generalized Semantic Segmentation (DGSS) aims to utilize segmentation model training on known source domains to make predictions on unknown target domains. Currently, there are two network architectures: one based on Convolutional Neural Networks (CNNs) and the other based on Visual Transformers (ViTs). However, both CNN-based and ViT-based DGSS methods face challenges: the former lacks a global receptive field, while the latter requires more computational demands. Drawing inspiration from State Space Models (SSMs), which not only possess a global receptive field but also maintain linear complexity, we propose SSM-based method for achieving DGSS. In this work, we first elucidate why does mask make sense in SSM-based DGSS and propose our mask learning mechanism. Leveraging this mechanism, we present our Mask Vision Mamba network (MaskViM), a model for SSM-based DGSS, and design our mask loss to optimize MaskViM. Our method achieves superior performance on four diverse DGSS setting, which demonstrates the effectiveness of our method.
Jiahao Li 0003, Yang Lu 0009, Yuan Xie 0006, Yanyun Qu
AAAI2
2025 Asynchronous Federated Clustering with Unknown Number of Clusters
abstract
Federated Clustering (FC) is crucial to mining knowledge from unlabeled non-Independent Identically Distributed (non-IID) data provided by multiple clients while preserving their privacy. Most existing attempts learn cluster distributions at local clients, then securely pass the desensitized information to the server for aggregation. However, some tricky but common FC problems are still relatively unexplored, including the heterogeneity in terms of clients' communication capacity and the unknown number of proper clusters. To further bridge the gap between FC and real application scenarios, this paper first shows that the clients' communication asynchrony and unknown proper cluster numbers are complex coupling problems, and then proposes an Asynchronous Federated Cluster Learning (AFCL) method accordingly. It spreads the excessive number of seed points to clients as a learning medium and coordinates them across clients to form a consensus. To alleviate the distribution imbalance cumulated due to the unforeseen asynchronous uploading from the heterogeneous clients, we also design a balancing mechanism for seeds updating. As a result, the seeds gradually adapt to each other to reveal a proper number of clusters. Extensive experiments demonstrate the efficacy of AFCL.
Yiqun Zhang 0006, Yang Lu 0009, Mengke Li 0001, Yiu-Ming Cheung
AAAI3
2025 Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch
abstract
Federated Semi-Supervised Learning (FSSL) aims to leverage unlabeled data across clients with limited labeled data to train a global model with strong generalization ability. Most FSSL methods rely on consistency regularization with pseudo-labels, converting predictions from local or global models into hard pseudo-labels as supervisory signals. However, we discover that the quality of pseudo-label is largely deteriorated by data heterogeneity, an intrinsic facet of federated learning. In this paper, we study the problem of FSSL in-depth and show that (1) heterogeneity exacerbates pseudo-label mismatches, further degrading model performance and convergence, and (2) local and global models’ predictive tendencies diverge as heterogeneity increases. Motivated by these findings, we propose a simple and effective method called Semi-supervised Aggregation for Globally-Enhanced Ensemble (SAGE), that can flexibly correct pseudo-labels based on confidence discrepancies. This strategy effectively mitigates performance degradation caused by incorrect pseudo-labels and enhances consensus between local and global models. Experimental results demonstrate that SAGE outperforms existing FSSL methods in both performance and convergence. Our code is available at https://github.com/Jay-Codeman/SAGE.
Xinyi Shang, Yiqun Zhang 0006, Yang Lu 0009, Chen Gong 0002, Jing-Hao Xue, Hanzi Wang
CVPR4
2025 Weighted Density for The Win: Accurate Subspace Density Clustering
abstract
k-clustering typically struggles with the detection of irregular-distributed clusters due to the natural bias, while density clustering usually cannot well-adapt to different datasets and clustering tasks as it is not an oriented optimization process. This paper, therefore, proposes to perform density clustering in dynamically learned subspaces. To exploit the irregular-distributed clusters obtained by density clustering for the subspace determination, we design a new strategy to appropriately evaluate the importance of attributes. It turns out that the proposed Weighted Density-based Subspace Clustering (WDSC) algorithm inherits the unbiased merits of density clustering, and also upgrades the unlearning density clustering to be learnable under the subspace learning paradigm of k-clustering. A comprehensive evaluation including significance tests, ablation studies, qualitative comparisons, etc., shows the superiority of WDSC.
Maixuan Peng, Yuyang Wu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Yiu-Ming Cheung
ICASSP3
2025 PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt Relocation
abstract
Visual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the importance of each block varying depending on the task. In this paper, we introduce adaptive distribution optimization (ADO) by tackling two key questions: (1) How to appropriately and formally define ADO, and (2) How to design an adaptive distribution strategy guided by this definition? Through empirical analysis, we first confirm that properly adjusting the distribution significantly improves VPT performance, and further uncover a key insight that a nested relationship exists between ADO and VPT. Based on these findings, we propose a new VPT framework, termed PRO-VPT (iterative Prompt RelOcation-based VPT), which adaptively adjusts the distribution built upon a nested optimization formulation. Specifically, we develop a prompt relocation strategy derived from this formulation, comprising two steps: pruning idle prompts from prompt-saturated blocks, followed by allocating these prompts to the most prompt-needed blocks. By iteratively performing prompt relocation and VPT, our proposal can adaptively learn the optimal prompt distribution in a nested optimization-based manner, thereby unlocking the full potential of VPT. Extensive experiments demonstrate that our proposal significantly outperforms advanced VPT methods, e.g., PRO-VPT surpasses VPT by 1.6 pp and 2.0 pp average accuracy, leading prompt-based methods to state-of-the-art performance on VTAB-1k and FGVC benchmarks. The code is available at https://github.com/ckshang/PRO-VPT.
Chikai Shang, Mengke Li 0001, Yiqun Zhang 0006, Zhen Chen 0018, Jinlin Wu, Fangqing Gu, Yang Lu 0009, Yiu-Ming Cheung
ICCV7
2025 You are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-Tailed Data
abstract
Data heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main problems and proposed neural-collapse-inspired synthetic simplex ETF to help representations be closer to neural collapse optima. However, we find that the neural-collapse-inspired methods are not strong enough to reach neural collapse and still have huge gaps to centralized training. In this paper, we rethink this issue from a self-bootstrap perspective and propose FedYoYo (You Are Your Own Best Teacher), introducing Augmented Self-bootstrap Distillation (ASD) to improve representation learning by distilling knowledge between weakly and strongly augmented local samples, without needing extra datasets or models. We further introduce Distribution-aware Logit Adjustment (DLA) to balance the self-bootstrap process and correct biased feature representations. FedYoYo nearly eliminates the performance gap, achieving centralized-level performance even under mixed heterogeneity. It enhances local representation learning, reducing model drift and improving convergence, with feature prototypes closer to neural collapse optimality. Extensive experiments show FedYoYo achieves state-of-the-art results, even surpassing centralized logit adjustment methods by 5.4\% under global long-tailed settings.
Shanshan Yan, Zexi Li 0001, Chao Wu 0001, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
ICCV5
2025 Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation (OVSS) conducts pixel-level classification via text-driven alignment, where the domain discrepancy between base category training and open-vocabulary inference poses challenges in discriminative modeling of latent unseen category. To address this challenge, existing vision-language model (VLM)-based approaches demonstrate commendable performance through pre-trained multi-modal representations. However, the fundamental mechanisms of latent semantic comprehension remain underexplored, making the bottleneck for OVSS. In this work, we initiate a probing experiment to explore distribution patterns and dynamics of latent semantics in VLMs under inductive learning paradigms. Building on these insights, we propose X-Agent, an innovative OVSS framework employing latent semantic-aware ''agent'' to orchestrate cross-modal attention mechanisms, simultaneously optimizing latent semantic dynamic and amplifying its perceptibility. Extensive benchmark evaluations demonstrate that X-Agent achieves state-of-the-art performance while effectively enhancing the latent semantic saliency.
Jiahao Li 0003, Yang Lu 0009, Yachao Zhang 0001, Fangyong Wang, Yuan Xie 0006, Yanyun Qu
ACM Multimedia2
2025 FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled Data
abstract
Semi-supervised learning (SSL) has achieved significant progress by leveraging both labeled data and unlabeled data. Existing SSL methods overlook a common real-world scenario when labeled data is extremely scarce, potentially as limited as a single labeled sample in the dataset. General SSL approaches struggle to train effectively from scratch under such constraints, while methods utilizing pre-trained models often fail to find an optimal balance between leveraging limited labeled data and abundant unlabeled data. To address this challenge, we propose Firstly Adapt, Then catEgorize (FATE), a novel SSL framework tailored for scenarios with extremely limited labeled data. At its core, the two-stage prompt tuning paradigm FATE exploits unlabeled data to compensate for scarce supervision signals, then transfers to downstream tasks. Concretely, FATE first adapts a pre-trained model to the feature distribution of downstream data using volumes of unlabeled samples in an unsupervised manner. It then applies an SSL method specifically designed for pre-trained models to complete the final classification task. FATE is designed to be compatible with both vision and vision-language pre-trained models. Extensive experiments demonstrate that FATE effectively mitigates challenges arising from the scarcity of labeled samples in SSL, achieving an average performance improvement of 33.74% across seven benchmarks compared to state-of-the-art SSL methods. Code is available at https://github.com/ganchi-huanggua/FATE.git.
Hezhao Liu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Shreyank N. Gowda, Chen Gong 0002, Hanzi Wang
ACM Multimedia2
2025 Progressive Data Dropout: An Embarrassingly Simple Approach to Train Faster
abstract
The success of the machine learning field has reliably depended on training on large datasets. While effective, this trend comes at an extraordinary cost. This is due to two deeply intertwined factors: the size of models and the size of datasets. While promising research efforts focus on reducing the size of models, the other half of the equation remains fairly mysterious. Indeed, it is surprising that the standard approach to training remains to iterate over and over, uniformly sampling the training dataset. In this paper we explore a series of alternative training paradigms that leverage insights from hard-data-mining and dropout, simple enough to implement and use that can become the new training standard. The proposed Progressive Data Dropout reduces the number of effective epochs to as little as 12.4\% of the baseline. This savings actually do not come at any cost for accuracy. Surprisingly, the proposed method improves accuracy by up to 4.82\%. Our approach requires no changes to model architecture or optimizer, and can be applied across standard training pipelines, thus posing an excellent opportunity for wide adoption. Code can be found here: \url{https://github.com/bazyagami/LearningWithRevision}.
Shriram M. S, Xinyue Hao 0001, Shihao Hou, Yang Lu 0009, Laura Sevilla-Lara, Anurag Arnab, Shreyank N. Gowda
NeurIPS4
2025 Unlocker: Disentangle the Deadlock of Learning between Label-noisy and Long-tailed Data
abstract
In real world, the observed label distribution of a dataset often mismatches its true distribution due to noisy labels. In this situation, noisy labels learning (NLL) methods directly integrated with long-tail learning (LTL) methods tend to fail due to a dilemma: NLL methods normally rely on unbiased model predictions to recover true distribution by selecting and correcting noisy labels; while LTL methods like logit adjustment depends on true distributions to adjust biased predictions, leading to a deadlock of mutual dependency defined in this paper. To address this, we propose \texttt{Unlocker}, a bilevel optimization framework that integrates NLL methods and LTL methods to iteratively disentangle this deadlock. The inner optimization leverages NLL to train the model, incorporating LTL methods to fairly select and correct noisy labels. The outer optimization adaptively determines an adjustment strength, mitigating model bias from over- or under-adjustment. We also theoretically prove that this bilevel optimization problem is convergent by transferring the outer optimization target to an equivalent problem with a closed-form solution. Extensive experiments on synthetic and real-world datasets demonstrate the effectiveness of our method in alleviating model bias and handling long-tailed noisy label data. Code is available at \url{https://anonymous.4open.science/r/neurips-2025-anonymous-1015/}.
Chen Shu, Ruichi Zhang, Mengke Li 0001, Yonggang Zhang 0003, Yang Lu 0009, Bo Han 0003, Yiu-Ming Cheung, Hanzi Wang
NeurIPS6
2025 Dual-Imbalance Mitigation in Semi-Supervised Federated Learning Through Candidate-Aware Prototype Aggregation
Shanshan Yan, Hezhao Liu, Shihao Hou, Yang Lu 0009
PRCV (2)5
2025 Data Enhancement for Long-tailed Tasks: Diffusion Model with Optimized Quality Filter
abstract
In the field of long-tail learning, the uneven distribution of datasets leads to a significant decrease in the model’s accuracy for tail classes. The data augmentation method is an effective way to alleviate the long-tail problem. The diversity of most existing augmentation methods is insufficient, leading to the limited improvement of tail class information. Due to the diffusion model’s ability to generate images with high quality, rich details, and diversity by gradually denoising images, we propose a augmentation method based on the diffusion model called Diffusion-VagMix (DVM). Based on the diffusion model, this method uses an Optimized Quality Filter (OptiFilter) to scalp low-quality generated images and perform further data augmentation. Our approach effectively addresses the issue of insufficient samples in tail classes and uneven quality of the resulting images by augmenting the dataset with high-quality generated images. This strategy leads to improvements in the accuracy of tail classes and enhances the model’s overall performance. Our DVM method can be used on other long-tail learning methods at will, which can make further improvements. The source code is available at https://github.com/Alert-M/Diffusion-VagMix.
Jiawei You, Yichuan Zhai, Mengke Li 0001, Yang Lu 0009
SMC4
2025 Learning unified distance metric for heterogeneous attribute data clustering
Yiqun Zhang 0006, Mingjie Zhao 0003, Yang Lu 0009, Yiu-Ming Cheung
Expert Syst. Appl.4
2025 IMC-Det: Intra-Inter Modality Contrastive Learning for Video Object Detection
Qiang Qi, Zhenyu Qiu, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
Int. J. Comput. Vis.4
2025 Adaptive Middle Modality Alignment Learning for Visible-Infrared Person Re-identification
Yan Yan 0001, Yang Lu 0009, Hanzi Wang
Int. J. Comput. Vis.3
2025 CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
Mingquan Lin, Gregory Holste, Song Wang 0026, Yiliang Zhou, Yishu Wei, Imon Banerjee, Pengyi Chen, Tianjie Dai, Yuexi Du, Nicha C. Dvornek, Yuyan Ge, Zuwei Guo, Shohei Hanaoka, Dongkyun Kim, Pablo Messina, Yang Lu 0009, Denis Parra, Donghyun Son, Alvaro Soto, Aisha Urooj Khan, René Vidal, Yosuke Yamagishi, Pingkun Yan, Zefan Yang, Ruichi Zhang, Yang Zhou 0019, Leo A. Celi, Ronald M. Summers, Zhiyong Lu, Hao Chen 0011, Adam E. Flanders, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.16
2025 FediOS: decoupling orthogonal subspaces for personalization in feature-skew federated learning
Lingzhi Gao, Zexi Li 0001, Xinyi Shang, Yang Lu 0009, Chao Wu 0001
Mach. Learn.4
2025 GradToken: Decoupling tokens with class-aware gradient for visual explanation of Transformer network
Yang Lu 0009, Yiu-Ming Cheung
Neural Networks3
2025 Categorical Data Clustering via Value Order Estimated Distance Metric Learning
abstract
Clustering is a popular machine learning technique for data mining that can process and analyze datasets to automatically reveal sample distribution patterns. Since the ubiquitous categorical data naturally lack a well-defined metric space such as the Euclidean distance space of numerical data, the distribution of categorical data is usually under-represented, and thus valuable information can be easily twisted in clustering. This paper, therefore, introduces a novel order distance metric learning approach to intuitively represent categorical attribute values by learning their optimal order relationship and quantifying their distance in a line similar to that of the numerical attributes. Since subjectively created qualitative categorical values involve ambiguity and fuzziness, the order distance metric is learned in the context of clustering. Accordingly, a new joint learning paradigm is developed to alternatively perform clustering and order distance metric learning with low time complexity and a guarantee of convergence. Due to the clustering-friendly order learning mechanism and the homogeneous ordinal nature of the order distance and Euclidean distance, the proposed method achieves superior clustering accuracy on categorical and mixed datasets. More importantly, the learned order distance metric greatly reduces the difficulty of understanding and managing the non-intuitive categorical data. Experiments with ablation studies, significance tests, case studies, etc., have validated the efficacy of the proposed method. The source code is available at https://github.com/csmjzhao/OCL_Source_Code.
Yiqun Zhang 0006, Mingjie Zhao 0003, Hong Jia, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung
Proc. ACM Manag. Data5
2025 Uncertainty-Aware Label Refinement on Hypergraphs for Personalized Federated Facial Expression Recognition
abstract
Most facial expression recognition (FER) models are trained on large-scale expression data with centralized learning. Unfortunately, collecting a large amount of centralized expression data is difficult in practice due to privacy concerns of facial images. In this paper, we investigate FER under the framework of personalized federated learning, which is a valuable and practical decentralized setting for real-world applications. To this end, we develop a novel uncertainty-Aware label refineMent on hYpergraphs (AMY) method. For local training, each local model consists of a backbone, an uncertainty estimation (UE) block, and an expression classification (EC) block. In the UE block, we leverage a hypergraph to model complex high-order relationships between expression samples and incorporate these relationships into uncertainty features. A personalized uncertainty estimator is then introduced to estimate reliable uncertainty weights of samples in the local client. In the EC block, we perform label propagation on the hypergraph, obtaining high-quality refined labels for retraining an expression classifier. Based on the above, we effectively alleviate heterogeneous sample uncertainty across clients and learn a robust personalized FER model in each client. Experimental results on two challenging real-world facial expression databases show that our proposed method consistently outperforms several state-of-the-art methods. This indicates the superiority of hypergraph modeling for uncertainty estimation and label refinement on the personalized federated FER task. The source code will be released athttps://github.com/mobei1006/AMY.
Hu Ding 0005, Yan Yan 0001, Yang Lu 0009, Jing-Hao Xue, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2025 CGATracker: Correlation-Aware Graph Alignment for Referring Multi-Object Tracking
abstract
Referring multi-object tracking (RMOT) aims to identify specific targets based on sentence descriptions. To enhance multi-modal learning, previous works typically relied on a simple fusion module at early or late stages. However, those methods frequently underutilize textual semantics and struggle to model the relationships between region-level features and word-level features. To address these limitations, we propose CGATracker, a correlation-aware graph alignment method for RMOT, which facilitates precise relationship modeling through relational scoring. Specifically, we design a Language-driven Relational Alignment (LRA) module, which establishes two connection graphs to generate positive and negative samples for the visual-textual alignment. Additionally, to effectively leverage referring information, we introduce a Semantic Clarify Booster (SCBooster) module based on a semantic infusion mechanism and a bias-aware verification mechanism for interactions with different modalities. Moreover, by designing a Multi-level Cross-modal Fusion (MCF) module, our method aggregates contextual features at multiple depths to enable the creation of the enriched correlation-aware graph. Extensive experiments conducted on the Refer-KITTI and Refer-KITTI-V2 datasets demonstrate the effectiveness of CGATracker.
Siping Zhuang, Qiangqiang Wu, Yang Lu 0009, Hai-Miao Hu, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Augmentation Matters: A Mix-Paste Method for X-Ray Prohibited Item Detection Under Noisy Annotations
abstract
Automatic X-ray prohibited item detection is vital for public safety. Existing deep learning-based methods all assume that the annotations of training X-ray images are correct. However, obtaining correct annotations is extremely hard if not impossible for large-scale X-ray images, where item overlapping is ubiquitous. As a result, X-ray images are easily contaminated with noisy annotations, leading to performance deterioration of existing methods. In this paper, we address the challenging problem of training a robust prohibited item detector under noisy annotations (including both category noise and bounding box noise) from a novel perspective of data augmentation, and propose an effective label-aware mixed patch paste augmentation method (Mix-Paste). Specifically, for each item patch, we mix several item patches with the same category label from different images and replace the original patch in the image with the mixed patch. In this way, the probability of containing the correct prohibited item within the generated image is increased. Meanwhile, the mixing process mimics item overlapping, enabling the model to learn the characteristics of X-ray images. Moreover, we design an item-based large-loss suppression (LLS) strategy to suppress the large losses corresponding to potentially positive predictions of additional items due to the mixing operation. We show the superiority of our method on X-ray datasets under noisy annotations. In addition, we evaluate our method on the noisy MS-COCO dataset to showcase its generalization ability. These results clearly indicate the great potential of data augmentation to handle noise annotations. The source code is released athttps://github.com/wscds/Mix-Paste.
Ruikang Chen, Yan Yan 0001, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang
IEEE Trans. Inf. Forensics Secur.4
2025 A Unified Multi-Domain Face Normalization Framework for Cross-Domain Prototype Learning and Heterogeneous Face Recognition
abstract
Face normalization is a critical technique for improving the robustness and generalizability of face recognition systems by reducing intra-personal variations arising from expressions, poses, occlusions, illuminations, and domain shifts. Existing normalization methods, however, often lack the flexibility to handle multi-factorial variations and exhibit limited cross-domain adaptability. To address these challenges, we propose a Unified Multi-Domain Face Normalization Network (UMFN), which is designed to process facial images with diverse variations from various domains and reconstruct frontal, neutralized facial prototypes in the target domain. As an unsupervised domain adaptation model, the UMFN facilitates concurrent training across multiple cross-domain datasets and demonstrates robust prototype reconstruction capabilities. Notably, the UMFN functions as a joint prototype and feature learning framework, extracting domain-agnostic identity features through a decoupling mapping network and adversarial training with a feature domain classifier. Furthermore, we design an efficient Heterogeneous Face Recognition (HFR) network that integrates these domain-agnostic features and the identity-discriminative features extracted from normalized prototypes, enhanced by contrastive learning to improve identity recognition accuracy. Empirical evaluation on multiple cross-domain benchmark datasets validate the effectiveness of the UMFN for face normalization and the superiority of the HFR network for heterogeneous face recognition.
Yang Lu 0009, Yiu-Ming Cheung, Nanrun Zhou
IEEE Trans. Inf. Forensics Secur.3
2025 Image-Attribute and Frequency-Spatial Dual Collaborative Learning for Pedestrian Attribute Recognition
Xinwen Fan, Yang Lu 0009, Hanzi Wang
IEEE Trans. Inf. Forensics Secur.4
2025 Frequency Domain Nuances Mining for Visible-Infrared Person Re-Identification
abstract
This paper focuses on the visible-infrared person re-identification (VIReID) task, which is essential for information forensics and security as it enables accurate person re-identification across low-light or nighttime conditions. The primary challenge in the VIReID task is to reduce the modality discrepancy between visible and infrared images. Current methods mainly utilize the spatial information, often neglecting the discriminative potential of frequency information. To address this issue, this paper aims to mitigate the modality discrepancy from a frequency domain perspective. Specifically, we propose a novel Frequency Domain Nuances Mining (FDNM) method, which mainly includes a Salience-guided Phase Enhancement (SPE) module and an Amplitude Nuances Mining (ANM) module, to effectively explore the cross-modality frequency domain information. These two modules are mutually beneficial to jointly explore frequency-domain visible-infrared nuances, thereby significantly reducing the modality discrepancy in the frequency domain. Additionally, we propose a Center-guided Nuances Mining (CNM) loss to ensure that the ANM module retains discriminative identity information while discovering diverse cross-modality nuances. Extensive experiments show that the proposed FDNM has significant advantages in improving the performance of VIReID. For instance, our method respectively outperforms the second-best method by 5.2% in Rank-1 accuracy and 5.8% in mAP on the SYSU-MM01 dataset under the indoor search mode. Furthermore, we also demonstrate the effectiveness and generalization of the proposed FDNM method in the challenging visible-infrared face recognition task.
Hanzi Wang, Yang Lu 0009, Yan Yan 0001, Xuelong Li 0001
IEEE Trans. Inf. Forensics Secur.3
2025 MOOD: Leveraging Out-of-Distribution Data to Enhance Imbalanced Semi-Supervised Learning
abstract
The imbalanced semi-supervised learning (SSL) has emerged as a critical research area due to the prevalence of class imbalanced and partially labeled data in real-world scenarios. As the requirement for data volume increases, naturally collected datasets inevitably contain out-of-distribution (OOD) samples. However, the performance of existing imbalanced SSL methods experiences a marked deterioration with OOD data. In this article, we propose an imbalanced SSL method called mixup-OOD (MOOD) to address this issue. The core idea is to "turn waste into treasure," exploring the potential of leveraging seemingly detrimental OOD data to expand the feature space, particularly for tail classes. Specifically, we first filter OOD data from unlabeled data, and then fuse it with labeled data to boost feature diversity for the tail classes. To avoid feature overlapping with OOD data, we develop a push-and-pull (PaP) loss to attract in-distribution (ID) instances toward respective class centroids while repelling OOD samples from them. Extensive experiments show that MOOD achieves superior performance compared with other state-of-the-art methods and exhibits robustness across data with different imbalanced ratios and OOD proportions. The source code is available at: https://github.com/xlhuang132/MOODv2.
Yang Lu 0009, Xiaolin Huang, Mengke Li 0001, Yan Yan 0001, Chen Gong 0002, Hanzi Wang
IEEE Trans. Neural Networks Learn. Syst.1
2024 Feature Fusion from Head to Tail for Long-Tailed Visual Recognition
abstract
The imbalanced distribution of long-tailed data presents a considerable challenge for deep learning models, as it causes them to prioritize the accurate classification of head classes but largely disregard tail classes. The biased decision boundary caused by inadequate semantic information in tail classes is one of the key factors contributing to their low recognition accuracy. To rectify this issue, we propose to augment tail classes by grafting the diverse semantic information from head classes, referred to as head-to-tail fusion (H2T). We replace a portion of feature maps from tail classes with those belonging to head classes. These fused features substantially enhance the diversity of tail classes. Both theoretical analysis and practical experimentation demonstrate that H2T can contribute to a more optimized solution for the decision boundary. We seamlessly integrate H2T in the classifier adjustment stage, making it a plug-and-play module. Its simplicity and ease of implementation allow for smooth integration with existing long-tailed recognition methods, facilitating a further performance boost. Extensive experiments on various long-tailed benchmarks demonstrate the effectiveness of the proposed H2T. The source code is available at https://github.com/Keke921/H2T.
Mengke Li 0001, Zhikai Hu, Yang Lu 0009, Weichao Lan, Yiu-Ming Cheung, Hui Huang 0004
AAAI3
2024 Federated Learning with Extremely Noisy Clients via Negative Distillation
abstract
Federated learning (FL) has shown remarkable success in cooperatively training deep models, while typically struggling with noisy labels. Advanced works propose to tackle label noise by a re-weighting strategy with a strong assumption, i.e., mild label noise. However, it may be violated in many real-world FL scenarios because of highly contaminated clients, resulting in extreme noise ratios, e.g., >90%. To tackle extremely noisy clients, we study the robustness of the re-weighting strategy, showing a pessimistic conclusion: minimizing the weight of clients trained over noisy data outperforms re-weighting strategies. To leverage models trained on noisy clients, we propose a novel approach, called negative distillation (FedNed). FedNed first identifies noisy clients and employs rather than discards the noisy clients in a knowledge distillation manner. In particular, clients identified as noisy ones are required to train models using noisy labels and pseudo-labels obtained by global models. The model trained on noisy labels serves as a ‘bad teacher’ in knowledge distillation, aiming to decrease the risk of providing incorrect information. Meanwhile, the model trained on pseudo-labels is involved in model aggregation if not identified as a noisy client. Consequently, through pseudo-labeling, FedNed gradually increases the trustworthiness of models trained on noisy clients, while leveraging all clients for model aggregation through negative distillation. To verify the efficacy of FedNed, we conduct extensive experiments under various settings, demonstrating that FedNed can consistently outperform baselines and achieve state-of-the-art performance.
Yang Lu 0009, Yonggang Zhang 0003, Yiliang Zhang, Bo Han 0003, Yiu-Ming Cheung, Hanzi Wang
AAAI1
2024 CLIP-Guided Federated Learning on Heterogeneity and Long-Tailed Data
abstract
Federated learning (FL) provides a decentralized machine learning paradigm where a server collaborates with a group of clients to learn a global model without accessing the clients' data. User heterogeneity is a significant challenge for FL, which together with the class-distribution imbalance further enhances the difficulty of FL. Great progress has been made in large vision-language models, such as Contrastive Language-Image Pre-training (CLIP), which paves a new way for image classification and object recognition. Inspired by the success of CLIP on few-shot and zero-shot learning, we use CLIP to optimize the federated learning between server and client models under its vision-language supervision. It is promising to mitigate the user heterogeneity and class-distribution balance due to the powerful cross-modality representation and rich open-vocabulary prior knowledge. In this paper, we propose the CLIP-guided FL (CLIP2FL) method on heterogeneous and long-tailed data. In CLIP2FL, the knowledge of the off-the-shelf CLIP model is transferred to the client-server models, and a bridge is built between the client and server. Specifically, for client-side learning, knowledge distillation is conducted between client models and CLIP to improve the ability of client-side feature representation. For server-side learning, in order to mitigate the heterogeneity and class-distribution imbalance, we generate federated features to retrain the server model. A prototype contrastive learning with the supervision of the text encoder of CLIP is introduced to generate federated features depending on the client-side gradients, and they are used to retrain a balanced server classifier. Extensive experimental results on several benchmarks demonstrate that CLIP2FL achieves impressive performance and effectively deals with data heterogeneity and long-tail distribution. The code is available at https://github.com/shijiangming1/CLIP2FL.
Jiangming Shi, Shanshan Zheng, Xiangbo Yin, Yang Lu 0009, Yuan Xie 0006, Yanyun Qu
AAAI4
2024 GMAE: Representation Learning on Graph via Masked Graph Autoencoders
abstract
In recent years, self-supervised learning, particularly autoencoders, has garnered extensive research interest, emerging as one of the most exciting learning paradigms and achieving success in computer vision and other artificial intelligence domains. However, autoencoders have yet to demonstrate their potential in processing graph data fully. To effectively learn from graph data, we propose a novel Masked Graph Autoencoder (MGAE) framework. MGAE utilizes the pretext task of masked graph reconstruction to learn high-quality representations from unlabeled graph data, incorporating two core designs in its structure. Firstly, we randomly masked numerous node features and edges during training and attempted to reconstruct this missing information. This approach creates a challenging yet meaningful self-supervised task, beneficial for applying the model to downstream tasks. Secondly, to aid the model in reconstructing the extensively masked information, we introduce a new module for matching the latent representation space, utilizing a neural network with the same structure as the encoder to generate predictive targets from the unmasked graph. This design mitigates the impact of graph masking and enhances the model’s capability to capture high-level semantic information. Our framework outperforms traditional graph autoencoders in handling graph data by offering a more robust and reliable self-supervised mechanism. Extensive experiments conducted on various graph benchmarks have demonstrated MGAE’s effectiveness across multiple tasks, including node classification, link prediction, and graph classification.
Chengbin Zheng, Yang Lu 0009
CSCWD3
2024 Learning Order Forest for Qualitative-Attribute Data Clustering
abstract
Clustering is a fundamental approach to understanding data patterns, wherein the intuitive Euclidean distance space is commonly adopted. However, this is not the case for implicit cluster distributions reflected by qualitative attribute values, e.g., the nominal values of attributes like symptoms, marital status, etc. This paper, therefore, discovered a tree-like distance structure to flexibly represent the local order relationship among intra-attribute qualitative values. That is, treating a value as the vertex of the tree allows to capture rich order relationships among the vertex value and the others. To obtain the trees in a clustering-friendly form, a joint learning mechanism is proposed to iteratively obtain more appropriate tree structures and clusters. It turns out that the latent distance space of the whole dataset can be well-represented by a forest consisting of the learned trees. Extensive experiments demonstrate that the joint learning adapts the forest to the clustering task to yield accurate results. Comparisons of 10 counterparts on 12 real benchmark datasets with significance tests verify the superiority of the proposed method. Source code of the proposed method is available at [39].
Mingjie Zhao 0003, Sen Feng, Yiqun Zhang 0006, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung
ECAI5
2024 Semantic-Guided Network with Contrastive Learning for Video Caption
abstract
Video captioning is a challenging task, which aims at generating a sentence to describe the content of a video using the natural language. Many existing methods model visual features (2D/3D) extracted from videos to generate captions, but they neglect semantic guidance. Empirically, visual features contain fine-grained information such as color and shape, while the generation of captions requires more emphasis on the semantic and syntax clues that cannot be studied adequately only from the caption loss. To alleviate this problem, we propose a semantic-guided network based on contrastive learning (SNCL), which makes use of both vision and text information to enrich the features with contextual guidance. Based on the meaningful features, hierarchical reasoning modules are employed to perform the key phrase prediction task in order to enhance our model with specific semantic guidance. Experimental results on the MSVD and MSR-VTT datasets show that our SNCL outperforms recent state-of-the-art methods.
Kaixuan Chen 0006, Qianji Di, Yang Lu 0009, Hanzi Wang
ICASSP3
2024 Visual-Linguistic Representation Learning with Deep Cross-Modality Fusion for Referring Multi-Object Tracking
abstract
Referring multi-object tracking is a new rising research topic that aims at detecting and tracking the referred objects in a video sequence based on a natural language expression. Compared with traditional multi-object tracking, this setting guides object tracking with high-level semantic information, which may bring more flexible and robust tracking performance in practical scenarios. However, existing methods perform cross-modal fusion in only one phase. The limitation of visual-linguistic representation is prone to causing visionlanguage mismatching and producing poor tracking results. To effectively fuse vision and language modalities, we propose DeepRMOT with deep cross-modality fusion, including an enhanced early-fusion module, a bidirectional crossmodality encoder, and a cross-modality decoder. Therefore, DeepRMOT can boost object detection and data association by the enhanced visual-linguistic representation. Extensive experiments on the Refer-KITTI demonstrate the effectiveness of our method.
Wenyan He, Yajun Jian, Yang Lu 0009, Hanzi Wang
ICASSP3
2024 Spatio-Temporal Correlation Learning for Multiple Object Tracking
abstract
Multi-object tracking (MOT) has gained remarkable progress in recent years, while due to the complexity of real-world environments, there are still many challenges that remain unsolved, such as object occlusion and deformation. To effectively alleviate this problem, we propose a simple yet effective Transformer-based tracker, named CLNet, consisting of an Instance-Aware Localization (IAL) module and a Temporal Context Aggregation (TCA) module. Specifically, the former learns the correlation of object positions for potential location estimation, and the latter learns the correlation of background contexts to obtain robust re-ID features for data association. Experimental results show that CLNet outperforms the baseline method by +2.2 MOTA and +2.3 IDF1 on MOT17 and +6.3 MOTA and +3.6 IDF1 on MOT20 respectively, which demonstrate the effectiveness of the proposed method.
Yajun Jian, Chihui Zhuang, Wenyan He, Kaiwen Du, Yang Lu 0009, Hanzi Wang
ICASSP5
2024 Proposal Distillation of Multi-Modal Feature Aggregation Network for Video Object Detection
abstract
Video object detection is a challenging task due to deteriorated object appearances. In order to bolster per-frame feature representations, one way is to aggregate features from relevant frames. However, relying exclusively on RGB modal for feature aggregation may limit the detection performance for lacking of motion robustness. We propose a novel proposal distillation of multi-modal feature aggregation network (PDMAN). Specially, it initially aligns the feature domain and flow domain via a lightweight flow module (LFM) and then facilities frame-level feature aggregation. Subsequently, a global-based semantic embedding module (GSEM) is designed to incorporate global semantic features into instance features and introduce a global multi-label classification loss to guide encoding with high class-wise responsiveness. Finally, to alleviate the presence of insufficient and redundant information in multi-modal instance-level feature aggregation, a proposal distilled aggregation module (PDAM) is employed. By distilling the instance set, this approach realizes a fine-grained feature aggregation, ultimately boosting the detection performance. Experimental results demonstrate that the proposed PDMAN achieves a favorable result on the most representative large-scale ImageNet VID dataset.
Zhenyu Qiu, Qiang Qi, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
ICASSP3
2024 Dynamically Anchored Prompting for Task-Imbalanced Continual Learning
Chenxing Hong, Zhiqi Kang, Mengke Li 0001, Yang Lu 0009, Hanzi Wang
IJCAI6
2024 Clustering by Learning the Ordinal Relationships of Qualitative Attribute Values
abstract
In many real-world clustering tasks, data objects are described by both quantitative and qualitative attributes. Attributes with semantically ordered qualitative values are very common and are usually coded according to their order (i.e., consecutive integers) for clustering. However, semantic order is not always globally interdependent with a certain clustering task. An intuitive case is that level of income (attribute) is not always positively correlated with the level of mental health (label). Using mismatched order surely forms a bottleneck to clustering performance, and conversely, the unsupervised clustering process prevents understanding of "true" order. Therefore, we proposed a novel learning paradigm to tune the value order. More specifically, we adjust the intra-attribute orders, and let this process learn mutually with object clustering, thus bridging the gap between value order and clustering task. To the best of our knowledge, this is the first attempt to learn ordinal relationships among qualitative attribute values. Extensive experiments with significance tests show that our method outperforms the existing relevant clustering approaches on qualitative attribute data.
Yiqun Zhang 0006, Yang Lu 0009, Mengke Li 0001, Yiu-Ming Cheung
IJCNN4
2024 Diverse Consensuses Paired with Motion Estimation-Based Multi-Model Fitting
abstract
Multi-model fitting aims to robustly estimate the parameters of various model instances in data contaminated by noise and outliers. Most previous works employ only a single type of consensus or implicit fusion model to represent the correlation between data points and model hypotheses. This approach often results in unrealistic and incorrect model fitting in the presence of noise and uncertainty. In this paper, we propose a novel method of diverse Consensuses paired with Motion estimation-based multi-Model Fitting (CMMF), which leverages three types of diverse consensuses along with inter-model collaboration to enhance the effectiveness of multi-model fusion. We design a Tangent Consensus Residual Reconstruction (TCRR) module to capture motion structure information of two points at the pixel level. Additionally, we introduce a Cross Consensus Affinity (CCA) framework to strengthen the correlation between data points and model hypotheses. To address the challenge of multi-body motion estimation, we propose a Nested Consensus Clustering (NCC) strategy, which formulates multi-model fitting as a motion estimation problem. It explicitly establishes motion collaboration between models and ensures that multiple models are well-fitted. Extensive quantitative and qualitative experiments are conducted on four public datasets (i.e., AdelaideRMF-F, Hopkins155, KITTI, MTPV62), and the results demonstrate that our proposed method outperforms several state-of-the-art methods.
Wenyu Yin, Shuyuan Lin, Yang Lu 0009, Hanzi Wang
ACM Multimedia3
2024 Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised Visible-Infrared Person Re-identification (USVI-ReID) presents a formidable challenge, which aims to match pedestrian images across visible and infrared modalities without any annotations. Recently, clustered pseudo-label methods have become predominant in USVI-ReID, although the inherent noise in pseudo-labels presents a significant obstacle. Most existing works primarily focus on shielding the model from the harmful effects of noise, neglecting to calibrate noisy pseudo-labels usually associated with hard samples, which will compromise the robustness of the model. To address this issue, we design a Robust Pseudo-label Learning with Neighbor Relation (RPNR) framework for USVI-ReID. To be specific, we first introduce a straightforward yet potent Noisy Pseudo-label Calibration module to correct noisy pseudo-labels. Due to the high intra-class variations, noisy pseudo-labels are difficult to calibrate completely. Therefore, we introduce a Neighbor Relation Learning module to reduce high intra-class variations by modeling potential interactions between all samples. Subsequently, we devise an Optimal Transport Prototype Matching module to establish reliable cross-modality correspondences. On that basis, we design a Memory Hybrid Learning module to jointly learn modality-specific and modality-invariant information. Comprehensive experiments conducted on two widely recognized benchmarks, SYSU-MM01 and RegDB, demonstrate that RPNR outperforms the current state-of-the-art GUR with an average Rank-1 improvement of 10.3%. The code is available at https://github.com/XiangboYin/RPNR.
Xiangbo Yin, Jiangming Shi, Yachao Zhang 0001, Yang Lu 0009, Zhizhong Zhang 0001, Yuan Xie 0006, Yanyun Qu
ACM Multimedia4
2024 Semi-supervised Visible-Infrared Person Re-identification via Modality Unification and Confidence Guidance
abstract
Semi-supervised visible-infrared person re-identification (SSVI-ReID) aims to match pedestrian images of the same identity from different modalities (visible and infrared) while only annotating visible images, which is highly related to multimedia and multi-modal processing. Existing works primarily focus on assigning accurate pseudo-labels to infrared images, but overlook the two key challenges: erroneous pseudo-labels and large modality discrepancy. To alleviate these issues, this paper proposes a novel Modality-Unified and Confidence-Guided (MUCG) semi-supervised learning method. Specifically, we first propose a Dynamic Intermediate Modality Generation (DIMG) module, which transfers knowledge from labeled visible images to unlabeled infrared images, enhancing the pseudo-label quality and bridging the modality discrepancy. Meanwhile, we propose a Weighted Identification Loss (WIL) that can reduce the model's dependence on erroneous labels by using confidence weighting. Moreover, an effective Modality Consistency Loss (MCL) is proposed to narrow the distribution of visible and infrared features, further narrowing the modality discrepancy and enabling the learning of modality-unified features. Extensive experiments show that the proposed MUCG has significant advantages in improving the performance of the SSVI-ReID task, surpassing the current state-of-the-art methods by a significant margin.
Xiying Zheng, Yang Lu 0009, Hanzi Wang
ACM Multimedia3
2024 Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual Recognition
abstract
Long-tailed visual recognition has received increasing attention recently. Despite fine-tuning techniques represented by visual prompt tuning (VPT) achieving substantial performance improvement by leveraging pre-trained knowledge, models still exhibit unsatisfactory generalization performance on tail classes. To address this issue, we propose a novel optimization strategy called Gaussian neighborhood minimization prompt tuning (GNM-PT), for VPT to address the long-tail learning problem. We introduce a novel Gaussian neighborhood loss, which provides a tight upper bound on the loss function of data distribution, facilitating a flattened loss landscape correlated to improved model generalization. Specifically, GNM-PT seeks the gradient descent direction within a random parameter neighborhood, independent of input samples, during each gradient update. Ultimately, GNM-PT enhances generalization across all classes while simultaneously reducing computational overhead. The proposed GNM-PT achieves state-of-the-art classification accuracies of 90.3%, 76.5%, and 50.1% on benchmark datasets CIFAR100-LT (IR 100), iNaturalist 2018, and Places-LT, respectively. The source code is available at https://github.com/Keke921/GNM-PT.
Mengke Li 0001, Yang Lu 0009, Yiqun Zhang 0006, Yiu-Ming Cheung, Hui Huang 0004
NeurIPS3
2024 Relationship Prompt Learning is Enough for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation (OVSS) aims to segment unseen classes without corresponding labels. Existing Vision-Language Model (VLM)-based methods leverage VLM's rich knowledge to enhance additional explicit segmentation-specific networks, yielding competitive results, but at the cost of extensive training cost. To reduce the cost, we attempt to enable VLM to directly produce the segmentation results without any segmentation-specific networks. Prompt learning offers a direct and parameter-efficient approach, yet it falls short in guiding VLM for pixel-level visual classification. Therefore, we propose the ${\bf R}$elationship ${\bf P}$rompt ${\bf M}$odule (${\bf RPM}$), which generates the relationship prompt that directs VLM to extract pixel-level semantic embeddings suitable for OVSS. Moreover, RPM integrates with VLM to construct the ${\bf R}$elationship ${\bf P}$rompt ${\bf N}$etwork (${\bf RPN}$), achieving OVSS without any segmentation-specific networks. RPN attains state-of-the-art performance with merely about ${\bf 3M}$ trainable parameters (2\% of total parameters).
Jiahao Li 0003, Yang Lu 0009, Yuan Xie 0006, Yanyun Qu
NeurIPS2
2024 Wavelet-domain feature decoupling for weakly supervised multi-object tracking
Yu-Lei Li, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
Sci. China Inf. Sci.3
2024 Personalized Federated Learning on long-tailed data via knowledge distillation and generated features
Fengling Lv, Pinxin Qian, Yang Lu 0009, Hanzi Wang
Pattern Recognit. Lett.3
2024 Label-noise learning via uncertainty-aware neighborhood sample selection
Yiliang Zhang, Yang Lu 0009, Hanzi Wang
Pattern Recognit. Lett.2
2024 PARFormer: Transformer-Based Multi-Task Network for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition (PAR) has received increasing attention because of its wide application in video surveillance and pedestrian analysis. Extracting robust feature representation is one of the key challenges in this task. The existing methods primarily rely on convolutional neural networks (CNNs) as the backbone network for feature extraction. However, these methods mainly focus on small discriminative regions while ignoring the global perspective. To overcome these limitations, we propose PARFormer, a pure transformer-based multi-task PAR network consisting of four modules. In the feature extraction module, we build a transformer-based strong baseline for feature extraction, which achieves competitive results on several PAR benchmarks compared with the existing CNN-based baseline methods. Since the PAR task is vulnerable to environmental factors, we enhance feature robustness in the feature processing module and propose an effective data augmentation strategy named batch random mask (BRM) block to reinforce the attentive feature learning of random patches. Furthermore, we propose a multi-attribute center loss (MACL) to augment the inter-attribute discriminability of feature representations. As viewpoints can affect some specific attributes, in the viewpoint perception module, we propose a multi-view contrastive loss (MVCL) that enables the network to exploit the viewpoint information. In the attribute recognition module, we alleviate the negative-positive imbalance problem to generate the attribute predictions. These modules interact and jointly learn a highly discriminative feature space and supervise the generation of the final features. Extensive experimental results show that the proposed PARFormer network performs well compared to the state-of-the-art methods on several public datasets, including PETA, RAP, and PA100K. Code will be released athttps://github.com/xwf199/PARFormer.
Xinwen Fan, Yang Lu 0009, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Few-Shot Action Recognition via Multi-View Representation Learning
abstract
Few-shot action recognition aims to recognize novel action classes with limited labeled samples and has recently received increasing attention. The core objective of few-shot action recognition is to enhance the discriminability of feature representations. In this paper, we propose a novel multi-view representation learning network (MRLN) to model intra-video and inter-video relations for few-shot action recognition. Specifically, we first propose a spatial-aware aggregation refinement module (SARM), which mainly consists of a spatial-aware aggregation sub-module and a spatial-aware refinement sub-module to explore the spatial context of samples at the frame level. Then, we design a temporal-channel enhancement module (TCEM), which can capture the temporal-aware and channel-aware features of samples with the elaborately designed temporal-aware enhancement sub-module and channel-aware enhancement sub-module. Third, we introduce a cross-video relation module (CVRM), which can explore the relations across videos by utilizing the self-attention mechanism. Moreover, we design a prototype-centered mean absolute error loss to improve the feature learning capability of the proposed MRLN. Extensive experiments on four prevalent few-shot action recognition benchmarks show that the proposed MRLN can significantly outperform a variety of state-of-the-art few-shot action recognition methods. Especially, on the 5-way 1-shot setting, our MRLN respectively achieves 75.7%, 86.9%, 65.5% and 45.9% on the Kinetics, UCF101, HMDB51 and SSv2 datasets.
Xiao Wang 0072, Yang Lu 0009, Wanchuan Yu, Yanwei Pang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.2
2023 Long-Tailed Visual Recognition via Self-Heterogeneous Integration with Knowledge Excavation
abstract
Deep neural networks have made huge progress in the last few decades. However, as the real-world data often exhibits a long-tailed distribution, vanilla deep models tend to be heavily biased toward the majority classes. To address this problem, state-of-the-art methods usually adopt a mixture of experts (MoE) to focus on different parts of the long-tailed distribution. Experts in these methods are with the same model depth, which neglects the fact that different classes may have different preferences to be fit by models with different depths. To this end, we propose a novel MoE-based method called Self-Heterogeneous Integration with Knowledge Excavation (SHIKE). We first propose Depth-wise Knowledge Fusion (DKF) to fuse features between different shallow parts and the deep part in one network for each expert, which makes experts more diverse in terms of representation. Based on DKF, we further propose Dynamic Knowledge Transfer (DKT) to reduce the influence of the hardest negative class that has a non-negligible impact on the tail classes in our MoEframework. As a result, the classification accuracy of long-tailed data can be significantly improved, especially for the tail classes. SHIKE achieves the state-of-the-art performance of 56.3%, 60.3%, 75.4% and 41.9% on CIFAR100-LT (IF100), ImageNet-LT, iNaturalist 2018, and Places-LT, respectively. The source code is available at https://github.com/jinyan-06/SHIKE.
Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung, Hanzi Wang
CVPR3
2023 Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object Tracking
abstract
Recently, some weakly supervised multi-object tracking (MOT) methods learn identity embedding features with pseudo identity labels rather than the high-cost manual ones. However, these pseudo identity labels may contain many false or missing identities, which adversely affect the optimization of tracking networks, resulting in interrupted trajectories of occluded targets. To effectively reconnect the interrupted trajectories caused by noisy pseudo labels, we propose a novel weakly supervised MOT method based on a Trajectory-Reconnecting Transformer (TRTMOT). TRT-MOT performs feature decoupling to extract discriminative embedding features for reconnecting trajectories of occluded targets. Experimental results show that TRTMOT outperforms previous weakly supervised MOT methods by at least +3.6 and +5.6 on MOTA for the MOT17 and MOT20 datasets, respectively.
Yu-Lei Li, Yang Lu 0009, Jie Li 0001, Hanzi Wang
ICASSP2
2023 Personalized Federated Learning on Long-Tailed Data via Adversarial Feature Augmentation
abstract
Personalized Federated Learning (PFL) aims to learn personalized models for each client based on the knowledge across all clients in a privacy-preserving manner. Existing PFL methods generally assume that the underlying global data across all clients are uniformly distributed without considering the long-tail distribution. The joint problem of data heterogeneity and long-tail distribution in the FL environment is more challenging and severely affects the performance of personalized models. In this paper, we propose a PFL method called Federated Learning with Adversarial Feature Aug-mentation (FedAFA) to address this joint problem in PFL. FedAFA optimizes the personalized model for each client by producing a balanced feature set to enhance the local minority classes. The local minority class features are generated by transferring the knowledge from the local majority class features extracted by the global model in an adversarial example learning manner. The experimental results on benchmarks under different settings of data heterogeneity and long-tail distribution demonstrate that FedAFA significantly improves the personalized performance of each client compared with the state-of-the-art PFL algorithm. The code is available at https://github.com/pxqian/FedAFA.
Yang Lu 0009, Pinxin Qian, Gang Huang 0004, Hanzi Wang
ICASSP1
2023 DQFORMER: Dynamic Query Transformer for Lane Detection
abstract
Lane detection is one of the most important tasks in self-driving. The critical purpose of lane detection is the prediction of lane shapes. Meanwhile, it is challenging and difficult to determine lane instance positions before predicting lane shapes in an image. In this paper, we propose a top-down method called Dynamic Query Transformer (DQFormer), which uses a Dynamic Lane Queries (DLQs) module to predict lane shapes. Specifically, to accurately predict lane shapes, we propose a new framework for generating dynamic weights based on DLQs, which can focus on the context of lane shapes dynamically. Unlike existing transformer-based methods, the proposed DQFormer does not require setting a fixed number of lane queries, so it is suitable for various scenes. In addition, we further propose a Line Voting Module (LVM) which collects votes from other lanes to enhance lane features, to determine lane instance positions. Extensive experiments demonstrate that DQFormer outperforms several state-of-the-art methods on two popular lane detection benchmarks (i.e., CULane and TuSimple).
Shuyuan Lin, Runqing Jiang, Yang Lu 0009, Hanzi Wang
ICASSP4
2023 Label-Noise Learning with Intrinsically Long-Tailed Data
abstract
Label noise is one of the key factors that lead to the poor generalization of deep learning models. Existing label-noise learning methods usually assume that the ground-truth classes of the training data are balanced. However, the real-world data is often imbalanced, leading to the inconsistency between observed and intrinsic class distribution with label noises. In this case, it is hard to distinguish clean samples from noisy samples on the intrinsic tail classes with the unknown intrinsic class distribution. In this paper, we propose a learning frame-work for label-noise learning with intrinsically long-tailed data. Specifically, we propose two-stage bi-dimensional sample selection (TABASCO) to better separate clean samples from noisy samples, especially for the tail classes. TABASCO consists of two new separation metrics that complement each other to compensate for the limitation of using a single metric in sample separation. Extensive experiments on benchmarks demonstrate the effectiveness of our method. Our code is available at https://github.com/Wakings/TABASCO.
Yang Lu 0009, Yiliang Zhang, Bo Han 0003, Yiu-Ming Cheung, Hanzi Wang
ICCV1
2023 Semantic Learning Network for Controllable Video Captioning
abstract
Video captioning is fundamental for visual understanding, which aims at describing the content of a video using natural language. Previous works make great efforts to visual representation learning by comparing the generated sentences and the ground truth in a supervised way. However, they neglect to explore linguistic semantics adequately due to the insufficient learning of visual words. In this paper, we propose the Semantic Learning Network (SLN) that explicitly learns specific semantics of all the visual words by aggregating static and dynamic features. Besides, to achieve controllable video captioning and alleviate the problem that one identical video is mapped to multiple annotations during training, we propose the predicate-based feature selection approach to convert the video to different textual features under the guidance of the predicate in the captions. Experimental results on the MSVD and MSR-VTT datasets show that our SLN outperforms recent state-of-the-art methods.
Kaixuan Chen 0006, Qianji Di, Yang Lu 0009, Hanzi Wang
ICIP3
2023 Long-Tailed Federated Learning Via Aggregated Meta Mapping
abstract
One major problem concerned in federated learning is data non-IIDness. Existing federated learning methods to deal with non-IID data generally assume that the data is globally balanced. However, real-world multi-class data tends to exhibit long-tail distribution. Therefore, we propose a new federated learning method called Federated Aggregated Meta Mapping (FedAMM) to address the joint problem of non-IID and global long-tailed data in a federated learning scenario. FedAMM assigns different weights to the local training samples by trainable loss-weight mapping in a meta-learning manner. To deal with data non-IIDness and global long-tail, the meta loss-weight mappings are aggregated on the server to acquire global long-tail distribution knowledge implicitly. We further propose an asynchronous meta updating mechanism to reduce the communication cost for meta-learning training. Experiments show that FedAMM outperforms the state-of-the-art federated learning methods.
Pinxin Qian, Yang Lu 0009, Hanzi Wang
ICIP2
2023 DeCAB: Debiased Semi-supervised Learning for Imbalanced Open-Set Data
Xiaolin Huang, Mengke Li 0001, Yang Lu 0009, Hanzi Wang
PRCV (9)3
2023 An Effective Visible-Infrared Person Re-identification Network Based on Second-Order Attention and Mixed Intermediate Modality
Haiyun Tao, Yang Lu 0009, Hanzi Wang
PRCV (9)3
2023 Cascaded-Scoring Tracklet Matching for Multi-object Tracking
Yixian Xie, Hanzi Wang, Yang Lu 0009
PRCV (10)3
2023 Unsupervised Concept Drift Detection via Imbalanced Cluster Discriminator Learning
Mingjie Zhao 0003, Yiqun Zhang 0006, Yuzhu Ji, Yang Lu 0009
PRCV (3)4
2023 TCNet: A Novel Triple-Cooperative Network for Video Object Detection
abstract
Video object detection aims at accurately localizing the objects in videos and correctly recognizing their categories. Off-the-shelf video object detection methods have made some progress in recent years but they still suffer from the problems of inaccurate object localization, incorrect object recognition or insufficient relation learning, resulting in limited detection performance. In this paper, we propose a novel triple-cooperative network (TCNet) for high-performance video object detection, with three substantial improvements to ameliorate the problems of existing methods. First, we develop a context-aware proposal refinement module to generate high-quality proposals, enabling our TCNet to achieve more accurate object localization. Second, we present a similarity-aware semantic distillation module that innovatively leverages the semantic knowledge of class labels as additional supervisory signals to enhance the object recognition ability of our TCNet. Third, we design a structure-aware relation learning module to effectively model the structural relations between features with an adaptive-pruning residual graph convolutional network, making our TCNet perform more effective feature aggregation. We conduct extensive experiments on the challenging ImageNet VID dataset and the experimental results demonstrate that our TCNet outperforms current state-of-the-art methods. More remarkably, our TCNet achieves 85.2% mAP and 86.3% mAP with ResNet-101 and ResNeXt-101, respectively.
Qiang Qi, Tianxiang Hou, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Drop Loss for Person Attribute Recognition With Imbalanced Noisy-Labeled Samples
abstract
Person attribute recognition (PAR) aims to simultaneously predict multiple attributes of a person. Existing deep learning-based PAR methods have achieved impressive performance. Unfortunately, these methods usually ignore the fact that different attributes have an imbalance in the number of noisy-labeled samples in the PAR training datasets, thus leading to suboptimal performance. To address the above problem of imbalanced noisy-labeled samples, we propose a novel and effective loss called drop loss for PAR. In the drop loss, the attributes are treated differently in an easy-to-hard way. In particular, the noisy-labeled candidates, which are identified according to their gradient norms, are dropped with a higher drop rate for the harder attribute. Such a manner adaptively alleviates the adverse effect of imbalanced noisy-labeled samples on model learning. To illustrate the effectiveness of the proposed loss, we train a simple ResNet-50 model based on the drop loss and term it DropNet. Experimental results on two representative PAR tasks (including facial attribute recognition and pedestrian attribute recognition) demonstrate that the proposed DropNet achieves comparable or better performance in terms of both balanced accuracy and classification accuracy over several state-of-the-art PAR methods.
Yan Yan 0001, Youze Xu, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang, Wentao Zhu 0002
IEEE Trans. Cybern.4
2023 DGRNet: A Dual-Level Graph Relation Network for Video Object Detection
abstract
Video object detection is a fundamental and important task in computer vision. One mainstay solution for this task is to aggregate features from different frames to enhance the detection on the current frame. Off-the-shelf feature aggregation paradigms for video object detection typically rely on inferring feature-to-feature (Fea2Fea) relations. However, most existing methods are unable to stably estimate Fea2Fea relations due to the appearance deterioration caused by object occlusion, motion blur or rare poses, resulting in limited detection performance. In this paper, we study Fea2Fea relations from a new perspective, and propose a novel dual-level graph relation network (DGRNet) for high-performance video object detection. Different from previous methods, our DGRNet innovatively leverages the residual graph convolutional network to simultaneously model Fea2Fea relations at two different levels including frame level and proposal level, which facilitates performing better feature aggregation in the temporal domain. To prune unreliable edge connections in the graph, we introduce a node topology affinity measure to adaptively evolve the graph structure by mining the local topological information of pairwise nodes. To the best of our knowledge, our DGRNet is the first video object detection method that leverages dual-level graph relations to guide feature aggregation. We conduct experiments on the ImageNet VID dataset and the results demonstrate the superiority of our DGRNet against state-of-the-art methods. Especially, our DGRNet achieves 85.0% mAP and 86.2% mAP with ResNet-101 and ResNeXt-101, respectively.
Qiang Qi, Tianxiang Hou, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
IEEE Trans. Image Process.3
2022 Embedding Adaptation Network with Transformer for Few-Shot Action Recognition
Rongrong Jin, Xiao Wang 0072, Guangge Wang, Yang Lu 0009, Hai-Miao Hu, Hanzi Wang
ACML4
2022 Long-tailed Visual Recognition via Gaussian Clouded Logit Adjustment
abstract
Long-tailed data is still a big challenge for deep neural networks, even though they have achieved great success on balanced data. We observe that vanilla training on longtailed data with crossentropy loss makes the instance-rich head classes severely squeeze the spatial distribution of the tail classes, which leads to difficulty in classifying tail class samples. Furthermore, the original crossentropy loss can only propagate gradient short-lively because the gradient in softmax form rapidly approaches zero as the logit difference increases. This phenomenon is called softmax saturation. It is unfavorable for training on balanced data, but can be utilized to adjust the validity of the samples in long-tailed data, thereby solving the distorted embedding space of long-tailed problems. To this end, this paper proposes the Gaussian clouded logit adjustment by Gaussian perturbation of different class logits with varied amplitude. We define the amplitude of perturbation as cloud size and set relatively large cloud sizes to tail classes. The large cloud size can reduce the softmax saturation and thereby making tail class samples more active as well as enlarging the embedding space. To alleviate the bias in a classifier, we therefore propose the class-based effective number sampling strategy with classifier retraining. Extensive experiments on benchmark datasets validate the superior performance of the proposed method. Source code is available at https://github.com/Keke921/GCLLoss.
Mengke Li 0001, Yiu-Ming Cheung, Yang Lu 0009
CVPR3
2022 Multi-Focus Guided Semantic Aggregation for Video Object Detection
abstract
For the task of video object detection, it is useful to aggregate semantic information from supporting frames. However, existing methods only focus on the current frame during the semantic aggregation, called Single-Focus methods. They neglect semantic information among supporting frames and deteriorate overall performance. In this work, we propose a method called Multi-Focus guided Semantic Aggregation (MFSA) for video object detection. We introduce a novel Relation Propagation Module (RPM) to capture and propagate proposal-to-proposal semantic dependencies. Moreover, we propose a simple yet effective Multi-Focus strategy to leverage captured dependencies to guide feature enhancement at a batch level. Aided by this strategy, our method can greatly improve aggregation efficiency of Single-Focus methods and enhance the accuracy of a per-frame detector significantly with negligible computing overhead. We perform extensive experiments on the ImageNet VID dataset. The results show that MFSA achieves excellent performance and a superior speed-accuracy tradeoff among the competing methods.
Haihui Ye, Guangge Wang, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
ICASSP3
2022 Bounding Box Distribution Learning and Center Point Calibration for Robust Visual Tracking
abstract
Visual tracking aims at both robust target classification and accurate localization. However, the reliability of the target bounding box and classification score are not properly addressed by most existing trackers, resulting in inaccurate tracking performance. In this paper, we propose to learn bounding box distribution in training and calibrate the center point response in inference for robust online tracking. Specifically, we propose a simple yet effective bounding box distribution learning (BDL) module to model the target bounding box distribution and enhance the localization ability of our network. Furthermore, we propose a center point calibration (CPC) module to calibrate the origin classification score with the predicted localization uncertainty and generate an accurate target center point. The proposed tracking method is referred to as DLPC. The experimental results on four challenging datasets (i.e., OTB100, VOT2019, LaSOT, and TrackingNet) show that DLPC performs favorably against several state-of-the-art trackers while running in real-time at 60 fps.
Chihui Zhuang, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
ICASSP4
2022 Egnet: A Novel Edge Guided Network for Instance Segmentation
abstract
Edge information plays a significant role in instance segmentation. However, many instance segmentation methods directly perform pixel-wise classification via fully convolutional networks, which may ignore object edges. In this paper, we propose a novel Edge Guided Network (EGNet), which exploits edge information to improve the mask accuracy, for instance segmentation. Specifically, we propose an edge branch to extract edge information. Then, we use edge information as guidance and fuse it with mask features, in order to enrich the mask features. Furthermore, we propose a Spatial Attention (SA) module and add it to the backbone of our EGNet, enabling the network to focus more on foreground objects. In addition, we incorporate a Semantic Enhancement (SE) module into the edge branch, aiming to obtain additional global context information. Experimental results on the COCO 2017 dataset show the effectiveness of the proposed EGNet.
Kaiwen Du, Xiao Wang 0072, Yan Yan 0009, Yang Lu 0009, Hanzi Wang
ICIP4
2022 Guided Sampling Based Feature Aggregation for Video Object Detection
abstract
Video object detection is a challenging task due to the presence of appearance deterioration in video frames. Recently, feature aggregation based methods which aggregate context information from object proposals in different frames to improve the performance, have dominated the task. However, much invalid information may be introduced during feature aggregation since frames and proposals are usually selected at random. In this paper, we propose a guided sampling based feature aggregation network (GSFA) to perform more effective feature aggregation. Specifically, we introduce a frame-level sampling module and a proposal-level sampling module to sample informative frames and proposals from a video sequence adaptively. As a result, the proposed GSFA can effectively aggregate context information from the semantically rich frames and proposals to boost the performance. Experimental results on the ImageNet VID dataset show the proposed GSFA achieves the state-of-the-art performance of 84.8% mAP with ResNet-101 and 85.8% mAP with ResNeXt-101.
Haosheng Chen 0001, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
ICIP4
2022 SCINet: Semantic Cue Infusion Network for Lane Detection
abstract
Nowadays, lane detection plays an important role in autonomous driving. However, the task of lane detection still faces many challenges, such as external no-visual-clue and internal sparse supervisory signals. In this work, we propose a novel Semantic Cue Infusion Network (SCINet) that uses semantic cues to aid lane detection. Specifically, in order to overcome the external no-visual-clue condition, we introduce a strategy that utilizes semantic cues as additional supervisory signals, which facilitate SCINet to collect region-aware features in the shared layer. We also design a hypernetwork with the additional signals used as a critical part to generate dynamic weights for downstream output heads. Furthermore, we design two Slice Attention Modules (SAMs) based on the interdependencies between slices to improve the robustness of SCINet in distinguishing features between lanes and background. Experiments on two popular lane detection benchmarks (i.e., TuSimple and CULane) show that SCINet significantly outperforms several state-of-the-art methods.
Shuyuan Lin, Yang Lu 0009, Hanzi Wang
ICIP4
2022 Dual Selection Network for Video Object Detection
abstract
Some off-the-shelf video object detection methods usually enhance the degraded proposal features of target frames by aggregating the proposal features from support frames. However, the proposals generated by region proposal network may not be accurate, resulting in inaccurate proposal features and limited performance. To mitigate this, we propose a novel dual selection network (DSNet) for video object detection, which contains two successive stages: selecting proposals that fit objects more closely, and selecting proposal features that are more conducive to feature aggregation. Correspondingly, the proposal selection module (PSM) aims to select better proposals by exploiting their boundary information, and the selective aggregation module (SAM) aims to select better proposal features for aggregation. Consequently, DSNet can generate more robust proposal features through the novel dual selection mechanism implemented by PSM and SAM. Extensive experiments show that our DSNet obtains 83.7% mAP and achieves superior performance over several state-of-the-art methods.
Tianxiang Hou, Qiang Qi, Yang Lu 0009, Kaiwen Du, Hanzi Wang
ICME3
2022 FEDIC: Federated Learning on Non-IID and Long-Tailed Data via Calibrated Distillation
abstract
Federated learning provides a privacy guarantee for generating good deep learning models on distributed clients with different kinds of data. Nevertheless, dealing with non-IID data is one of the most challenging problems for federated learning. Researchers have proposed a variety of methods to eliminate the negative influence of non-IIDness. However, they only focus on the non-IID data provided that the universal class distribution is balanced. In many real-world applications, the universal class distribution is long-tailed, which causes the model seriously biased. Therefore, this paper studies the joint problem of non-IID and long-tailed data in federated learning and proposes a corresponding solution called Federated Ensemble Distillation with Imbalance Calibration (FEDIC). To deal with non-IID data, FEDIC uses model ensemble to take advantage of the diversity of models trained on non-IID data. Then, a new distillation method with logit adjustment and calibration gating network is proposed to solve the long-tail problem effectively. We evaluate FEDIC on CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT with a highly non-IID experimental setting, in comparison with the state-of-the-art methods of federated learning and long-tail learning. Our code is available at https://github.com/shangxinyi/FEDIC.
Xinyi Shang, Yang Lu 0009, Yiu-Ming Cheung, Hanzi Wang
ICME2
2022 Federated Learning on Heterogeneous and Long-Tailed Data via Classifier Re-Training with Federated Features
abstract
Federated learning (FL) provides a privacy-preserving solution for distributed machine learning tasks. One challenging problem that severely damages the performance of FL models is the co-occurrence of data heterogeneity and long-tail distribution, which frequently appears in real FL applications. In this paper, we reveal an intriguing fact that the biased classifier is the primary factor leading to the poor performance of the global model. Motivated by the above finding, we propose a novel and privacy-preserving FL method for heterogeneous and long-tailed data via Classifier Re-training with Federated Features (CReFF). The classifier re-trained on federated features can produce comparable performance as the one re-trained on real data in a privacy-preserving manner without information leakage of local data or class distribution. Experiments on several benchmark datasets show that the proposed CReFF is an effective solution to obtain a promising FL model under heterogeneous and long-tailed data. Comparative results with the state-of-the-art FL methods also validate the superiority of CReFF. Our code is available at https://github.com/shangxinyi/CReFF-FL.
Xinyi Shang, Yang Lu 0009, Gang Huang 0004, Hanzi Wang
IJCAI2
2022 TRL: Transformer based refinement learning for hybrid-supervised semantic segmentation
Pengfei Fang, Yan Yan 0001, Yang Lu 0009, Hanzi Wang
Pattern Recognit. Lett.4
2022 Triplet Relationship Guided Sampling Consensus for Robust Model Estimation
abstract
RANSAC (RANdom SAmple Consensus) is a widely used robust estimator for estimating a geometric model from feature matches in an image pair. Unfortunately, it becomes less effective when initial input feature matches (i.e., input data) are corrupted by a large number of outliers. In this paper, we propose a new robust estimator (called TRESAC) for model estimation, where data subsets are sampled with the guidance of the triplet relationships, which involve high relevance and local geometric consistency. Each triplet consists of three data, whose relationships satisfy spatial consistency constraints. Therefore, the triplet relationships can be used to effectively initialize and refine the sampling process. With the advantage of the triplet relationships, TRESAC significantly alleviates the influence of outliers and also improves the computational efficiency of model estimation. Experimental results on four challenging datasets show that TRESAC can achieve superior performance on both estimation accuracy and computational efficiency against several other state-of-the-art methods.
Hanlin Guo, Yang Lu 0009, Guobao Xiao, Shuyuan Lin, Hanzi Wang
IEEE Signal Process. Lett.2
2021 Towards a Unified Middle Modality Learning for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to search identities of pedestrians across different spectra. In this task, one of the major challenges is the modality discrepancy between the visible (VIS) and infrared (IR) images. Some state-of-the-art methods try to design complex networks or generative methods to mitigate the modality discrepancy while ignoring the highly non-linear relationship between the two modalities of VIS and IR. In this paper, we propose a non-linear middle modality generator (MMG), which helps to reduce the modality discrepancy. Our MMG can effectively project VIS and IR images into a unified middle modality image (UMMI) space to generate middle-modality (M-modality) images. The generated M-modality images and the original images are fed into the backbone network to reduce the modality discrepancy.Furthermore, in order to pull together the two types of M-modality images generated from the VIS and IR images in the UMMI space, we propose a distribution consistency loss (DCL) to make the modality distribution of the generated M-modalities images as consistent as possible. Finally, we propose a middle modality network (MMN) to further enhance the discrimination and richness of features in an explicit manner. Extensive experiments have been conducted to validate the superiority of MMN for VI-ReID over some state-of-the-art methods on two challenging datasets. The gain of MMN is more than 11.1% and 8.4% in terms of Rank-1 and mAP, respectively, even compared with the latest state-of-the-art methods on the SYSU-MM01 dataset.
Yan Yan 0001, Yang Lu 0009, Hanzi Wang
ACM Multimedia3
2021 Small-Vote Sample Selection for Label-Noise Learning
Youze Xu, Yan Yan 0001, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang
ECML/PKDD (3)4
2021 Recurrent Context Aggregation Network for Single Image Dehazing
abstract
Existing learning-based dehazing methods are prone to cause excessive dehazing and failure to dense haze, mainly because that the global features of hazy images are not fully utilized, while the local features of hazy images are not enough discriminative. In this letter, we propose a Recurrent Context Aggregation Network (RCAN) to effectively dehaze images and restore color fidelity. In RCAN, an efficient and generic module, called Context Aggression Block (CAB), is designed to improve the feature representation by taking advantage of both global and local features, which are complementary for robust dehazing because that local features can capture different levels of haze, and global features can focus on textures and object edges of a whole image. In addition, RCAN adopts a deep recurrent mechanism to improve the dehazing performance without introducing additional network parameters. Extensive experimental results on both synthetic and real-world datasets show that the proposed RCAN performs better than other state-of-the-art dehazing methods.
Runqing Chen, Yang Lu 0009, Yan Yan 0001, Hanzi Wang
IEEE Signal Process. Lett.3
2021 Self-Adaptive Multiprototype-Based Competitive Learning Approach: A k-Means-Type Algorithm for Imbalanced Data Clustering
abstract
Class imbalance problem has been extensively studied in the recent years, but imbalanced data clustering in unsupervised environment, that is, the number of samples among clusters is imbalanced, has yet to be well studied. This paper, therefore, studies the imbalanced data clustering problem within the framework of k -means-type competitive learning. We introduce a new method called self-adaptive multiprototype-based competitive learning (SMCL) for imbalanced clusters. It uses multiple subclusters to represent each cluster with an automatic adjustment of the number of subclusters. Then, the subclusters are merged into the final clusters based on a novel separation measure. We also propose a new internal clustering validation measure to determine the number of final clusters during the merging process for imbalanced clusters. The advantages of SMCL are threefold: 1) it inherits the advantages of competitive learning and meanwhile is applicable to the imbalanced data clustering; 2) the self-adaptive multiprototype mechanism uses a proper number of subclusters to represent each cluster with any arbitrary shape; and 3) it automatically determines the number of clusters for imbalanced clusters. SMCL is compared with the existing counterparts for imbalanced clustering on the synthetic and real datasets. The experimental results show the efficacy of SMCL for imbalanced clusters.
Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang
IEEE Trans. Cybern.1
2020 Global and local feature alignment for video object detection
abstract
Extending image-based object detectors into video domain suffers from immense inadaptability due to the deteriorated frames caused by motion blur, partial occlusion or strange poses. Therefore, the generated features of deteriorated frames encounter the poor quality of misalignment, which degrades the overall performance of video object detectors. How to capture valuable information locally or globally is of importance to feature alignment but remains quite challenging. In this paper, we propose a Global and Local Feature Alignment (abbreviated as GLFA) module for video object detection, which can distill both global and local information to excavate the deep relationship between features for feature alignment. Specifically, GLFA can model the spatial-temporal dependencies over frames based on propagating global information and capture the interactive correspondences within the same frame based on aggregating valuable local information. Moreover, we further introduce a Self-Adaptive Calibration (SAC) module to strengthen the semantic representation of features and distill valuable local information in a dual local-alignment manner. Experimental results on the ImageNet VID dataset show that the proposed method achieves high performance as well as a good trade-off between real-time speed and competitive accuracy.
Haihui Ye, Qiang Qi, Yang Lu 0009, Hanzi Wang
MMAsia4
2020 Adaptive Chunk-Based Dynamic Weighted Majority for Imbalanced Data Streams With Concept Drift
abstract
One of the most challenging problems in the field of online learning is concept drift, which deeply influences the classification stability of streaming data. If the data stream is imbalanced, it is even more difficult to detect concept drifts and make an online learner adapt to them. Ensemble algorithms have been found effective for the classification of streaming data with concept drift, whereby an individual classifier is built for each incoming data chunk and its associated weight is adjusted to manage the drift. However, it is difficult to adjust the weights to achieve a balance between the stability and adaptability of the ensemble classifiers. In addition, when the data stream is imbalanced, the use of a size-fixed chunk to build a single classifier can create further problems; the data chunk may contain too few or even no minority class samples (i.e., only majority class samples). A classifier built on such a chunk is unstable in the ensemble. In this article, we propose a chunk-based incremental learning method called adaptive chunk-based dynamic weighted majority (ACDWM) to deal with imbalanced streaming data containing concept drift. ACDWM utilizes an ensemble framework by dynamically weighting the individual classifiers according to their classification performance on the current data chunk. The chunk size is adaptively selected by statistical hypothesis tests to access whether the classifier built on the current data chunk is sufficiently stable. ACDWM has four advantages compared with the existing methods as follows: 1) it can maintain stability when processing nondrifted streams and rapidly adapt to the new concept; 2) it is entirely incremental, i.e., no previous data need to be stored; 3) it stores a limited number of classifiers to ensure high efficiency; and 4) it adaptively selects the chunk size in the concept drift environment. Experiments on both synthetic and real data sets containing concept drift show that ACDWM outperforms both state-of-the-art chunk-based and online methods.
Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang
IEEE Trans. Neural Networks Learn. Syst.1
2020 Bayes Imbalance Impact Index: A Measure of Class Imbalanced Data Set for Classification Problem
abstract
Recent studies of imbalanced data classification have shown that the imbalance ratio (IR) is not the only cause of performance loss in a classifier, as other data factors, such as small disjuncts, noise, and overlapping, can also make the problem difficult. The relationship between the IR and other data factors has been demonstrated, but to the best of our knowledge, there is no measurement of the extent to which class imbalance influences the classification performance of imbalanced data. In addition, it is also unknown which data factor serves as the main barrier for classification in a data set. In this article, we focus on the Bayes optimal classifier and examine the influence of class imbalance from a theoretical perspective. We propose an instance measure called the Individual Bayes Imbalance Impact Index (IBI3) and a data measure called the Bayes Imbalance Impact Index (BI3). IBI3and BI3reflect the extent of influence using only the imbalance factor, in terms of each minority class sample and the whole data set, respectively. Therefore, IBI3can be used as an instance complexity measure of imbalance and BI3as a criterion to demonstrate the degree to which imbalance deteriorates the classification of a data set. We can, therefore, use BI3to access whether it is worth using imbalance recovery methods, such as sampling or cost-sensitive methods, to recover the performance loss of a classifier. The experiments show that IBI3is highly consistent with the increase of the prediction score obtained by the imbalance recovery methods and that BI3is highly consistent with the improvement in the F1 score obtained by the imbalance recovery methods on both synthetic and real benchmark data sets.
Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang
IEEE Trans. Neural Networks Learn. Syst.1
2018 k-Times Markov Sampling for SVMC
abstract
Support vector machine (SVM) is one of the most widely used learning algorithms for classification problems. Although SVM has good performance in practical applications, it has high algorithmic complexity as the size of training samples is large. In this paper, we introduce SVM classification (SVMC) algorithm based on -times Markov sampling and present the numerical studies on the learning performance of SVMC with -times Markov sampling for benchmark data sets. The experimental results show that the SVMC algorithm with -times Markov sampling not only have smaller misclassification rates, less time of sampling and training, but also the obtained classifier is more sparse compared with the classical SVMC and the previously known SVMC algorithm based on Markov sampling. We also give some discussions on the performance of SVMC with -times Markov sampling for the case of unbalanced training samples and large-scale training samples.
Bin Zou 0002, Chen Xu 0007, Yang Lu 0009, Yuan Yan Tang, Jie Xu 0006, Xinge You
IEEE Trans. Neural Networks Learn. Syst.3
2017 Dynamic Weighted Majority for Incremental Learning of Imbalanced Data Streams with Concept Drift
abstract
Concept drifts occurring in data streams will jeopardize the accuracy and stability of the online learning process. If the data stream is imbalanced, it will be even more challenging to detect and cure the concept drift. In the literature, these two problems have been intensively addressed separately, but have yet to be well studied when they occur together. In this paper, we propose a chunk-based incremental learning method called Dynamic Weighted Majority for Imbalance Learning (DWMIL) to deal with the data streams with concept drift and class imbalance problem. DWMIL utilizes an ensemble framework by dynamically weighting the base classifiers according to their performance on the current data chunk. Compared with the existing methods, its merits are four-fold: (1) it can keep stable for non-drifted streams and quickly adapt to the new concept; (2) it is totally incremental, i.e. no previous data needs to be stored; (3) it keeps a limited number of classifiers to ensure high efficiency; and (4) it is simple and needs only one thresholding parameter. Experiments on both synthetic and real data sets with concept drift show that DWMIL performs better than the state-of-the-art competitors, with less computational cost.
Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang
IJCAI1
2016 Hybrid Sampling with Bagging for Class Imbalance Learning
Yang Lu 0009, Yiu-Ming Cheung, Yuan Yan Tang
PAKDD (1)1
2015 A Fractal Dimension and Wavelet Transform Based Method for Protein Sequence Similarity Analysis
abstract
One of the key tasks related to proteins is the similarity comparison of protein sequences in the area of bioinformatics and molecular biology, which helps the prediction and classification of protein structure and function. It is a significant and open issue to find similar proteins from a large scale of protein database efficiently. This paper presents a new distance based protein similarity analysis using a new encoding method of protein sequence which is based on fractal dimension. The protein sequences are first represented into the 1-dimensional feature vectors by their biochemical quantities. A series of Hybrid method involving discrete Wavelet transform, Fractal dimension calculation (HWF) with sliding window are then applied to form the feature vector. At last, through the similarity calculation, we can obtain the distance matrix, by which, the phylogenic tree can be constructed. We apply this approach by analyzing the ND5 (NADH dehydrogenase subunit 5) protein cluster data set. The experimental results show that the proposed model is more accurate than the existing ones such as Su's model, Zhang's model, Yao's model and MEGA software, and it is consistent with some known biological facts.
Yuan Yan Tang, Yang Lu 0009, Huiwu Luo
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 The Generalization Ability of SVM Classification Based on Markov Sampling
abstract
UNLABELLED: The previously known works studying the generalization ability of support vector machine classification (SVMC) algorithm are usually based on the assumption of independent and identically distributed samples. In this paper, we go far beyond this classical framework by studying the generalization ability of SVMC based on uniformly ergodic Markov chain (u.e.M.c.) samples. We analyze the excess misclassification error of SVMC based on u.e.M.c. samples, and obtain the optimal learning rate of SVMC for u.e.M.c. SAMPLES: We also introduce a new Markov sampling algorithm for SVMC to generate u.e.M.c. samples from given dataset, and present the numerical studies on the learning performance of SVMC based on Markov sampling for benchmark datasets. The numerical studies show that the SVMC based on Markov sampling not only has better generalization ability as the number of training samples are bigger, but also the classifiers based on Markov sampling are sparsity when the size of dataset is bigger with regard to the input dimension.
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009, Baochang Zhang 0001
IEEE Trans. Cybern.6
2015 Hyperspectral Image Classification Based on Three-Dimensional Scattering Wavelet Transform
abstract
Recent research has shown that utilizing the spectral-spatial information can improve the performance of hyperspectral image (HSI) classification. Since HSI is a 3-D cube datum, 3-D spatial filtering becomes a simple and effective method for extracting the spectral-spatial information. In this paper, we propose a 3-D scattering wavelet transform, which filters the HSI cube data with a cascade of wavelet decompositions, complex modulus, and local weighted averaging. The scattering feature can adequately capture the spectral-spatial information for classification. In the classification step, a support vector machine based on Gaussian kernel is used as a classifier due to its capability to deal with high-dimensional data. Our method is fully evaluated on four classic HSIs, i.e., Indian Pines, Pavia University, Botswana, and Kennedy Space Center. The classification results show that our method achieves as high as 94.46%, 99.30%, 97.57%, and 95.20% accuracies, respectively, when only 5% of the total samples per class is labeled.
Yuan Yan Tang, Yang Lu 0009
IEEE Trans. Geosci. Remote. Sens.2
2015 The Generalization Ability of Online SVM Classification Based on Markov Sampling
abstract
In this paper, we consider online support vector machine (SVM) classification learning algorithms with uniformly ergodic Markov chain (u.e.M.c.) samples. We establish the bound on the misclassification error of an online SVM classification algorithm with u.e.M.c. samples based on reproducing kernel Hilbert spaces and obtain a satisfactory convergence rate. We also introduce a novel online SVM classification algorithm based on Markov sampling, and present the numerical studies on the learning ability of online SVM classification based on Markov sampling for benchmark repository. The numerical studies show that the learning performance of the online SVM classification algorithm based on Markov sampling is better than that of classical online SVM classification based on random sampling as the size of training samples is larger.
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009
IEEE Trans. Neural Networks Learn. Syst.6
2014 A novel method for protein structure retrieval using tableau representation and sparse coding
abstract
Protein retrieval is a difficult task and has become a hot issue recently due to the complex structure and large data size of proteins. This work helps biologists investigate the link between structure and function of a protein in a deeper level and can be used in lots of biomedical applications. The retrieval system gives scores to all proteins in the database, e.g. SCOP or PDB, by given a query protein to compare with them. In this paper, we propose a novel algorithm based on sparse coding to retrieve proteins in the database using tableau representation. Both unsupervised and supervised methods are studied in the proposed algorithm where the sparse coefficient is regarded as similarity measurement. Experiments are conducted on ASTRAL 1.73 95% database and show that the proposed algorithms can improve the original feature extraction method which only uses cosine similarity.
Yang Lu 0009, Yulong Wang 0002, Huiwu Luo, Yuan Yan Tang
SMC1
2014 Spectral-spatial hyperspectral image destriping using low-rank representation and Huber-Markov random fields
abstract
This paper presents a novel spectral-spatial destriping method for hyperspectral images. The ubiquitous striping noise in hyperspectral images might degrade the quality of the imagery and bring difficulties in hyperspectral data processing. Although numerous methods have been proposed for striping noise reduction recently, most of them fail to consider the spectral correlation and spatial information of the hyperspectral images simultaneously. In order to remedy this drawback, the proposed method integrates the spectral and spatial information to remove the striping noise in the hyperspectral images. To this end, firstly, the low-rank representation (LRR) is used to take advantage of the spectral information. Then, the spatial information is included using a Huber-Markov random field (MRF) prior model, which is convex and can well preserve the edge and texture information while removing the noise. The experimental results on simulated and real hyperspectral data sets demonstrate the effectiveness of the proposed method.
Yulong Wang 0002, Yuan Yan Tang, Huiwu Luo, Yang Lu 0009
SMC6
2014 Protein sequence analysis based on fractal-wavelet scheme
abstract
It is a significant issue to find similar proteins from a large scale of protein database efficiently. This paper presents a new algorithm of protein sequence which is based on fractal dimension and wavelet transform. A hybrid method consisting fractal dimension calculation, discrete wavelet transform and sliding window are applied to generate a new encoding feature. Through the computation between the feature vectors, we can obtain the distance matrix and the phylogenic tree can be constructed.We apply this approach by analyzing the ND5 (NADH dehydrogenase subunit 5) protein dataset. The experimental results show that the proposed model is more accurate than the existing ones such as Su's model, Zhang's model and Yao's model, and it is consistent with the result generated from MEGA software and some known facts.
Yuan Yan Tang, Yang Lu 0009, Huiwu Luo, Yulong Wang 0002
SMC3
2014 Feature extraction based on kernel sparse representation for hyperspectral image classification
abstract
Feature extraction is a promising technique for hyperspectral image classification. Recent research has shown that the criterion of sparse representation classification (SRC) can help to design a feature extraction method. This method is called the SRC steered discriminative projection (SRCDP). Motivated by the fact that kernel trick can exploit the nonlinear case of features, this paper generalizes SRCDP to its kernel case named KSRCDP. Extensive experiments show that KSRCDP can obtain excellent classification performance on two classic hyperspectral images.
Huiwu Luo, Yang Lu 0009, Yulong Wang 0002, Yuan Yan Tang
SMC4
2014 Generalization performance of Gaussian kernels SVMC based on Markov sampling
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009
Neural Networks6
2014 The Generalization Performance of Regularized Regression Algorithms Based on Markov Sampling
abstract
This paper considers the generalization ability of two regularized regression algorithms [least square regularized regression (LSRR) and support vector machine regression (SVMR)] based on non-independent and identically distributed (non-i.i.d.) samples. Different from the previously known works for non-i.i.d. samples, in this paper, we research the generalization bounds of two regularized regression algorithms based on uniformly ergodic Markov chain (u.e.M.c.) samples. Inspired by the idea from Markov chain Monto Carlo (MCMC) methods, we also introduce a new Markov sampling algorithm for regression to generate u.e.M.c. samples from a given dataset, and then, we present the numerical studies on the learning performance of LSRR and SVMR based on Markov sampling, respectively. The experimental results show that LSRR and SVMR based on Markov sampling can present obviously smaller mean square errors and smaller variances compared to random sampling.
Bin Zou 0002, Yuan Yan Tang, Zongben Xu, Luoqing Li, Jie Xu 0006, Yang Lu 0009
IEEE Trans. Cybern.6