Lei Wang 0095

dblp:w/LeiWang95 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-3860-5139ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Learning paradigms · 51% Deep learning architectures and training · 21% Transfer learning and domain adaptation · 13%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
multi-label classification
3.442025
SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image Classification · IEEE Trans. Multim. 2025
Towards Space and Semantics: Object-Purified Representation Learning for Multi-Label Image Classification · ACM Multimedia 2025
Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning · CVPR 2025
Machine learning › Deep learning architectures and training
data augmentation
0.912025
SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image Classification · IEEE Trans. Multim. 2025
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
visual prompt tuning
0.912025
Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning · CVPR 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.312025
Towards Space and Semantics: Object-Purified Representation Learning for Multi-Label Image Classification · ACM Multimedia 2025
Machine learning › Deep learning architectures and training › regularization
consistency training
0.312025
SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image Classification · IEEE Trans. Multim. 2025
Natural language and speech › Language models and text generation
prompt tuning
0.212024
Text-Region Matching for Multi-Label Image Recognition with Missing Labels · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

vision transformer · 0.9representation purification · 0.9prompt tuning · 0.9mixup · 0.9mixture of experts · 0.9cutmix · 0.9consistency learning · 0.9attention mechanism · 0.9multimodal contrastive learning · 0.8category prototype · 0.8
YearPublicationVenuePosition
2025 Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning
abstract
Modeling label correlations has always played a pivotal role in multi-label image classification (MLC), attracting significant attention from researchers. However, recent studies have overemphasized co-occurrence relationships among labels, which can lead to overfitting risk on this overemphasis, resulting in suboptimal models. To tackle this problem, we advocate for balancing correlative and discriminative relationships among labels to mitigate the risk of overfitting and enhance model performance. To this end, we propose the Multi-Label Visual Prompt Tuning framework, a novel and parameter-efficient method that groups classes into multiple class subsets according to label co-occurrence and mutual exclusivity relationships, and then models them respectively to balance the two relationships. In this work, since each group contains multiple classes, multiple prompt tokens are adopted within Vision Transformer (ViT) to capture the correlation or discriminative label relationship within each group, and effectively learn correlation or discriminative representations for class subsets. On the other hand, each group contains multiple group-aware visual representations that may correspond to multiple classes, and the mixture of experts (MoE) model can cleverly assign them from the group-aware to the label-aware, adaptively obtaining label-aware representation, which is more conducive to classification. Experiments on multiple benchmark datasets show that our proposed approach achieves competitive results and outperforms SOTA methods on multiple pre-trained models.
Leilei Ma 0002, Ming-Kun Xie, Lei Wang 0095, Dengdi Sun, Haifeng Zhao 0001
CVPR4
2025 Towards Space and Semantics: Object-Purified Representation Learning for Multi-Label Image Classification
abstract
Multi-label image classification requires simultaneously recognizing multiple objects with complex interdependencies. While existing attention-based methods are prominent, their performance is hampered by two forms of representation entanglement: 1) Spatial entanglement, where contextual interference from backgrounds and co-occurring objects confuses specific object representations; 2) Semantic entanglement, where models overfit label co-occurrence priors, thereby impairing a genuine semantic understanding of the image. To address these challenges, we propose an Object-Purified Representation Learning framework. Concretely, for spatial entanglement, we propose the Spatial-wise Representation Purification Module that employs Spatial-Purified Attention to eliminate object-irrelevant feature activations for contextual interference reduction, combined with Spatial-Aware Supervision to enhance object perception capability. For semantic entanglement, we develop the Semantic-wise Association Purification Module that synergistically integrates our proposed average message with the original co-occurrence-based message. This design effectively models co-occurrence relationships while preventing their overemphasis. Furthermore, we design the Bidirectional Representation Refinement Module to efficiently enhance representations, further boosting classification performance. Extensive experiments on multiple benchmark datasets with different configurations demonstrate that our proposed method achieves state-of-the-art performance.
Haifeng Zhao 0001, Leilei Ma 0002, Lei Wang 0095, Dengdi Sun
ACM Multimedia5
2025 SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image Classification
abstract
Recently, Mix-style data augmentation methods (e.g., Mixup and CutMix) have shown promising performance in various visual tasks. However, these methods are primarily designed for single-label images, ignoring the considerable discrepancies between single- and multi-label images,i.e., a multi-label image involves multiple co-occurred categories and fickle object scales. On the other hand, previous multi-label image classification (MLIC) methods tend to design elaborate models, bringing expensive computation. In this article, we introduce a simple but effective augmentation strategy for multi-label image classification, namely SpliceMix. The “splice” in our method is two-fold:1)Each mixed image is a splice of several downsampled images in the form of a grid, where the semantics of images attending to mixing are blended without object deficiencies for alleviating co-occurred bias;2)We splice mixed images and the original mini-batch to form a new SpliceMixed mini-batch, which allows an image with different scales to contribute to training together. Furthermore, such splice in our SpliceMixed mini-batch enables interactions between mixed images and original regular images. We also provide a simple and non-parametric extension based on consistency learning (SpliceMix-CL) to show the potential of extending our SpliceMix. Extensive experiments on various tasks demonstrate that only using SpliceMix with a baseline model (e.g., ResNet) achieves better performance than state-of-the-art methods. Moreover, the generalizability of our SpliceMix is further validated by the improvements in current MLIC methods when married with our SpliceMix.
Lei Wang 0095, Yibing Zhan, Leilei Ma 0002, Dapeng Tao, Liang Ding 0006, Chen Gong 0002
IEEE Trans. Multim.1
2024 Text-Region Matching for Multi-Label Image Recognition with Missing Labels
abstract
Recently, large-scale visual language pre-trained (VLP) models have demonstrated impressive performance across various downstream tasks. Motivated by these advancements, pioneering efforts have emerged in multi-label image recognition with missing labels, leveraging VLP prompt-tuning technology. However, they usually cannot match text and vision features well, due to complicated semantics gaps and missing labels in a multi-label image. To tackle this challenge, we propose Text-Region Matching for optimizing Multi-Label prompt tuning, namely TRM-ML, a novel method for enhancing meaningful cross-modal matching. Compared to existing methods, we advocate exploring the information of category-aware regions rather than the entire image or pixels, which contributes to bridging the semantic gap between textual and visual representations in a one-to-one matching manner. Concurrently, we further introduce multimodal contrastive learning to narrow the semantic gap between textual and visual modalities and establish intra-class and inter-class relationships. Additionally, to deal with missing labels, we propose a multimodal category prototype that leverages intra- and inter-category semantic relationships to estimate unknown labels, facilitating pseudo-label generation. Extensive experiments on the MS-COCO, PASCAL VOC, Visual Genome, NUS-WIDE, and CUB-200-211 benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art methods by a significant margin. Our code is available here.
Leilei Ma 0002, Hongxing Xie, Lei Wang 0095, Yanping Fu, Dengdi Sun, Haifeng Zhao 0001
ACM Multimedia3
2024 Attention-Aware Sobel Graph Convolutional Network for Remote Sensing Image Change Detection
abstract
In the study of remote sensing images, the problem of change detection (CD) is crucial. Convolutional neural networks (CNNs) are well-liked feature extraction structures that are frequently used in CD. On the other hand, graph convolutional networks (GCNs) are effective in building contextual structure information. Compared with CNN, GCN can make full use of the graph structure information to capture the changing features between different areas in the graph by learning the connections and interactions between nodes. In contrast, traditional pixel-based CNNs may have difficulty modeling semantic relationships and temporal variations among features and are susceptible to noise interference. So in this article, we extract optimization information using a GCN structure. Due to the particularity of remote sensing images, edge information is often ignored, which is useful in the field of CD. In this article, we propose an attention-aware Sobel GCN (ASGCN) for remote sensing image CD. First, we use a Siamese CNN to extract primary multilevel features. Then, a dual-branch attention module (DAM) including coordinate attention and multiscale local attention module (MLAM) is proposed to focus on informative pixels, we use Sobel operator to construct graph, and the graph convolutional module can expand receptive field and extract edge information. Attention fusion module (AFM) is adopted at decoder to perform effective feature fusion. Extensive comparative experiments on three CD datasets, LEVIR-CD, WHU-CD, and DSIFN-CD, verify the effectiveness of the proposed ASGCN.
Lei Wang 0095, Zhi-Hui You, Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Semantic-Aware Dual Contrastive Learning for Multi-Label Image Classification
abstract
Extracting image semantics effectively and assigning corresponding labels to multiple objects or attributes for natural images is challenging due to the complex scene contents and confusing label dependencies. Recent works have focused on modeling label relationships with graph and understanding object regions using class activation maps (CAM). However, these methods ignore the complex intra- and inter-category relationships among specific semantic features, and CAM is prone to generate noisy information. To this end, we propose a novel semantic-aware dual contrastive learning framework that incorporates sample-to-sample contrastive learning (SSCL) as well as prototype-to-sample contrastive learning (PSCL). Specifically, we leverage semantic-aware representation learning to extract category-related local discriminative features and construct category prototypes. Then based on SSCL, label-level visual representations of the same category are aggregated together, and features belonging to distinct categories are separated. Meanwhile, we construct a novel PSCL module to narrow the distance between positive samples and category prototypes and push negative samples away from the corresponding category prototypes. Finally, the discriminative label-level features related to the image content are accurately captured by the joint training of the above three parts. Experiments on five challenging large-scale public datasets demonstrate that our proposed method is effective and outperforms the state-of-the-art methods. Code and supplementary materials are released on https://github.com/yu-gi-oh-leilei/SADCL.
Leilei Ma 0002, Dengdi Sun, Lei Wang 0095, Haifeng Zhao 0001, Bin Luo 0001
ECAI3
2020 Weighted discriminative collaborative competitive representation for robust image classification
Jianping Gou, Lei Wang 0095, Zhang Yi 0001, Yun-Hao Yuan 0001, Weihua Ou, Qirong Mao
Neural Networks2
2019 Discriminative Group Collaborative Competitive Representation for Visual Classification
abstract
In pattern recognition, the representation-based classification (RBC) has attracted much attention recently. As a representative one of RBC, collaborative representation-based classification (CRC) and its variants have achieved promising classification performance in many visual classification tasks. However, most of the CRC methods cannot directly consider the class discrimination information of data that is very important for classification. To fully use the class discrimination information, we propose a novel discriminative group collaborative competitive representation-based classification method (DGCCR) in this paper. In the designed DGCCR model, the discriminative competitive relationships of classes, the discriminative decorrelations among classes and the weighted class-specific group constraints are simultaneously taken into account for strengthening the power of pattern discrimination. Experiments on three visual classification data sets demonstrate that the proposed DGCCR out-performs state-of-the-art RBC methods.
Jianping Gou, Lei Wang 0095, Zhang Yi 0001, Yun-Hao Yuan 0001, Weihua Ou, Qirong Mao
ICME2
2019 Two-phase probabilistic collaborative representation-based classification
Jianping Gou, Lei Wang 0095, Bing Hou, Jiancheng Lv 0001, Yun-Hao Yuan 0001, Qirong Mao
Expert Syst. Appl.2