Wei Lu 0026

dblp:98/6613-26 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-6566-775XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LUFormer : A luminance-informed localized transformer with frequency augmentation for nighttime flare removal
Wei Lu 0026, Dubuke Ma, Peiguang Jing
Neural Networks1
2025 A Novel Pruning Algorithm Based on Data-Driven and Knowledge-Driven Global Channel Pruning
abstract
Pruning reduces the number of parameters in CNN models, making them lighter while preserving accuracy. Nevertheless, current pruning algorithms do not fit well with the input data, making it hard to recover the accuracy of the network after pruning. To this end, we propose a Data-driven and Knowledge-driven Global Channel Pruning algorithm (DKGCP), which leverages a data-driven channel self-attention module to align channel importance with the input data, enhancing the performance of pruned models, while incorporating knowledge-driven modules for local reconstruction and global feature alignment to improve generalization. Specifically, 1) During pre-training, the data-driven channel self-attention module uses all input data to determine channel importance aligned with the dataset. 2) During model optimization, the local reconstruction feature knowledge-driven module fits intermediate features between the pruned and original networks, while the global warming feature knowledge-driven module aligns global features and combines the cross-entropy loss to guide the feature learning process of the pruned network. In comparative experiments on VGGNet, GoogLeNet, and ResNet, our method outperformed state-of-the-art approaches. Notably, it achieved 94.55% accuracy on the pruned ResNet-110 with a 50% pruning rate on the CIFAR-10 dataset.
Wei Lu 0026, Xinjie Ding, Jinghui Chu
IEEE Signal Process. Lett.1
2025 PFPRNet: A Phase-Wise Feature Pyramid With Retention Network for Polyp Segmentation
abstract
Early detection of colonic polyps is crucial for the prevention and diagnosis of colorectal cancer. Currently, deep learning-based polyp segmentation methods have become mainstream and achieved remarkable results. Acquiring a large number of labeled data is time-consuming and labor-intensive, and meanwhile the presence of numerous similar wrinkles in polyp images also hampers model prediction performance. In this paper, we propose a novel approach called Phase-wise Feature Pyramid with Retention Network (PFPRNet), which leverages a pre-trained Transformer-based Encoder to obtain multi-scale feature maps. A Phase-wise Feature Pyramid with Retention Decoder is designed to gradually integrate global features into local features and guide the model's attention towards key regions. Additionally, our custom Enhance Perception module enables capturing image information from a broader perspective. Finally, we introduce an innovative Low-layer Retention module as an alternative to Transformer for more efficient global attention modeling. Evaluation results on several widely-used polyp segmentation datasets demonstrate that our proposed method has strong learning ability and generalization capability, and outperforms the state-of-the-art approaches.
Jinghui Chu, Wangtao Liu, Wei Lu 0026
IEEE J. Biomed. Health Informatics4
2025 Multimodal Dual-Graph Collaborative Network With Serial Attentive Aggregation Mechanism for Micro-Video Multi-Label Classification
abstract
The increasing commercial value of micro-videos has spurred a rising demand for grasping their contents. The abundant multimodal cues in micro-videos exhibit substantial potential in enhancing content comprehension. However, effectively harnessing the collaborative characteristics across different modalities remains a significant challenge, especially in multi-label scenarios due to inconsistent behaviors regarding label correlations. To better tackle this issue, in this paper, we first introduce a multimodal dual-graph collaborative network with serial attentive aggregation mechanism (MDGCN) for micro-video multi-label classification. In MDGCN, we exploit an asymmetric encoder-decoder framework, which incorporates multiple parallel encoders with complementary representations and a decoder to ensure the completeness of encoded results. Meanwhile, an adversarial constraint is used to ensure individual differences prominently featured within each modality. Furthermore, considering the inconsistency of label correlations across various modalities, we then construct a serial attentive graph convolutional network that employs an interactive dual-graph attention paradigm to sequentially integrate multimodal representations and dynamically explore label correlations. The experiments conducted on two datasets demonstrate that our proposed method outperforms state-of-the-art approaches.
Wei Lu 0026, Peiguang Jing, Weiming Wang 0002, Yuting Su 0001
IEEE Trans. Multim.2
2024 HiEval: A scheduling performance estimation approach for spatial accelerators via hierarchical abstraction
Yue Hu 0004, Wei Lu 0026, Yu Liu 0004
J. Syst. Archit.4
2024 VMemNet: A Deep Collaborative Spatial-Temporal Network With Attention Representation for Video Memorability Prediction
abstract
Video memorability measures the degree to which a video is remembered by different viewers and has shown great potential in various contexts, including advertising, education, and health care. While extensive research has been conducted on image memorability, the study of video memorability is still in its early stages. Existing methods in this field primarily focus on coarse-grained spatial feature representation and decision fusion strategies, overlooking the crucial interactions between spatial and temporal domains. Therefore, we propose an end-to-end collaborative spatial-temporal network called VMemNet, which incorporates targeted attention mechanisms and intermediation fusion strategies. This enables VMemNet to capture the intricate relationships between spatial and temporal information and uncover more elements of memorability within video visual features. VMemNet integrates spatially and semantically guided attention modules into a dual-stream network architecture, allowing it to simultaneously capture static local cues and dynamic global cues in videos. Specifically, the spatial attention module is used to aggregate more memorable elements from spatial locations, and the semantically guided attention module is used to achieve semantic alignment and intermediate fusion of the local and global cues. In addition, two types of loss functions with complementary decision rules are associated with the corresponding attention modules to guide the training process of the proposed network. Experimental results obtained on a publicly available dataset verify that the proposed VMemNet approach outperforms all current single- and multi-modal methods in terms of video memorability prediction.
Wei Lu 0026, Jiaze Han, Peiguang Jing, Yu Liu 0004, Yuting Su 0001
IEEE Trans. Multim.1
2023 A Novel Channel Pruning Approach based on Local Attention and Global Ranking for CNN Model Compression
abstract
Channel pruning facilitates the acceleration and deployment of convolutional neural networks on resource-constrained devices. Nevertheless, existing related methods mainly focus on the importance of an individual channel, neglecting the intra-layer relationship and inter-layer influence. In this paper, we propose a novel local attention and global ranking (LAGR) method for channel pruning. Specifically, we first introduce the attention mechanism to explore the local correlation between channels of the intra-layer. On this basis, we evaluate the global ranking of all channels across the network by the normalization operation. Besides, we introduce a noisy training strategy in the pre-training stage to ensure a balanced weight distribution. Extensive experiments conducted on three representative networks, including VGGNet, GoogLeNet, and ResNet, have demonstrated the superior performance of the proposed method in comparison with several state-of-the-art methods.
Wei Lu 0026, Peiguang Jing, Jinghui Chu, Fugui Fan
ICME1
2023 A Multimodal Aggregation Network With Serial Self-Attention Mechanism for Micro-Video Multi-Label Classification
abstract
Currently, micro-videos have attracted increasing attention due to their unique properties and great commercial value. Considering that micro-videos naturally incorporate multimodal information, a powerful representation method for distinct joint multimodal representations is essential for real applications. Inspired by the potential of attention neural network architectures over various tasks, we propose a multimodal aggregation network (MANET) with a serial self-attention mechanism to perform tasks of micro-video multi-label classification. Specifically, we first propose a parallel content-dependent graph neural networks (CDGNN) module, which explores category-related embeddings of micro-videos by disentangling category relations into modality-specific and modality-shared category dependency patterns. Then we introduce a serial self-attention (SSA) module to transmit the multimodal information in sequential order, in which an aggregation bottleneck is incorporated to better collect and condense the significant information. Experiments conducted on a large-scale multi-label micro-video dataset demonstrate that our proposed method has achieved competitive results compared with several state-of-the-art methods.
Wei Lu 0026, Peiguang Jing, Yuting Su 0001
IEEE Signal Process. Lett.1
2023 Learning Dual Low-Rank Representation for Multi-Label Micro-Video Classification
abstract
Currently, with the rapid development of mobile Internet, micro-video has become a prevailing format of user-generated contents (UGCs) on various social media platforms. Several studies have been conducted towards to understanding high-level micro-video semantics, such as venue categorization, memorability, and popularity. However, these approaches supported tasks with only a single output, which exhibited limitations when attempting to use them to resolve tasks with multiple outputs, especially the multi-label micro-video classification. To tackle this problem, in this paper, we propose a dual multi-modal low-rank decomposition (DMLRD) method for multi-label micro-video classification tasks. To learn more comprehensive micro-video representations, we first learn the low-rank-regularized modality-specific and modality-shared components by considering the consistency and the complementarity among modalities simultaneously. Meanwhile, the less descriptive power of each modality aroused by inherent properties can be solved to a certain extent. To obtain unseen label representations, we next construct a sparsity-regularized multi-matrix normal estimation term to jointly encode the latent relationship structures among labels and dimensions. Experiments on two datasets demonstrate the effectiveness of our proposed method over the state-of-art methods.
Wei Lu 0026, Liqiang Nie, Peiguang Jing, Yuting Su 0001
IEEE Trans. Multim.1
2021 Joint Co-Attention And Co-Reconstruction Representation Learning For One-Shot Object Detection
abstract
One-shot object detection aims to detect all candidate instances in a target image whose label class is unavailable in training, and only one labeled query image is given in testing. Nevertheless, insufficient utilization of the only known sample is one significant reason causing the performance degradation of current one-shot object detection models. To tackle the problem, we develop joint co-attention and co-reconstruction (CoAR) representation learning for one-shot object detection. The main contributions are described as follows. First, we propose a high-order feature fusion operation to exploit the deep co-attention of each target-query pair, which aims to enhance the correlation of the same class. Second, we use a low-rank structure to reconstruct the target-query feature in channel level, which aims to remove the irrelevant noise and enhance the latent similarity between the region proposals in target image and the query image. Experiments on both PASCAL VOC and MS COCO datasets demonstrate that our method outperforms previous state-of-the-art algorithms.
Jinghui Chu, Peiguang Jing, Wei Lu 0026
ICIP4
2021 Illumination-based adaptive saliency detection network through fusion of multi-source features
Chunxu Jiang, Yu Liu 0004, Jinglin Sun, Jichang Guo, Wei Lu 0026
J. Vis. Commun. Image Represent.5
2021 A Dual Rank-Constrained Filter Pruning Approach for Convolutional Neural Networks
abstract
Filter pruning has attracted increasing attentions to compress and accelerate the convolutional neural networks (CNNs) on computationally restricted devices. Existing related methods mainly focus on independently leveraging the spatial information of individual filters while ignoring the inner correlation among filters. In this letter, we propose a dual rank-constrained filter pruning approach for convolutional neural networks, in which the representation, clustering, and identification of representative filters are integrated into an adaptive graph regularization framework. Particularly, the proposed approach utilizes the low-rank constraint to capture the low-dimensional intrinsic representations of filters for the adaptive affinity graph construction and clustering. It is noteworthy that the original filters are projected as points on Grassmann manifold for geometrical structure preserving. Meanwhile, the high-rank constraint is employed to select the most informative filter for representing the cluster. Experimental results on CIFAR-10 dataset show that the proposed approach achieves competitive results compared with the state-of-the-art methods by 87.1% in VGGNet-16, 52.4% in GoogLeNet and 72.9% in ResNet-56 in terms of model parameters compression.
Fugui Fan, Yuting Su 0001, Peiguang Jing, Wei Lu 0026
IEEE Signal Process. Lett.4
2020 Super-resolution using multi-channel merged convolutional network
Jinghui Chu, Wei Lu 0026
Neurocomputing4
2019 Wearable Computing for Internet of Things: A Discriminant Approach for Human Activity Recognition
abstract
With the rapid development of the wireless sensor network and the continuous improvement of its key technologies, the concept of Internet of Things has been encouraged and extended due to its wide applications in scenarios, such as smart homes and healthcare. Under the background, human activity recognition has drawn great attention in recent years. In this paper, we present a discriminant approach to recognize daily human activities recorded through accelerometer sensor. In the proposed approach, we first use S transform (ST) to extract features, and then introduce a supervised regularization-based robust subspace (SRRS) learning method to learn low-dimensional intrinsic feature representation from the original feature subspace. Particularly, ST has been described as a joint time-frequency representation, which is insensitive to noise. SRRS can learn more robust and discriminative features to reinforce the descriptions of samples while removing noise and redundancy. Experiments are conducted on three publicly available datasets, i.e., wireless sensor data mining, SCUT-NAA, and mHealth demonstrating the superior performance of our proposed scheme compared with state-of-the-art methods.
Wei Lu 0026, Fugui Fan, Jinghui Chu, Peiguang Jing, Yuting Su 0001
IEEE Internet Things J.1
2018 Effective pattern recognition and find-density-peaks clustering based blind identification for underdetermined speech mixing systems
Xiangdong Huang 0002, Runan Song, Wei Lu 0026
Multim. Tools Appl.4
2018 A Novel Multiconnected Convolutional Network for Super-Resolution
abstract
Convolutional neural networks exhibit superior performance for single image super-resolution (SISR) tasks. However, as the network grows deeper, features from the earlier layers are impeded or less used in later layers. In SISR, the earlier layers are mainly composed of local features that are essential to the task. In this letter, we present a novel multiconnected convolutional network for SISR tasks by enhancing the combination of both low- and high-level features. We design a structure built on multiconnected blocks to extract diversified and complicated features via the concatenation of low-level features to high-level features. In addition to stacking multiconnected blocks, a long skip-connection is implemented to further aggregate features of the first layer and a specific later layer. Furthermore, we employ a flexible two-parameter loss function to optimize the training process. The proposed method yields state-of-the-art performance both in terms of quantitative metrics and visual quality. The method also outperforms existing methods on datasets via unknown degrading operators, indicating an excellent generalization ability.
Jinghui Chu, Wei Lu 0026, Xiangdong Huang 0002
IEEE Signal Process. Lett.3
2017 Adaptive Ensemble Undersampling-Boost: A novel learning framework for imbalanced data
Wei Lu 0026, Zhe Li 0013, Jinghui Chu
J. Syst. Softw.1
2016 Dynamic Hand Gesture Recognition With Leap Motion Controller
abstract
Dynamic hand gesture recognition is a crucial but challenging task in the pattern recognition and computer vision communities. In this paper, we propose a novel feature vector which is suitable for representing dynamic hand gestures, and presents a satisfactory solution to recognizing dynamic hand gestures with a Leap Motion controller (LMC) only. These have not been reported in other papers. The feature vector with depth information is computed and fed into the Hidden Conditional Neural Field (HCNF) classifier to recognize dynamic hand gestures. The systematic framework of the proposed method includes two main steps: feature extraction and classification with the HCNF classifier. The proposed method is evaluated on two dynamic hand gesture datasets with frames acquired with a LMC. The recognition accuracy is 89.5% for the LeapMotion-Gesture3D dataset and 95.0% for the Handicraft-Gesture dataset. Experimental results show that the proposed method is suitable for certain dynamic hand gesture recognition tasks.
Wei Lu 0026, Zheng Tong, Jinghui Chu
IEEE Signal Process. Lett.1