Longsheng Wei

dblp:38/8613 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-2492-5510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PDNet: Pluralistic depth-aware network for RGB-D salient object detection
Longsheng Wei, Zhiqiang Zhu, Yichen Mi
Signal Process.1
2025 RAAG:Redundancy-adaptive and attention-guided token pruning for efficient video action detection
Jun Chen 0019, Sailong Deng, Wei Yu 0018, Longsheng Wei
Neurocomputing4
2025 CCAF-Net: Cascade Complementarity-Aware Fusion Network for traffic accident prediction in dashcam videos
Wei Liu 0005, Yixiang Gao, Longsheng Wei, Jun Chen 0019
Neurocomputing5
2025 Attention-enhanced re-activation CAMs for weakly supervised semantic segmentation
Longsheng Wei, Tangqiang Li
Neurocomputing1
2025 CCINet: A cascaded consensus interaction network for co-saliency object detection
Longsheng Wei, Xu Pei, Jiu Huang
Neurocomputing1
2025 A pavement crack segmentation method based on deformable convolution and enhanced perceive network
Longsheng Wei
Multim. Tools Appl.2
2025 VRTNet: Vector Rectifier Transformer for Two-View Correspondence Learning
abstract
Finding reliable correspondences in two-view image and recovering the camera poses are key problems in photogrammetry and image signal processing. Multilayer perceptron (MLP) has a wide application in two-view correspondence learning for which is good at learning disordered sparse correspondences, but it is susceptible to the dominant outliers and requires additional functional blocks to capture context information. CNN can naturally extract local context information, but it cannot handle disordered data and extract global context and channel information. In order to overcome the shortcomings of MLP and CNN, we design a correspondence learning network based on Transformer, named Vector Rectifier Transformer (VRTNet). Transformer is an encoder-decoder structure which can handle disordered sparse correspondences and output sequences of arbitrary length. Therefore, we design two sub-Transformers in VRTNet to achieve the mutual conversion between disordered and ordered correspondences. The self-attention and cross-attention mechanisms in them allow VRTNet to focus on the global context relations of all correspondences. To capture local context and channel information, we propose rectifier network (including CNN and channel attention block) as the backbone of VRTNet, which avoids the complex design of additional blocks. Rectifier network can correct the errors of ordered correspondences to obtain rectified correspondences. Finally, outliers are removed by comparing original and rectified correspondences. VRTNet performs better than the state-of-the-art methods in the tasks of relative pose estimation, outlier removal and image registration.
Meng Yang 0031, Jun Chen 0019, Xin Tian 0006, Longsheng Wei, Jiayi Ma 0001
IEEE Trans. Multim.4
2024 Semi-supervised anomaly detection in video surveillance by inpainting
Wei Liu 0005, Mingqiang Gao, Shuaidong Duan, Longsheng Wei
Multim. Tools Appl.4
2024 KRRNet: Keypoint Relational Regression Network for Bottom-Up Anchor-Free Object Detection
abstract
Anchor-free detection methods identify different objects by perceiving bounding box keypoints without predefined anchor boxes, which have attracted much attention due to their straightforward design and comparable performance. Currently, most anchor-free methods detect bounding box corners to regress object locations. In clutter environments, the bounding box corners may lie in background regions, which have limited relation with the object itself. In addition, the relationships between object keypoints are always neglected, potentially affecting the perceptibility of the detector for high-precision object detection. In this paper, we propose the Keypoint Relational Regression Network (KRRNet) to detect object keypoints with semantic relations instead of bounding box corners. The relational regression head is designed to enhance the keypoint relationship exploration capability and reason accuracy object locations. Moreover, the random background sampling strategy is proposed to sample negative background points around foreground object regions and form point pairs with object keypoints. Then, KRRNet can explicitly learn discriminative feature embedding from contrastive learning to pull close the positive pairs and push apart the negative pairs, resisting the influence of surrounding complex environments. KRRNet can be trained on one Nvidia RTX 3090 GPU and achieves a single-scale test AP of 48.9% and multi-scale test AP of 50.6% on the MS-COCO test-dev with the backbone of Hourglass-104, surpassing state-of-the-art bottom-up anchor-free detector using the same backbone.
Yinyuan Wang, Haowen Du, Changxin Gao, Longsheng Wei, Dapeng Luo
IEEE Trans. Circuits Syst. Video Technol.5
2023 A cascaded refined rgb-d salient object detection network based on the attention mechanism
Guanyu Zong, Longsheng Wei, Yongtao Wang
Appl. Intell.2
2023 THAT-Net: Two-layer hidden state aggregation based two-stream network for traffic accident prediction
Wei Liu 0005, Yisheng Lu, Jun Chen 0019, Longsheng Wei
Inf. Sci.5
2023 EGA-Net: Edge feature enhancement and global information attention network for RGB-D salient object detection
abstract
With the supplement of texture and geometry cues in depth maps, salient object detection (SOD) shifts from 2D to 3D, aiming to detect the most attractive object in a pair of color and depth images. Previous work primarily focused on regional integrity. Few methods are used to improve the edge quality of prediction results, resulting in a final prediction with a complete structure but blurred edges. Moreover, due to the complexity of real-life scenarios, the problem of effectively separating the salient object from complex background has become a hot potato. Aiming to address these issues, we propose a novel network, EGA-Net, to improve the edge quality and highlight the main features of the salient object. Specifically, in the EGA-Net, we propose a feature interaction (FI) module and an edge feature enhancement (EFE) module, respectively. Among them, the FI module is used to remove unimodal feature redundancy, capture multi-modal feature complementarity, and reduce the contamination of low-quality depth maps. The EFE is used to improve the edge quality of the final salient object prediction results. Furthermore, a Global Information Guide Integration (GIGI) module has been proposed to suppress the background noise and effectively highlight the salient objects’ main features. It uses interleaving and fusion methods to automatically select and enhance the vital information in the original input features under the guidance of global features. We put the training of EGA-Net under the supervision of a new hybrid loss function that can simultaneously take global pixel point, foreground, and depth map quality into account. Quantitative and qualitative experiment results demonstrate that our method outperforms the 19 advanced methods on eight publicly available RGB-D salient object detection datasets with five evaluation metrics. You can find the code and results of our method at https://github.com/guanyuzong/EGA-Net .
Longsheng Wei, Guanyu Zong
Inf. Sci.1
2022 Multiscale face recognition in cluttered backgrounds based on visual attention
Guoqing Du, Longsheng Wei, Huaiying Lu, Changxin Gao, Dapeng Luo
Neurocomputing3
2022 Perceptual quality assessment for no-reference image via optimization-based meta-learning
Longsheng Wei, Qingqing Yan, Wei Liu 0005, Dapeng Luo
Inf. Sci.1
2022 UAV Image Stitching Using Shape-Preserving Warp Combined With Global Alignment
abstract
In this letter, we propose a strategy for unmanned aerial vehicle (UAV) image stitching to generate natural-looking panoramas. Traditional methods using homography to perform alignment cannot account for images with parallax, so they require that the input images should be taken from the same viewpoint or the scene should be near the planar. However, remote sensing images obtained by UAVs usually do not satisfy such an ideal situation, and the stitching results always suffer from artifacts. To overcome these challenges and obtain natural-looking panoramas, a global alignment strategy is proposed to better align the input images. Combined with a shape-preserving warp, the stitching results can achieve better alignment accuracy while maintaining the shape. Meanwhile, locality preserving matching (LPM) is used to eliminate mismatches during feature detection and matching for accurate alignment. In addition, to make the stitching results more natural-looking, we also use multiband blending to eliminate artifacts that may exist in the results due to unmodeled effects. Experiments show that our stitching strategy can effectively improve alignment accuracy and obtain natural-looking results compared to other state-of-the-art methods.
Donghai Guo, Jun Chen 0019, Linbo Luo 0002, Wenping Gong, Longsheng Wei
IEEE Geosci. Remote. Sens. Lett.5
2022 Learning scene-specific object detectors based on a generative-discriminative model with minimal supervision
Dapeng Luo, Siyuan Lei, Changxin Gao, Longsheng Wei
Pattern Recognit. Lett.7
2021 Unsupervised domain-adaptive scene-specific pedestrian detection for static video surveillance
Quanzheng Mou, Longsheng Wei, Conghao Wang, Dapeng Luo, Songze He, Changxin Gao
Pattern Recognit.2
2021 Drone Image Stitching Using Local Mesh-Based Bundle Adjustment and Shape-Preserving Transform
abstract
This article proposes a strategy for drone image stitching using local mesh-based bundle adjustment and shape-preserving transform, which aims to effectively stitch multiple overlapping drone images into a natural panoramic image. Existing traditional methods using a simple homography cannot handle the situation that the input drone images have parallax effect, and the image mosaic result always suffers from artifacts. In order to achieve natural-looking stitching results without the above limitation, we divide the proposed method into the following steps. Starting from initial feature sets obtained by off-the-shelf feature extraction methods, we incorporate the parallax errors into an energy minimum framework and construct a robust alignment energy. This energy can be minimized efficiently based on local bundle adjustment and robust$3\sigma $principle, which could eliminate parallax effects and achieve accurate alignment. Then the seamless panoramic image is obtained by warping the target image and the source images onto the mesh plane directly. An image patch can be transformed by projective transformation (e.g., homography), which provides good alignment but may cause distortions. Consequently, combined with mesh-based shape-preserving transform, our proposed strategy can improve the naturalness of the results flexibly. Experiments show that our stitching strategy can eliminate parallax effects more effectively and achieve natural-looking results compared to other state-of-the-art methods.
Qi Wan, Jun Chen 0019, Linbo Luo 0002, Wenping Gong, Longsheng Wei
IEEE Trans. Geosci. Remote. Sens.5
2020 Bio-inspired head detection framework based on online learning algorithm
Dapeng Luo, Quanzheng Mou, Zhipeng Zeng, Longsheng Wei, Xiangli Zhang
Multim. Tools Appl.5
2011 Building Recognition Based on Indirect Location of Planar Landmark in FLIR Image Sequences
abstract
A novel method is proposed to deal with the problem of building recognition in forward-looking infrared (FLIR) image sequences. Two computational models are deduced in this paper: one is the perspective transformation model, through which an image is transformed perspectively from downward-looking state to forward-looking state; the other is the indirect location model, through which the position of a building is computed in an FLIR image. In addition, in order to illustrate the application scope of our method, the error analysis is presented. The proposed approach is validated by extensive experiments, with images taken in different weather conditions, seasons and at different times. Superior recognition results were obtained on three FLIR image sequences we collected.
Dengwei Wang, Tianxu Zhang, Meijun Wan, Wenjun Shi, Longsheng Wei
Int. J. Pattern Recognit. Artif. Intell.5
2010 A Biologically-Inspired Top-Down Learning Model Based on Visual Attention
abstract
A biologically-inspired top-down learning model based on visual attention is proposed in this paper. Low-level visual features are extracted from learning object itself and do not depend on the background information. All the features are expressed as a feature vector, which is looked as a random variable following a normal distribution. So every learning object is represented as the mean and standard deviation. All the learning objects are combined as an object class, which is represented as class's mean and class's standard deviation stored in long-term memory (LTM). Then the learned knowledge is used to find the similar location in an attended image. Experimental results indicate that: when the attended object doesn't always appear in the background similar to that in the learning objects or their combinations change hugely between learning images and attended images, our model is excellent to other two top-down visual attention models.
Nong Sang, Longsheng Wei, Yuehuan Wang
ICPR2