Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shengxiang Qi

dblp:122/1504 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 50% Autonomous driving · 31% Representation and self-supervised learning · 19%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.522024
RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception
0.812024
RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection
0.812024
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024
Robotics › Autonomous driving › perception
LiDAR perception
0.812024
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.812024
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024
Computer vision › 3D vision › 3d scene understanding
multi-view 3d perception
0.812024
HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras · ECCV (50) 2024
Robotics › Autonomous driving
perception
0.812024
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked autoencoder
point cloud masked autoencoder
0.812024
BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024
Computer vision › 3D vision › 3d object detection › multimodal 3d object detection
radar-camera 3d object detection
0.812024
RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024
Computer vision › 3D vision › 3d scene understanding
multi-task 3d perception
0.212024
HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras · ECCV (50) 2024
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles
0.212024
RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024

Methods — techniques the papers use, named apart from their topics

transformer · 0.8masked autoencoder · 0.8hybrid encoding · 0.8deformable attention · 0.8cross-attention fusion · 0.8bird's-eye-view representation · 0.8
YearPublicationVenuePosition
2024 BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
abstract
Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive and time-consuming. Self-supervised pre-training is an effective and desirable way to alleviate this dependence on extensive annotated data. In this work, we present BEV-MAE, an efficient masked autoencoder pre-training framework for LiDAR-based 3D object detection in autonomous driving. Specifically, we propose a bird's eye view (BEV) guided masking strategy to guide the 3D encoder learning feature representation in a BEV perspective and avoid complex decoder design during pre-training. Furthermore, we introduce a learnable point token to maintain a consistent receptive field size of the 3D encoder with fine-tuning for masked point cloud inputs. Based on the property of outdoor point clouds in autonomous driving scenarios, i.e., the point clouds of distant objects are more sparse, we propose point density prediction to enable the 3D encoder to learn location information, which is essential for object detection. Experimental results show that BEV-MAE surpasses prior state-of-the-art self-supervised methods and achieves a favorably pre-training efficiency. Furthermore, based on TransFusion-L, BEV-MAE achieves new state-of-the-art LiDAR-based 3D object detection results, with 73.6 NDS and 69.6 mAP on the nuScenes benchmark. The source code will be released at https://github.com/VDIGPKU/BEV-MAE.
Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001
AAAI3
2024 RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
abstract
Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely on cameras is difficult to achieve highly accurate and robust 3D object detection. An effective solution to this issue is combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection. In this paper, we introduce RCBEVDet, a radar-camera fusion 3D object detection method in the bird's eye view (BEV). Specifically, we first design RadarBEVNet for radar BEV feature extraction. RadarBEVNet consists of a dual-stream radar backbone and a Radar Cross-Section (RCS) aware BEV encoder. In the dual-stream radar backbone, a point-based encoder and a transformer-based encoder are proposed to extract radar features, with an injection and extraction module to facilitate communication between the two encoders. The RCS-aware BEV encoder takes RCS as the object size prior to scattering the point feature in BEV. Besides, we present the Cross-Attention Multi-layer Fusion module to automatically align the multi-modal BEV feature from radar and camera with the deformable attention mechanism, and then fuse the feature with channel and spatial fusion layers. Experimental results show that RCBEVDet achieves new state-of-the-art radar-camera fusion results on nuScenes and view-of-delft (VoD) 3D object detection benchmarks. Furthermore, RCBEVDet achieves better 3D detection results than all real-time camera-only and radar-camera 3D object detectors with a faster inference speed at 21∼28 FPS. The source code will be released at https://github.com/VDIGPKU/RCBEVDet.
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ce Zhu
CVPR6
2024 HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001
ECCV (50)6
2020 Efficient Single Shot Object Detector Towards More Accurate and Faster Prediction
Shengxiang Qi, Jiarong Yang
PRCV (3)1
2015 Kernel regression in mixed feature spaces for spatio-temporal saliency detection
Yansheng Li 0001, Yihua Tan, Jin-Gang Yu, Shengxiang Qi, Jinwen Tian
Comput. Vis. Image Underst.4
2015 Salient object detection via contrast information and object vision organization cues
Shengxiang Qi, Jin-Gang Yu, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian
Neurocomputing1
2015 Built-Up Area Detection From Satellite Images Using Multikernel Learning, Multifield Integrating, and Multihypothesis Voting
abstract
This letter proposes a novel supervised approach for accurate built-up area detection from high-resolution remote sensing images. In existing supervised built-up area detection approaches based on block-based image interpretation, the determination of the block size and the pursuit of the pixel-level result are not well addressed. Concerning these issues, this letter proposes a complete and systematic approach. It first utilizes multikernel learning to incorporate multiple features to implement the block-level image interpretation. Then, multifield integrating (i.e., the image interpretation results using different block sizes are fused) is proposed to obtain the block-level result. On the basis of the achieved result of the second step, multihypothesis voting is finally presented for working toward the pixel-level built-up area detection result through multihypothesis superpixel representation and graph smoothing. The proposed approach has been validated in the ZY-3 and GF-1 satellite images, and experimental results show that the proposed approach can outperform the state-of-the-art approaches.
Yansheng Li 0001, Yihua Tan, Shengxiang Qi, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.4
2015 Unsupervised Ship Detection Based on Saliency and S-HOG Descriptor From Optical Satellite Images
abstract
With the development of high-resolution imagery, ship detection in optical satellite images has attracted a lot of research interest because of the broad applications in fishery management, vessel salvage, etc. Major challenges for this task include cloud, wave, and wake clutters, and even the variability of ship sizes. In this letter, we propose an unsupervised ship detection method toward overcoming these existing issues. Visual saliency, which focuses on highlighting salient signals from scenes, is applied to extract candidate regions followed by a homogeneous filter presented to confirm suspected ship targets with complete profiles. Then, a novel descriptor, ship histogram of oriented gradient, which characterizes the gradient symmetry of ship sides, is provided to discriminate real ships. Experimental results on numerous panchromatic satellite images demonstrate the good performance of our method compared to state-of-the-art methods.
Shengxiang Qi, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.1
2014 Visual saliency detection using feature activity weighted decorrelation cues
abstract
In this paper, a novel model based on feature activity weighted decorrelation cues is proposed for visual saliency detection in natural images. It consists of two parts: the feature decorrelation and feature information-activity. For the first part, Laplacian sparse coding and low-rank decomposition are used to extract decorrelated features from the scenes. For the second part, Incremental Coding Length is applied to measure the information-activity contained in features, which is then employed to weight the decorrelated features. Finally, visual saliency is estimated through a max pooling strategy. Experimental results on a publicly available benchmark demonstrate the effectiveness of our proposed model with good performance against the state-of-the-art methods.
Shengxiang Qi, Jin-Gang Yu, Ji Zhao 0001, Jie Ma 0003, Jinwen Tian
ICIP1
2013 A Robust Directional Saliency-Based Method for Infrared Small-Target Detection Under Various Complex Backgrounds
abstract
Infrared small-target detection plays an important role in image processing for infrared remote sensing. In this letter, different from traditional algorithms, we formulate this problem as salient region detection, which is inspired by the fact that a small target can often attract attention of human eyes in infrared images. This visual effect arises from the discrepancy that a small target resembles isotropic Gaussian-like shape due to the optics point spread function of the thermal imaging system at a long distance, whereas background clutters are generally local orientational. Based on this observation, a new robust directional saliency-based method is proposed incorporating with visual attention theory for infrared small-target detection. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods for real infrared images with various typical complex backgrounds.
Shengxiang Qi, Jie Ma 0003, Chao Tao 0001, Changcai Yang, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.1