EDBT 2026 Demo / reviewers in the wild / expert
Shengxiang Qi
dblp:122/1504
· DBLP profile ↗
10ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 50% Autonomous driving · 31% Representation and self-supervised learning · 19% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d object detection |
1.5 | 2 | 2024 | RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024 BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024 |
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception |
0.8 | 1 | 2024 | RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024 |
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection |
0.8 | 1 | 2024 | BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024 |
Robotics › Autonomous driving › perception
LiDAR perception |
0.8 | 1 | 2024 | BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.8 | 1 | 2024 | BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024 |
Computer vision › 3D vision › 3d scene understanding
multi-view 3d perception |
0.8 | 1 | 2024 | HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras · ECCV (50) 2024 |
Robotics › Autonomous driving
perception |
0.8 | 1 | 2024 | BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked autoencoder
point cloud masked autoencoder |
0.8 | 1 | 2024 | BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios · AAAI 2024 |
Computer vision › 3D vision › 3d object detection › multimodal 3d object detection
radar-camera 3d object detection |
0.8 | 1 | 2024 | RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024 |
Computer vision › 3D vision › 3d scene understanding
multi-task 3d perception |
0.2 | 1 | 2024 | HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras · ECCV (50) 2024 |
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles |
0.2 | 1 | 2024 | RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.8masked autoencoder · 0.8hybrid encoding · 0.8deformable attention · 0.8cross-attention fusion · 0.8bird's-eye-view representation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving ScenariosabstractExisting LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive and time-consuming. Self-supervised pre-training is an effective and desirable way to alleviate this dependence on extensive annotated data. In this work, we present BEV-MAE, an efficient masked autoencoder pre-training framework for LiDAR-based 3D object detection in autonomous driving. Specifically, we propose a bird's eye view (BEV) guided masking strategy to guide the 3D encoder learning feature representation in a BEV perspective and avoid complex decoder design during pre-training. Furthermore, we introduce a learnable point token to maintain a consistent receptive field size of the 3D encoder with fine-tuning for masked point cloud inputs. Based on the property of outdoor point clouds in autonomous driving scenarios, i.e., the point clouds of distant objects are more sparse, we propose point density prediction to enable the 3D encoder to learn location information, which is essential for object detection. Experimental results show that BEV-MAE surpasses prior state-of-the-art self-supervised methods and achieves a favorably pre-training efficiency. Furthermore, based on TransFusion-L, BEV-MAE achieves new state-of-the-art LiDAR-based 3D object detection results, with 73.6 NDS and 69.6 mAP on the nuScenes benchmark. The source code will be released at https://github.com/VDIGPKU/BEV-MAE. Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001 |
AAAI | 3 |
| 2024 | RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object DetectionabstractThree-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely on cameras is difficult to achieve highly accurate and robust 3D object detection. An effective solution to this issue is combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection. In this paper, we introduce RCBEVDet, a radar-camera fusion 3D object detection method in the bird's eye view (BEV). Specifically, we first design RadarBEVNet for radar BEV feature extraction. RadarBEVNet consists of a dual-stream radar backbone and a Radar Cross-Section (RCS) aware BEV encoder. In the dual-stream radar backbone, a point-based encoder and a transformer-based encoder are proposed to extract radar features, with an injection and extraction module to facilitate communication between the two encoders. The RCS-aware BEV encoder takes RCS as the object size prior to scattering the point feature in BEV. Besides, we present the Cross-Attention Multi-layer Fusion module to automatically align the multi-modal BEV feature from radar and camera with the deformable attention mechanism, and then fuse the feature with channel and spatial fusion layers. Experimental results show that RCBEVDet achieves new state-of-the-art radar-camera fusion results on nuScenes and view-of-delft (VoD) 3D object detection benchmarks. Furthermore, RCBEVDet achieves better 3D detection results than all real-time camera-only and radar-camera 3D object detectors with a faster inference speed at 21∼28 FPS. The source code will be released at https://github.com/VDIGPKU/RCBEVDet. Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ce Zhu |
CVPR | 6 |
| 2024 | HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001 |
ECCV (50) | 6 |
| 2020 | Efficient Single Shot Object Detector Towards More Accurate and Faster Prediction
Shengxiang Qi, Jiarong Yang |
PRCV (3) | 1 |
| 2015 | Kernel regression in mixed feature spaces for spatio-temporal saliency detection
Yansheng Li 0001, Yihua Tan, Jin-Gang Yu, Shengxiang Qi, Jinwen Tian |
Comput. Vis. Image Underst. | 4 |
| 2015 | Salient object detection via contrast information and object vision organization cues
Shengxiang Qi, Jin-Gang Yu, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian |
Neurocomputing | 1 |
| 2015 | Built-Up Area Detection From Satellite Images Using Multikernel Learning, Multifield Integrating, and Multihypothesis VotingabstractThis letter proposes a novel supervised approach for accurate built-up area detection from high-resolution remote sensing images. In existing supervised built-up area detection approaches based on block-based image interpretation, the determination of the block size and the pursuit of the pixel-level result are not well addressed. Concerning these issues, this letter proposes a complete and systematic approach. It first utilizes multikernel learning to incorporate multiple features to implement the block-level image interpretation. Then, multifield integrating (i.e., the image interpretation results using different block sizes are fused) is proposed to obtain the block-level result. On the basis of the achieved result of the second step, multihypothesis voting is finally presented for working toward the pixel-level built-up area detection result through multihypothesis superpixel representation and graph smoothing. The proposed approach has been validated in the ZY-3 and GF-1 satellite images, and experimental results show that the proposed approach can outperform the state-of-the-art approaches. Yansheng Li 0001, Yihua Tan, Shengxiang Qi, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Unsupervised Ship Detection Based on Saliency and S-HOG Descriptor From Optical Satellite ImagesabstractWith the development of high-resolution imagery, ship detection in optical satellite images has attracted a lot of research interest because of the broad applications in fishery management, vessel salvage, etc. Major challenges for this task include cloud, wave, and wake clutters, and even the variability of ship sizes. In this letter, we propose an unsupervised ship detection method toward overcoming these existing issues. Visual saliency, which focuses on highlighting salient signals from scenes, is applied to extract candidate regions followed by a homogeneous filter presented to confirm suspected ship targets with complete profiles. Then, a novel descriptor, ship histogram of oriented gradient, which characterizes the gradient symmetry of ship sides, is provided to discriminate real ships. Experimental results on numerous panchromatic satellite images demonstrate the good performance of our method compared to state-of-the-art methods. Shengxiang Qi, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Visual saliency detection using feature activity weighted decorrelation cuesabstractIn this paper, a novel model based on feature activity weighted decorrelation cues is proposed for visual saliency detection in natural images. It consists of two parts: the feature decorrelation and feature information-activity. For the first part, Laplacian sparse coding and low-rank decomposition are used to extract decorrelated features from the scenes. For the second part, Incremental Coding Length is applied to measure the information-activity contained in features, which is then employed to weight the decorrelated features. Finally, visual saliency is estimated through a max pooling strategy. Experimental results on a publicly available benchmark demonstrate the effectiveness of our proposed model with good performance against the state-of-the-art methods. Shengxiang Qi, Jin-Gang Yu, Ji Zhao 0001, Jie Ma 0003, Jinwen Tian |
ICIP | 1 |
| 2013 | A Robust Directional Saliency-Based Method for Infrared Small-Target Detection Under Various Complex BackgroundsabstractInfrared small-target detection plays an important role in image processing for infrared remote sensing. In this letter, different from traditional algorithms, we formulate this problem as salient region detection, which is inspired by the fact that a small target can often attract attention of human eyes in infrared images. This visual effect arises from the discrepancy that a small target resembles isotropic Gaussian-like shape due to the optics point spread function of the thermal imaging system at a long distance, whereas background clutters are generally local orientational. Based on this observation, a new robust directional saliency-based method is proposed incorporating with visual attention theory for infrared small-target detection. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods for real infrared images with various typical complex backgrounds. Shengxiang Qi, Jie Ma 0003, Chao Tao 0001, Changcai Yang, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 1 |