VLDB 2026 Research / reviewers in the wild / expert
Shaoqing Xu
dblp:81/7774
· DBLP profile ↗
18ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Computer networks · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy RobustnessabstractThe safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents' future trajectories. This direct-editing approach fully leverages the VLM's powerful generalization capabilities to create a diverse curriculum of plausible yet challenging scenarios that extend beyond the scope of traditional methods. We demonstrate that our approach substantially enhances the safety and robustness of the resulting AD policy, particularly in its ability to navigate critical long-tail events. Qimao Chen, Shaoqing Xu, Zhiyi Lai, Zixun Xie, Yuechen Luo, Shengyin Jiang, Hanbing Li, Long Chen 0005 |
AAAI | 3 |
| 2026 | Think before Go: Hierarchical Reasoning for Image-goal NavigationabstractPengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Lin Zhao, Long Chen, Zhi-Xin Yang, Nanning Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengna Li, Kangyi Wu, Shaoqing Xu, Long Chen 0005, Nanning Zheng 0001 |
ACL (1) | 3 |
| 2026 | FPC-VLA: A vision-language-action framework with a supervisor for failure prediction and correction
Zhixiang Duan, Tianshi Xie, Fuyu Cao, Pinxi Shen, Peili Song, Chenyang Zhao 0009, Piaopiao Jin, Guokang Sun, Shaoqing Xu, Yangwei You, Jingtai Liu |
Expert Syst. Appl. | 10 |
| 2026 | DGFusion: Dual-Guided Fusion for Robust Multi-Modal 3D Object DetectionabstractAs a critical task in autonomous driving perception systems, 3D object detection is used to identify and track key objects, such as vehicles and pedestrians. However, detecting distant, small, or occluded objects (hard instances) remains a challenge, which directly compromises the safety of autonomous driving systems. We observe that existing multi-modal 3D object detection methods often follow a single-guided paradigm, failing to account for the differences in information density of hard instances between modalities. In this work, we propose DGFusion, based on the Dual-guided paradigm, which fully inherits the advantages of the Point-guide-Image paradigm and integrates the Image-guide-Point paradigm to address the limitations of the single paradigms. The core of DGFusion, the Difficulty-aware Instance Pair Matcher (DIPM), performs instance-level feature matching based on difficulty to generate easy and hard instance pairs, while the Dual-guided Modules exploit the advantages of both pair types to enable effective multi-modal feature fusion. Experimental results demonstrate that our DGFusion outperforms the baseline methods, with respective improvements of +1.0% mAP, +0.8% NDS, and +1.3% average recall on nuScenes. Extensive experiments demonstrate consistent robustness gains for hard instance detection across ego-distance, size, visibility, and small-scale training scenarios. Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lei Yang 0060, Ziying Song |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | TiGDistill-BEV: Multi-View BEV 3D Object Detection via Target Inner-Geometry Learning DistillationabstractAccurate multi-view 3D object detection is essential for applications such as autonomous driving. Researchers have consistently aimed to leverage LiDAR’s precise spatial information to enhance camera-based detectors through methods like depth supervision and bird-eye-view (BEV) feature distillation. However, existing approaches often face challenges due to the inherent differences between LiDAR and camera data representations. In this paper, we introduce the TiGDistill-BEV, a novel approach that effectively bridges this gap by leveraging the strengths of both sensors. Our method distills knowledge from diverse modalities(e.g., LiDAR) as the teacher model to a camera-based student detector, utilizing the Target Inner-Geometry learning scheme to enhance camera-based BEV detectors through both depth and BEV features by leveraging diverse modalities. Specially, we propose two key modules: an inner-depth supervision module to learn the low-level relative depth relations within objects which equips detectors with a deeper understanding of object-level spatial structures, and an inner-feature BEV distillation module to transfer high-level semantics of different keypoints within foreground targets. To further alleviate the domain gap, we incorporate both inter-channel and inter-keypoint distillation to model feature similarity. Extensive experiments on the nuScenes benchmark demonstrate that TiGDistill-BEV significantly boosts camera-based only detectors achieving a state-of-the-art with 62.8% NDS and surpassing previous methods by a significant margin. The codes is available at: https://github.com/Public-BOTs/TiGDistill-BEV.git. Shaoqing Xu, Peixiang Huang, Ziying Song, Zhi-Xin Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving frameworks enable seamless integration of perception and planning but often rely on one-shot trajectory prediction, which may lead to unstable control and vulnerability to occlusions in single-frame perception. To address this, we propose the Momentum-Aware Driving (MomAD) framework, which introduces trajectory momentum and perception momentum to stabilize and refine trajectory predictions. MomAD comprises two core components: (1) Topological Trajectory Matching (TTM) employs Hausdorff Distance to select the optimal planning query that aligns with prior paths to ensure coherence; (2) Momentum Planning Interactor (MPI) cross-attends the selected planning query with historical queries to expand static and dynamic perception files. This enriched query, in turn, helps regenerate long-horizon trajectory and reduce collision risks. To mitigate noise arising from dynamic environments and detection errors, we introduce robust instance denoising during training, enabling the planning model to focus on critical signals and improve its robustness. We also propose a novel Trajectory Prediction Consistency (TPC) metric to quantitatively assess planning stability. Experiments on the nuScenes dataset demonstrate that MomAD achieves superior long-term consistency (≥ 3s) compared to SOTA methods. Moreover, evaluations on the curated Turning-nuScenes shows that MomAD reduces the collision rate by 26% and improves TPC by 0.97m (33.45%) over a 6s prediction horizon, while closed- loop on Bench2Drive demonstrates an up to 16.3% improvement in success rate. The source code is available at https://github.com/adept-thu/MomAD. Ziying Song, Caiyan Jia, Hongyu Pan, Shaoqing Xu, Lei Yang 0060, Yadan Luo |
CVPR | 8 |
| 2025 | Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyabstractWe present Genesis, a unified world model for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-represented LiDAR generator with NeRF-based rendering and adaptive sampling. Both modalities are directly coupled through a shared condition input, enabling coherent evolution across visual and geometric domains. To guide the generation with structured semantics, we introduce DataCrafter, a captioning module built on vision-language models that provides scene-level and instance-level captions. Extensive experiments on the nuScenes benchmark demonstrate that Genesis achieves state-of-the-art performance across video and LiDAR metrics (FVD 16.95, FID 4.24, Chamfer 0.611), and benefits downstream tasks including segmentation and 3D detection, validating the semantic fidelity and practical utility of the synthetic data. Zhanqian Wu, Kaixin Xiong, Gangwei Xu, Shaoqing Xu, Hangjun Ye, Wenyu Liu 0001, Xinggang Wang |
NeurIPS | 7 |
| 2024 | GraphBEV: Towards Robust BEV Feature Alignment for Multi-modal 3D Object Detection
Ziying Song, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092 |
ECCV (26) | 3 |
| 2024 | RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
Ziying Song, Guoxing Zhang, Lei Yang 0060, Shaoqing Xu, Caiyan Jia, Feiyang Jia, Li Wang 0092 |
IJCAI | 5 |
| 2024 | SparseInteraction: Sparse Semantic Guidance for Radar and Camera 3D Object DetectionabstractMulti-modal fusion techniques, such as radar and images, enable a complementary and cost-effective perception of the surrounding environment regardless of lighting and weather conditions. However, existing fusion methods for surround-view images and radar are challenged by the inherent noise and positional ambiguity of radar, which leads to significant performance losses. To address this limitation effectively, our paper presents a robust, end-to-end fusion framework dubbed SparseInteraction. First, we introduce the Noisy Radar Filter (NRF) module to extract foreground features by creatively using queried semantic features from the image to filter out noisy radar features. Furthermore, we implement the Sparse Cross-Attention Encoder (SCAE) to effectively blend foreground radar features and image features to address positional ambiguity issues at a sparse level. Ultimately, to facilitate model convergence and performance, the foreground prior queries containing position information of the foreground radar are concatenated with predefined queries and fed into the subsequent transformer-based decoder. The experimental results demonstrate that the proposed fusion strategies markedly enhance detection performance and achieve new state-of-the-art results on the nuScenes benchmark. Source code is available at https://github.com/GG-Bonds/SparseInteraction. Shaoqing Xu, Shengyin Jiang, Li Liu 0069, Ziying Song, Zhi-Xin Yang 0001 |
ACM Multimedia | 1 |
| 2024 | Multi-Sem Fusion: Multimodal Semantic Fusion for 3-D Object DetectionabstractLIDAR and camera fusion techniques are promising for achieving 3D object detection in autonomous driving. Most multi-modal 3D object detection frameworks integrate semantic knowledge from 2D images into 3D LiDAR point clouds to enhance detection accuracy. Nevertheless, the restricted resolution of 2D feature maps impedes accurate re-projection and often induces a pronounced boundary-blurring effect, which is primarily attributed to erroneous semantic segmentation. To address these limitations, we present theMulti-Sem Fusion (MSF)framework, a versatile multi-modal fusion approach that employs 2D/3D semantic segmentation methods to generate parsing results for both modalities. Subsequently, the 2D semantic information undergoes re-projection into 3D point clouds utilizing calibration parameters. To tackle misalignment challenges between the 2D and 3D parsing results, we introduce an Adaptive Attention-based Fusion (AAF) module to fuse them by learning an adaptive fusion score. Then the point cloud with the fused semantic label is sent to the following 3D object detectors. Furthermore, we propose a Deep Feature Fusion (DFF) module to aggregate deep features at different levels to boost the final detection performance. The effectiveness of the framework has been verified on two public large-scale 3D object detection benchmarks by comparing them with different baselines. And the experimental results show that the proposed fusion strategies can significantly improve the detection performance compared to the methods using only point clouds and the methods using only 2D semantic information. Moreover, our approach seamlessly integrates as a plug-in within any detection framework. Shaoqing Xu, Ziying Song, Sifen Wang, Zhi-Xin Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object DetectionabstractLiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when fusing sparse voxel features with dense image features in a one-to-one manner, resulting in the loss of the advantages of images, including semantic and continuity information, leading to sub-optimal detection performance, especially at long distances. In this paper, we present VoxelNextFusion, a multi-modal 3D object detection framework specifically designed for voxel-based methods, which effectively bridges the gap between sparse point clouds and dense images. In particular, we propose a voxel-based image pipeline that involves projecting point clouds onto images to obtain both pixel- and patch-level features. These features are then fused using a self-attention to obtain a combined representation. Moreover, to address the issue of background features present in patches, we propose a feature importance module that effectively distinguishes between foreground and background features, thus minimizing the impact of the background features. Extensive experiments were conducted on the widely used KITTI and nuScenes 3D object detection benchmarks. Notably, our VoxelNextFusion achieved around +3.20% in [email protected] improvement for car detection in hard level compared to the Voxel R-CNN baseline on the KITTI test dataset. Ziying Song, Jun Xie 0003, Caiyan Jia, Shaoqing Xu, Zhepeng Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Scalable 3D Beam-Steering for Directional Millimeter Wave Wireless NetworksabstractMulti-Gbps 60 GHz millimeter wave (mmWave) networks, are considered as the enabling technology for emerging applications such as untethered VR and 4K/8K Miracast. However, user motion, and even orientation change, can cause mis-alignment between mmWave transceivers’ directional beams and thus severe link outage. Within the practical 3D spaces, the combination of location and orientation dynamics leads to the exponential growth of beam searching complexity, which substantially exacerbates the outage. In this paper, we first measure the impact of 3D motion on 60 GHz link performance in the context of VR and Miracast applications. We find that 3D motion exhibits inherent non-predictability, so conventional beam steering solutions are no longer effective. Therefore, we propose a model-driven 3D beam-steering mechanism called Parallel Scanner (PSCAN), which can maintain high performance for mobile 60 GHz links. To enable PSCAN, we first discover and prove a hidden interaction between 3D beams and the spatial channel profile of 60 GHz radios. Leveraging on which, PSCAN strategically scans the 3D space to reduce the search latency by more than one order of magnitude. Experiment results based on a custom-built 60 GHz platform demonstrate PSCAN’s remarkable throughput gain, up to$5\times $, compared with the state-of-the-art. Yi Yang 0035, Anfu Zhou, Leilei Wu, Shaoqing Xu, Huadong Ma, Teng Wei, Xinyu Zhang 0003 |
IEEE Trans. Wirel. Commun. | 4 |
| 2020 | Robotic Millimeter-Wave Wireless NetworksabstractThe emerging millimeter-wave (mmWave) networking technology promises to unleash a new wave of multi-Gbps wireless applications. However, due to high directionality of the mmWave radios, maintaining stable link connection remains an open problem. Users' slight orientation change, coupled with motion and blockage, can easily disconnect the link. In this paper, we propose RoMil, a robotic mmWave relay that optimizes network coverage through wireless sensing and autonomous motion/rotation planning. The robot relay automatically constructs the geometry/reflectivity of the environment, by estimating the geometries of all signal paths. It then navigates itself along an optimal moving trajectory, and ensures continuous connectivity for the client despite environment/human dynamics. We have prototyped RoMil on a programmable robot carrying a commodity 60 GHz radio. Our field trials demonstrate that RoMil can achieve nearly full coverage in dynamic environment, even with constrained speed and mobility region. Anfu Zhou, Shaoqing Xu, Jingqi Huang, Shaoyuan Yang, Teng Wei, Xinyu Zhang 0003, Huadong Ma |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | An Unmanned Aerial Vehicle Navigation Mechanism with Preserving PrivacyabstractVisual-based Unmanned Aerial Vehicles (UAVs) (e.g. equipped with an optical camera) have become more popular in daily life, because of their flexibility and convenience in capturing images/videos. Existing works mainly focus on the image/video capturing efficiency, but overlook the privacy violations that may be caused by the misuse of UAVs. In this paper, we study the Privacy Preserving Navigation (PPN) problem of the path planning of a UAV in 3D space to cover a 2D Target Area (TA), so that TA is covered while the privacy of sub-areas within TA is violated. We prove the PPN is NP-hard and propose a heuristic solution to PPN. The real word data trace driven emulation results show our solution is effective. Yan Pan 0003, ShiNing Li, Juan Luque Chang, Yan Yan 0025, Shaoqing Xu, Yinghai An, Ting Zhu 0001 |
ICC | 5 |
| 2019 | Robot Navigation in Radio Beam Space: Leveraging Robotic Intelligence for Seamless mmWave Network CoverageabstractThe emerging millimeter-wave (mmWave) networking technology promises to unleash a new wave of multi-Gbps wireless applications. However, due to high directionality of the mmWave radios, maintaining stable link connection remains an open problem. Users' slight orientation change, coupled with motion and blockage, can easily disconnect the link. In this paper, we propose miDroid, a robotic mmWave relay that optimizes network coverage through wireless sensing and autonomous motion/rotation planning. The robot relay automatically constructs the geometry/reflectivity of the environment, by estimating the geometries of all signal paths. It then navigates itself along an optimal moving trajectory, and ensures continuous connectivity for the client despite environment/human dynamics. We have prototyped miDroid on a programmable robot carrying a commodity 60 GHz radio. Our field trials demonstrate that miDroid can achieve nearly full coverage in dynamic environment, even with constrained speed and mobility region. Anfu Zhou, Shaoqing Xu, Jingqi Huang, Shaoyuan Yang, Teng Wei, Xinyu Zhang 0003, Huadong Ma |
MobiHoc | 2 |
| 2018 | Following the Shadow: Agile 3-D Beam-Steering for 60 GHz Wireless Networksabstract60 GHz networks, with multi-Gbps bitrate, are considered as the enabling technology for emerging applications such as wireless Virtual Reality (VR) and 4K/8K real-time Miracast. However, user motion, and even orientation change, can cause mis-alignment between 60 GHz transceivers' directional beams, thus causing severe link outage. Within the practical 3D spaces, the combination of location and orientation dynamics leads to exponential growth of beam searching complexity, which substantially exacerbates the outage and hinders fast recovery. In this paper, we first conduct an extensive measurement to analyze the impact of 3D motion on 60 GHz link performance, in the context of VR and Miracast applications. We find that 3D motion exhibits inherent non-predictability, so conventional beam steering solutions, which targets 2D scenarios with lower search space and short-term motion coherence, fail in practical 3D setup. Motivated by these observations, we propose a model-driven 3D beam-steering mechanism called Orthogonal Scanner (OScan), which can maintain high performance for mobile 60 GHz links in 3D space. OScan discovers and leverages a hidden interaction between 3D beams and the spatial channel profile of 60 GHz radios, and strategically scans the 3D space so as to reduce the search latency by more than one order of magnitude. Experiment results based on a custom-built 60 GHz platform along with a trace-driven emulator demonstrate OScan's remarkable throughput gain, up to 5×, compared with the state-of-the-art. Anfu Zhou, Leilei Wu, Shaoqing Xu, Huadong Ma, Teng Wei, Xinyu Zhang 0003 |
INFOCOM | 3 |
| 2009 | Asymmetrical interval regression using extended epsilon-SVM with robust algorithm
Shaoqing Xu, Qiang-Yi Luo |
Fuzzy Sets Syst. | 1 |