Xiaokai Bai

dblp:139/1989 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0002-8382-2976ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 69% Vision and language · 20% Autonomous driving · 9%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.922026
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025
Computer vision › 3D vision › visual localization
cross-view geo-localization
1.012026
Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework · AAAI 2026
Computer vision › Vision and language
multimodal prompt learning
1.012026
Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework · AAAI 2026
Computer vision › Vision and language
vision-language dataset
1.012026
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision
3d scene understanding
0.912025
Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025
Computer vision › 3D vision
depth estimation
0.912025
Structure-Aware Radar-Camera Depth Estimation · ICRA 2025
Computer vision › 3D vision › 3d object detection
multi-view indoor 3d object detection
0.912025
Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025
Robotics › Autonomous driving
perception
0.912025
Structure-Aware Radar-Camera Depth Estimation · ICRA 2025
Computer vision › 3D vision › depth estimation › depth completion
radar-camera depth estimation
0.912025
Structure-Aware Radar-Camera Depth Estimation · ICRA 2025
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.312026
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Robotics › Robot navigation and mapping › localization › vehicle localization
UAV localization
0.312026
Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework · AAAI 2026
Computer vision › 3D vision
3d reconstruction
0.312025
Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025
Image and video processing › image enhancement
feature enhancement
0.312025
Structure-Aware Radar-Camera Depth Estimation · ICRA 2025

Methods — techniques the papers use, named apart from their topics

radar depth enhancement · 1.7multi-scale structure guided network · 1.7transformer · 1.0surround-view camera · 1.0cross-attention · 1.0contrastive learning · 1.0LiDAR · 1.04d imaging radar · 1.0sparse volume construction · 0.9geometry and context aware aggregation · 0.9
YearPublicationVenuePosition
2026 Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework
abstract
We present MoP-UAV, a new benchmark for UAV-based cross-view object geo-localization guided by multi-modal prompts. MoP-UAV supports fine-grained object-level cross-view localization under diverse prompt modalities, including natural language, bounding boxes, and click points. It offers potential for incorporating large foundation models like large language models (LLMs) and promotes the building of more flexible and intelligent UAV agents. Based on the benchmark, we propose MoPT, a multi-modal-prompt-guided tansformer that embeds prompts as token sequences and extract object location from UAV and satellite features via cross-attention. To enhance semantic consistency and performance, we further adopt a cross-view contrastive loss and propose a RefCOCOg-based pre-training strategy. Extensive experiments show that MoPT achieves robust localization under arbitrary prompt combinations. Notably, multi-modal-prompt training significantly boosts unimodal-prompt inference performance, highlighting the generalization benefits of multi-modal learning. MoPT trained with multi-modal prompts outperforms prior unimodal prompt works under the same setting.
Zhangkai Shen, Si-Yuan Cao, Xiaokai Bai, Zheheng Han, Qi Ming
AAAI4
2026 OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
abstract
The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions. Next-generation autonomous driving datasets must be multimodal, incorporating data from advanced sensors that feature extensive data coverage, detailed annotations, and diverse scene representation. To address this need, we present OmniHD-Scenes, a large-scale multimodal dataset that provides comprehensive omnidirectional high-definition data. The OmniHD-Scenes dataset combines data from 128-beam LiDAR, six cameras, and six 4D imaging radar systems to achieve full environmental perception. The dataset comprises 1501 clips, each approximately 30-s long, totaling more than 450 K synchronized frames and more than 5.85 million synchronized sensor data points. We also propose a novel 4D annotation pipeline. To date, we have annotated 200 clips with more than 514 K precise 3D bounding boxes. These clips also include semantic segmentation annotations for static scene elements. Additionally, we introduce a novel automated pipeline for generation of the dense occupancy ground truth, which effectively leverages information from non-key frames. Alongside the proposed dataset, we establish comprehensive evaluation metrics, baseline models, and benchmarks for 3D detection and semantic occupancy prediction. These benchmarks utilize surround-view cameras and 4D imaging radar to explore cost-effective sensor solutions for autonomous driving applications. Extensive experiments demonstrate the effectiveness of our low-cost sensor configuration and its robustness under adverse conditions.
Lianqing Zheng, Qunshu Lin, Wenjin Ai, Minghao Liu 0021, Shouyi Lu, Hongze Ren, Jingyue Mo, Xiaokai Bai, Zhixiong Ma, Xichan Zhu
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 SADet: A semantic-aware tiny object detection network against missed detection
Dongqi Yuan, Yihan Hu 0006, Beinan Yu, Xiaokai Bai, Si-Yuan Cao, Yuncheng Jin, Bailin Yang
Pattern Recognit.7
2026 Doracamom: Joint 3D Detection and Occupancy Prediction With Multi-View 4D Radars and Cameras for Omnidirectional Perception
Lianqing Zheng, Runwei Guan, Shouyi Lu, Xiaokai Bai, Zhixiong Ma, Xichan Zhu
IEEE Trans. Circuits Syst. Video Technol.7
2026 Remote PPG Measurement Using a Synergistic Time-Frequency Network
abstract
Remote photoplethysmography (rPPG) aims to estimate the blood volume pulse (BVP) signal from facial videos. Existing rPPG approaches still suffer from limitations. We attribute this issue to two primary problems: (1) the reliance solely on time-domain processing that makes the signal susceptible to interference, and (2) the presence of a phase discrepancy between the supervision signal and the ground-truth PPG. To address these problems, we propose TFSNet, a novel time-frequency synergy network for rPPG signal estimation and heart rate prediction. Specifically, we leverage time-frequency fusion (TFF) module, which integrates frequency-domain information into the learning process to enrich the feature representations. Additionally, we introduce the amplitude-phase decoupling (APD) module, which apply phase compensation in frequency domain to mitigate the adverse effects of incorrect phase supervision. Extensive experiments demonstrate that TFSNet achieves state-of-the-art performance, significantly outperforming current approaches in both accuracy and robustness.
Qinglin He, Yuguang Chu, Yuanhui Hu, Xiaokai Bai
IEEE J. Biomed. Health Informatics7
2025 Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction
abstract
This work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate geometric and contextual information within adaptive regions in each image and dynamically adjust the contributions from different views, enhancing the representation capability of voxel features. Furthermore, we propose a sparse volume construction strategy that adaptively identifies and selects voxels with high occupancy probabilities for feature refinement, minimizing redundant computation in free space. Benefiting from the above designs, our framework achieves effective and efficient volume construction in an adaptive way. Better still, our network can be supervised using only 3D bounding boxes, eliminating the dependence on ground-truth scene geometry. Experimental results demonstrate that SGCDet achieves state-of-the-art performance on the ScanNet, ScanNet200 and ARKitScenes datasets. The source code is available at https://github.com/RM-Zhang/SGCDet.
Runmin Zhang, Zhu Yu 0001, Si-Yuan Cao, Lingyu Zhu 0006, Guangyi Zhang 0005, Xiaokai Bai
ICCV6
2025 Structure-Aware Radar-Camera Depth Estimation
abstract
Radar has gained much attention in autonomous driving due to its accessibility and robustness. However, its standalone application for depth perception is constrained by issues of sparsity and noise. Radar-camera depth estimation offers a more promising complementary solution. Despite significant progress, current approaches fail to produce satisfactory dense depth maps, due to the unsatisfactory processing of the sparse and noisy radar data. They constrain the regions of interest for radar points in rigid rectangular regions, which may introduce unexpected errors and confusions. To address these issues, we develop a structure-aware strategy for radar depth enhancement, which provides more targeted regions of interest by leveraging the structural priors of RGB images. Furthermore, we design a Multi-Scale Structure Guided Network to enhance radar features and preserve detailed structures, achieving accurate and structure-detailed dense metric depth estimation. Building on these, we propose a structure-aware radar-camera depth estimation framework, named SA-RCD. Extensive experiments demonstrate that our SA-RCD achieves state-of-the-art performance on the nuScenes dataset. Our code will be available at https://github.com/FreyZhangYeh/SA-RCD.
Fuyi Zhang, Zhu Yu 0001, Chunhao Li, Runmin Zhang, Xiaokai Bai, Si-Yuan Cao
ICRA5
2025 LGDD: Local-Global Synergistic Dual-Branch 3D Object Detection Using 4D Radar
abstract
4D millimeter-wave radar plays a pivotal role in autonomous driving due to its cost-effectiveness and robustness in adverse weather. However, the application of 4D radar point cloud in 3D perception tasks is hindered by its inherent sparsity and noise. To address these challenges, we propose LGDD, a novel local-global synergistic dual-branch 3D object detection framework using 4D radar. Specifically, we first introduce a point-based branch, which utilize a voxel-attended point feature extractor (VPE) to integrate semantic segmentation with cluster voting, thereby mitigating radar noise and extracting local-clustered instances features. Then, for the conventional pillar-based branch, we design a query-based feature pre-fusion (QFP) to address the sparsity and enhance global context representation. Additionally, we devise a proposal mask to filter out noisy points, enabling more focused clustering on regions of interest. Finally, we align the local instances with global context through semantics-geometry aware fusion (SGF) module to achieve comprehensive scene understanding. Extensive experiments demonstrate that LGDD achieves state-of-the-art performance on the public View-of-Delft and TJ4DRadSet datasets. Source code is available at https://github.com/shawnnnkb/LGDD.
Xiaokai Bai, Fuyi Zhang, Si-Yuan Cao, Lianqing Zheng, Beinan Yu
IROS1
2025 Beyond Registration: Self-supervised Unknown Border Completion
Xiaokai Bai, Leyuan Yu, Runmin Zhang, Beinan Yu, Si-Yuan Cao
PRCV (8)1
2025 Boosting Conditional Diffusion Models Using Intermediate Segmentation Map for Infrared Small Target Detection
Wenkai Zhao, Xiaokai Bai, Yicheng Tong, Si-Yuan Cao
PRCV (4)2
2025 TEFormer: Thermal Infrared Image Enhancement by Preserving Spatial Consistency and Details
abstract
Thermal infrared (TIR) images suffer from low contrast due to the atmospheric thermal radiation effect, especially under extreme conditions like low temperature. TIR image enhancement aims to improve image contrast, but previous enhancement approaches usually produce enhanced results with two limitations: spatial inconsistency and detail blurring. To deal with the limitations, we propose a novel TIR image enhancement method, named TEFormer, to preserve spatial consistency and restore fine-grained details during image enhancement. To preserve spatial consistency, we devise the global enhancement module (GEM) to enhance the low-resolution representation. The GEM performs long-range interactions across spatial dimensions and channel dimensions to condition the enhancement curve fitting. To keep details clear, we design the local enhancement module (LEM) as the decoding unit. The LEM injects additional detail structures into the enhanced low-resolution representation for high-resolution reconstruction. Besides, we further apply histogram-based supervision to facilitate learning in intensity distribution of clear images. Extensive experimental results on three challenging benchmarks demonstrate that the proposed method outperforms other state-of-the-art approaches.
Yunxin Li, Runmin Zhang, Si-Yuan Cao, Jiacheng Ying, Xiaokai Bai, Shujie Chen 0001, Bailin Yang
IEEE Trans. Geosci. Remote. Sens.7
2025 GeoFormer: Boosting Object Distinguishing and Prompt Understanding for Cross-View Object Geo-Localization
abstract
Cross-view object geo-localization (CVOGL) determines the geographic location of an object on the satellite view reference image. The object is indicated by a point prompt in a ground- or drone-view query image. Despite wide applications, the current CVOGL method still suffers from limited performance, and the reason for this limitation remains unknown. In this work, we analyze CVOGL and find two primary challenges that hinder its performance,i.e., 1) point prompt understanding and 2) similar-appearance object distinguishing. Therefore, we propose a novel end-to-end framework called object geo-localization transformer (GeoFormer). Specifically, we leverage the knowledge of the segment anything model (SAM) through two settings for accurate point prompt understanding,i.e., using SAM during training and inference (GeoFormer) or using SAM only in training via knowledge distillation (GeoFormer-KD). Additionally, we devise an information aggregation module (IAM) to leverage local and global perception for similar-appearance object distinguishing. Except for the public dataset, we manually annotated a new dataset that contains 1642 image pairs for further comparison. Experiments show that our method significantly outperforms the previous work. Notably, the tiny versions of our method (GeoFormer-t and GeoFormer-t-KD) maintain state-of-the-art performance while substantially reducing parameter costs. Our code and the new dataset will be made available at https://github.com/Temperature-ai/GeoFormer.
Si-Yuan Cao, Zhu Yu 0001, Xiaokai Bai, Beinan Yu
IEEE Trans. Geosci. Remote. Sens.6