EDBT 2026 Demo / reviewers in the wild / expert
Xiaokai Bai
dblp:139/1989
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0002-8382-2976ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 69% Vision and language · 20% Autonomous driving · 9% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d object detection |
1.9 | 2 | 2026 | OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026 Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025 |
Computer vision › 3D vision › visual localization
cross-view geo-localization |
1.0 | 1 | 2026 | Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework · AAAI 2026 |
Computer vision › Vision and language
multimodal prompt learning |
1.0 | 1 | 2026 | Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework · AAAI 2026 |
Computer vision › Vision and language
vision-language dataset |
1.0 | 1 | 2026 | OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | Structure-Aware Radar-Camera Depth Estimation · ICRA 2025 |
Computer vision › 3D vision › 3d object detection
multi-view indoor 3d object detection |
0.9 | 1 | 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025 |
Robotics › Autonomous driving
perception |
0.9 | 1 | 2025 | Structure-Aware Radar-Camera Depth Estimation · ICRA 2025 |
Computer vision › 3D vision › depth estimation › depth completion
radar-camera depth estimation |
0.9 | 1 | 2025 | Structure-Aware Radar-Camera Depth Estimation · ICRA 2025 |
Computer vision › 3D vision › 3d scene understanding
semantic scene completion |
0.3 | 1 | 2026 | OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Robotics › Robot navigation and mapping › localization › vehicle localization
UAV localization |
0.3 | 1 | 2026 | Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework · AAAI 2026 |
Computer vision › 3D vision
3d reconstruction |
0.3 | 1 | 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume Construction · ICCV 2025 |
Image and video processing › image enhancement
feature enhancement |
0.3 | 1 | 2025 | Structure-Aware Radar-Camera Depth Estimation · ICRA 2025 |
Methods — techniques the papers use, named apart from their topics
radar depth enhancement · 1.7multi-scale structure guided network · 1.7transformer · 1.0surround-view camera · 1.0cross-attention · 1.0contrastive learning · 1.0LiDAR · 1.04d imaging radar · 1.0sparse volume construction · 0.9geometry and context aware aggregation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT FrameworkabstractWe present MoP-UAV, a new benchmark for UAV-based cross-view object geo-localization guided by multi-modal prompts. MoP-UAV supports fine-grained object-level cross-view localization under diverse prompt modalities, including natural language, bounding boxes, and click points. It offers potential for incorporating large foundation models like large language models (LLMs) and promotes the building of more flexible and intelligent UAV agents. Based on the benchmark, we propose MoPT, a multi-modal-prompt-guided tansformer that embeds prompts as token sequences and extract object location from UAV and satellite features via cross-attention. To enhance semantic consistency and performance, we further adopt a cross-view contrastive loss and propose a RefCOCOg-based pre-training strategy. Extensive experiments show that MoPT achieves robust localization under arbitrary prompt combinations. Notably, multi-modal-prompt training significantly boosts unimodal-prompt inference performance, highlighting the generalization benefits of multi-modal learning. MoPT trained with multi-modal prompts outperforms prior unimodal prompt works under the same setting. Zhangkai Shen, Si-Yuan Cao, Xiaokai Bai, Zheheng Han, Qi Ming |
AAAI | 4 |
| 2026 | OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous DrivingabstractThe rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions. Next-generation autonomous driving datasets must be multimodal, incorporating data from advanced sensors that feature extensive data coverage, detailed annotations, and diverse scene representation. To address this need, we present OmniHD-Scenes, a large-scale multimodal dataset that provides comprehensive omnidirectional high-definition data. The OmniHD-Scenes dataset combines data from 128-beam LiDAR, six cameras, and six 4D imaging radar systems to achieve full environmental perception. The dataset comprises 1501 clips, each approximately 30-s long, totaling more than 450 K synchronized frames and more than 5.85 million synchronized sensor data points. We also propose a novel 4D annotation pipeline. To date, we have annotated 200 clips with more than 514 K precise 3D bounding boxes. These clips also include semantic segmentation annotations for static scene elements. Additionally, we introduce a novel automated pipeline for generation of the dense occupancy ground truth, which effectively leverages information from non-key frames. Alongside the proposed dataset, we establish comprehensive evaluation metrics, baseline models, and benchmarks for 3D detection and semantic occupancy prediction. These benchmarks utilize surround-view cameras and 4D imaging radar to explore cost-effective sensor solutions for autonomous driving applications. Extensive experiments demonstrate the effectiveness of our low-cost sensor configuration and its robustness under adverse conditions. Lianqing Zheng, Qunshu Lin, Wenjin Ai, Minghao Liu 0021, Shouyi Lu, Hongze Ren, Jingyue Mo, Xiaokai Bai, Zhixiong Ma, Xichan Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | SADet: A semantic-aware tiny object detection network against missed detection
Dongqi Yuan, Yihan Hu 0006, Beinan Yu, Xiaokai Bai, Si-Yuan Cao, Yuncheng Jin, Bailin Yang |
Pattern Recognit. | 7 |
| 2026 | Doracamom: Joint 3D Detection and Occupancy Prediction With Multi-View 4D Radars and Cameras for Omnidirectional Perception
Lianqing Zheng, Runwei Guan, Shouyi Lu, Xiaokai Bai, Zhixiong Ma, Xichan Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Remote PPG Measurement Using a Synergistic Time-Frequency NetworkabstractRemote photoplethysmography (rPPG) aims to estimate the blood volume pulse (BVP) signal from facial videos. Existing rPPG approaches still suffer from limitations. We attribute this issue to two primary problems: (1) the reliance solely on time-domain processing that makes the signal susceptible to interference, and (2) the presence of a phase discrepancy between the supervision signal and the ground-truth PPG. To address these problems, we propose TFSNet, a novel time-frequency synergy network for rPPG signal estimation and heart rate prediction. Specifically, we leverage time-frequency fusion (TFF) module, which integrates frequency-domain information into the learning process to enrich the feature representations. Additionally, we introduce the amplitude-phase decoupling (APD) module, which apply phase compensation in frequency domain to mitigate the adverse effects of incorrect phase supervision. Extensive experiments demonstrate that TFSNet achieves state-of-the-art performance, significantly outperforming current approaches in both accuracy and robustness. Qinglin He, Yuguang Chu, Yuanhui Hu, Xiaokai Bai |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionabstractThis work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate geometric and contextual information within adaptive regions in each image and dynamically adjust the contributions from different views, enhancing the representation capability of voxel features. Furthermore, we propose a sparse volume construction strategy that adaptively identifies and selects voxels with high occupancy probabilities for feature refinement, minimizing redundant computation in free space. Benefiting from the above designs, our framework achieves effective and efficient volume construction in an adaptive way. Better still, our network can be supervised using only 3D bounding boxes, eliminating the dependence on ground-truth scene geometry. Experimental results demonstrate that SGCDet achieves state-of-the-art performance on the ScanNet, ScanNet200 and ARKitScenes datasets. The source code is available at https://github.com/RM-Zhang/SGCDet. Runmin Zhang, Zhu Yu 0001, Si-Yuan Cao, Lingyu Zhu 0006, Guangyi Zhang 0005, Xiaokai Bai |
ICCV | 6 |
| 2025 | Structure-Aware Radar-Camera Depth EstimationabstractRadar has gained much attention in autonomous driving due to its accessibility and robustness. However, its standalone application for depth perception is constrained by issues of sparsity and noise. Radar-camera depth estimation offers a more promising complementary solution. Despite significant progress, current approaches fail to produce satisfactory dense depth maps, due to the unsatisfactory processing of the sparse and noisy radar data. They constrain the regions of interest for radar points in rigid rectangular regions, which may introduce unexpected errors and confusions. To address these issues, we develop a structure-aware strategy for radar depth enhancement, which provides more targeted regions of interest by leveraging the structural priors of RGB images. Furthermore, we design a Multi-Scale Structure Guided Network to enhance radar features and preserve detailed structures, achieving accurate and structure-detailed dense metric depth estimation. Building on these, we propose a structure-aware radar-camera depth estimation framework, named SA-RCD. Extensive experiments demonstrate that our SA-RCD achieves state-of-the-art performance on the nuScenes dataset. Our code will be available at https://github.com/FreyZhangYeh/SA-RCD. Fuyi Zhang, Zhu Yu 0001, Chunhao Li, Runmin Zhang, Xiaokai Bai, Si-Yuan Cao |
ICRA | 5 |
| 2025 | LGDD: Local-Global Synergistic Dual-Branch 3D Object Detection Using 4D Radarabstract4D millimeter-wave radar plays a pivotal role in autonomous driving due to its cost-effectiveness and robustness in adverse weather. However, the application of 4D radar point cloud in 3D perception tasks is hindered by its inherent sparsity and noise. To address these challenges, we propose LGDD, a novel local-global synergistic dual-branch 3D object detection framework using 4D radar. Specifically, we first introduce a point-based branch, which utilize a voxel-attended point feature extractor (VPE) to integrate semantic segmentation with cluster voting, thereby mitigating radar noise and extracting local-clustered instances features. Then, for the conventional pillar-based branch, we design a query-based feature pre-fusion (QFP) to address the sparsity and enhance global context representation. Additionally, we devise a proposal mask to filter out noisy points, enabling more focused clustering on regions of interest. Finally, we align the local instances with global context through semantics-geometry aware fusion (SGF) module to achieve comprehensive scene understanding. Extensive experiments demonstrate that LGDD achieves state-of-the-art performance on the public View-of-Delft and TJ4DRadSet datasets. Source code is available at https://github.com/shawnnnkb/LGDD. Xiaokai Bai, Fuyi Zhang, Si-Yuan Cao, Lianqing Zheng, Beinan Yu |
IROS | 1 |
| 2025 | Beyond Registration: Self-supervised Unknown Border Completion
Xiaokai Bai, Leyuan Yu, Runmin Zhang, Beinan Yu, Si-Yuan Cao |
PRCV (8) | 1 |
| 2025 | Boosting Conditional Diffusion Models Using Intermediate Segmentation Map for Infrared Small Target Detection
Wenkai Zhao, Xiaokai Bai, Yicheng Tong, Si-Yuan Cao |
PRCV (4) | 2 |
| 2025 | TEFormer: Thermal Infrared Image Enhancement by Preserving Spatial Consistency and DetailsabstractThermal infrared (TIR) images suffer from low contrast due to the atmospheric thermal radiation effect, especially under extreme conditions like low temperature. TIR image enhancement aims to improve image contrast, but previous enhancement approaches usually produce enhanced results with two limitations: spatial inconsistency and detail blurring. To deal with the limitations, we propose a novel TIR image enhancement method, named TEFormer, to preserve spatial consistency and restore fine-grained details during image enhancement. To preserve spatial consistency, we devise the global enhancement module (GEM) to enhance the low-resolution representation. The GEM performs long-range interactions across spatial dimensions and channel dimensions to condition the enhancement curve fitting. To keep details clear, we design the local enhancement module (LEM) as the decoding unit. The LEM injects additional detail structures into the enhanced low-resolution representation for high-resolution reconstruction. Besides, we further apply histogram-based supervision to facilitate learning in intensity distribution of clear images. Extensive experimental results on three challenging benchmarks demonstrate that the proposed method outperforms other state-of-the-art approaches. Yunxin Li, Runmin Zhang, Si-Yuan Cao, Jiacheng Ying, Xiaokai Bai, Shujie Chen 0001, Bailin Yang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | GeoFormer: Boosting Object Distinguishing and Prompt Understanding for Cross-View Object Geo-LocalizationabstractCross-view object geo-localization (CVOGL) determines the geographic location of an object on the satellite view reference image. The object is indicated by a point prompt in a ground- or drone-view query image. Despite wide applications, the current CVOGL method still suffers from limited performance, and the reason for this limitation remains unknown. In this work, we analyze CVOGL and find two primary challenges that hinder its performance,i.e., 1) point prompt understanding and 2) similar-appearance object distinguishing. Therefore, we propose a novel end-to-end framework called object geo-localization transformer (GeoFormer). Specifically, we leverage the knowledge of the segment anything model (SAM) through two settings for accurate point prompt understanding,i.e., using SAM during training and inference (GeoFormer) or using SAM only in training via knowledge distillation (GeoFormer-KD). Additionally, we devise an information aggregation module (IAM) to leverage local and global perception for similar-appearance object distinguishing. Except for the public dataset, we manually annotated a new dataset that contains 1642 image pairs for further comparison. Experiments show that our method significantly outperforms the previous work. Notably, the tiny versions of our method (GeoFormer-t and GeoFormer-t-KD) maintain state-of-the-art performance while substantially reducing parameter costs. Our code and the new dataset will be made available at https://github.com/Temperature-ai/GeoFormer. Si-Yuan Cao, Zhu Yu 0001, Xiaokai Bai, Beinan Yu |
IEEE Trans. Geosci. Remote. Sens. | 6 |