Xuanbin Wang

dblp:334/3099 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0004-0849-556XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 59% Generative modeling · 41%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.722025
No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models · NeurIPS 2025
Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation · ICCV 2025
Computer vision › Segmentation and scene understanding
3d semantic segmentation
0.912025
No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models · NeurIPS 2025
Computer vision › Segmentation and scene understanding › 3d semantic segmentation
cross-domain 3d semantic segmentation
0.912025
No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features
0.912025
No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models · NeurIPS 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.912025
Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation · ICCV 2025
Computer vision › Segmentation and scene understanding
semantic segmentation
0.912025
Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation · ICCV 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.7object agent queries · 0.9cross-modal learning · 0.9
YearPublicationVenuePosition
2026 GREAT Dataset: A Multi-Sensor Raw Observation Dataset for High-Precision Urban Navigation
Chunxi Xia, Xuanbin Wang, Yuxuan Zhou 0001
IEEE Trans Autom. Sci. Eng.4
2025 Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic Segmentation
Xuanbin Wang, Yuelei Xu
ICCV2
2025 No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion Models
abstract
Enhancing the cross-domain generalization of 3D semantic segmentation is a pivotal task in computer vision that has recently gained increasing attention. Most existing methods, whether using consistency regularization or cross-modal feature fusion, focus solely on individual objects while overlooking implicit semantic dependencies among them, resulting in the loss of useful semantic information. Inspired by the diffusion model's ability to flexibly compose diverse objects into high-quality images across varying domains, we seek to harness its capacity for capturing underlying contextual distributions and spatial arrangements among objects to address the challenging task of cross-domain 3D semantic segmentation. In this paper, we propose a novel cross-modal learning framework based on diffusion models to enhance the generalization of 3D semantic segmentation, named XDiff3D. XDiff3D comprises three key ingredients: (1) constructing object agent queries from diffusion features to aggregate instance semantic information; (2) decoupling fine-grained local details from object agent queries to prevent interference with 3D semantic representation; (3) leveraging object agent queries as an interface to enhance the modeling of object semantic dependencies in 3D representations. Extensive experiments validate the effectiveness of our method, achieving state-of-the-art performance across multiple benchmarks in different task settings. Code is available at \url{https://github.com/FanLiHub/XDiff3D}.
Xuanbin Wang, Yuelei Xu
NeurIPS3
2025 Accurate and Capable GNSS-Inertial-Visual Vehicle Navigation via Tightly Coupled Multiple Homogeneous Sensors
abstract
Continuous and reliable estimation of navigation states is of paramount importance in ensuring the safe operation of intelligent vehicles. The conventional global navigation satellite system (GNSS)-Inertial-Visual navigation systems have demonstrated the capability to achieve locally accurate and globally drift-free pose estimation. However, challenges such as frequent satellite signal interference, limited visual features, and sudden sensor failures can severely degrade performance or even lead to complete system collapse when utilizing minimal sensor configurations. For this reason, we propose a novel and globally drift-free tightly coupled (TC) system that can integrate any number of GNSS, inertial measurement units (IMUs), and cameras to enable accurate and robust vehicle navigation. Specifically, a stacked state estimator centered on the IMU is designed to fuse information from all sensors at the raw measurement level. The pseudorange and carrier phase measurements from all GNSS terminals are directly correlated with the core IMU, ensuring accurate and fast positioning of the system in a global frame. The feature measurements from multiple independent cameras are also used to update the states by exploiting inter-epoch geometric constraints. In addition, the multiple homogeneous IMUs can not only further improve the state estimation of the system by imposing rigid constraints on the core IMU, but also switch over in time for smooth state estimation when the core IMU are faulty. We comprehensively evaluate the state estimation accuracy and robustness of the proposed approach through a series of in-vehicle experiments and simulation experiments in real urban scenarios. The results indicate that the proposed system can achieve 93.2% availability with a horizontal position error less than 0.5 m and 97.8% availability with a heading error less than 0.2 deg in typical urban environments, significantly outperforming both conventional and state-of-the-art approaches. Note to Practitioners—This study focuses on the tight integration of multiple homogeneous and heterogeneous sensors with the goal of addressing frequent interference and degradation challenges in wide-area vehicle navigation applications. We propose a general-purpose GNSS-Inertial-Visual tight coupled framework capable of integrating any number of GNSS, IMUs, and cameras at the raw measurement level. It maximizes the use of as much sensor information as possible to achieve accurate and robust state estimation and is resilient to anomalous measurements and sensor unavailability. This solution holds practical and effective for autonomous vehicles that are now commonly equipped with multiple sensors.
Zhiheng Shen, Yuxuan Zhou 0001, Zongzhou Wu, Xuanbin Wang
IEEE Trans Autom. Sci. Eng.6
2024 A Novel Factor Graph Framework for Tightly Coupled GNSS/INS Integration With Carrier-Phase Ambiguity Resolution
abstract
Accurate position, velocity and orientation are essential for the autonomous navigation of unmanned vehicles. The integration of GNSS and INS that can deliver continuous navigation states is widely used for intelligent vehicle systems. In this paper, we propose a tightly coupled GNSS/INS positioning framework with carrier-phase ambiguity resolution based on factor graph optimization (FGO). In this approach, a sliding window optimizer is employed to fuse the multi-GNSS pseudorange and carrier-phase observations with inertial measurements. The same ambiguity within the window is considered as a state node, and the constraint on the ambiguity is continuously preserved by marginalization. To further improve the accuracy and reliability of precise positioning, the carrier-phase ambiguity resolution is introduced to the FGO-based GNSS/INS framework. When the vehicle is detected to be stationary, a zero-velocity constraint and an attitude invariant constraint will be imposed. Several experimental results indicate that the proposed method can accomplish the centimeter-level position estimation performance with beyond 90% positioning availability (horizontal$<$10 cm and vertical$<$10 cm) and outperforms the current state-of-the-art filter-based tightly coupled method.
Zhiheng Shen, Xuanbin Wang, Zongzhou Wu, Xin Li 0117, Yuxuan Zhou 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Ground-VIO: Monocular Visual-Inertial Odometry With Online Calibration of Camera-Ground Geometric Parameters
abstract
Monocular visual-inertial odometry (VIO) is a low-cost solution to provide high-accuracy, low-drifting pose estimation. However, it encounters challenges in vehicular scenarios, as the restricted motion of a ground vehicle could lead to degraded observability, and a lack of stable features might occur in dynamic road environments. In this paper, we propose Ground-VIO, which utilizes ground features and the specific camera-ground geometry to enhance monocular VIO performance in realistic road environments. In the method, the camera-ground geometry is modeled with vehicle-centered parameters and integrated into an optimization-based VIO framework. These parameters could be calibrated online and simultaneously improve the odometry accuracy by providing stable scale-awareness. Besides, a specially designed visual front-end is developed to stably extract and track ground features via the inverse perspective mapping (IPM) technique. Both real-world experiments and tests on public datasets are conducted to verify the effectiveness of the proposed method. The results show that our implementation could dramatically improve monocular VIO accuracy in vehicular scenarios, achieving comparable performance to state-of-art stereo VIO solutions and showing good robustness in challenging conditions. The system can also be used for the auto-calibration of IPM which is widely used in vehicle perception. A toolkit for ground feature processing, together with the experimental datasets, has been made open-source.
Yuxuan Zhou 0001, Xuanbin Wang, Zhiheng Shen
IEEE Trans. Intell. Transp. Syst.4