Xinggang Hu

dblp:280/0534 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-9607-3973ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DyGS-SLAM: Real-Time Accurate Localization and Gaussian Reconstruction for Dynamic Scenes
Xinggang Hu, Chenyangguang Zhang, Yuanze Gui, Xiangkui Zhang, Xiangyang Ji
ICCV1
2025 Dy3DGS-SLAM: Monocular 3D Gaussian Splatting SLAM for Dynamic Environments
abstract
Current Simultaneous Localization and Mapping (SLAM) methods based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting excel in reconstructing static 3D scenes but struggle with tracking and reconstruction in dynamic environments, such as real-world scenes with moving elements. Existing NeRF-based SLAM approaches addressing dynamic challenges typically rely on RGB-D inputs, with few methods accommodating pure RGB input. To overcome these limitations, we propose Dy3DGS-SLAM, the first 3D Gaussian Splatting (3DGS) SLAM method for dynamic scenes using monocular RGB input. To address dynamic interference, we fuse optical flow masks and depth masks through a probabilistic model to obtain a fused dynamic mask. With only a single network iteration, this can constrain tracking scales and refine rendered geometry. Based on the fused dynamic mask, we designed a novel motion loss to constrain the pose estimation network for tracking. In mapping, we use the rendering loss of dynamic pixels, color, and depth to eliminate transient interference and occlusion caused by dynamic objects. Experimental results demonstrate that Dy3DGS-SLAM achieves state-of-the-art tracking and rendering in dynamic environments, outperforming or matching existing RGB-D methods.
Hongxing Zhou, Xinggang Hu, Florian Roemer, Hongyu Wang 0001, Ahmad Osman
ICRA4
2025 DPGS-SLAM: Gaussian Splatting SLAM for Dynamic Scenes with Planar Constraints
Yuanze Gui, Xinggang Hu
PRCV (10)2
2025 DYO-SLAM: Visual Localization and Object Mapping in Dynamic Scenes
abstract
Addressing the impact of dynamic factors on localization accuracy and constructing a long-term consistent map containing only static elements are two crucial tasks in visual simultaneous localization and mapping (SLAM) for dynamic scenes. The introduction of dynamic elements can compromise the geometric constraints essential for visual SLAM, leading to a decrease in localization accuracy. Existing related research faces challenges in simultaneously ensuring localization accuracy in both low-dynamic and high-dynamic scenarios, while also maintaining the system’s real-time performance. To address this issue, we propose a two-stage, coarse-to-fine static-probability-based localization scheme. The construction of object-level maps offers strong support for tasks involving higher-level intelligent agent manipulation as well as augmented reality (AR). However, current research is inadequate for dynamic scenes where the objects to be modeled are frequently and irregularly obscured by dynamic objects, and where there are significant challenges such as severe image and point cloud noise, semantic noise, and lack of observational perspectives. To overcome these challenges, we first propose an object parameter estimation algorithm that combines clustering, weighted Principal Component Analysis (PCA) based on an energy function, and a minimum bounding rectangle. Then, we design a multi-modal object data association strategy based on appearance, semantic, and spatial features. The proposed object parameter estimation algorithm and data association strategy demonstrate improved accuracy and robustness in dynamic scenes with the aforementioned challenges. Finally, based on the entire system, we further develop a dynamic object tracking algorithm and construct an AR system to demonstrate the system’s application prospects. A series of public datasets and real-world scene results have been used to evaluate the effectiveness of the proposed system.
Xinggang Hu, Yanmin Wu, Zhenzhong Cao, Xiangkui Zhang, Xiangyang Ji
IEEE Trans. Circuits Syst. Video Technol.1
2025 PAS-SLAM: A Visual SLAM System for Planar-Ambiguous Scenes
abstract
Visual SLAM (Simultaneous Localization and Mapping) systems based on planar features have been widely applied in fields such as environmental structure perception and augmented reality (AR). However, current research still faces challenges in accurate localization and map construction in planar ambiguous scenes, primarily due to the insufficient accuracy of the planar features and data association methods employed. In this paper, we propose a visual SLAM system based on planar features designed for ambiguous planar scenes, including planar analysis and processing, data association, and multi-constraint factor graph optimization. Initially, we introduce a planar analysis and processing strategy that integrates semantic information to analyze the structure of planes and further refine the selection of planes, providing accurate planar information for subsequent association and optimization processes. Then, we integrate various planar data to propose a multimodal fusion data association strategy, achieving accurate and robust planar data association in ambiguous planar scenes. Finally, based on accurate and rich planar information along with related constraints, we design a set of multi-constraint factor graphs for camera pose optimization. Public datasets and real-world experiments demonstrate that, compared to state-of-the-art related research, our proposed system shows significant competitive advantages in terms of accuracy and robustness for both map construction and camera localization. Regarding quantifiable localization accuracy, our system achieves an average improvement in Absolute Trajectory Error (ATE) of approximately 57% in planar ambiguous scenes and about 25% in non-planar ambiguous scenes. Additionally, the system exhibits great application potential in fields such as augmented reality.
Xinggang Hu, Yanmin Wu, Linghao Yang, Xiangkui Zhang, Xiangyang Ji
IEEE Trans. Circuits Syst. Video Technol.1
2025 Bidirectional Patch-Based Correlations With Local Rigidity for Global Nonrigid Registration
abstract
The registration of time-varying 3D shapes with high degrees of freedom remains a challenging task. Most existing techniques attempt to address this issue by solving an optimization problem defined on deformation graph with as-rigid-as-possible smoothness prior, which usually struggle to capture large scale displacements. Motivated by the insight that a set of points tends to collectively undergo significant rigid motion accompanied by slight nonrigid deformation, we propose a two-step approach to address nonrigid registration in a coarse-to-fine manner. In the first step, coarse correlations between source and target points are constructed by estimating a set of rigid transformations for local patches which are regional clusters of points. To leverage more contextual information, a bidirectional registration module is introduced that estimates both the forward and backward patch-wise rigid transformation fields (PRTFs). Subsequently, in the second step, the source point set is warped by blending both forward and backward PRTFs and fed into a deformation optimization module. Here, unidirectional point-based correspondences are sought to refine the global nonrigid transformation fields (GNTFs) while adhering to local rigidity constraints. To illustrate the efficacy of our method, we conduct tests on challenging scenarios involving human datasets, including large displacements resulting from fast inter-frame motions or pose changes. Both qualitative and quantitative results demonstrate that our approach outperforms several state-of-the-art methods in terms of robustness and registration accuracy.
Xuexin Yu, Xinggang Hu, Long Xu 0001, Xiangyang Ji
IEEE Trans. Circuits Syst. Video Technol.3
2025 Combating Voice Spoofing Attacks on Wearables via Speech Movement Sequences
abstract
Voice assistants, increasingly integrated into wearable devices with limited human-computer interaction modalities, are susceptible to voice spoofing attacks. Such attacks exploit pre-recorded or synthesized voice commands to trick the assistants into executing actions unauthorized by legitimate users. In this work, we propose GyroTalk, a novel approach extracts individual and reliable features from speech movement sequences of users, using built-in gyroscopes in wearables, to differentiate between legitimate users and malicious attackers. GyroTalk is inspired by two critical insights. First, speech, as a highly intricate motor task, necessitates the synchronized coordination of multiple respiratory, laryngeal, lingual and mandibular muscles. These collective muscle movements propagate throughout the body, providing unique movement signatures. Second, the distinctive speech movement sequences of individual speakers, essential for generating specific words, can be grabbed by embedded IMU of wearables. We conduct a comprehensive evaluation of GyroTalk across various COTS Android devices, including smart phones, watches and glasses. Our experimental results demonstrate that GyroTalk can achieve a mean FAR of 2.23% and a FRR of 2.48%, even in the face of complicated voice spoofing attacks.
Shan Chang, Luo Zhou, Wei Liu 0138, Hongzi Zhu, Xinggang Hu, Lei Yang 0025
IEEE Trans. Dependable Secur. Comput.5
2025 UniQuadric: A SLAM Backend for Unknown Rigid Object 3-D Tracking and Light-Weight Modeling
abstract
Tracking and modeling unknown rigid objects in the environment play a crucial role in autonomous uncrewed systems and virtual-real interactive applications. However, many existing Simultaneous Localization, Mapping and Moving Object Tracking (SLAMMOT) methods focus solely on estimating specific object poses and lack estimation of object scales and are unable to effectively track unknown objects. In this paper, we propose a novel SLAM backend that unifies ego-motion tracking, rigid object motion tracking, and modeling within a joint optimization framework. In the perception part, we designed a pixel-level asynchronous object tracker (AOT) based on the Segment Anything Model (SAM) and DeAOT, enabling the tracker to effectively track target unknown objects guided by various predefined tasks and prompts. In the modeling part, we present a novel object-centric quadric parameterization to unify both static and dynamic object initialization and optimization. Subsequently, in the part of object state estimation, we propose a tightly coupled optimization model for object pose and scale estimation, incorporating hybrids constraints into a novel dual sliding window optimization framework for joint estimation. To our knowledge, we are the first to tightly couple object pose tracking with light-weight modeling of dynamic and static objects using quadric. We conduct qualitative and quantitative experiments on simulation datasets and real-world datasets, demonstrating the state-of-the-art robustness and accuracy in motion estimation and modeling. This showcases the significant potential of our method for object perception in complex dynamic environments.
Linghao Yang, Yanmin Wu, Rui Tian 0002, Xinggang Hu
IEEE Trans. Intell. Transp. Syst.5
2022 CFP-SLAM: A Real-time Visual SLAM Based on Coarse-to-Fine Probability in Dynamic Environments
abstract
The dynamic factors in the environment will lead to the decline of camera localization accuracy due to the violation of the static environment assumption of SLAM algorithm. Recently, some related works generally use the combination of semantic constraints and geometric constraints to deal with dynamic objects, but problems can still be raised, such as poor real-time performance, easy to treat people as rigid bodies, and poor performance in low dynamic scenes. In this paper, a dynamic scene-oriented visual SLAM algorithm based on object detection and coarse-to-fine static probability named CFP-SLAM is proposed. The algorithm combines semantic constraints and geometric constraints to calculate the static probability of objects, keypoints and map points, and takes them as weights to participate in camera pose estimation. Extensive evaluations show that our approach can achieve almost the best results in high dynamic and low dynamic scenarios compared to the state-of-the-art dynamic SLAM methods, and shows quite high real-time ability.
Xinggang Hu, Yunzhou Zhang, Zhenzhong Cao, Yanmin Wu, Zhiqiang Deng, Wenkai Sun
IROS1
2022 VOGUE: Secure User Voice Authentication on Wearable Devices using Gyroscope
abstract
Voice assistants are popular to wearable devices with limited input and output capabilities, however vulnerable to voice attacks, which cheat a voice assistant by playing forged voice commands without user awareness. In this paper, we propose VOGUE, which captures unique yet stable pattern of speech movement sequences of speakers with embedded gyroscope in wearable devices, to distinguish between registered legal user and malicious attackers (human or machines). The design of VOGUE is based on two key observations. First, speech, as a type of highly complex motor task, inherently requires coordinated actions of many orofacial, laryngeal, pharyngeal, and respiratory muscles, and the collective movements of muscles propagate to distant body segments. Second, to generate a certain word, the speech movement sequence of a speaker is known to be distinctive, and can be captured by inertial sensors. We implement VOGUE on three kinds of COTS android devices including smart glasses, watches and phones, and conduct comprehensive evaluation on the performances. Experimental results show that VOGUE achieves a mean false-acceptance rate (FAR) and false- rejection rate (FRR) of 2.23% and 2.48%, respectively, even under sophisticated voice impersonation attacks.
Shan Chang, Xinggang Hu, Hongzi Zhu, Wei Liu 0138, Lei Yang 0025
SECON2
2021 Object SLAM-Based Active Mapping and Robotic Grasping
abstract
This paper presents the first active object mapping framework for complex robotic manipulation and autonomous perception tasks. The framework is built on an object SLAM system integrated with a simultaneous multi-object pose estimation process that is optimized for robotic grasping. Aiming to reduce the observation uncertainty on target objects and increase their pose estimation accuracy, we also design an object-driven exploration strategy to guide the object mapping process, enabling autonomous mapping and high-level perception. Combining the mapping module and the exploration strategy, an accurate object map that is compatible with robotic grasping can be generated. Additionally, quantitative evaluations also indicate that the proposed framework has a very high mapping accuracy. Experiments with manipulation (including object grasping and placement) and augmented reality significantly demonstrate the effectiveness and advantages of our proposed framework.
Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Sonya A. Coleman, Wenkai Sun, Xinggang Hu, Zhiqiang Deng
3DV7
2021 CoSafe: Securing Mobile Devices through Mutual Mobility Consistency Verification
abstract
As mobile devices play increasingly important roles in our daily lives, it is of great significance to protect personal mobile devices from being lost. Noticing the trend that one person normally carries more than one mobile device, we propose an innovative scheme, calledCoSafe, to detect device loss by verifying the motion consistency between a pair of devices. The rationale is that the vibrations perceived on devices carried by the same person should be tightly coupled whereas a lost device would show distinct mobility characteristics from others. Specifically, CoSafe compares the mobility consistency between a pair of devices on three levels, where coarse features (i.e., the mobility state and motion periodicity) are first compared to give fast response and more complex comparison on subtle feature (i.e., the relative phase) is conducted only when needed. In this way, CoSafe can instantly respond and introduce very low computation and communication costs. We implement CoSafe on a Commercial-Off-The-Shelf Android smartphone and a smartwatch, and conduct both trace-driven simulations and real-world experiments to evaluate the performance of CoSafe. The results show that CoSafe achieves a mean false negative ratio and false positive ratio of 1.46 and 3.12 percent, respectively, even under sophisticated stealing attacks.
Shan Chang, Hongzi Zhu, Xinggang Hu
IEEE Trans. Mob. Comput.4