Xiaoqiang Teng

dblp:147/1358 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-5876-7605ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Deep Learning for 3-D Lane Detection in Autonomous Driving: A Survey
abstract
3D lane detection has become a critical component in the perception task of autonomous vehicles. Unlike 2D lane detection, which operates in the image plane, 3D lane detection estimates the spatial layout of lanes in real-world coordinates, enabling fine-grained localization, map construction, and planning. However, the task remains challenging due to depth ambiguity, sensor limitations, and diverse road conditions. Existing surveys mostly focus on 2D or organize 3D lane detection by sensor modality, lacking a systematic treatment of algorithmic designs. In this paper, we present a comprehensive survey of deep learning-based 3D lane detection methods. We introduce a dual-axis taxonomy that jointly considers modeling paradigms and representation spaces. Based on this framework, we categorize existing methods into four primary paradigms: geometry-based, end-to-end, query-based, and implicit field-based. We analyze how each paradigm interacts with spatial representations such as image, BEV, 3D, and topological spaces. For each category, we review representative frameworks, architectural principles, and performance trade-offs. We also provide an extensive summary of public datasets, evaluation metrics, and state-of-the-art results across multiple benchmarks. Finally, we identify current limitations and outline future research directions toward robust, scalable, and interpretable 3D lane detection.
Xiaoqiang Teng, Zuo Chen, Shunpeng Chen, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002
IEEE Internet Things J.1
2026 Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout Adapter
abstract
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.7
2025 DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion Refinement
abstract
Inertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty capturing long-range temporal dependencies, limiting their robustness and accuracy. To address these challenges, we propose DiffusionIMU, a novel diffusion-based framework for inertial navigation. DiffusionIMU enhances direct velocity regression from IMU data through an iterative generative denoising process, progressively refining motion state estimation. It integrates the noise-adaptive feature modulation for sensor variability handling, the feature alignment mechanism for representation consistency, and the diffusion-based temporal modeling to decrease accumulated drift. Experiments show that DiffusionIMU consistently outperforms existing methods, demonstrating superior generalization to unseen users while alleviating the impact of the sensor noise.
Xiaoqiang Teng, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002, Weiliang Meng, Xiaopeng Zhang 0001
IJCAI1
2025 ARPDR++: Exploiting local-global temporal modeling for smartphone-based indoor pedestrian localization
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Pengfei Xu 0013, Runbo Hu
Comput. Networks1
2025 VANE-IN: Velocity Auto-Encoder for Inertial Navigation
abstract
Data-driven inertial navigation is crucial for mobile computing applications, such as navigation, augmented reality, and robotics. It typically depends on a trained velocity regression network (VRN) to estimate velocities from inertial measurement unit (IMU) data, enabling position determination through integration. However, using prior velocity information for feature representation in inertial navigation remains underexplored. This work introduces a framework called velocity auto-encoder for inertial navigation (VANE-IN), which employs a Teacher-Student scheme to enhance VRN performance by encoding the velocity. Specifically, a velocity auto-encoder (VANE) is proposed as a student model to distill prior velocity insights from the training dataset, which is then guided by the VRN acting as the teacher model. Additionally, an attention mechanism is introduced to fuse these insights into the features of the VRN. To this end, the VANE-IN achieves a state-of-the-art position accuracy on the RoNIN benchmarks. Our experimental results demonstrate that the VANE-IN achieves approximately 5% performance improvements over existing methods regarding position accuracy.
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Hai-Sheng Li 0002
IEEE Internet Things J.1
2024 MIM-HD: Making Smaller Masked Autoencoder Better with Efficient Distillation
abstract
Self-supervised learning and knowledge distillation intersect to achieve exceptional performance on downstream tasks across diverse network capacities. This paper introduces MIM-HD, which implements enhancements for masked image modeling (MIM) distillation, in two key aspects. First, a vision transformer head-level relation adaptive distillation approach is proposed, allowing the student to dynamically draw multi-source knowledge from the teacher based on its evolving state, compatible with scenarios where teacher-student transformer block head count differs. Second, to address the overemphasis on the encoder and neglect of the decoder role in maintaining representation consistency in previous MIM distillations, a dual-view decoding strategy for latent visual representations is introduced, reusing the teacher’s decoder to alleviate MIM burdens on smaller networks. MIM-HD effectiveness is demonstrated through evaluations on ADE20K (mIoU) and ImageNet-1K (Acc), achieving +1.4% and +0.5% improved performance, respectively, compared to state-of-the-art methods, with substantial advantages on smaller pre-training datasets. Moreover, MIM-HD achieves superior efficiency, reducing pre-training epochs from 300 to 100.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Li Guo 0004, Jiguang Zhang, Xiaoqiang Teng, Wenbo Xu 0003
ECAI8
2024 HCF-Net: Hierarchical Context Fusion Network for Infrared Small Object Detection
abstract
Infrared small object detection is an important computer vision task involving the recognition and localization of tiny objects in infrared images, which usually contain only a few pixels. However, it encounters difficulties due to the diminutive size of the objects and the generally complex backgrounds in infrared images. In this paper, we propose a deep learning method, HCF-Net, that significantly improves infrared small object detection performance through multiple practical modules. Specifically, it includes the parallelized patch-aware attention (PPA) module, dimension-aware selective integration (DASI) module, and multi-dilated channel refiner (MDCR) module. The PPA module uses a multi-branch feature extraction strategy to capture feature information at different scales and levels. The DASI module enables adaptive channel selection and fusion. The MDCR module captures spatial features of different receptive field ranges through multiple depth-separable convolutional layers. Extensive experimental results on the SIRST infrared single-frame image dataset show that the proposed HCF-Net performs well, surpassing other traditional and deep learning models. Code is available at https://github.com/zhengshuchen/HCFNet.
Shibiao Xu, ShuChen Zheng, Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Xiaoqiang Teng, Ao Li 0002, Li Guo 0004
ICME7
2024 DTTCNet: Time-to-Collision Estimation With Autonomous Emergency Braking Using Multi-Scale Transformer Network
abstract
The rapid advancement of autonomous driving technologies has brought the significance of Autonomous Emergency Braking (AEB) systems, which are paramount in mitigating collision risk and elevating road safety by preemptively applying brakes when a potential collision is detected. Within the core mechanisms of AEB systems, the Time-to-Collision (TTC) estimation plays a pivotal role, in quantitatively determining the criticality and timing for initiating braking interventions. However, existing TTC estimation approaches exhibit sensitivity to diverse driving scenarios, compromising the performance of AEB systems, especially in instantaneous situations. To address these issues, this paper presents DTTCNet, a novel supervised deep learning model for TTC estimation that leverages multi-scale transformer architectures and multi-task losses, thereby enhancing precision and boosting system performance. The DTTCNet first extracts spatiotemporal features from raw sensor data and utilizes a supervised training strategy. The multi-scale transformer architecture effectively captures variations across different scales, while the multi-task loss function optimizes the network training performance. Our experimental results on a challenging dataset demonstrate that DTTCNet achieves approximately 20% performance improvements over existing methods in terms of accuracy. This signifies a promising approach to augmenting the safety of autonomous driving systems with the integration of aftermarket mobile devices (e.g., Mobileye and Bosch products).
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Mob. Comput.1
2021 SiFi: Self-Updating of Indoor Semantic Floorplans for Annotated Objects
abstract
Due to the rapid development of indoor location-based services, automatically deriving an indoor semantic floorplan becomes a highly promising technique for ubiquitous applications. To make an indoor semantic floorplan fully practical, it is essential to handle the dynamics of semantic information. Despite several methods proposed for automatic construction and semantic labeling of indoor floorplans, this problem has not been well studied and remains open. In this article, we present a system called SiFi to provide accurate and automatic self-updating service. It updates semantics with instant videos acquired by mobile devices in indoor scenes. First, a crowdsourced-based task model is designed to attract users to contribute semantic-rich videos. Second, we use the maximum likelihood estimation method to solve the text inferring problem as the sequential relationship of texts provides additional geometrical constraints. Finally, we formulate the semantic update as an inference problem to accurately label semantics at correct locations on the indoor floorplans. Extensive experiments have been conducted across 9 weeks in a shopping mall with more than 250 stores. Experimental results show that SiFi achieves 84.5% accuracy of semantic update.
Deke Guo, Xiaoqiang Teng, Yulan Guo, Xiaolei Zhou 0001, Zhong Liu 0002
ACM Trans. Internet Things2
2020 ARPDR: An Accurate and Robust Pedestrian Dead Reckoning System for Indoor Localization on Handheld Smartphones
abstract
The proliferation of mobile computing has prompted Pedestrian Dead Reckoning (PDR) to be one of the most attractive and promising indoor localization techniques for ubiquitous applications. The existing PDR approaches either suffer position drifts caused by accumulative errors or are sensitive to various users. This paper presents ARPDR, an accurate and robust PDR approach to improve the accuracy and robustness of indoor localization methods. Particularly, we propose a novel step counting algorithm based on motion models by deeply exploiting inertial sensor data. We then combine step counting with adaptive thresholding to personalize the PDR system for different users. Furthermore, we propose a novel stride-heading model with a deep neural network to predict stride lengths and walking orientations, thus the displacement errors are significantly reduced. Extensive experiments on public datasets demonstrate that ARPDR outperforms the state-of-the-art PDR methods.
Xiaoqiang Teng, Pengfei Xu 0013, Deke Guo, Yulan Guo, Runbo Hu, Didi Chuxing
IROS1
2019 Enabling entity discovery in indoor commercial environments without pre-deployed infrastructure
Xiaolei Zhou 0001, Xiaoqiang Teng, Deke Guo
Frontiers Comput. Sci.3
2019 CloudNavi: Toward Ubiquitous Indoor Navigation Service with 3D Point Clouds
abstract
The rapid development of mobile computing has prompted indoor navigation to be one of the most attractive and promising applications. Conventional designs of indoor navigation systems depend on either infrastructures or indoor floor maps. This article presents CloudNavi, a ubiquitous indoor navigation solution, which relies on the point clouds acquired by the 3D camera embedded in a mobile device. Particularly, CloudNavi first efficiently infers the walking trace of each user from captured point clouds and inertial data. Many shared walking traces and associated point clouds are combined to generate the point cloud traces, which are then used to generate a 3D path-map. Accordingly, CloudNavi can accurately estimate the location of a user by fusing point clouds and inertial data using a particle filter algorithm and then guiding the user to its destination from its current location. Extensive experiments are conducted on office building and shopping mall datasets. Experimental results indicate that CloudNavi exhibits outstanding navigation performance in both office buildings and shopping malls and obtains around 34% improvement compared with the state-of-the-art method.
Xiaoqiang Teng, Deke Guo, Yulan Guo, Xiaolei Zhou 0001, Zhong Liu 0002
ACM Trans. Sens. Networks1
2018 From one to crowd: a survey on crowdsourcing-based wireless indoor localization
Xiaolei Zhou 0001, Tao Chen 0013, Deke Guo, Xiaoqiang Teng
Frontiers Comput. Sci.4
2018 SISE: Self-Updating of Indoor Semantic Floorplans for General Entities
abstract
Indoor semantic floorplan is important for a range of location based service (LBS) applications, attracting many research efforts in several years. In many cases, the out-of-date indoor semantic floorplans would gradually deteriorate and even break down the LBS performance. Thus, it is important to automatically update changed semantics of indoor floorplans caused by environmental variation. However, few research has been focused on the continuous semantic updating problem. This paper presents SISE as a mobile crowdsourcing system that uses a new abstraction for indoor general entities and their semantics, enGraph, to automatically update changed semantics of indoor floorplans using images and inertial data. We first propose efficient methods to generate enGraph. Thus, an image can be associated with an indoor semantic floorplan. Accordingly, we formulate the enGraph matching problem and then propose a quality-based maximum common subgraph matching algorithm so that entities extracted from an image can be corresponded to entities in the indoor semantic floorplan. Furthermore, we propose a quadrant comparison algorithm and a region shrink based localization algorithm to detect and localize changed entities. Thus, the new semantics can be labeled and out-of-date semantics can be removed. Extensive experiments have been conducted on real and synthetic data. Experimental results show that 80 percent of out-of-date semantics of indoor general entities can be updated by SISE.
Xiaoqiang Teng, Deke Guo, Yulan Guo, Xiang Zhao 0002, Zhong Liu 0002
IEEE Trans. Mob. Comput.1
2017 Source selection problem in multi-source multi-destination multicasting
Deke Guo, Xiaoqiang Teng, Zhiyao Hu, Bangbang Ren
Comput. Networks2
2017 IONavi: An Indoor-Outdoor Navigation Service via Mobile Crowdsensing
abstract
The proliferation of mobile computing has prompted navigation to be one of the most attractive and promising applications. Conventional designs of navigation systems mainly focus on either indoor or outdoor navigation. However, people have a strong need for navigation from a large open indoor environment to an outdoor destination in real life. This article presents IONavi, a joint navigation solution, which can enable passengers to easily deploy indoor-outdoor navigation service for subway transportation systems in a crowdsourcing way. Any self-motivated passenger records and shares individual walking traces from a location inside a subway station to an uncertain outdoor destination within a given range, such as one kilometer. IONavi further extracts navigation traces from shared individual traces, each of which is not necessary to be accurate. A subsequent following user achieves indoor-outdoor navigation services by tracking a recommended navigation trace. Extensive experiments are conducted on a subway transportation system. The experimental results indicate that IONavi exhibits outstanding navigation performance from an uncertain location inside a subway station to an outdoor destination. Although IONavi is to enable indoor-outdoor navigation for subway transportation systems, the basic idea can naturally be extended to joint navigation from other open indoor environments to outdoor environments.
Xiaoqiang Teng, Deke Guo, Yulan Guo, Xiaolei Zhou 0001, Zeliu Ding, Zhong Liu 0002
ACM Trans. Sens. Networks1
2015 Poster: An Indoor-Outdoor Navigation Service for Subway Transportation Systems
abstract
The proliferation of mobile computing has prompted navigation to be one of the most attractive and promising applications. Conventional designs of navigation systems mainly focus either indoor or outdoor navigation. However, people have a strong need for navigation from a large open indoor environment to an outdoor destination in real life. In this poster, we present a joint navigation system, named ioNavi. It can enable passengers to easily deploy indoor-outdoor navigation service for subway transportation systems in a crowdsourcing way, without comprehensive indoor localization systems. Any self-motivated passenger records and shares its individual walking trace and associated rich set of sensor readings, from a location inside a subway station to an uncertain outdoor destination within a given range, such as one kilometer. ioNavi further extracts navigation traces from shared individual traces, each of which is not necessary to be accurate and useful. A subsequent following user achieves indoor-outdoor navigation services by tracking a recommended navigation trace.
Xiaoqiang Teng, Deke Guo, Xiaolei Zhou 0001, Zhong Liu 0002
SenSys1