EDBT 2026 Demo / reviewers in the wild / expert
Xinhu Zheng
dblp:175/5373
· DBLP profile ↗
34ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-9898-5543ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZeRCP: Towards Communication-Efficient Collaborative Perception and Future Scene Prediction via Request-Free Spatial FilteringabstractMulti-Agent collaboration addresses inherent limitations of individual agent systems, including limited sensing range and occlusion-induced blind spots. Despite significant progress, persistent challenges such as constrained communication bandwidth and under-explored subsequent extensions still hinder real-time deployment and further developments of collaborative autonomous driving systems. In this work, we propose ZeRCP, a unified communication-efficient framework that bridges collaborative perception with future scene prediction. Specifically, (i) we devise a plug-and-play request-free spatial filtering module (ZeroR) that eliminates the reliance on request maps while preserving inter-agent spatial complementarity modeling. This approach further reduce communication latency and bandwidth consumptions. (ii) We design a multi-scale pyramidal prediction network anchored by a novel Spatial-Temporal Deformable Attention (STDA) module, extending frame-wise detection to multi-frame predictions. This method adeptly models spatiotemporal dynamics without relying on auto-regressive recursion. We evaluate our method on a large-scale dataset in challenging semantic segmentation and scene prediction tasks. Extensive experiments demonstrate the superiority and effectiveness of ZeRCP in bandwidth-constrained collaboration scenarios and spatiotemporal prediction applications. Yuzhe Ji, Xiaoyun Qiu, Ying-Cong Chen, Xinhu Zheng |
AAAI | 6 |
| 2026 | RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object DetectionabstractMulti-view 3D detection with bird’s eye view (BEV) is crucial for autonomous driving and robotics, but its robustness in real-world is limited as it struggles to predict accurate depth values. A mainstream solution, cross-modal distillation, transfers depth information from LiDAR to camera models but also unintentionally transfers depth-irrelevant information (e.g. LiDAR density). To mitigate this issue, we propose RayD3D, which transfers crucial depth knowledge along the ray: a line projecting from the camera to true location of an object. It is based on the fundamental imaging principle that predicted location of this object can only vary along this ray, which is finally determined by predicted depth value. Therefore, distilling along the ray enables more effective depth information transfer. More specifically, we design two ray-based distillation modules. Ray-based Contrastive Distillation (RCD) incorporates contrastive learning into distillation by sampling along the ray to learn how LiDAR accurately locates objects. Ray-based Weighted Distillation (RWD) adaptively adjusts distillation weight based on the ray to minimize the interference of depth-irrelevant information in LiDAR. For validation, we widely apply RayD3D into three representative types of BEV-based models, including BEVDet, BEVDepth4D, and BEVFormer. Our method is trained on clean NuScenes, and tested on both clean NuScenes and RoboBEV with a variety types of data corruptions. Our method significantly improves the robustness of all the three base models in all scenarios without increasing inference costs, and achieves the best when compared to recently released multi-view and distillation models. Zhaonian Kuang, Zongwei Zhou, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
AAAI | 5 |
| 2026 | Towards Robust Event-Based Depth Estimation: Bridging Synthetic and Real Domains with Motion AdaptationabstractEvent cameras provide microsecond latency and high dynamic range, making them ideal for 3D perception tasks in traffic scenes with challenging lighting conditions. Yet existing methods often struggle to generalize to out-of-domain environments due to the limited availability of diverse training data. While synthetic data offers an easily accessible alternative, it introduces a significant sim-to-real gap, particularly in motion patterns. We tackle this challenge by introducing Motion-Adaptation Mamba (MA-Mamba), a dual-track framework that advances both architecture and data augmentation. At the architectural level, we introduce a lightweight Spatio-Temporal Association module that captures motion-induced appearance variations at arbitrary scales, and an Adaptive Memory Balancing module, built on the Mamba state-space framework, that adaptively filters memory updates to maintain stable scene context under diverse dynamics. At the data level, we design event-oriented augmentations that simulate varied motion patterns and apply priority-based masked sequence modeling to strengthen long-range spatio-temporal reasoning. Trained solely on synthetic data, MA-Mamba delivers substantial zero-shot gains on multiple real-world benchmarks, demonstrating strong robustness and generalizability. Yuzhe Ji, Xiang Cheng 0001, Liuqing Yang 0001, Xinhu Zheng |
AAAI | 6 |
| 2026 | Adaptive-Smooth LiDAR-Camera Knowledge Distillation with Heterogeneous Fusion for Multi-View 3D Object DetectionabstractMulti-view 3D object detection has garnered increasing attention, particularly due to its success in autonomous driving systems. Although multi-view systems possess rich semantic information, their spatial-geometric reasoning capabilities remain limited. Recent studies employ simulated point cloud generation mechanisms to facilitate LiDAR-camera multi-modal knowledge distillation, achieving formal structural consistency. Despite advancements, these methods still face two main issues: i) alignment challenges caused by discrepancies between LiDAR and camera data, and ii) prediction errors from simulated point clouds that compromise the semantic information extracted from images during fusion. To address these problems, we propose adaptive-smooth distillation to optimize alignment granularity based on feature discrepancies for improved LiDAR-camera knowledge distillation. Specifically, this work considers both LIDAR-to-camera cross-modal distillation and LiDAR-camera fusion to simulated point cloud-camera fusion multi-modal distillation. Then, we introduce a heterogeneous fusion module to strategically bias the fusion process toward the extracted camera features, thereby enhancing the robustness of the fusion feature. Additionally, soft-weighted response distillation is proposed to facilitate the student model to selectively mimic the high-quality output of the teacher model. Extensive experiments have demonstrated the superiority of our method, achieving statistically significant improvements of 4.9% in mean Average Precision (mAP) and 4.5% in NuScenes Detection Score (NDS) over the benchmark. Rui Zhao 0029, Shuoyao Wang, Xinhu Zheng, Shijian Gao |
AAAI | 3 |
| 2026 | ClustView: Point clustering and depth view fusion for point cloud analysis
Xiaoyang Xiao, Yuanbo Chen, Runzhao Yao, Jue Jiang, Xinhu Zheng, Shaoyi Du, Long Guo |
Expert Syst. Appl. | 6 |
| 2026 | Object-Scene-Camera Decomposition and Recomposition for Data Efficient Monocular 3D Object Detection
Zhaonian Kuang, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
Mengqi Lei, Siqi Li 0001, Xinhu Zheng, Shaoyi Du, Yue Gao 0002 |
Int. J. Comput. Vis. | 4 |
| 2026 | Multi-Modal Decouple and Recouple Network for Robust 3D Object DetectionabstractMulti-modal 3D object detection with bird’s eye view (BEV) has achieved desired advances on benchmarks. Nonetheless, the accuracy may drop significantly in the real world due to data corruption such as sensor configurations for LiDAR and scene conditions for camera. One design bottleneck of previous models resides in the tightly coupling of multi-modal BEV features during fusion, which may degrade the overall system performance if one modality or both is corrupted. To mitigate, we propose a Multi-Modal Decouple and Recouple Network for robust 3D object detection under data corruption. Different modalities commonly share some high-level invariant features. We observe that these invariant features across modalities do not always fail simultaneously, because different types of data corruption affect each modality in distinct ways. These invariant features can be recovered across modalities for robust fusion under data corruption. To this end, we explicitly decouple Camera/LiDAR BEV features into modality-invariant and modality-specific parts. It allows invariant features to compensate each other while mitigates the negative impact of a corrupted modality on the other. We then recouple these features into three experts to handle different types of data corruption, respectively, i.e., LiDAR, camera, and both. For each expert, we use modality-invariant features as robust information, while modality-specific features serve as a complement. Finally, we adaptively fuse the three experts to exact robust features for 3D object detection. For validation, we collect a benchmark with a large quantity of data corruption for LiDAR, camera, and both based on nuScenes. Our model is trained on clean nuScenes and tested on all types of data corruption. Our model consistently achieves the best accuracy on both corrupted and clean data compared to recent models. Zhaonian Kuang, Yuzhe Ji, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Spatial-Aware and Viewpoint-Robust Vision-Language NavigationabstractVision-language navigation (VLN) requires an agent to follow visual observations and language instructions to navigate within a 3D environment. Most prior VLN approaches utilize recurrent units, topological graphs, or grid maps to represent the agent’s previously explored environment. However, these methods face two key limitations: (1) difficulty in accurately and adequately representing multi-level spatial information, and (2) a lack of robustness in semantic feature extraction against viewpoint variations. To overcome these challenges, we design a novel spatial-aware scene representation (SSR) as well as a viewpoint-robust feature extraction method. Specifically, for SSR, we first construct a multi-layer map that models the scene at different granularities, incorporating waypoints, objects, rooms, and floors. We then propose to extract spatial features from this multi-layer map based on a heterogeneous graph transformer and align them with the instructions, effectively guiding the agent’s planning process. As to viewpoint-robust feature extraction, we integrate 3D Gaussian splatting to capture the 3D geometry and visual texture of the environment, enabling novel view object observation and feature extraction. This improves observation coverage and enhances the viewpoint robustness of feature extraction. Extensive experiments demonstrated that our approach achieves state-of-the-art performance on two VLN benchmarks and shows strong performance in real-world scenarios. Source codes will be publicly available upon paper acceptance. Zhide Zhong, Xiangchen Liu, Xinhu Zheng, Zhe Liu 0022, Hesheng Wang 0001, Haoang Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Fusion Filtering for RIS-Assisted Vehicle Localization With Unknown Inputs: A Privacy-Preserving Output-Mask StrategyabstractThis article investigates the privacy-preserving fusion filtering problem for vehicle localization subject to unknown inputs. High-accuracy and privacy-preserving localization is essential for the safe and reliable operation of intelligent transportation systems. In practice, vehicle localization is often challenged by measurement bias caused by nonline-of-sight (NLOS) propagation, unknown inputs arising from uncertainties or acceleration/deceleration maneuvers, and risks of signal and location privacy leakage. To address these issues, a privacy-preserving fusion filtering framework is proposed by integrating the reconfigurable intelligent surface (RIS) technique, an unknown-input estimation method, and an output-mask mechanism. A unified measurement model is first developed to represent both line-of-sight (LOS) and NLOS scenarios, and RISs are employed to construct virtual LOS paths to mitigate NLOS effects. An output-mask-based privacy-preserving strategy is then designed to prevent eavesdroppers from inferring vehicle locations or signal characteristics while maintaining the required filtering performance. Based on this model, an unknown-input estimator and a privacy-preserving filter are constructed, and the impacts of NLOS propagation, unknown inputs, and masking on filtering performance are analyzed. The associated gain matrix parameters are obtained by solving the corresponding optimization problems. Furthermore, an RIS-assisted privacy-preserving fusion filtering algorithm is developed to exploit multisource measurements and enhance localization robustness and accuracy. Simulation results demonstrate the effectiveness of the proposed method. Kaiqun Zhu, Zidong Wang 0001, Xinhu Zheng, Zhiyong Cui, Zhenning Li 0001, Keqiang Li 0002 |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known AttributesabstractAdapting pre-trained models has become the dominant approach for anomalous sound detection (ASD), where classifying the attributes of machine working status is commonly chosen as the deputy task for fine-tuning. However, attributes might be intractable to collect for some machines, causing the label to bear mixed granularity and thus deprecating the ASD performance. Therefore, we propose an adaptive proto-type learning scheme for fine-tuning pre-trained models, which adaptively scales coarse-grained labels to sub-centers so as to keep consistency with fine-grained labels. To deal with domain shift, we employ SMOTE to over-sample the prototypes of the target domain. The experiment on the DCASE 2024 ASD dataset demonstrates the efficacy of the proposed scheme, setting up a new milestone of 65.01% on both sets and outperforming the best system of the challenge. A detailed ablation study is also conducted to validate the effectiveness. Anbai Jiang, Xinhu Zheng, Yihong Qiu, Pingyi Fan, Cheng Lu 0007, Jia Liu 0001 |
ICASSP | 2 |
| 2025 | SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban EnvironmentsabstractUnmanned Aerial Vehicles (UAVs) have emerged as versatile tools across various sectors, driven by their mobility and adaptability. This paper introduces SkyVLN, a novel framework integrating vision-and-language navigation (VLN) with Nonlinear Model Predictive Control (NMPC) to enhance UAV autonomy in complex urban environments. Unlike traditional navigation methods, SkyVLN leverages Large Language Models (LLMs) to interpret natural language instructions and visual observations, enabling UAVs to navigate through dynamic 3D spaces with improved accuracy and robustness. We present a multimodal navigation agent equipped with a fine-grained spatial verbalizer and a history path memory mechanism. These components allow the UAV to disambiguate spatial contexts, handle ambiguous instructions, and backtrack when necessary. The framework also incorporates an NMPC module for dynamic obstacle avoidance, ensuring precise trajectory tracking and collision prevention. To validate our approach, we developed a high-fidelity 3D urban simulation environment using AirSim, featuring realistic imagery and dynamic urban elements. Extensive experiments demonstrate that SkyVLN significantly improves navigation success rates and efficiency, particularly in new and unseen environments. Tianshun Li, Tianyi Huai, Yichun Gao, Haoang Li, Xinhu Zheng |
IROS | 6 |
| 2025 | Deployment-Friendly Lane-Changing Intention Prediction Powered by Brain-Inspired Spiking Neural NetworksabstractAccurate and real-time prediction of surrounding vehicles' lane-changing intentions is a critical challenge in deploying safe and efficient autonomous driving systems in open-world scenarios. Existing high-performing methods remain hard to deploy due to their high computational cost, long training times, and excessive memory requirements. Here, we propose an efficient lane-changing intention prediction approach based on brain-inspired Spiking Neural Networks (SNN). By leveraging the event-driven nature of SNN, the proposed approach enables us to encode the vehicle's states in a more efficient manner. Comparison experiments conducted on HighD and NGSIM datasets demonstrate that our method significantly improves training efficiency and reduces deployment costs while maintaining comparable prediction accuracy. Particularly, compared to the baseline, our approach reduces training time by 75% and memory usage by 99.9%. These results validate the efficiency and reliability of our method in lane-changing predictions, highlighting its potential for safe and efficient autonomous driving systems while offering significant advantages in deployment, including reduced training time, lower memory usage, and faster inference. Shuqi Shen, Xinhu Zheng, Hai Yang 0003 |
IV | 5 |
| 2025 | Predictive Target-to-User Association in Complex Scenarios via Hybrid-Field ISAC SignalingabstractThis paper presents a novel and robust target-to-user (T2U) association framework to support reliable vehicle-to-infrastructure (V2I) networks that potentially operate within the hybrid field (near-field and far-field). To address the challenges posed by complex vehicle maneuvers and user association ambiguity, an interacting multiple-model filtering scheme is developed, which combines coordinated turn and constant velocity models for predictive beamforming. Building upon this foundation, a lightweight association scheme leverages user-specific integrated sensing and communication (ISAC) signaling while employing probabilistic data association to manage clutter measurements in dense traffic. Numerical results validate that the proposed framework significantly outperforms conventional methods in terms of both tracking accuracy and association reliability. Yifeng Yuan, Miaowen Wen, Xinhu Zheng, Shuoyao Wang, Shijian Gao |
VTC2025-Spring | 3 |
| 2025 | Incomplete multi-view subspace clustering based on robust matrix completion
Lei Xing 0003, Xinhu Zheng, Badong Chen |
Neurocomputing | 2 |
| 2025 | Scale Propagation Network for Generalizable Depth CompletionabstractDepth completion, inferring dense depth maps from sparse measurements, is crucial for robust 3D perception. Although deep learning based methods have made tremendous progress in this problem, these models cannot generalize well across different scenes that are unobserved in training, posing a fundamental limitation that yet to be overcome. A careful analysis of existing deep neural network architectures for depth completion, which are largely borrowing from successful backbones for image analysis tasks, reveals that a key design bottleneck actually resides in the conventional normalization layers. These normalization layers are designed, on one hand, to make training more stable, on the other hand, to build more visual invariance across scene scales. However, in depth completion, the scale is actually what we want to robustly estimate in order to better generalize to unseen scenes. To mitigate, we propose a novel scale propagation normalization (SP-Norm) method to propagate scales from input to output, and simultaneously preserve the normalization operator for easy convergence. More specifically, we rescale the input using learned features of a single-layer perceptron from the normalized input, rather than directly normalizing the input as conventional normalization layers. We then develop a new network architecture based on SP-Norm and the ConvNeXt V2 backbone. We explore the composition of various basic blocks and architectures to achieve superior performance and efficient inference for generalizable depth completion. Extensive experiments are conducted on six unseen datasets with various types of sparse depth maps, i.e., randomly sampled 0.1%/1%/10% valid pixels, 4/8/16/32/64-line LiDAR points, and holes from Structured-Light. Our model consistently achieves the best accuracy with faster speed and lower memory when compared to state-of-the-art methods. Haotian Wang 0009, Meng Yang 0002, Xinhu Zheng, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | E²BA: Environment Exploration and Backtracking Agent for Visual Language Object NavigationabstractRobot navigation in an unknown environment is a challenging task, due to the lack of spatial awareness and semantic understanding of the environment. Previous works predominantly relied on prior scene knowledge and semantic information, lacking generalization and transferability. This paper proposes an environment exploration and backtracking agent (E2BA) for visual language object navigation, which leverages the rich semantic prior knowledge and commonsense reasoning of large language models (LLMs) to explore the environment and find the object. By fusing LLM scores and spatial geometric costs using particle filters, we select a redefined optimal frontier as sub-goal for environment exploration. To avoid redundant exploration and paths, we design a backtracking discriminator to evaluate the state of the agent and determine the timing of backtracking triggering through a double-level cascade mechanism. Additionally, we design a random instruction fuzzy semantic guessing task to verify the application diversity of this method. Comprehensive experiments on the Habitat-Matterport 3D dataset show that our method achieves a success rate of 0.704, which is higher than the existing baseline method. This study explores the potential application of LLMs in environment exploration without the need for additional training and semantic supplementation. Lihang Sun, Xinhu Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | EditFollower: Tunable Car Following Models for Customizable Driving BehaviorabstractIn the realm of driving technologies, fully autonomous vehicles have not been widely adopted yet, making advanced driver assistance systems (ADAS) crucial for enhancing driving experiences. Among these, car-following behavior modeling plays a pivotal role, forming the foundation for systems that ensure safe and efficient vehicle interactions. However, current approaches often rely on fixed parameters, failing to capture the diverse social preferences and driving styles of individuals. To overcome these limitations, we propose the Editable Behavior Generation (EBG) model, a data-driven car-following model that allows for adjusting driving discourtesy levels. The framework integrates diverse courtesy calculation methods into long short-term memory (LSTM) and Transformer architectures, offering a comprehensive approach to capture nuanced driving dynamics. By integrating various discourtesy values during the training process, our model generates realistic agent trajectories with different levels of courtesy in car-following behavior. Experimental results on the naturalistic datasets showcase a reduction in Mean Squared Error (MSE) of spacing and MSE of speed compared to baselines, establishing style controllability. To the best of our knowledge, this work represents the first data-driven car-following model capable of dynamically adjusting discourtesy levels. Our model provides valuable insights for the development of ADAS that take into account drivers’ social preferences. Xianda Chen, Xu Han 0017, Meixin Zhu, Xiaowen Chu 0001, PakHin Tiu, Xinhu Zheng, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Depth Map Super-Resolution via Deep Cross-Modality and Cross-Scale GuidanceabstractGuided depth super-resolution is essential in many applications, which enhances low-resolution (LR) depth maps using high-resolution (HR) RGB images from the same scene. However, the challenge lies in avoiding the texture-copy artifacts issue caused by structural inconsistencies between two modalities. To mitigate, we propose a cross-modality and cross-scale guided depth super-resolution network (D2CNet). We first design a novel two-stage feature integration module to effectively fuse multi-modal RGB and depth while minimizing texture-copy artifacts. That is, a cross-modality fusion stage transfers consistent structures from RGB to depth in a multi-scale manner, and a cross-scale refinement stage mitigates inconsistent structures across modalities. In addition, we design a convolution group as the basic module to well extract high-frequency features and an LR and HR domain projection strategy to enrich features between the fusion and refinement stages. We then develop a new network architecture by progressively repeating the feature integration module and the convolution group, which is flexibly controllable to strike a balance between accuracy and cost for easy implementation in real world. Extensive experiments on multiple benchmarks demonstrate that our D2CNet consistently achieves superior accuracy and generalization ability across sampling scales in both qualitative and quantitative evaluations, when compared to state-of-the-art baselines. Shuzhe Liu, Delong Suzhang, Meng Yang 0002, Xinhu Zheng, Ce Zhu |
IEEE Trans. Multim. | 4 |
| 2025 | CCPoint: Contrasting Corrupted Point Clouds for Self-Supervised Representation LearningabstractSelf-supervised Learning (SSL), including mainstream contrastive learning, has achieved significant success in learning visual representations without the need for data annotations in 3D vision. While most contrastive learning methods focus on instance-level information through random affine transformations, they pay limited attention to the intrinsic structures within point clouds. In this work, we propose a novel SSL paradigm for point cloud representation learning, called CCPoint, which incorporates a novel form of data corruption as a negative augmentation strategy. Specifically, we degrade the input point cloud with various corruptions and conduct contrastive learning among the augmented, raw, and corrupted points to learn robust and discriminative representations. To preserve the semantic structure of the point cloud even under heavy degradation, an auxiliary reconstruction decoder is introduced into the corruption branch to provide an additional supervision signal. We explore four families of corruptions—affine, noise, masking, and combined transformations. Different from previous methods that rely on multi-modal data or complex network architectures, CCPoint achieves state-of-the-art performance on three widely used datasets (ModelNet40, ScanObjectNN, and ShapeNetPart) with a lightweight and efficient structure, reaching top linear accuracies of 92.4% and 86.2% on ModelNet40 and ScanObjectNN, respectively. Xiaoyang Xiao, Shaoyi Du, Meiqin Liu 0001, Xinhu Zheng |
IEEE Trans. Multim. | 5 |
| 2024 | Improving Car-Following Control in Mixed Traffic: A Deep Reinforcement Learning Framework with Aggregated Human-Driven VehiclesabstractTraffic oscillations pose safety and efficiency challenges in mixed scenarios involving connected and automated vehicles (CAVs) and human-driven vehicles (HDVs). Existing control strategies fail to handle the unpredictability of HDV behaviors, resulting in disruptive "stop-and-go" traffic patterns. This study proposes a novel algorithm that uses Deep Reinforcement Learning (DRL) integrated into a distinctive "CAV-AHDV-CAV" structure for car-following events. The consecutive HDVs are treated as an aggregated unit called Aggregated HDVs (AHDVs) to eliminate stochasticity and leverage collective traffic features as inputs, addressing the driver heterogeneity issue. Our training and testing were conducted using a dataset of 9,200 car-following events extracted from the HighD dataset. In these events, the lead vehicle serves as our CAV in front, while the following vehicle represents the AHDV. We simulated our controlled vehicle to follow the AHDV, aiming to achieve the vehicle equilibrium state with respect to both the AHDV and the CAV in front. The results demonstrate a reduction in the impact of HDVs and an enhancement of equilibrium states compared to baseline models. Specifically, we achieved a speed mean square error (MSE) of 3.151 and spacing MSE values of 50.484 (with respect to the AHDV) and 47.855 (with respect to the CAV). These findings offer robust and adaptable control strategies for efficient and safe mixed traffic dominated by CAVs. Xianda Chen, PakHin Tiu, Yihuai Zhang, Meixin Zhu, Xinhu Zheng, Yinhai Wang |
IV | 5 |
| 2024 | Driving Style-aware Car-following Considering Cut-in Tendencies of Adjacent Vehicles with Inverse Reinforcement LearningabstractDespite the widespread implementation, the Adaptive Cruise Control (ACC) systems still fall short in delivering a satisfactory human-likely experience, primarily due to the heterogeneity in driving experience preferences and the unpredictable, heterogeneous nature of human driving behaviors. To address this critical gap, we introduce an innovative driving style-aware car-following model that effectively captures the varying cut-in tendencies of adjacent vehicles by utilizing the Maximum Entrop Inverse Reinforcement Learning (Max-Ent IRL) method. A distinct reward function is developed to replicate human driving behavior, which can achieve a harmonious equilibrium between efficiency, safety, and comfort. The efficacy of this model is rigorously evaluated through a comprehensive analysis on car-following episodes extracted from the Next Generation Simulation (NGSIM) I-80 dataset. A novel human-likely metric is utilized for evaluating the performance of the proposed model in comparison to standard benchmarks. The results demonstrably favor our approach, showing notable enhancements in efficiency, safety, and comfort. Additionally, the model’s versatility is confirmed by its ability to accommodate a wide spectrum of driving styles, as evidenced by the diverse weights learned from different driving styles. These findings highlight the significant potential of our model in advancing ACC technology for more human-oriented vehicular systems that align closely with the natural driving instincts and preferences of humans. Xiaoyun Qiu, Meixin Zhu, Liuqing Yang 0001, Xinhu Zheng |
IV | 5 |
| 2024 | VeXKD: The Versatile Integration of Cross-Modal Fusion and Knowledge Distillation for 3D PerceptionabstractRecent advancements in 3D perception have led to a proliferation of network architectures, particularly those involving multi-modal fusion algorithms. While these fusion algorithms improve accuracy, their complexity often impedes real-time performance. This paper introduces VeXKD, an effective and Versatile framework that integrates Cross-Modal Fusion with Knowledge Distillation. VeXKD applies knowledge distillation exclusively to the Bird's Eye View (BEV) feature maps, enabling the transfer of cross-modal insights to single-modal students without additional inference time overhead. It avoids volatile components that can vary across various 3D perception tasks and student modalities, thus improving versatility. The framework adopts a modality-general cross-modal fusion module to bridge the modality gap between the multi-modal teachers and single-modal students. Furthermore, leveraging byproducts generated during fusion, our BEV query guided mask generation network identifies crucial spatial locations across different BEV feature maps in a data-driven manner, significantly enhancing the effectiveness of knowledge distillation. Extensive experiments on the nuScenes dataset demonstrate notable improvements, with up to 6.9\%/4.2\% increase in mAP and NDS for 3D detection tasks and up to 4.3\% rise in mIoU for BEV map segmentation tasks, narrowing the performance gap with multi-modal models. Yuzhe Ji, Liuqing Yang 0001, Xinhu Zheng |
NeurIPS | 6 |
| 2024 | Improving Anomalous Sound Detection Via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio ModelsabstractAnomalous Sound Detection (ASD) has gained significant interest through the application of various Artificial Intelligence (AI) technologies in industrial settings. Though possessing great potential, ASD systems can hardly be readily deployed in real production sites due to the generalization problem, which is primarily caused by the difficulty of data collection and the complexity of environmental factors. This paper introduces a robust ASD model that leverages audio pre-trained models. Specifically, we fine-tune these models using machine operation data, employing SpecAug as a data augmentation strategy. Additionally, we investigate the impact of utilizing Low-Rank Adaptation (LoRA) tuning instead of full fine-tuning to address the problem of limited data for fine-tuning. Our experiments on the DCASE2023 Task 2 dataset establish a new benchmark of 77.75% on the evaluation set, with a significant improvement of 6.48% compared with previous state-of-the-art (SOTA) models, including top-tier traditional convolutional networks and speech pre-trained models, which demonstrates the effectiveness of audio pre-trained models with LoRA tuning. Ablation studies are also conducted to showcase the efficacy of the proposed scheme. Xinhu Zheng, Anbai Jiang, Bing Han 0008, Yanmin Qian, Pingyi Fan, Jia Liu 0001, Weiqiang Zhang 0001 |
SLT | 1 |
| 2023 | Confidence Evaluation for Machine Learning Schemes in Vehicular Sensor NetworksabstractIn this paper, we study a cooperative perception scheme in a vehicular sensor network, attempting to fuse the semantic information provided by different sensors at multiple vehicles, so as to expand the vehicle’s perception range, eliminate blind spots, improve the ability to handle environmental interference and enhance the accuracy and robustness of the perception results. The key to guide the fusion process is the evaluation of the confidence levels of the outputs provided by various machine learning schemes implemented at individual sensors in the vehicular sensor network. We first propose an evaluation criterion termed as Environmental Sensitivity (ES), which is used to measure the sensitivity of the network to environmental changes. Based on the ES, we further evaluate the confidence of the perception output of neural networks and quantify the confidence level considering the abnormal level of the input data, the general performance of the perception algorithm, the detection performance and the ES of the network. Semantic information fusion algorithm is then developed based upon the confidence levels. Experiment results are provided to validate the proposed fusion method in various scenarios. Xinhu Zheng, Sijiang Li, Yuru Li, Dongliang Duan, Liuqing Yang 0001, Xiang Cheng 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | Real-time driving style classification based on short-term observationsabstractAbstract Vehicle behaviour prediction provides important information for decision‐making in modern intelligent transportation systems. People with different driving styles have considerably different driving behaviours and hence exhibit different behaviour tendency. However, most existing prediction methods do not consider the different tendencies in driving styles and apply the same model to all vehicles. Furthermore, most of the existing driver classification methods rely on offline learning that requires a long observation of driving history and hence are not suitable for real‐time driving behaviour analysis. To facilitate personalised models that can potentially improve vehicle behaviour prediction, the authors propose an algorithm that classifies drivers into different driving styles. The algorithm only requires data from a short observation window and it is more applicable for real‐time online applications compared with existing methods that require a long term observation. Experiment results demonstrate that the proposed algorithm can achieve consistent classification results and provide intuitive interpretation and statistical characteristics of different driving styles, which can be further used for vehicle behaviour prediction. Xinhu Zheng, Pengtao Yang, Dongliang Duan, Xiang Cheng 0001, Liuqing Yang 0001 |
IET Commun. | 1 |
| 2022 | Supervised assisted deep reinforcement learning for emergency voltage control of power systems
Xiaoshuang Li, Xiao Wang 0002, Xinhu Zheng, Yuxin Dai, Zhihong Yu, Jun Jason Zhang, Guangquan Bu, Fei-Yue Wang 0001 |
Neurocomputing | 3 |
| 2022 | SADRL: Merging human experience with machine intelligence via supervised assisted deep reinforcement learning
Xiaoshuang Li, Xiao Wang 0002, Xinhu Zheng, Junchen Jin, Yanhao Huang, Jun Jason Zhang, Fei-Yue Wang 0001 |
Neurocomputing | 3 |
| 2022 | Multivehicle Multisensor Occupancy Grid Maps (MVMS-OGM) for Autonomous DrivingabstractIn autonomous driving, environment perception is the fundamental task for intelligent vehicles which provides the necessary environment information for other applications. The main issues in existing environment perception can be categorized into two aspects. On the one hand, all sensors are prone to measurement errors and failures. On the other hand, in complex driving environments, vehicles may encounter a variety of blind spots caused by vehicle occlusions, overlaps, and harsh weather conditions, which will cause sensors to experience low-quality data or to miss crucial environmental information. To cope with these issues, a multivehicle and multisensor (MVMS) cooperative perception method is presented to construct the occupancy grid map (OGM) of vehicles in a global view for the environment perception of autonomous driving. Distinct from existing environment perception methods, our proposed MVMS-OGM not only provides continuous geographical information but also captures and fuses continuous information with soft occupancy probabilities, resulting in more comprehensive and raw environmental information. Simulations and real-world experiments demonstrate that the proposed approach not only expands the perception range in comparison with single-vehicle sensing but also better captures the uncertainty of sensor data by fusing the occupancy probabilities with soft information. Xinhu Zheng, Yuru Li, Dongliang Duan, Liuqing Yang 0001, Chen Chen 0002, Xiang Cheng 0001 |
IEEE Internet Things J. | 1 |
| 2021 | Visual Human-Computer Interactions for Intelligent Vehicles and Intelligent Transportation Systems: The State of the Art and Future DirectionsabstractResearch on intelligent vehicles has been popular in the past decade. To fill the gap between automatic approaches and man-machine control systems, it is indispensable to integrate visual human-computer interactions (VHCIs) into intelligent vehicles systems. In this article, we review existing studies on VHCI in intelligent vehicles from three aspects: 1) visual intelligence; 2) decision making; and 3) macro deployment. We discuss how VHCI evolves in intelligent vehicles and how it enhances the capability of intelligent vehicles. We present several simulated scenarios and cases for future intelligent transportation system. Xumeng Wang, Xinhu Zheng, Wei Chen 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | Capturing Car-Following Behaviors by Deep LearningabstractIn this paper, we propose a deep neural network-based car-following model that has two distinctive properties. First, unlike most existing car-following models that take only the instantaneous velocity, velocity difference, and position difference as inputs, this new model takes the velocities, velocity differences, and position differences that were observed in the last few time intervals as inputs. That is, we assume that drivers’ actions are temporally dependent in this model and try to embed prediction capability or memory effect of human drivers in a natural and efficient way. Second, this car-following model is built in a data-driven way, in which we reduce human interference to the minimum degree. Specially, we use recently developing deep neural networks rather than conventional neural networks to establish the model, since deep learning technique provides us more flexibility and accuracy to describe complicated human actions. Tests on empirical trajectory records show that this deep neural network-based car-following model yield significantly higher simulation accuracy than existing car-following models. All these findings provide a novel way to study traffic flow theory and traffic simulations. Xiao Wang 0013, Li Li 0013, Yilun Lin 0002, Xinhu Zheng, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | Analysis of Cyber Interactive Behaviors Using Artificial Community and Computational ExperimentsabstractAn artificial community is constructed for studying cyber interactive behavior of message publishing by considering impacts from individual's psychological state and social environment. The bottom-up multiagent approach is applied to model the interactions between individuals, and the evolution and organization of collective cyber behaviors. The cyber message publishing behaviors in response to online topics is demonstrated in the artificial community, where each agent perceives the state of the topics and make decisions to operate on specific themes with particular message publishing mechanisms. The mechanisms are designed in line with the analysis of statistical results of 10 years data from Tianya.cn, which are further validated through the artificial community using computational experimental leverages. Experimental results illuminate that the individual's post publishing behavior is dominated by agent's psychological states, while the comment publishing behavior is more susceptible to the agent's social environment expressing the collective efforts of topics and other agents in the community. The artificial community has been used to test the impact of several factors (e.g., individual's psychological state, number of initial agents, and number of new added agents) to individual's message publishing mechanism and to cyber collective behavior patterns. It provides a software-defined social laboratory for social and cyber behavior studies with computational experimental measures. Xiao Wang 0002, Xinhu Zheng, Xinzhan Zhang, Ke Zeng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2016 | Crowdsourcing in ITS: The State of the Work and the NetworkingabstractIn the last decade, crowdsourcing has emerged as a novel mechanism for accomplishing temporal and spatial critical tasks in transportation with the collective intelligence of individuals and organizations. This paper presents a timely literature review of crowdsourcing and its applications in intelligent transportation systems (ITS). We investigate the ITS services enabled by crowdsourcing, the keyword co-occurrence and coauthorship networks formed by ITS publications, and identify the problems and challenges that need further research. Finally, we briefly introduce our future works focusing on using geospatial tagged data to analyze real-time traffic conditions and the management of traffic flow in urban environment. This review aims to help ITS practitioners and researchers build a state-of-the-art understanding of crowdsourcing in ITS, as well as to call for more research on the application of crowdsourcing in transportation systems. Xiao Wang 0002, Xinhu Zheng, Qingpeng Zhang, Tao Wang 0172, Dayong Shen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Big Data for Social TransportationabstractBig data for social transportation brings us unprecedented opportunities for resolving transportation problems for which traditional approaches are not competent and for building the next-generation intelligent transportation systems. Although social data have been applied for transportation analysis, there are still many challenges. First, social data evolve with time and contain abundant information, posing a crucial need for data collection and cleaning. Meanwhile, each type of data has specific advantages and limitations for social transportation, and one data type alone is not capable of describing the overall state of a transportation system. Systematic data fusing approaches or frameworks for combining social signal data with different features, structures, resolutions, and precision are needed. Second, data processing and mining techniques, such as natural language processing and analysis of streaming data, require further revolutions in effective utilization of real-time traffic information. Third, social data are connected to cyber and physical spaces. To address practical problems in social transportation, a suite of schemes are demanded for realizing big data in social transportation systems, such as crowdsourcing, visual analysis, and task-based services. In this paper, we overview data sources, analytical approaches, and application systems for social transportation, and we also suggest a few future research directions for this new social transportation field. Xinhu Zheng, Wei Chen 0001, Dayong Shen, Songhang Chen, Xiao Wang 0002, Qingpeng Zhang, Liuqing Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |