VLDB 2026 Research / reviewers in the wild / expert
Runwei Guan
dblp:338/6709
· DBLP profile ↗
28ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0003-4013-2107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AerialMind: Towards Referring Multi-Object Tracking in UAV ScenariosabstractReferring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrains their ability to capture broad-scale scene contexts and perform comprehensive tracking and path planning. In contrast, Unmanned Aerial Vehicles (UAVs) leverage their expansive aerial perspectives and superior maneuverability to enable wide-area surveillance. Moreover, UAVs have emerged as critical platforms for Embodied Intelligence, which has given rise to an unprecedented demand for intelligent aerial systems capable of natural language interaction. To this end, we introduce AerialMind, the first large-scale RMOT benchmark in UAV scenarios, which aims to bridge this research gap. To facilitate its construction, we develop an innovative semi-automated collaborative agent-based labeling assistant (COALA) framework that significantly reduces labor costs while maintaining annotation quality. Furthermore, we propose HawkEyeTrack (HETrack), a novel method that collaboratively enhances vision-language representation learning and improves the perception of UAV scenarios. Comprehensive experiments validated the challenging nature of our dataset and the effectiveness of our method. Chenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun, Haocheng Zhao, Haiyun Jiang, Tao Huang 0008, Henghui Ding, Qing-Long Han |
AAAI | 3 |
| 2026 | RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation SystemabstractCurrent roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introduce RoadSceneVQA, a large-scale and richly annotated visual question answering (VQA) dataset specifically tailored for roadside scenarios. The dataset comprises 34,736 diverse QA pairs collected under varying weather, illumination, and traffic conditions, targeting not only object attributes but also the intent, legality, and interaction patterns of traffic participants. RoadSceneVQA challenges models to perform both explicit recognition and implicit commonsense reasoning, grounded in real-world traffic rules and contextual dependencies. To fully exploit the reasoning potential of Multi-modal Large Language Models (MLLMs), we further propose CogniAnchor Fusion (CAF), a vision-language fusion module inspired by human-like scene anchoring mechanisms. CAF enables precise and efficient cross-modal interaction. Moreover, we propose the Assisted Decoupled Chain-of-Thought (AD-CoT) to enhance the reasoned thinking via CoT prompting and multi-task learning. Experimental results on RoadSceneVQA and CODA-LM benchmark show that the pipeline consistently improves both reasoning accuracy and computational efficiency, allowing the MLLM to achieve state-of-the-art performance in structural traffic perception and reasoning tasks. Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningyuan Xiao, Ziren Tang, Ningwei Ouyang, Shaofeng Liang, Yuxuan Fan, Wanjie Sun, Yutao Yue |
AAAI | 1 |
| 2026 | Generating transferable attacks across large vision-language models using adversarial deformation learning
Daizong Liu, Wangqin Liu, Xiaowen Cai 0001, Pan Zhou 0001, Runwei Guan, Xiaoye Qu, Bo Du 0001 |
Pattern Recognit. | 5 |
| 2026 | Da Yu: Toward ASV-Based Image Captioning for Waterway Surveillance and Scene UnderstandingabstractAutomated waterway environment perception is crucial for enabling unmanned surface vessels (USVs) to understand their surroundings and make informed decisions. Most existing waterway perception models primarily focus on instance-level object perception paradigms (e.g., detection, segmentation). However, due to the complexity of waterway environments, current perception datasets and models fail to achieve global semantic understanding of waterways, limiting large-scale monitoring and structured log generation. With the advancement of vision-language models (VLMs), we leverage image captioning to introduce WaterCaption, the first captioning dataset specifically designed for waterway environments. WaterCaption focuses on fine-grained, multi-region long-text descriptions, providing a new research direction for visual geo-understanding and spatial scene cognition. Exactly, it includes 20.2k image-text pair data with 1.8 million vocabulary size. Additionally, we propose Da Yu, an edge-deployable multi-modal large language model for USVs, where we propose a novel vision-to-language projector called Nano Transformer Adaptor (NTA). NTA effectively balances computational efficiency with the capacity for both global and fine-grained local modeling of visual features, thereby significantly enhancing the model’s ability to generate long-form textual outputs. Da Yu achieves an optimal balance between performance and efficiency, surpassing state-of-the-art models on WaterCaption and several other captioning benchmarks. The project is available at https://github.com/GuanRunwei/WaterCaption. Runwei Guan, Ningwei Ouyang, Tianhao Xu, Shaofeng Liang, Yafeng Sun, Shang Gao 0012, Songning Lai, Shanliang Yao, Xuming Hu, Ryan Wen Liu, Yutao Yue, Hui Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Doracamom: Joint 3D Detection and Occupancy Prediction With Multi-View 4D Radars and Cameras for Omnidirectional Perception
Lianqing Zheng, Runwei Guan, Shouyi Lu, Xiaokai Bai, Zhixiong Ma, Xichan Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | ScopeDrive: Text-Anchored Cross-Modal Calibration and Density-Aware Modulation for Autonomous DrivingabstractVision–language models are emerging as a unified paradigm for perception, prediction, and planning in autonomous driving. However, most existing approaches still rely on shallow fusion between visual and textual features, leading to weak cross-modal interaction, feature drift, and poor adaptability in complex scenes. This work present ScopeDrive, an end-to-end VLM framework that enables deep semantic alignment and adaptive reasoning through two novel components. The text-anchored calibrator transforms textual semantics into multiscale anchors that progressively calibrate visual features across layers, strengthening cross-modal correspondence. The density-aware agent modulator estimates scene complexity and dynamically adjusts attention distribution, allowing the model to focus on dense or dynamic regions when needed. Evaluated on the DriveLM, DriveBench, and NuScenes-QA benchmarks, ScopeDrive surpasses both lightweight and large-scale baselines, achieving a BLEU-4 of 53.27 and METEOR of 38.75 on DriveLM while maintaining only 328 M parameters. It also delivers state-of-the-art performance on perception and planning tasks under both clean and corrupted conditions. These results demonstrate that ScopeDrive effectively breaks the shallow-fusion barrier, offering a lightweight yet semantically aligned foundation for interpretable autonomous-driving intelligence. Minghui Hou, Runwei Guan, Tao Huang 0008, Qing-Long Han |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Imperceptible Transfer Attack on Large Vision-Language ModelsabstractIn spite of achieving significant progress in recent years, Large Vision-Language Models (LVLMs) are proven to be vulnerable to adversarial examples. Therefore, there is an urgent need for an effective adversarial attack to identify the deficiencies of LVLMs in security-sensitive applications. However, existing LVLM attackers generally optimize adversarial samples against a specific textual prompt with a certain LVLM model, tending to overfit the target prompt/network and hardly remain malicious once they are transferred to attack a different prompt/model. To this end, in this paper, we propose a novel Imperceptible Transfer Attack (ITA) against LVLMs to generate prompt/model-agnostic adversarial samples to enhance such adversarial transferability while further improving the imperceptibility. Specifically, we learn to apply appropriate visual transformations on image inputs to create diverse input patterns by selecting the optimal combination of operations from a pool of candidates, consequently improving adversarial transferability. We conceptualize the selection of optimal transformation combinations as an adversarial learning problem and employ a gradient approximation strategy with noise budget constraints to effectively generate imperceptible transferable samples. Extensive experiments on three LVLM models and two widely used datasets with three tasks demonstrate the superior performance of our ITA. Xiaowen Cai 0001, Daizong Liu, Runwei Guan, Pan Zhou 0001 |
ICASSP | 3 |
| 2025 | PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised LearningabstractFine-grained image classification has witnessed significant advancements with the advent of deep learning and computer vision technologies. However, the scarcity of detailed annotations remains a major challenge, especially in scenarios where obtaining high-quality labeled data is costly or time-consuming. To address this limitation, we introduce Precision-Enhanced Pseudo-Labeling (PEPL) approach specifically designed for fine-grained image classification within a semi-supervised learning framework. Our method leverages the abundance of unlabeled data by generating high-quality pseudo-labels that are progressively refined through two key phases: initial pseudo-label generation and semantic-mixed pseudo-label generation. These phases utilize Class Activation Maps (CAMs) to accurately estimate the semantic content and generate refined labels that capture the essential details necessary for fine-grained classification. By focusing on semantic-level information, our approach effectively addresses the limitations of standard data augmentation and image-mixing techniques in preserving critical fine-grained features. We achieve state-of-the-art performance on benchmark datasets, demonstrating significant improvements over existing semi-supervised strategies, with notable boosts in accuracy and robustness. Songning Lai, Lujundong Li, Zhihao Shuai, Runwei Guan, Yutao Yue |
ICASSP | 5 |
| 2025 | Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression ComprehensionabstractEmbodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20, 558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar. Runwei Guan, Ruixiao Zhang 0001, Ningwei Ouyang, Ka Lok Man, Xiaohao Cai, Ming Xu 0011, Jeremy S. Smith, Eng Gee Lim, Yutao Yue, Hui Xiong 0001 |
ICRA | 1 |
| 2025 | DRIVE: Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous DrivingabstractRecent advancements in autonomous driving have seen a paradigm shift towards end-to-end learning paradigms, which map sensory inputs directly to driving actions, thereby enhancing the robustness and adaptability of autonomous vehicles. However, these models often sacrifice interpretability, posing significant challenges to trust, safety, and regulatory compliance. To address these issues, we introduce DRIVE – Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous Driving, a comprehensive framework designed to improve the dependability and stability of explanations in end-to-end unsupervised autonomous driving models. Our work specifically targets the inherent instability problems observed in the Driving through the Concept Gridlock (DCG) model, which undermine the trustworthiness of its explanations and decisionmaking processes. We define four key attributes of DRIVE: consistent interpretability, stable interpretability, consistent output, and stable output. These attributes collectively ensure that explanations remain reliable and robust across different scenarios and perturbations. Through extensive empirical evaluations, we demonstrate the effectiveness of our framework in enhancing the stability and dependability of explanations, thereby addressing the limitations of current models. Our contributions include an in-depth analysis of the dependability issues within the DCG model, a rigorous definition of DRIVE with its fundamental properties, a framework to implement DRIVE, and novel metrics for evaluating the dependability of concept-based explainable autonomous driving models. These advancements lay the groundwork for the development of more reliable and trusted autonomous driving systems, paving the way for their broader acceptance and deployment in real-world applications. “We can only see a short distance ahead, but we can see plenty there that needs to be done.” – Alan Turing Songning Lai, Tianlang Xue, Hongru Xiao, Lijie Hu, Jiemin Wu, Ninghui Feng, Runwei Guan, Haicheng Liao, Zhenning Li 0001, Yutao Yue |
ICRA | 7 |
| 2025 | Dynamic Compact Consensus Tracking for Aerial RobotsabstractExisting one-stream trackers have attracted widespread attention. However, they are not applicable in real-time aerial robot tracking systems due to substantial computational overhead, especially when dynamic templates are introduced. To address this issue, we propose a novel Dynamic Compact Consensus Tracker (DC2T), constructed by stacking blocks that each consists of a Compact Token Encoder (CTE) and Dynamic Consensus Attention (DCA). Unlike traditional methods that convert images into a large number of tokens, the CTE, inspired by “superpixel”, extracts a compact set of representative tokens from both initial and dynamic templates, eliminating the need for a large token set. This strategic reduction in the number of compact tokens markedly decreases the computational load of CTE, enhancing the efficiency of subsequent attention operations. To achieve linear complexity of the DCA, compact dynamic template tokens (as keys) are requeried by search tokens (as queries) to perform dynamic consensus on the aggregated tokens (as values). This arrangement seamlessly incorporates dynamic spatio-temporal features into the DCA while avoiding the computational burden typically associated with dynamic templates. With the aim of further enhancing the system's responsiveness and accuracy, a direct control network is crafted to seamlessly incorporate the prediction of high-level control values into the tracking network, ensuring a cohesive and efficient interaction with the controller. Comprehensive experiments and real-world evaluations have proven DC2T's superior performance, accompanied by a significant reduction in FLOPs. Furthermore, we have conducted experiments that demonstrate the tracker's ability to integrate seamlessly with other technologies such as SLAM and detection, enabling precise tracking of arbitrary objects. The tracker code will be released in the github. com/xiaolousun/ refine-pytracking. Xiaolou Sun, Zhibin Quan, Yuntian Li, Wufei Si, Wenhui Ni, Runwei Guan |
ICRA | 8 |
| 2025 | UniBEVFusion: Unified Radar-Vision Bevfusion for 3D Object Detectionabstract4D millimeter-wave (MMW) radar, which provides both height information and dense point cloud data over 3D MMW radar, has become increasingly popular in 3D object detection. In recent years, radar-vision fusion models have demonstrated performance close to that of LiDAR-based models, offering advantages in terms of lower hardware costs and better resilience in extreme conditions. However, many radar-vision fusion models treat radar as a sparse LiDAR, underutilizing radar-specific information. Additionally, these multi-modal networks are often sensitive to the failure of a single modality, particularly vision. To address these challenges, we propose the Radar Depth Lift-Splat-Shoot (RDL) module, which integrates radar-specific data into the depth prediction process, enhancing the quality of visual Bird's-Eye View (BEV) features. We further introduce a Unified Feature Fusion (UFF) approach that extracts BEV features across different modalities using shared module. To assess the robustness of multimodal models, we develop a novel Failure Test (FT) ablation experiment, which simulates vision modality failure by injecting Gaussian noise. We conduct extensive experiments on the View-of-Delft (VoD) and TJ4D datasets. The results demonstrated that our proposed Unified BEVFusion (UniBEVFusion) network significantly outperforms state-of-the-art models on the TJ4D dataset, with improvements of 3.96% in 3D and 4.17% in BEV object detection accuracy. Haocheng Zhao, Runwei Guan, Taoyu Wu, Ka Lok Man, Limin Yu, Yutao Yue |
ICRA | 2 |
| 2025 | MonoAttack: A Strong Attack Framework with Depth-Migration and Attribute-Tampering for Monocular 3D Object DetectionabstractAlthough many efforts have been made into attacks on deep neural networks (DNNs) in recent years, no research explores the vulnerability of monocular 3D object detection (M3D) models. This M3D task is fundamental but essential in safety-critical 3D applications, potentially bringing hazards to autonomous driving. In this paper, we thoroughly investigate the sensitivity of current M3D models to adversarial noise and propose a novel M3D adversarial attack method called MonoAttack. The key insight of our method is exploring both depth-migration and attribute-tampering for generating M3D adversarial samples. Specifically, in addition to the general misleading of the detection model, we deceive the M3D model by changing the potential object depth into its opposite position. We also guide the M3D model to mis-recognize the class attribute of its detected object for generating low-confidence bounding boxes. Moreover, we further disentangle the depth knowledge from the geometric and semantic perspectives to auxiliary correlate the detection and attribute information for jointly generating the latent perturbation. In this manner, our attack framework is strong and can effectively attack M3D models with trivial perturbations. Experimental results on the KITTI dataset demonstrate that our attack achieves high adversarial ability against current monocular 3D detection models. Xiayue Zhang, Huashuo Lei, Daizong Liu, Xiaoye Qu, Runwei Guan, Keyan Jin |
IJCNN | 6 |
| 2025 | Manipulating the Bounding Box: Multimodal Controlled Backdoor Attacks on 3D Visual Grounding Modelsabstract3D visual grounding models, pivotal in interpreting and aligning objects within 3D spaces with textual descriptions, have become integral to the advancement of the multimedia community. As these models are widely used in daily life as real-world applications, they become more susceptible to be attacked. Backdoor attacks are designed to corrupt a model in such a way that it responds with adversary-wanted outputs when specific trigger patterns are introduced, while responding normally to clean inputs. Unlike traditional backdoor attack methods that focus on attacking simple classification models, attacking 3D visual grounding models presents unique challenges due to their multi-modal inputs and the nature of their output, which is the localization box of objects described by the text within the 3D scene. This necessitates distinct attack strategies and trigger designs, adding complexity to executing successful attacks. To this end, in this paper, we present a novel multimodal controlled backdoor attack aimed at manipulating the positioning and size of bounding boxes in the challenging multi-modal 3D visual grounding models. Specifically, we design triggers for both point cloud and textual modalities, along with specialized placement strategies for each, to enhance the stealth and precision of the attack. Furthermore, we develop optimization strategies to enhance the efficacy of the point cloud trigger. Experimental results across various standard models confirm the effectiveness of our backdoor attack method, with negligible impact on performance in clean datasets. Xiayue Zhang, Huashuo Lei, Daizong Liu, Xiaoye Qu, Runwei Guan, Keyan Jin |
IJCNN | 6 |
| 2025 | NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave RadarabstractRecently, visual grounding and multi-sensors setting have been incorporated into perception system for terrestrial autonomous driving systems and Unmanned Surface Vessels (USVs), yet the high complexity of modern learning-based visual grounding model using multi-sensors prevents such model to be deployed on USVs in the real-life. To this end, we design a low-power multi-task model named NanoMVG for waterway embodied perception, guiding both camera and 4D millimeter-wave radar to locate specific object(s) through natural language. NanoMVG can perform both box-level and mask-level visual grounding tasks simultaneously. Compared to other visual grounding models, NanoMVG achieves highly competitive performance on the WaterVG dataset, particularly in harsh environments. Moreover, the real-world experiments with deployment of NanoMVG on embedded edge device of USV demonstrates its fast inference speed for real-time perception and capability of boasting ultra-low power consumption for long endurance. Runwei Guan, Liye Jia, Haocheng Zhao, Shanliang Yao, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Yutao Yue |
IROS | 1 |
| 2025 | USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
Shanliang Yao, Runwei Guan, Yi Ni, Yong Yue 0001, Ryan Wen Liu |
IROS | 2 |
| 2025 | Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsabstractAlthough Multimodal Large Language Models (MLLMs) have demonstrated remarkable achievements in recent years, they remain vulnerable to adversarial examples that result in harmful responses. Existing attacks typically focus on optimizing adversarial perturbations for a certain multimodal image-prompt pair or fixed training dataset, which often leads to overfitting. Consequently, these perturbations fail to remain malicious once transferred to attack unseen image-prompt pairs, suffering from significant resource costs to cover the diverse multimodal inputs in complicated real-world scenarios. To alleviate this issue, this paper proposes a novel adversarial attack on MLLMs based on distribution approximation theory, which models the potential image-prompt input distribution and adds the same distribution-fitting adversarial perturbation on multimodal input pairs to achieve effective cross-image/prompt transfer attacks. Specifically, we exploit the Laplace approximation to model the Gaussian distribution of the image and prompt inputs for the MLLM, deriving an estimate of the mean and covariance parameters. By sampling from this approximated distribution with Monte Carlo mechanism, we efficiently optimize and fit a single input‑agnostic perturbation over diverse image‑prompt pairs, yielding strong universality and transferability. Extensive experiments are conducted to verify the strong adversarial capabilities of our proposed attack against prevalent MLLMs spanning a spectrum of images/prompts. Hai Yan, Haijian Ma, Xiaowen Cai 0001, Daizong Liu, Zenghui Yuan, Xiaoye Qu, Jianfeng Dong, Runwei Guan, Hongyang He, Yulai Xie 0002, Pan Zhou 0001 |
NeurIPS | 8 |
| 2025 | Referring flexible image restoration
Runwei Guan, Rongsheng Hu, Zhuhao Zhou, Tianlang Xue, Ka Lok Man, Jeremy S. Smith, Eng Gee Lim, Weiping Ding 0001, Yutao Yue |
Expert Syst. Appl. | 1 |
| 2025 | RePaIR: Repaired pruning at initialization resilienceabstractOver the past decade, the size of neural network models has gradually increased in both breadth and depth, leading to a growing interest in the application of neural network pruning. Unstructured pruning provides fine-grained sparsity and achieves better inference acceleration under specific hardware support. Unstructured Pruning at Initialization (PaI) optimizes the iterative pruning pipeline, but sparse weights increase the risk of underfitting during training. More importantly, almost all PaI algorithms focus only on obtaining the best pruning mask without considering whether the retained weights are suitable for training. Introducing Lipschitz constants during model initialization can reduce the risk of model underfitting and overfitting. As a result, we firstly analyze the impact of Lipschitz initialization on model training and propose the Repaired Initialization (ReI) algorithm for common modules with BatchNorm. Then, we utilize the same idea to repair the weight of unstructured pruned model, and name it Repaired Pruning at Initialization Resilience (RePaIR) algorithm. Extensive experiments and demonstrate that our proposed ReI and RePaIR can improve the training robustness of unpruned and pruned models, respectively, and achieve up to 1.7% accuracy gain with the same sparse pruning mask on TinyImageNet. Furthermore, we provide an improved SynFlow algorithm called Repair SynFlow (ReSynFlow), which employs Lipschitz scaling to overcome the problem of score computation in deeper models. ReSynFlow can effectively improve the maximum compression rate and is suitable for deeper models, with an accuracy improvement of up to 1.3% compared to the SynFlow algorithm on TinyImageNet. Haocheng Zhao, Runwei Guan, Ka Lok Man, Limin Yu, Yutao Yue |
Neural Networks | 2 |
| 2025 | OptiPMB: Enhancing 3D Multi-Object Tracking With Optimized Poisson Multi-Bernoulli FilteringabstractAccurate 3D multi-object tracking (MOT) is crucial for autonomous driving, as it enables robust perception, navigation, and planning in complex environments. While deep learning-based solutions have demonstrated impressive 3D MOT performance, model-based approaches remain appealing for their simplicity, interpretability, and data efficiency. Conventional model-based trackers typically rely on random vector-based Bayesian filters within the tracking-by-detection (TBD) framework but face limitations due to heuristic data association and track management schemes. In contrast, random finite set (RFS)-based Bayesian filtering handles object birth, survival, and death in a theoretically sound manner, facilitating interpretability and parameter tuning. In this paper, we present OptiPMB, a novel RFS-based 3D MOT method that employs an optimized Poisson multi-Bernoulli (PMB) filter while incorporating several key innovative designs within the TBD framework. Specifically, we propose a measurement-driven hybrid adaptive birth model for improved track initialization, employ adaptive detection probability parameters to effectively maintain tracks for occluded objects, and optimize density pruning and track extraction modules to further enhance overall tracking performance. Extensive evaluations on nuScenes and KITTI datasets show that OptiPMB achieves superior tracking accuracy compared with state-of-the-art methods, thereby establishing a new benchmark for model-based 3D MOT and offering valuable insights for future research on RFS-based trackers in autonomous driving. Guanhua Ding, Yuxuan Xia, Runwei Guan, Qinchen Wu, Tao Huang 0008, Weiping Ding 0001, Jinping Sun, Guoqiang Mao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | WaterVG: Waterway Visual Grounding Based on Text-Guided Vision and mmWave RadarabstractWaterway perception is critical for the special operations and autonomous navigation of Unmanned Surface Vessels (USVs), but current perception schemes are sensor-based, neglecting the interaction between humans and USVs for embodied perception in various operations. Therefore, inspired by visual grounding, we present WaterVG, the inaugural visual grounding dataset tailored for USV-based waterway perception guided by human prompts. WaterVG contains a wealth of prompts describing multiple targets, with instance-level annotations, including bounding boxes and masks. Specifically, WaterVG comprises 11,568 samples and 34,987 referred targets, integrating both visual and radar characteristics. The text-guided two-sensor pattern provides a fine granularity of text prompts aligned with the visual and radar features of the referent targets, containing both qualitative and numeric descriptions. To enhance the endurance and maintain the normal operations of USVs in open waterways, we propose Potamoi, a low-power visual grounding model. Potamoi is a multi-task model employing a sophisticated Phased Heterogeneous Modality Fusion (PHMF) mechanism, which includes Adaptive Radar Weighting (ARW) and Multi-Head Slim Cross Attention (MHSCA). The ARW module utilizes a gating mechanism to adaptively extract essential radar features for fusion with visual inputs, ensuring prompt alignment. MHSCA, characterized by its low parameter count and computational efficiency (FLOPs), effectively integrates contextual information from both sensors with linguistic features, delivering outstanding performance in visual grounding tasks. Comprehensive experiments and evaluations on WaterVG demonstrate that Potamoi achieves state-of-the-art results compared to existing methods. The project is available athttps://github.com/GuanRunwei/WaterVG. Runwei Guan, Liye Jia, Shanliang Yao, Fengyufan Yang, Erick Purwanto, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Xuming Hu, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Exploring Radar Data Representations in Autonomous Driving: A Comprehensive ReviewabstractWith the rapid advancements of sensor technology and deep learning, autonomous driving systems are providing safe and efficient access to intelligent vehicles as well as intelligent transportation. Among these equipped sensors, the radar sensor plays a crucial role in providing robust perception information in diverse environmental conditions. This review focuses on exploring different radar data representations utilized in autonomous driving systems. Firstly, we introduce the capabilities and limitations of the radar sensor by examining the working principles of radar perception and signal processing of radar measurements. Then, we delve into the generation process of five radar representations, including the ADC signal, radar tensor, point cloud, grid map, and micro-Doppler signature. For each radar representation, we examine the related datasets, methods, advantages and limitations. Furthermore, we discuss the challenges faced in these data representations and propose potential research directions. Above all, this comprehensive review offers an in-depth insight into how these representations enhance autonomous system capabilities, providing guidance for radar perception researchers. To facilitate retrieval and comparison of different data representations, datasets and methods, we provide an interactive website at https://radar-camera-fusion.github.io/radar. Shanliang Yao, Runwei Guan, Zitian Peng, Chenhang Xu, Yilu Shi, Weiping Ding 0001, Eng Gee Lim, Yong Yue 0001, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | radarODE: An ODE-Embedded Deep Learning Model for Contactless ECG Reconstruction From Millimeter-Wave RadarabstractRadar-based cardiac monitoring has become a popular research direction recently, but the fine-grained electrocardiogram (ECG) signal is still hard to reconstruct from millimeter-wave radar signal. The key obstacle is to decouple cardiac activities in the electrical domain (i.e., ECG) from that in the mechanical domain (i.e., heartbeat), and most existing research only uses purely data-driven methods to map such domain transformation as a black box. Therefore, this work first proposes a signal model that considers the fine-grained cardiac feature sensed by radar, and a novel deep learning framework called radarODE is designed to extract both temporal and morphological features for generating ECG. In addition, ordinary differential equations are embedded in radarODE as a decoder to provide morphological prior, helping the convergence of the model training and improving the robustness under body movements. After being validated on the dataset, the proposed radarODE achieves better performance compared with the benchmark in terms of missed detection rate, root mean square error, Pearson correlation coefficient with improvements of 9%, 16% and 19%, respectively. The validation results imply that radarODE is capable of recovering ECG signals from radar signals with high fidelity and can potentially be implemented in real-life scenarios Runwei Guan, Rui Yang 0007, Yutao Yue, Eng Gee Lim |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave RadarabstractPanoptic Driving Perception (PDP) is critical for the autonomous navigation of Unmanned Surface Vehicles (USVs). A PDP model typically integrates multiple tasks, necessitating the simultaneous and robust execution of various perception tasks to facilitate downstream path planning. The fusion of visual and radar sensors is currently acknowledged as a robust and cost-effective approach. However, most existing research has primarily focused on fusing visual and radar features dedicated to object detection or utilizing a shared feature space for multiple tasks, neglecting the individual representation differences between various tasks. To address this gap, we propose a pair of Asymmetric Fair Fusion (AFF) modules with favorable explainability designed to efficiently interact with independent features from both visual and radar modalities, tailored to the specific requirements of object detection and semantic segmentation tasks. The AFF modules treat image and radar maps as irregular point sets and transform these features into a crossed-shared feature space for multitasking, ensuring equitable treatment of vision and radar point cloud features. Leveraging AFF modules, we propose a novel and efficient PDP model, ASY-VRNet, which processes image and radar features based on irregular super-pixel point sets. Additionally, we propose an effective multi-task learning method specifically designed for PDP models. Compared to other lightweight models, ASY-VRNet achieves state-of-the-art performance in object detection, semantic segmentation, and drivable-area segmentation on the WaterScenes benchmark. Our project is publicly available at https://github.com/GuanRunwei/ASY-VRNet. Runwei Guan, Shanliang Yao, Ka Lok Man, Yong Yue 0001, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
IROS | 1 |
| 2024 | FindVehicle and VehicleFinder: a NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval systemabstractAbstract Natural language (NL) based vehicle retrieval is a task aiming to retrieve a vehicle that is most consistent with a given NL query from among all candidate vehicles. Because NL query can be easily obtained, such a task has a promising prospect in building an interactive intelligent traffic system (ITS). Current solutions mainly focus on extracting both text and image features and mapping them to the same latent space to compare the similarity. However, existing methods usually use dependency analysis or semantic role-labelling techniques to find keywords related to vehicle attributes. These techniques may require a lot of pre-processing and post-processing work, and also suffer from extracting the wrong keyword when the NL query is complex. To tackle these problems and simplify, we borrow the idea from named entity recognition (NER) and construct FindVehicle, a NER dataset in the traffic domain. It has 42.3k labelled NL descriptions of vehicle tracks, containing information such as the location, orientation, type and colour of the vehicle. FindVehicle also adopts both overlapping entities and fine-grained entities to meet further requirements. To verify its effectiveness, we propose a baseline NL-based vehicle retrieval model called VehicleFinder. Our experiment shows that by using text encoders pre-trained by FindVehicle, VehicleFinder achieves 87.7% precision and 89.4% recall when retrieving a target vehicle by text command on our homemade dataset based on UA-DETRAC [1]. From loading the command into VehicleFinder to identifying whether the target vehicle is consistent with the command, the time cost is 279.35 ms on one ARM v8.2 CPU and 93.72 ms on one RTX A4000 GPU, which is much faster than the Transformer-based system. The dataset is open-source via the link https://github.com/GuanRunwei/FindVehicle , and the implementation can be found via the link https://github.com/GuanRunwei/VehicleFinder-CTIM . Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
Multim. Tools Appl. | 1 |
| 2024 | WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water SurfacesabstractAutonomous driving on water surfaces plays an essential role in executing hazardous and time-consuming missions, such as maritime surveillance, survivor rescue, environmental monitoring, hydrography mapping and waste cleaning. This work presents WaterScenes, the first multi-task 4D radar-camera fusion dataset for autonomous driving on water surfaces. Equipped with a 4D radar and a monocular camera, our Unmanned Surface Vehicle (USV) proffers all-weather solutions for discerning object-related information, including color, shape, texture, range, velocity, azimuth, and elevation. Focusing on typical static and dynamic objects on water surfaces, we label the camera images and radar point clouds at pixel-level and point-level, respectively. In addition to basic perception tasks, such as object detection, instance segmentation and semantic segmentation, we also provide annotations for free-space segmentation and waterline segmentation. Leveraging the multi-task and multi-modal data, we conduct benchmark experiments on the uni-modality of radar and camera, as well as the fused modalities. Experimental results demonstrate that 4D radar-camera fusion can considerably improve the accuracy and robustness of perception on water surfaces, especially in adverse lighting and weather conditions. WaterScenes dataset is public onhttps://waterscenes.github.io. Shanliang Yao, Runwei Guan, Zhaodong Wu, Yi Ni, Zile Huang, Ryan Wen Liu, Yong Yue 0001, Weiping Ding 0001, Eng Gee Lim, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | PLSR: Unstructured Pruning with Layer-Wise Sparsity RatioabstractIn the current era of multi-modal and large models gradually revealing their potential, neural network pruning has emerged as a crucial means of model compression. It is widely recognized that models tend to be over-parameterized, and pruning enables the removal of unimportant weights, leading to improved inference speed while preserving accuracy. From early methods such as gradient-based, and magnitude-based pruning to modern algorithms like iterative magnitude pruning, lottery ticket hypothesis, and pruning at initialization, researchers have strived to increase the compression ratio of model parameters while maintaining high accuracy. Currently, mainstream algorithms focus on the global pruning of neural networks using various scoring functions, followed by different pruning strategies to enhance the accuracy of sparse model. Recent studies have shown that random pruning with varying layer-wise sparsity ratio has achieved robust results for large models and out-of-distribution data. Based on this discovery, we propose a new score called FeatIO, which is based on module input and output feature map sizes. As a score function used in PaI, FeatIO surpasses the performance of other PaI score functions. Additionally, we propose a novel pruning strategy called Pruning with Layer-wise Sparsity Ratio (PLSR), which conbines the layer-wise sparsity ratios and magnitude-based score function, resulting in optimal evaluation performance. Almost all algorithms exhibit improved performance when using our novel pruning strategy. The combination of PLSR and FeatIO consistently outperforms other algorithms in testing, demonstrating the significant potential of our proposed approach. Our code will be available here. Haocheng Zhao, Limin Yu, Runwei Guan, Liye Jia, Junqing Zhang, Yutao Yue |
ICMLA | 3 |
| 2023 | MAN and CAT: mix attention to nn and concatenate attention to YOLO
Runwei Guan, Ka Lok Man, Haocheng Zhao, Ruixiao Zhang 0001, Shanliang Yao, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
J. Supercomput. | 1 |