EDBT 2026 Demo / reviewers in the wild / expert
Yutao Yue
dblp:205/3143
· DBLP profile ↗
34ranked-venue papers
0as first author
34since 2021 · last 2026
0000-0003-4532-0924ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Systems, architecture and hardware · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation SystemabstractCurrent roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introduce RoadSceneVQA, a large-scale and richly annotated visual question answering (VQA) dataset specifically tailored for roadside scenarios. The dataset comprises 34,736 diverse QA pairs collected under varying weather, illumination, and traffic conditions, targeting not only object attributes but also the intent, legality, and interaction patterns of traffic participants. RoadSceneVQA challenges models to perform both explicit recognition and implicit commonsense reasoning, grounded in real-world traffic rules and contextual dependencies. To fully exploit the reasoning potential of Multi-modal Large Language Models (MLLMs), we further propose CogniAnchor Fusion (CAF), a vision-language fusion module inspired by human-like scene anchoring mechanisms. CAF enables precise and efficient cross-modal interaction. Moreover, we propose the Assisted Decoupled Chain-of-Thought (AD-CoT) to enhance the reasoned thinking via CoT prompting and multi-task learning. Experimental results on RoadSceneVQA and CODA-LM benchmark show that the pipeline consistently improves both reasoning accuracy and computational efficiency, allowing the MLLM to achieve state-of-the-art performance in structural traffic perception and reasoning tasks. Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningyuan Xiao, Ziren Tang, Ningwei Ouyang, Shaofeng Liang, Yuxuan Fan, Wanjie Sun, Yutao Yue |
AAAI | 13 |
| 2026 | CoRe: Contrast and reconstruction combination self-supervised point cloud representation learning
Changyu Zeng, Jimin Xiao, Anh Nguyen 0003, Xuming Hu, Wei Wang 0042, Yutao Yue |
Expert Syst. Appl. | 6 |
| 2026 | Da Yu: Toward ASV-Based Image Captioning for Waterway Surveillance and Scene UnderstandingabstractAutomated waterway environment perception is crucial for enabling unmanned surface vessels (USVs) to understand their surroundings and make informed decisions. Most existing waterway perception models primarily focus on instance-level object perception paradigms (e.g., detection, segmentation). However, due to the complexity of waterway environments, current perception datasets and models fail to achieve global semantic understanding of waterways, limiting large-scale monitoring and structured log generation. With the advancement of vision-language models (VLMs), we leverage image captioning to introduce WaterCaption, the first captioning dataset specifically designed for waterway environments. WaterCaption focuses on fine-grained, multi-region long-text descriptions, providing a new research direction for visual geo-understanding and spatial scene cognition. Exactly, it includes 20.2k image-text pair data with 1.8 million vocabulary size. Additionally, we propose Da Yu, an edge-deployable multi-modal large language model for USVs, where we propose a novel vision-to-language projector called Nano Transformer Adaptor (NTA). NTA effectively balances computational efficiency with the capacity for both global and fine-grained local modeling of visual features, thereby significantly enhancing the model’s ability to generate long-form textual outputs. Da Yu achieves an optimal balance between performance and efficiency, surpassing state-of-the-art models on WaterCaption and several other captioning benchmarks. The project is available at https://github.com/GuanRunwei/WaterCaption. Runwei Guan, Ningwei Ouyang, Tianhao Xu, Shaofeng Liang, Yafeng Sun, Shang Gao 0012, Songning Lai, Shanliang Yao, Xuming Hu, Ryan Wen Liu, Yutao Yue, Hui Xiong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 12 |
| 2026 | From High-SNR Radar Signal to ECG: A Transfer Learning Model With Cardio-Focusing Algorithm for Scenarios With Limited DataabstractElectrocardiogram (ECG), as a crucial fine-grained cardiac feature, has been successfully recovered from radar signals in the literature, but the performance heavily relies on the high-quality radar signal and numerous radar-ECG pairs for training, restricting the applications in new scenarios due to data scarcity. Therefore, this work focuses on radar-based ECG recovery in new scenarios with limited data and proposes a cardio-focusing and-tracking (CFT) algorithm to precisely track the cardiac location to ensure an efficient acquisition of high quality radar signals. Furthermore, a transfer learning model (RFcardi) is proposed to extract cardio-related information from the radar signal without ECG ground truth based on the intrinsic sparsity of cardiac features, and only a few synchronous radar ECG pairs are required to fine-tune the pre-trained model for ECG recovery. The experimental results reveal that the proposed CFT can dynamically identify the cardiac location, and the RFcardi model can effectively generate faithful ECG recoveries after using a small number of radar-ECG pairs for training. The code and dataset will be made available after publication. Haocheng Zhao, Sijie Xiong, Rui Yang 0007, Eng Gee Lim, Yutao Yue |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Can GRPO Boost Complex Multimodal Table Understanding?abstractXiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu, Xiaobo Jin, Kaizhu Huang, Wei Wang, Yutao Yue, Xiaowei Huang, Qiufeng Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xiaoqiang Kang, Shengen Wu, Xiao-Bo Jin, Kaizhu Huang, Wei Wang 0042, Yutao Yue |
EMNLP | 8 |
| 2025 | Spatial-Temporal Perception with Causal Inference for Naturalistic Driving Action RecognitionabstractNaturalistic driving action recognition is essential for vehicle cabin monitoring systems. However, the complexity of real-world backgrounds presents significant challenges for this task, and previous approaches have struggled with practical implementation due to their limited ability to observe subtle behavioral differences and effectively learn inter-frame features from video. In this paper, we propose a novel Spatial-Temporal Perception (STP) architecture that emphasizes both temporal information and spatial relationships between key objects, incorporating a causal decoder to perform behavior recognition and temporal action localization. Without requiring multimodal input, STP directly extracts temporal and spatial distance features from RGB video clips. Subsequently, these dual features are jointly encoded by maximizing the expected likelihood across all possible permutations of the factorization order. By integrating temporal and spatial features at different scales, STP can perceive subtle behavioral changes in challenging scenarios. Additionally, we introduce a causal-aware module to explore relationships between video frame features, significantly enhancing detection efficiency and performance. We validate the effectiveness of our approach using two publicly available driver distraction detection benchmarks. The results demonstrate that our framework achieves state-of-the-art performance. Zhihao Shuai, Limin Yu, Yutao Yue |
ICASSP | 5 |
| 2025 | PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised LearningabstractFine-grained image classification has witnessed significant advancements with the advent of deep learning and computer vision technologies. However, the scarcity of detailed annotations remains a major challenge, especially in scenarios where obtaining high-quality labeled data is costly or time-consuming. To address this limitation, we introduce Precision-Enhanced Pseudo-Labeling (PEPL) approach specifically designed for fine-grained image classification within a semi-supervised learning framework. Our method leverages the abundance of unlabeled data by generating high-quality pseudo-labels that are progressively refined through two key phases: initial pseudo-label generation and semantic-mixed pseudo-label generation. These phases utilize Class Activation Maps (CAMs) to accurately estimate the semantic content and generate refined labels that capture the essential details necessary for fine-grained classification. By focusing on semantic-level information, our approach effectively addresses the limitations of standard data augmentation and image-mixing techniques in preserving critical fine-grained features. We achieve state-of-the-art performance on benchmark datasets, demonstrating significant improvements over existing semi-supervised strategies, with notable boosts in accuracy and robustness. Songning Lai, Lujundong Li, Zhihao Shuai, Runwei Guan, Yutao Yue |
ICASSP | 7 |
| 2025 | IMTS is Worth Time × Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction
Zhangyi Hu, Jiemin Wu, Mingqian Liao, Ninghui Feng, Songning Lai, Yutao Yue |
ICML | 8 |
| 2025 | Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression ComprehensionabstractEmbodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20, 558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar. Runwei Guan, Ruixiao Zhang 0001, Ningwei Ouyang, Ka Lok Man, Xiaohao Cai, Ming Xu 0011, Jeremy S. Smith, Eng Gee Lim, Yutao Yue, Hui Xiong 0001 |
ICRA | 10 |
| 2025 | DRIVE: Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous DrivingabstractRecent advancements in autonomous driving have seen a paradigm shift towards end-to-end learning paradigms, which map sensory inputs directly to driving actions, thereby enhancing the robustness and adaptability of autonomous vehicles. However, these models often sacrifice interpretability, posing significant challenges to trust, safety, and regulatory compliance. To address these issues, we introduce DRIVE – Dependable Robust Interpretable Visionary Ensemble Framework in Autonomous Driving, a comprehensive framework designed to improve the dependability and stability of explanations in end-to-end unsupervised autonomous driving models. Our work specifically targets the inherent instability problems observed in the Driving through the Concept Gridlock (DCG) model, which undermine the trustworthiness of its explanations and decisionmaking processes. We define four key attributes of DRIVE: consistent interpretability, stable interpretability, consistent output, and stable output. These attributes collectively ensure that explanations remain reliable and robust across different scenarios and perturbations. Through extensive empirical evaluations, we demonstrate the effectiveness of our framework in enhancing the stability and dependability of explanations, thereby addressing the limitations of current models. Our contributions include an in-depth analysis of the dependability issues within the DCG model, a rigorous definition of DRIVE with its fundamental properties, a framework to implement DRIVE, and novel metrics for evaluating the dependability of concept-based explainable autonomous driving models. These advancements lay the groundwork for the development of more reliable and trusted autonomous driving systems, paving the way for their broader acceptance and deployment in real-world applications. “We can only see a short distance ahead, but we can see plenty there that needs to be done.” – Alan Turing Songning Lai, Tianlang Xue, Hongru Xiao, Lijie Hu, Jiemin Wu, Ninghui Feng, Runwei Guan, Haicheng Liao, Zhenning Li 0001, Yutao Yue |
ICRA | 10 |
| 2025 | UniBEVFusion: Unified Radar-Vision Bevfusion for 3D Object Detectionabstract4D millimeter-wave (MMW) radar, which provides both height information and dense point cloud data over 3D MMW radar, has become increasingly popular in 3D object detection. In recent years, radar-vision fusion models have demonstrated performance close to that of LiDAR-based models, offering advantages in terms of lower hardware costs and better resilience in extreme conditions. However, many radar-vision fusion models treat radar as a sparse LiDAR, underutilizing radar-specific information. Additionally, these multi-modal networks are often sensitive to the failure of a single modality, particularly vision. To address these challenges, we propose the Radar Depth Lift-Splat-Shoot (RDL) module, which integrates radar-specific data into the depth prediction process, enhancing the quality of visual Bird's-Eye View (BEV) features. We further introduce a Unified Feature Fusion (UFF) approach that extracts BEV features across different modalities using shared module. To assess the robustness of multimodal models, we develop a novel Failure Test (FT) ablation experiment, which simulates vision modality failure by injecting Gaussian noise. We conduct extensive experiments on the View-of-Delft (VoD) and TJ4D datasets. The results demonstrated that our proposed Unified BEVFusion (UniBEVFusion) network significantly outperforms state-of-the-art models on the TJ4D dataset, with improvements of 3.96% in 3D and 4.17% in BEV object detection accuracy. Haocheng Zhao, Runwei Guan, Taoyu Wu, Ka Lok Man, Limin Yu, Yutao Yue |
ICRA | 6 |
| 2025 | CD^2: Constrained Dataset Distillation for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) receives significant attention from the public to perform classification continuously with a few training samples, which suffers from the key catastrophic forgetting problem. Existing methods usually employ an external memory to store previous knowledge and treat it with incremental classes equally, which cannot properly preserve previous essential knowledge. To solve this problem and inspired by recent distillation works on knowledge transfer, we propose a framework termed Constrained Dataset Distillation (CD^2) to facilitate FSCIL, which includes a dataset distillation module (DDM) and a distillation constraint module (DCM). Specifically, the DDM synthesizes highly condensed samples guided by the classifier, forcing the model to learn compacted essential class-related clues from a few incremental samples. The DCM introduces a designed loss to constrain the previously learned class distribution, which can preserve distilled knowledge more sufficiently. Extensive experiments on three public datasets show the superiority of our method against other state-of-the-art competitors. Kexin Bao, Daichi Zhang, Hansong Zhang 0003, Yutao Yue, Shiming Ge |
IJCAI | 5 |
| 2025 | NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave RadarabstractRecently, visual grounding and multi-sensors setting have been incorporated into perception system for terrestrial autonomous driving systems and Unmanned Surface Vessels (USVs), yet the high complexity of modern learning-based visual grounding model using multi-sensors prevents such model to be deployed on USVs in the real-life. To this end, we design a low-power multi-task model named NanoMVG for waterway embodied perception, guiding both camera and 4D millimeter-wave radar to locate specific object(s) through natural language. NanoMVG can perform both box-level and mask-level visual grounding tasks simultaneously. Compared to other visual grounding models, NanoMVG achieves highly competitive performance on the WaterVG dataset, particularly in harsh environments. Moreover, the real-world experiments with deployment of NanoMVG on embedded edge device of USV demonstrates its fast inference speed for real-time perception and capability of boasting ultra-low power consumption for long endurance. Runwei Guan, Liye Jia, Haocheng Zhao, Shanliang Yao, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Yutao Yue |
IROS | 10 |
| 2025 | BioD2C: A Dual-Level Semantic Consistency Constraint Framework for Biomedical VQA
Zhengyang Ji, Shang Gao 0012, Yifan Jia 0010, Yutao Yue |
MICCAI (10) | 5 |
| 2025 | ANT: Adaptive Neural Temporal-Aware Text-to-Motion ModelabstractWhile diffusion models advance text-to-motion generation, their static semantic conditioning ignores temporal-frequency demands: early denoising requires structural semantics for motion foundations while later stages need localized details for text alignment. This mismatch mirrors biological morphogenesis where developmental phases demand distinct genetic programs. Inspired by epigenetic regulation governing morphological specialization, we propose **(ANT)**, an **A**daptive **N**eural **T**emporal-Aware architecture. ANT orchestrates semantic granularity through: **(i) Semantic Temporally Adaptive (STA) Module:** Automatically partitions denoising into low-frequency structural planning and high-frequency refinement via spectral analysis. **(ii) Dynamic Classifier-Free Guidance scheduling (DCFG):** Adaptively adjusts conditional to unconditional ratio enhancing efficiency while maintaining fidelity. Extensive experiments show that ANT can be applied to various baselines, significantly improving model performance, and achieving state-of-the-art semantic alignment on StableMoFusion. Wenshuo Chen, Kuimou Yu, Haozhe Jia, Kaishen Yuan, Zexu Huang, Songning Lai, Hongru Xiao, Erhang Zhang, Lei Wang 0108, Yutao Yue |
ACM Multimedia | 11 |
| 2025 | Physics-Informed Representation Alignment for Sparse Radio-Map ReconstructionabstractWith the rapid development of wireless communication technology, the efficient utilization of spectrum resources, optimization of communication quality, and intelligent communication have become critical. Radio map reconstruction is essential for enabling advanced applications, yet challenges such as complex signal propagation and sparse observational data hinder accurate reconstruction in practical scenarios. Existing methods often fail to align physical constraints with data-driven features, particularly under sparse measurement conditions. To address these issues, we propose Physics-Aligned Radio Map Diffusion Model (PhyRMDM), a novel framework that establishes cross-domain representation alignment between physical principles and neural network features through dual learning pathways. The proposed model integrates Physics-Informed Neural Networks (PINNs) with a representation alignment mechanism that explicitly enforces consistency between Helmholtz equation constraints and environmental propagation patterns. Our architecture employs two synergistic U-Nets: the first ensures physical consistency by minimizing PDE residuals and boundary conditions through latent space alignment, while the second refines predictions via diffusion-based denoising with attention-guided feature fusion. This dual alignment strategy enables simultaneous satisfaction of wave propagation laws and data distribution characteristics. Experimental results demonstrate significant improvements over state-of-the-art methods, achieving NMSE of 0.0031 and RMSE of 0.0125 under Static Radio Map (SRM) conditions, and NMSE of 0.0047 with RMSE of 0.0146 in Dynamic Radio Map (DRM) scenarios. The proposed representation alignment paradigm provides 37.2% accuracy enhancement in ultra-sparse cases (1% sampling rate), confirming its effectiveness in bridging physics-based modeling and deep learning for radio map reconstruction. These advancements establish a new framework for sparse signal environment characterization, with direct applications in 5G/6G network optimization and intelligent spectrum management. The code can be found on the website: https://github.com/Hxxxz0/RMDM Haozhe Jia, Wenshuo Chen, Lei Wang 0108, Hongru Xiao, Nanqian Jia, Keming Wu, Songning Lai, Yutao Yue |
ACM Multimedia | 10 |
| 2025 | From Guesswork to Guarantee: Towards Faithful Multimedia Web Forecasting with TimeSieveabstractThe domain of time series forecasting has gained significant attention due to its critical applications in multimedia-rich web traffic (including video streaming workloads and dynamic content delivery) and cross-platform advertisement click predictions, which are essential for web operations planning. While models like TimeSieve have demonstrated strong capabilities in predicting web visitation metrics, they suffer from critical unfaithfulness issues, including sensitivity to random seeds, input noise, layer noise, and parametric perturbations. To address these limitations, we propose Faithful TimeSieve (FTS), an enhanced framework designed to improve prediction reliability and robustness. Our approach systematically detects and mitigates unfaithfulness in TimeSieve, significantly enhancing its stability and consistency. Experimental results demonstrate that FTS substantially improves the model's faithfulness, setting a new standard for temporal forecasting methods. This advancement not only increases TimeSieve's reliability but also contributes to more robust temporal modeling, particularly crucial for web traffic forecasting where prediction accuracy directly impacts operational decisions. Our work thus represents a significant step toward more dependable time series predictions in web-related applications. Songning Lai, Ninghui Feng, Jiechao Gao, Hao Wang 0220, Haochen Sui, Xin Zou 0001, Wenshuo Chen, Lijie Hu, Hang Zhao 0010, Xuming Hu, Yutao Yue |
ACM Multimedia | 12 |
| 2025 | Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck ModelsabstractConcept Bottleneck Models (CBMs) enhance the interpretability of AI systems, particularly by bridging visual input with human-understandable concepts, effectively acting as a form of multimodal interpretability model. However, existing CBMs typically assume static datasets, which fundamentally limits their adaptability to real-world, continuously evolving multimodal data streams. To address this, we define a novel continual learning task for CBMs: simultaneously handling concept-incremental and class-incremental learning. This task requires models to continuously acquire new concepts (often representing cross-modal attributes) and classes while robustly preserving previously learned knowledge. To tackle this challenging problem, we propose CONceptual Continual Incremental Learning (CONCIL), a novel framework that fundamentally re-imagines concept and decision layer updates as linear regression problems. This reformulation eliminates the need for gradient-based optimization, thereby effectively preventing catastrophic forgetting. Crucially, CONCIL relies solely on recursive matrix operations, rendering it highly computationally efficient and well-suited for real-time and large-scale multimodal data applications. Experimental results compellingly demonstrate that CONCIL achieves ''absolute knowledge memory'' and significantly surpasses the performance of traditional CBM methods in both concept- and class-incremental settings, thus establishing a new paradigm for continual learning in CBMs, particularly valuable for dynamic multimodal understanding. Songning Lai, Mingqian Liao, Zhangyi Hu, Wenshuo Chen, Hongru Xiao, Jianheng Tang 0001, Haicheng Liao, Yutao Yue |
ACM Multimedia | 9 |
| 2025 | CFSSeg: Closed-Form Solution for Class-Incremental Semantic Segmentation of 2D Images and 3D Point Cloudsabstract2D images and 3D point clouds are foundational data types for multimedia applications, including real-time video analysis, augmented reality (AR), and 3D scene understanding. Class-incremental semantic segmentation (CSS) requires incrementally learning new semantic categories while retaining prior knowledge. Existing methods typically rely on computationally expensive training based on stochastic gradient descent, employing complex regularization or exemplar replay. However, stochastic gradient descent-based approaches inevitably update the model's weights for past knowledge, leading to catastrophic forgetting, a problem exacerbated by pixel/point-level granularity. To address these challenges, we propose CFSSeg, a novel exemplar-free approach that leverages a closed-form solution, offering a practical and theoretically grounded solution for continual semantic segmentation tasks. This eliminates the need for iterative gradient-based optimization and storage of past data, requiring only a single pass through new samples per step. It not only enhances computational efficiency but also provides a practical solution for dynamic, privacy-sensitive multimedia environments. Extensive experiments on 2D and 3D benchmark datasets such as Pascal VOC2012, S3DIS, and ScanNet demonstrate CFSSeg's superior performance. Jianyu Qi, Songning Lai, Linpu Lv, Kejia Fan, Jianheng Tang 0001, Yutao Yue, Dongzhan Zhou, Yunhuai Liu, Huiping Zhuang |
ACM Multimedia | 8 |
| 2025 | Text2Weight: Bridging Natural Language and Neural Network Weight SpacesabstractHow far are we really from automatically generating neural networks? While neural network weight generation shows promise, current approaches struggle with generalization to unseen tasks and practical application exploration. To address this, we propose T2W, a diffusion transformer framework that generates task-specific weights conditioned on natural language descriptions. T2W hierarchically processes network parameters into uniform blocks, integrates text embeddings from CLIP via a prior attention mechanism, and employs adversarial training with weight-space augmentation to enhance generalization. Experiments on Cifar100, Caltech256, and TinyImageNet demonstrate T2W's ability to produce high-quality weights for unseen tasks, outperforming optimization-based initialization and enabling novel applications such as weight enhancement and text-guided model fusion. Our work bridges textual semantics with weight-space dynamics, supported by an open-source dataset of text-weight pairs, advancing the practicality of generative models in neural network parameter synthesis. Our code is available on https://github.com/TianSuya/T2W. Wenshuo Chen, Zexi Li 0001, Songning Lai, Jiemin Wu, Yutao Yue |
ACM Multimedia | 6 |
| 2025 | Referring flexible image restoration
Runwei Guan, Rongsheng Hu, Zhuhao Zhou, Tianlang Xue, Ka Lok Man, Jeremy S. Smith, Eng Gee Lim, Weiping Ding 0001, Yutao Yue |
Expert Syst. Appl. | 9 |
| 2025 | RePaIR: Repaired pruning at initialization resilienceabstractOver the past decade, the size of neural network models has gradually increased in both breadth and depth, leading to a growing interest in the application of neural network pruning. Unstructured pruning provides fine-grained sparsity and achieves better inference acceleration under specific hardware support. Unstructured Pruning at Initialization (PaI) optimizes the iterative pruning pipeline, but sparse weights increase the risk of underfitting during training. More importantly, almost all PaI algorithms focus only on obtaining the best pruning mask without considering whether the retained weights are suitable for training. Introducing Lipschitz constants during model initialization can reduce the risk of model underfitting and overfitting. As a result, we firstly analyze the impact of Lipschitz initialization on model training and propose the Repaired Initialization (ReI) algorithm for common modules with BatchNorm. Then, we utilize the same idea to repair the weight of unstructured pruned model, and name it Repaired Pruning at Initialization Resilience (RePaIR) algorithm. Extensive experiments and demonstrate that our proposed ReI and RePaIR can improve the training robustness of unpruned and pruned models, respectively, and achieve up to 1.7% accuracy gain with the same sparse pruning mask on TinyImageNet. Furthermore, we provide an improved SynFlow algorithm called Repair SynFlow (ReSynFlow), which employs Lipschitz scaling to overcome the problem of score computation in deeper models. ReSynFlow can effectively improve the maximum compression rate and is suitable for deeper models, with an accuracy improvement of up to 1.3% compared to the SynFlow algorithm on TinyImageNet. Haocheng Zhao, Runwei Guan, Ka Lok Man, Limin Yu, Yutao Yue |
Neural Networks | 5 |
| 2025 | Corrections to "A Bidirectional Deep Learning Approach for Designing MEMS Sensors"
Xiong Cheng, Pengfei Zhang 0018, Daying Sun, Wenhua Gu, Yutao Yue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | WaterVG: Waterway Visual Grounding Based on Text-Guided Vision and mmWave RadarabstractWaterway perception is critical for the special operations and autonomous navigation of Unmanned Surface Vessels (USVs), but current perception schemes are sensor-based, neglecting the interaction between humans and USVs for embodied perception in various operations. Therefore, inspired by visual grounding, we present WaterVG, the inaugural visual grounding dataset tailored for USV-based waterway perception guided by human prompts. WaterVG contains a wealth of prompts describing multiple targets, with instance-level annotations, including bounding boxes and masks. Specifically, WaterVG comprises 11,568 samples and 34,987 referred targets, integrating both visual and radar characteristics. The text-guided two-sensor pattern provides a fine granularity of text prompts aligned with the visual and radar features of the referent targets, containing both qualitative and numeric descriptions. To enhance the endurance and maintain the normal operations of USVs in open waterways, we propose Potamoi, a low-power visual grounding model. Potamoi is a multi-task model employing a sophisticated Phased Heterogeneous Modality Fusion (PHMF) mechanism, which includes Adaptive Radar Weighting (ARW) and Multi-Head Slim Cross Attention (MHSCA). The ARW module utilizes a gating mechanism to adaptively extract essential radar features for fusion with visual inputs, ensuring prompt alignment. MHSCA, characterized by its low parameter count and computational efficiency (FLOPs), effectively integrates contextual information from both sensors with linguistic features, delivering outstanding performance in visual grounding tasks. Comprehensive experiments and evaluations on WaterVG demonstrate that Potamoi achieves state-of-the-art results compared to existing methods. The project is available athttps://github.com/GuanRunwei/WaterVG. Runwei Guan, Liye Jia, Shanliang Yao, Fengyufan Yang, Erick Purwanto, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Xuming Hu, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 12 |
| 2025 | Exploring Radar Data Representations in Autonomous Driving: A Comprehensive ReviewabstractWith the rapid advancements of sensor technology and deep learning, autonomous driving systems are providing safe and efficient access to intelligent vehicles as well as intelligent transportation. Among these equipped sensors, the radar sensor plays a crucial role in providing robust perception information in diverse environmental conditions. This review focuses on exploring different radar data representations utilized in autonomous driving systems. Firstly, we introduce the capabilities and limitations of the radar sensor by examining the working principles of radar perception and signal processing of radar measurements. Then, we delve into the generation process of five radar representations, including the ADC signal, radar tensor, point cloud, grid map, and micro-Doppler signature. For each radar representation, we examine the related datasets, methods, advantages and limitations. Furthermore, we discuss the challenges faced in these data representations and propose potential research directions. Above all, this comprehensive review offers an in-depth insight into how these representations enhance autonomous system capabilities, providing guidance for radar perception researchers. To facilitate retrieval and comparison of different data representations, datasets and methods, we provide an interactive website at https://radar-camera-fusion.github.io/radar. Shanliang Yao, Runwei Guan, Zitian Peng, Chenhang Xu, Yilu Shi, Weiping Ding 0001, Eng Gee Lim, Yong Yue 0001, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 13 |
| 2025 | radarODE: An ODE-Embedded Deep Learning Model for Contactless ECG Reconstruction From Millimeter-Wave RadarabstractRadar-based cardiac monitoring has become a popular research direction recently, but the fine-grained electrocardiogram (ECG) signal is still hard to reconstruct from millimeter-wave radar signal. The key obstacle is to decouple cardiac activities in the electrical domain (i.e., ECG) from that in the mechanical domain (i.e., heartbeat), and most existing research only uses purely data-driven methods to map such domain transformation as a black box. Therefore, this work first proposes a signal model that considers the fine-grained cardiac feature sensed by radar, and a novel deep learning framework called radarODE is designed to extract both temporal and morphological features for generating ECG. In addition, ordinary differential equations are embedded in radarODE as a decoder to provide morphological prior, helping the convergence of the model training and improving the robustness under body movements. After being validated on the dataset, the proposed radarODE achieves better performance compared with the benchmark in terms of missed detection rate, root mean square error, Pearson correlation coefficient with improvements of 9%, 16% and 19%, respectively. The validation results imply that radarODE is capable of recovering ECG signals from radar signals with high fidelity and can potentially be implemented in real-life scenarios Runwei Guan, Rui Yang 0007, Yutao Yue, Eng Gee Lim |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language ModelabstractProjecting visual features into word embedding space has become a significant fusion strategy adopted by Multimodal Large Language Models (MLLMs).However, its internal mechanisms have yet to be explored.Inspired by multilingual research, we identify domain-specific neurons in multimodal large language models.Specifically, we investigate the distribution of domain-specific neurons and the mechanism of how MLLMs process features from diverse domains.Furthermore, we propose a threestage mechanism for language model modules in MLLMs when handling projected image features, and verify this hypothesis using logit lens.Extensive experiments indicate that while current MLLMs exhibit Visual Question Answering (VQA) capability, they may not fully utilize domain-specific information.Manipulating domain-specific neurons properly will result in a 10% change of accuracy at most, shedding light on the development of cross-domain, all-encompassing MLLMs in the future.The source code is available at this URL. Jiahao Huo, Boren Hu, Yutao Yue, Xuming Hu |
EMNLP | 4 |
| 2024 | ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave RadarabstractPanoptic Driving Perception (PDP) is critical for the autonomous navigation of Unmanned Surface Vehicles (USVs). A PDP model typically integrates multiple tasks, necessitating the simultaneous and robust execution of various perception tasks to facilitate downstream path planning. The fusion of visual and radar sensors is currently acknowledged as a robust and cost-effective approach. However, most existing research has primarily focused on fusing visual and radar features dedicated to object detection or utilizing a shared feature space for multiple tasks, neglecting the individual representation differences between various tasks. To address this gap, we propose a pair of Asymmetric Fair Fusion (AFF) modules with favorable explainability designed to efficiently interact with independent features from both visual and radar modalities, tailored to the specific requirements of object detection and semantic segmentation tasks. The AFF modules treat image and radar maps as irregular point sets and transform these features into a crossed-shared feature space for multitasking, ensuring equitable treatment of vision and radar point cloud features. Leveraging AFF modules, we propose a novel and efficient PDP model, ASY-VRNet, which processes image and radar features based on irregular super-pixel point sets. Additionally, we propose an effective multi-task learning method specifically designed for PDP models. Compared to other lightweight models, ASY-VRNet achieves state-of-the-art performance in object detection, semantic segmentation, and drivable-area segmentation on the WaterScenes benchmark. Our project is publicly available at https://github.com/GuanRunwei/ASY-VRNet. Runwei Guan, Shanliang Yao, Ka Lok Man, Yong Yue 0001, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
IROS | 8 |
| 2024 | Self-supervised learning for point cloud data: A surveyabstract3D point clouds are a crucial type of data collected by LiDAR sensors and widely used in transportation applications due to its concise descriptions and accurate localization. Deep neural networks (DNNs) have achieved remarkable success in processing large amount of disordered and sparse 3D point clouds, especially in various computer vision tasks, such as pedestrian detection and vehicle recognition. Among all the learning paradigms, Self-Supervised Learning (SSL), an unsupervised training paradigm that mines effective information from the data itself, is considered as an essential solution to solve the time-consuming and labor-intensive data labelling problems via smart pre-training task design. This paper provides a comprehensive survey of recent advances on SSL for point clouds. We first present an innovative taxonomy, categorizing the existing SSL methods into four broad categories based on the pretexts’ characteristics. Under each category, we then further categorize the methods into more fine-grained groups and summarize the strength and limitations of the representative methods. We also compare the performance of the notable SSL methods in literature on multiple downstream tasks on benchmark datasets both quantitatively and qualitatively. Finally, we propose a number of future research directions based on the identified limitations of existing SSL research on point clouds. Changyu Zeng, Wei Wang 0042, Anh Nguyen 0003, Jimin Xiao, Yutao Yue |
Expert Syst. Appl. | 5 |
| 2024 | FindVehicle and VehicleFinder: a NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval systemabstractAbstract Natural language (NL) based vehicle retrieval is a task aiming to retrieve a vehicle that is most consistent with a given NL query from among all candidate vehicles. Because NL query can be easily obtained, such a task has a promising prospect in building an interactive intelligent traffic system (ITS). Current solutions mainly focus on extracting both text and image features and mapping them to the same latent space to compare the similarity. However, existing methods usually use dependency analysis or semantic role-labelling techniques to find keywords related to vehicle attributes. These techniques may require a lot of pre-processing and post-processing work, and also suffer from extracting the wrong keyword when the NL query is complex. To tackle these problems and simplify, we borrow the idea from named entity recognition (NER) and construct FindVehicle, a NER dataset in the traffic domain. It has 42.3k labelled NL descriptions of vehicle tracks, containing information such as the location, orientation, type and colour of the vehicle. FindVehicle also adopts both overlapping entities and fine-grained entities to meet further requirements. To verify its effectiveness, we propose a baseline NL-based vehicle retrieval model called VehicleFinder. Our experiment shows that by using text encoders pre-trained by FindVehicle, VehicleFinder achieves 87.7% precision and 89.4% recall when retrieving a target vehicle by text command on our homemade dataset based on UA-DETRAC [1]. From loading the command into VehicleFinder to identifying whether the target vehicle is consistent with the command, the time cost is 279.35 ms on one ARM v8.2 CPU and 93.72 ms on one RTX A4000 GPU, which is much faster than the Transformer-based system. The dataset is open-source via the link https://github.com/GuanRunwei/FindVehicle , and the implementation can be found via the link https://github.com/GuanRunwei/VehicleFinder-CTIM . Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
Multim. Tools Appl. | 9 |
| 2024 | WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water SurfacesabstractAutonomous driving on water surfaces plays an essential role in executing hazardous and time-consuming missions, such as maritime surveillance, survivor rescue, environmental monitoring, hydrography mapping and waste cleaning. This work presents WaterScenes, the first multi-task 4D radar-camera fusion dataset for autonomous driving on water surfaces. Equipped with a 4D radar and a monocular camera, our Unmanned Surface Vehicle (USV) proffers all-weather solutions for discerning object-related information, including color, shape, texture, range, velocity, azimuth, and elevation. Focusing on typical static and dynamic objects on water surfaces, we label the camera images and radar point clouds at pixel-level and point-level, respectively. In addition to basic perception tasks, such as object detection, instance segmentation and semantic segmentation, we also provide annotations for free-space segmentation and waterline segmentation. Leveraging the multi-task and multi-modal data, we conduct benchmark experiments on the uni-modality of radar and camera, as well as the fused modalities. Experimental results demonstrate that 4D radar-camera fusion can considerably improve the accuracy and robustness of perception on water surfaces, especially in adverse lighting and weather conditions. WaterScenes dataset is public onhttps://waterscenes.github.io. Shanliang Yao, Runwei Guan, Zhaodong Wu, Yi Ni, Zile Huang, Ryan Wen Liu, Yong Yue 0001, Weiping Ding 0001, Eng Gee Lim, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 14 |
| 2023 | PLSR: Unstructured Pruning with Layer-Wise Sparsity RatioabstractIn the current era of multi-modal and large models gradually revealing their potential, neural network pruning has emerged as a crucial means of model compression. It is widely recognized that models tend to be over-parameterized, and pruning enables the removal of unimportant weights, leading to improved inference speed while preserving accuracy. From early methods such as gradient-based, and magnitude-based pruning to modern algorithms like iterative magnitude pruning, lottery ticket hypothesis, and pruning at initialization, researchers have strived to increase the compression ratio of model parameters while maintaining high accuracy. Currently, mainstream algorithms focus on the global pruning of neural networks using various scoring functions, followed by different pruning strategies to enhance the accuracy of sparse model. Recent studies have shown that random pruning with varying layer-wise sparsity ratio has achieved robust results for large models and out-of-distribution data. Based on this discovery, we propose a new score called FeatIO, which is based on module input and output feature map sizes. As a score function used in PaI, FeatIO surpasses the performance of other PaI score functions. Additionally, we propose a novel pruning strategy called Pruning with Layer-wise Sparsity Ratio (PLSR), which conbines the layer-wise sparsity ratios and magnitude-based score function, resulting in optimal evaluation performance. Almost all algorithms exhibit improved performance when using our novel pruning strategy. The combination of PLSR and FeatIO consistently outperforms other algorithms in testing, demonstrating the significant potential of our proposed approach. Our code will be available here. Haocheng Zhao, Limin Yu, Runwei Guan, Liye Jia, Junqing Zhang, Yutao Yue |
ICMLA | 6 |
| 2023 | A Bidirectional Deep Learning Approach for Designing MEMS SensorsabstractTo achieve the desired characteristics for MEMS sensors, the traditional design process obtains the geometrical parameters based on complex theoretical calculations and interactive finite-element (FE) simulations, which are time consuming and data consuming. To solve the above problems, a data-driven bidirectional design approach based on the deep learning (DL) method is introduced to improve the design efficiency of MEMS sensors in this work. By using the piezoresistive acceleration sensor as a design example, the forward artificial neural network (ANN) with the sensor geometrical parameters as the input and the sensor performance as the output is trained and realized by using 1000 groups of data collected through FE simulation. This forward ANN can accurately predict the sensor performance, including the measurement range, sensitivity, and resonant frequency. In addition, the inverse ANN with the sensor performance as the input and the sensor geometrical parameters as the output is also achieved by using a tandem network. This inverse ANN can provide the geometrical parameters directly and instantly according to the target performance. Both the forward and inverse networks cost only about 6 ms for each task and the mean relative errors are less than 3%. The high efficiency and low relative error indicate that DL is a promising approach to improve the design efficiency for MEMS sensors. Xiong Cheng, Pengfei Zhang 0018, Daying Sun, Wenhua Gu, Yutao Yue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | MAN and CAT: mix attention to nn and concatenate attention to YOLO
Runwei Guan, Ka Lok Man, Haocheng Zhao, Ruixiao Zhang 0001, Shanliang Yao, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
J. Supercomput. | 8 |