Yuxu Lu

dblp:272/5292 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-9845-7516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ShipTraj-R1: Reinforcing Ship Trajectory Prediction in Large Language Models via Group Relative Policy Optimization
Yang Zhan 0007, Yuxu Lu, Yan Li 0110
PAKDD (1)4
2026 Uncertainty-aware vessel trajectory prediction for heterogeneous data fusion in internet of things-driven smart waterways
Yuxu Lu, Kaisen Yang, Dong Yang 0003, Haifeng Ding, Jinxian Weng
Eng. Appl. Artif. Intell.1
2026 USRNet: Unified scene recovery network for image restoration under multiple adverse weather conditions
Yuxu Lu, Ai Chen, Dong Yang 0003, Ryan Wen Liu
Pattern Recognit.1
2026 Graph Learning-Driven Multi-Vessel Association: Fusing Multimodal Data for Maritime Intelligence
abstract
Ensuring maritime safety and optimizing traffic management in increasingly crowded and complex waterways require effective waterway monitoring. However, current methods struggle with challenges arising from multimodal data, such as dimensional disparities, mismatched target counts, vessel scale variations, occlusions, and asynchronous data streams from systems like the automatic identification system (AIS) and closed-circuit television (CCTV). Traditional multi-vessel association methods often struggle with these complexities, particularly in densely trafficked waterways. To overcome these issues, we propose a graph learning-driven multi-vessel association (named GMvA) method tailored for maritime multimodal data fusion. By integrating AIS and CCTV data, GMvA leverages time series learning and graph neural networks to capture the spatiotemporal features of vessel trajectories effectively. To enhance feature representation, the proposed method incorporates temporal graph attention and spatiotemporal attention, effectively capturing both local and global vessel interactions. Furthermore, a multi-layer perceptron-based uncertainty fusion module computes robust similarity scores, and the Hungarian algorithm is adopted to ensure globally consistent and accurate target matching. To validate the efficacy of our method, we have also constructed a new maritime multimodal dataset (termed MaritimeMmD) for data fusion. Extensive experiments demonstrate that GMvA delivers superior accuracy and robustness in multi-vessel association, outperforming existing methods even in challenging scenarios with high vessel density and incomplete or unevenly distributed AIS and CCTV data. The source dataset and code are available athttps://github.com/LouisYxLu/GMvA
Yuxu Lu, Kaisen Yang, Dong Yang 0003, Haifeng Ding, Jinxian Weng, Ryan Wen Liu
IEEE Trans. Intell. Transp. Syst.1
2025 VL-UR: Vision-Language-guided Universal Restoration of Images Degraded by Adverse Weather Conditions
abstract
Image restoration is critical for improving the quality of degraded images, which is vital for applications like autonomous driving, security surveillance, and digital content enhancement. However, existing methods are often tailored to specific degradation scenarios, limiting their adaptability to the diverse and complex challenges in real-world environments. Moreover, real-world degradations are typically non-uniform, highlighting the need for adaptive and intelligent solutions. To address these issues, we propose a novel vision-language-guided universal restoration (VL-UR) framework. VL-UR leverages a zero-shot contrastive language-image pre-training (CLIP) model to enhance image restoration by integrating visual and semantic information. A scene classifier is introduced to adapt CLIP, generating high-quality language embeddings aligned with degraded images while predicting degraded types for complex scenarios. Extensive experiments across eleven diverse degradation settings demonstrate VL-UR’s state-of-the-art performance, robustness, and adaptability. This positions VL-UR as a transformative solution for modern image restoration challenges in dynamic, real-world environments.
Yuxu Lu, Hushan Yu, Dong Yang 0003
ICME2
2025 Neptune-X: Active X-to-Maritime Generation for Universal Maritime Object Detection
abstract
Maritime object detection is essential for navigation safety, surveillance, and autonomous operations, yet constrained by two key challenges: the scarcity of annotated maritime data and poor generalization across various maritime attributes (e.g., object category, viewpoint, location, and imaging environment). To address these challenges, we propose Neptune-X, a data-centric generative-selection framework that enhances training effectiveness by leveraging synthetic data generation with task-aware sample selection. From the generation perspective, we develop X-to-Maritime, a multi-modality-conditioned generative model that synthesizes diverse and realistic maritime scenes. A key component is the Bidirectional Object-Water Attention module, which captures boundary interactions between objects and their aquatic surroundings to improve visual fidelity. To further improve downstream tasking performance, we propose Attribute-correlated Active Sampling, which dynamically selects synthetic samples based on their task relevance. To support robust benchmarking, we construct the Maritime Generation Dataset, the first dataset tailored for generative maritime learning, encompassing a wide range of semantic conditions. Extensive experiments demonstrate that our approach sets a new benchmark in maritime scene synthesis, significantly improving detection accuracy, particularly in challenging and previously underrepresented settings. The code is available at https://github.com/gy65896/Neptune-X.
Yu Guo 0008, Shengfeng He, Yuxu Lu, Haonan An 0001, Yihang Tao, Huilin Zhu, Jingxian Liu, Yuguang Fang
NeurIPS3
2025 Multi-view adaptive image enhancement with hierarchical attention for complex underground mining scenes
Wenyu Xu, Yuxu Lu, Dong Yang 0003
Expert Syst. Appl.3
2024 OneRestore: A Universal Restoration Framework for Composite Degradation
Yu Guo 0008, Yuan Gao 0015, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, Shengfeng He
ECCV (19)3
2024 AoSRNet: All-in-One Scene Recovery Networks via multi-knowledge integration
Yuxu Lu, Dong Yang 0003, Yuan Gao 0015, Ryan Wen Liu, Jun Liu 0012, Yu Guo 0008
Knowl. Based Syst.1
2024 Real-Time Multi-Scene Visibility Enhancement for Promoting Navigational Safety of Vessels Under Complex Weather Conditions
abstract
The visible-light camera, which is capable of environment perception and navigation assistance, has emerged as an essential imaging sensor for marine surface vessels in intelligent waterborne transportation systems (IWTS). However, the visual imaging quality inevitably suffers from several kinds of degradations (e.g., limited visibility, low contrast, color distortion, etc.) under complex weather conditions (e.g., haze, rain, and low-lightness). The degraded visual information will accordingly result in inaccurate environment perception and delayed operations for navigational risk. To promote the navigational safety of vessels, many computational methods have been presented to perform visual quality enhancement under poor weather conditions. However, most of these methods are essentially specific-purpose implementation strategies, only available for one specific weather type. To overcome this limitation, we propose to develop a general-purpose multi-scene visibility enhancement method, i.e., edge reparameterization- and attention-guided neural network (ERANet), to adaptively restore the degraded images captured under different weather conditions. In particular, our ERANet simultaneously exploits the channel attention, spatial attention, and reparameterization technology to enhance the visual quality while maintaining low computational cost. Extensive experiments conducted on standard and IWTS-related datasets have demonstrated that our ERANet could outperform several representative visibility enhancement methods in terms of both imaging quality and computational efficiency. The superior performance of IWTS-related object detection and scene segmentation could also be steadily obtained after ERANet-based visibility enhancement under complex weather conditions.
Ryan Wen Liu, Yuxu Lu, Yuan Gao 0015, Yu Guo 0008, Wenqi Ren, Fenghua Zhu, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.2
2023 Deep Network-Enabled Haze Visibility Enhancement for Visual IoT-Driven Intelligent Transportation Systems
abstract
The Internet of Things (IoT) has recently emerged as a revolutionary communication paradigm where a large number of objects and devices are closely interconnected to enable smart industrial environments. The tremendous growth of visual sensors can significantly promote the traffic situational awareness, traffic safety management, and intelligent vehicle navigation in intelligent transportation systems (ITSs). However, due to the absorption and scattering of light by the turbid medium in atmosphere, the visual IoT inevitably suffers from imaging quality degradation, e.g., contrast reduction, color distortion, etc. This negative impact can not only reduce the imaging quality, but also bring challenges for the deployment of several high-level vision tasks (e.g., object detection, tracking, recognition, etc.) in the ITS. To improve imaging quality under the hazy environment, we propose a deep network-enabled three-stage dehazing network (termed TSDNet) for promoting the visual IoT-driven ITS. In particular, the proposed TSDNet mainly contains three parts, i.e., multiscale attention module for estimating the hazy distribution in the RGB image domain, two-branch extraction module for learning the hazy features, and multifeature fusion module for integrating all characteristic information and reconstructing the haze-free image. Numerous experiments have been implemented on synthetic and real-world imaging scenarios. Dehazing results illustrated that our TSDNet remarkably outperformed several state-of-the-art methods in terms of both qualitative and quantitative evaluations. The high-accuracy object detection results have also demonstrated the superior dehazing performance of the TSDNet under hazy atmosphere conditions. The source code is available athttps://github.com/gy65896/TSDNet.
Ryan Wen Liu, Yu Guo 0008, Yuxu Lu, Kwok Tai Chui, Brij B. Gupta
IEEE Trans. Ind. Informatics3
2023 GradDT: Gradient-Guided Despeckling Transformer for Industrial Imaging Sensors
abstract
The speckle noise is a granular disturbance that often brings negative side effects on the detection and recognition of targets of interest in industrial imaging sensors. From the statistical point of view, this type of noise can be modeled as a multiplicative formula. The nonlinear multiplicative property makes despeckling more intractable with respect to noise reduction and details preservation. To blindly remove the undesirable speckle noise, we combine the gradient model and machine learning technology for despeckling. In particular, we first introduce the logarithmic transformation to transform the multiplicative speckle noise into an additive version. A gradient-guided despeckling transformer (termed GradDT) is then proposed to blindly reduce the additive noise in the transformed noisy images. To be specific, the proposed method mainly includes two modules, i.e., the spatial feature extraction module (SFEM) and the efficient transformer module (ETM). The SFEM can extract the spatial feature of speckle noise and the gradient maps corresponding to the noise-free image. The ETM module can calculate the spatial domain's cross-channel cross-covariance and produce global attention maps to reconstruct the sharp image. The proposed GradDT thus can effectively distinguish the speckle noise and vital image features (e.g., edge and texture) to balance the degree of noise suppression and details preservation. Extensive experiments have been implemented on both synthetic and realistic degraded images. Compared with several state-of-the-art speckle noise reduction methods, our GradDT could generate superior imaging performance in terms of both quantitative evaluation and visual quality.
Yuxu Lu, Yu Guo 0008, Ryan Wen Liu, Kwok Tai Chui, Brij B. Gupta
IEEE Trans. Ind. Informatics1
2023 Asynchronous Trajectory Matching-Based Multimodal Maritime Data Fusion for Vessel Traffic Surveillance in Inland Waterways
abstract
The automatic identification system (AIS) and video cameras have been widely exploited for vessel traffic surveillance in inland waterways. The AIS data could provide vessel identity and dynamic information on vessel position and movements. In contrast, the video data could describe the visual appearances of moving vessels without knowing the information on identity, position, movements, etc. To further improve vessel traffic surveillance, it becomes necessary to fuse the AIS and video data to simultaneously capture the visual features, identity, and dynamic information for the vessels of interest. However, the performance of AIS and video data fusion is susceptible to issues such as data spatial difference, message asynchronous transmission, visual object occlusion, etc. In this work, we propose a deep learning-based simple online and real-time vessel data fusion method (termed DeepSORVF). We first extract the AIS-and video-based vessel trajectories, and then propose an asynchronous trajectory matching method to fuse the AIS-based vessel information with the corresponding visual targets. In addition, by combining the AIS-and video-based movement features, we also present a prior knowledge-driven anti-occlusion method to yield accurate and robust vessel tracking results under occlusion conditions. To validate the efficacy of our DeepSORVF, we have also constructed a new benchmark dataset (termed FVessel) for vessel detection, tracking, and data fusion. It consists of many videos and the corresponding AIS data collected in various weather conditions and locations. The experimental results have demonstrated that our method is capable of guaranteeing high-reliable data fusion and anti-occlusion vessel tracking. The DeepSORVF code and FVessel dataset are publicly available at https://github.com/gy65896/DeepSORVF and https://github.com/gy65896/FVessel, respectively.
Yu Guo 0008, Ryan Wen Liu, Jingxiang Qu, Yuxu Lu, Fenghua Zhu
IEEE Trans. Intell. Transp. Syst.4
2022 Rep-Enhancer: Re-parameterizing Neural Network for Real-time Low-light Enhancement in Visual Maritime Surveillance
abstract
Vision-based maritime surveillance has become an essential part of the vessel traffic services system. The images collected in low-light maritime conditions often suffer from poor visibility. These images may significantly degenerate the performance of high-level visual tasks and increase the uncertainty in maritime surveillance. To address this problem, we propose a lightweight neural network (Rep-Enhancer) for low-light image enhancement. Specifically, we first design a re-parameterizable multi-branch edge extraction module, i.e., spatial domain-oriented convolution block (SDCB). Furthermore, skip connections and spatial attention operations are employed to strengthen the features. By exploiting these well-strengthened edge features, we can enhance the low-light images effectively with the encoder-decoder structure. The experimental results have shown that Rep-Enhancer can enhance the low-light image qualifiedly while maintaining great inference efficiency.
Xijing Li, Yuxu Lu, Yu Guo 0008, Jingxiang Qu, Ryan Wen Liu
EUC2
2022 MTRBNet: Multi-Branch Topology Residual Block-Based Network for Low-Light Enhancement
abstract
The learning-based low-light image enhancement methods have remarkable performance due to the robust feature learning and mapping capabilities. This paper proposes a multi-branch topology residual block (MTRB)-based network (MTRBNet), which can alleviate training difficulties and more efficiently use the parameters between neurons. Compared with the previous residual block, the proposed MTRB increases the width of the network and simultaneously transmits information along with the depth and width directions, which can effectively select network nodes to promote the network learning capacity. Meanwhile, the feature information of neighbor nodes is transferred to each other, thereby maximizing the information flow of the convolution unit. The proposed information connection and feedback mechanism can improve the network’s ability to capture the global and local features. We analyze the pros and cons of two multi-feature fusion strategies (i.e., addition and concatenation) and three normalization methods on the quantitative results. In addition, we embed our MTRB into traditional Encoder-Decoder structure to improve the image enhancement results under different low-light imaging conditions. Experiments on the LOL image dataset have demonstrated that our MTRBNet achieves superior performance compared with several state-of-the-art methods.
Yuxu Lu, Yu Guo 0008, Ryan Wen Liu, Wenqi Ren
IEEE Signal Process. Lett.1
2020 DSPNet: Deep Learning-Enabled Blind Reduction of Speckle Noise
abstract
Blind reduction of speckle noise has become a longstanding unsolved problem in several imaging applications, such as medical ultrasound imaging, synthetic aperture radar (SAR) imaging, and underwater sonar imaging, etc. The unwanted noise could lead to negative effects on the reliable detection and recognition of objects of interest. From a statistical point of view, speckle noise could be assumed to be multiplicative, significantly different from the common additive Gaussian noise. The purpose of this study is to blindly reduce the speckle noise under non-ideal imaging conditions. The multiplicative relationship between latent sharp image and random noise will be first converted into an additive version through a logarithmic transformation. To promote imaging performance, we introduced the feature pyramid network (FPN) and atrous spatial pyramid pooling (ASPP), contributing to a more powerful deep blind DeSPeckling Network (named as DSPNet). In particular, DSPNet is mainly composed of two subnetworks, i.e., Log-NENet (i.e., noise estimation network in logarithmic domain) and Log-DNNet (i.e., denoising network in logarithmic domain). Log-NENet and Log-DNNet are, respectively, proposed to estimate noise level map and reduce random noise in logarithmic domain. The multi-scale mixed loss function is further proposed to improve the robust generalization of DSPN et. The proposed deep blind despeckling network is capable of reducing random noise and preserving salient image details. Both synthetic and realistic experiments have demonstrated the superior performance of our DSPNet in terms of quantitative evaluations and visual image qualities.
Yuxu Lu, Meifang Yang, Ryan Wen Liu
ICPR1