VLDB 2026 Research / reviewers in the wild / expert
Jingzehua Xu
dblp:364/3079
· DBLP profile ↗
23ranked-venue papers
9as first author
23since 2021 · last 2026
0009-0007-2860-5203ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Computer networks · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is FISHER All You Need in the Multi-AUV Underwater Target Tracking Task?abstractIt is significant to employ multiple autonomous underwater vehicles (AUVs) to execute the underwater target tracking task collaboratively. However, it's pretty challenging to meet various prerequisites utilizing traditional control methods. Therefore, we propose an effective two-stage learning from demonstrations training framework, FISHER, to highlight the adaptability of reinforcement learning (RL) methods in the multi-AUV underwater target tracking task, while addressing its limitations such as extensive requirements for environmental interactions and the challenges in designing reward functions. The first stage utilizes imitation learning (IL) to realize policy improvement and generate offline datasets. To be specific, we introduce multi-agent discriminator-actor-critic based on improvements of the generative adversarial IL algorithm and multi-agent IL optimization objective derived from the Nash equilibrium condition. Then in the second stage, we develop multi-agent independent generalized decision transformer, which analyzes the latent representation to match the future states of high-quality samples rather than reward function, attaining further enhanced policies capable of handling various scenarios. Besides, we propose a simulation to simulation demonstration generation procedure to facilitate the generation of expert demonstrations in underwater environments, which capitalizes on traditional control methods and can easily accomplish the domain transfer to obtain demonstrations. Extensive simulation experiments from multiple scenarios showcase that FISHER possesses strong stability, multi-task performance and capability of generalization. Guanwen Xie, Jingzehua Xu, Xiangwang Hou, Dongfang Ma, Shuai Zhang 0015, Yong Ren 0001, Dusit Niyato |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Never Too Cocky to Cooperate: An FIM and RL-Based USV-AUV Collaborative System for Underwater Tasks in Extreme Sea ConditionsabstractThis paper develops a novel Unmanned Surface Vehicle (USV)–Autonomous Underwater Vehicle (AUV) collaborative system designed to enhance underwater task performance in extreme sea conditions. The system integrates a dual strategy: (1) high-precision multi-AUV localization enabled by Fisher Information Matrix (FIM)-optimized USV path planning, and (2) a Reinforcement Learning (RL)-based cooperative planning and control framework for multi-AUV task execution. Extensive experimental evaluations in the underwater data collection task demonstrate the system's operational feasibility, with quantitative results showing significant performance improvements over baseline methods. The proposed system exhibits robust coordination capabilities between USV and AUVs while maintaining stability in extreme sea conditions. To facilitate reproducibility and community advancement, Jingzehua Xu, Guanwen Xie, Jiwei Tang, Yimian Ding, Shuai Zhang 0015, Yi Li 0032 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | AoI-MDP: An AoI Optimized Markov Decision Process Dedicated in the Underwater Task (Student Abstract)abstractOcean exploration places high demands on autonomous underwater vehicles, especially when there's observation delay. We propose age of information optimized Markov decision process (AoI-MDP) to enhance underwater tasks by modeling observation delay as signal delay and including it in the state space. AoI-MDP also introduces wait time in the action space and integrates AoI with reward functions, optimizing information freshness and decision-making using reinforcement learning. Simulations show AoI-MDP outperforms the standard MDP, demonstrating superior performance, feasibility, and generalization in underwater tasks. To accelerate relevant research, we have made the codes available as open-source at https://github.com/Xiboxtg/AoI-MDP. Yimian Ding, Jingzehua Xu, Yiyuan Yang, Guanwen Xie, Xinqi Wang, Shuai Zhang 0015 |
AAAI | 2 |
| 2025 | ERFSL: An Efficient Reward Function Searcher via Large Language Models for Custom-Environment Multi-Objective Reinforcement Learning (Student Abstract)abstractWe propose ERFSL, an efficient reward function searcher using large language models (LLMs) for custom-environment, multi-objective reinforcement learning (RL). ERFSL generates reward components based on explicit user requirements and rectifies them, and iteratively optimizes the weights of these components based on textual context. Applied to an underwater data collection RL task, ERFSL corrects reward codes with only one feedback iteration per requirement, and acquires diverse reward functions within the Pareto set. ERFSL also presents robust capability for deviated weights and small-size LLMs such as GPT-4o mini. The full-text prompts, examples of LLM-generated answers, and source code are available at https://360zmem.github.io/LLMRsearcher/ . Guanwen Xie, Jingzehua Xu, Yiyuan Yang, Yimian Ding, Shuai Zhang 0015 |
AAAI | 2 |
| 2025 | UACOF: A USV-AUV Collaboration Framework for Underwater Tasks Under Extreme Sea Conditions (Student Abstract)abstractOcean exploration requires effective collaboration between the unmanned surface vehicle (USV) and autonomous underwater vehicles (AUVs). We propose UACOF, a USV-AUV collaboration framework that enhances multi-AUV performance under extreme sea conditions. The framework includes high-precision multi-AUV location via USV path planning with Fisher information matrix optimization and reinforcement learning training for cooperative tasks. Experimental results show UACOF's superior feasibility, performance, coordination and robustness in extreme conditions. Jingzehua Xu, Guanwen Xie, Yimian Ding, Yongming Zeng, Shuai Zhang 0015 |
AAAI | 1 |
| 2025 | LHQ-SVC: Lightweight and High Quality Singing Voice Conversion ModelingabstractSinging Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer’s voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional SVC methods have limitations in terms of audio quality, data requirements, and computational complexity. In this paper, we propose LHQ-SVC, a lightweight, CPU-compatible model based on the SVC framework and diffusion model, designed to reduce model size and computational demand without sacrificing performance. We incorporate features to improve inference quality, and optimize for CPU execution by using performance tuning tools and parallel computing frameworks. Our experiments demonstrate that LHQ-SVC maintains competitive performance, with significant improvements in processing speed and efficiency across different devices. The results suggest that LHQ-SVC can meet real-time performance requirements, even in resource-constrained environments. Muyang Ye, Anran Zhu, Jingzehua Xu, Shuai Zhang 0015, Weijie Niu |
ICASSP | 6 |
| 2025 | USV-AUV Collaboration Framework for Underwater Tasks under Extreme Sea ConditionsabstractAutonomous underwater vehicles (AUVs) are valuable for ocean exploration due to their flexibility and ability to carry communication and detection units. Nevertheless, AUVs alone often face challenges in harsh and extreme sea conditions. This study introduces a unmanned surface vehicle (USV)–AUV collaboration framework, which includes high-precision multi-AUV positioning using USV path planning via Fisher information matrix optimization and reinforcement learning for multi-AUV cooperative tasks. Applied to a multi-AUV underwater data collection task scenario, extensive simulations validate the framework’s feasibility and superior performance, highlighting exceptional coordination and robustness under extreme sea conditions. To accelerate relevant research in this field, we have made simulation code (demo version) available as open-source1. Jingzehua Xu, Guanwen Xie, Xinqi Wang, Yimian Ding, Shuai Zhang 0015 |
ICASSP | 1 |
| 2025 | Make Your AUV Adaptive: An Environment-Aware Reinforcement Learning Framework For Underwater TasksabstractThis study presents a novel environment-aware reinforcement learning (RL) framework designed to augment the operational capabilities of autonomous underwater vehicles (AUVs) in underwater environments. Departing from traditional RL architectures, the proposed framework integrates an environment-aware network module that dynamically captures flow field data, effectively embedding this critical environmental information into the state space. This integration facilitates real-time environmental adaptation, significantly enhancing the AUV’s situational awareness and decision-making capabilities. Furthermore, the framework incorporates AUV structure characteristics into the optimization process, employing a large language model (LLM)-based iterative refinement mechanism that leverages both environmental conditions and training outcomes to optimize task performance. Comprehensive experimental evaluations demonstrate the framework’s superior performance, robustness and adaptability. Yimian Ding, Jingzehua Xu, Guanwen Xie, Shuai Zhang 0015, Yi Li 0032 |
IROS | 2 |
| 2025 | Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea ConditionsabstractThe adaptivity and maneuvering capabilities of Autonomous Underwater Vehicles (AUVs) have drawn significant attention in oceanic research, due to the unpredictable disturbances and strong coupling among the AUV’s degrees of freedom. In this paper, we developed large language model (LLM)-enhanced reinforcement learning (RL)-based adaptive S-surface controller for AUVs. Specifically, LLMs are introduced for the joint optimization of controller parameters and reward functions in RL training. Using multi-modal and structured explicit task feedback, LLMs enable joint adjustments, balance multiple objectives, and enhance task-oriented performance and adaptability. In the proposed controller, the RL policy focuses on upper-level tasks, outputting task-oriented high-level commands that the S-surface controller then converts into control signals, ensuring cancellation of nonlinear effects and unpredictable external disturbances in extreme sea conditions. Under extreme sea conditions involving complex terrain, waves, and currents, the proposed controller demonstrates superior performance and adaptability in high-level tasks such as underwater target tracking and data collection, outperforming traditional PID and SMC controllers.3 Guanwen Xie, Jingzehua Xu, Yimian Ding, Shuai Zhang 0015, Yi Li 0032 |
IROS | 2 |
| 2025 | Multi-Modal Gradual Domain Osmosis: Stepwise Dynamic Learning with Batch Matching for Gradual Domain AdaptationabstractIn this paper, we propose a new method called Multi-Modal Gradual Domain Osmosis, which aims to solve the problem of smooth knowledge migration from the source domain to the target domain in Gradual Domain Adaptation (GDA). Traditional Gradual Domain Adaptation methods mitigate domain bias by introducing intermediate domains and self-training strategies but often face the challenges of inefficient knowledge migration or missing data in intermediate domains. In this paper, we design an optimization framework based on the hyperparameter łambda by dynamically balancing the loss weights of the source and target domains, which enables the model to progressively adjust the strength of knowledge migration (łambda incrementing from 0 to 1) during the training process, thus achieving cross-domain generalization more efficiently. Specifically, the method incorporates self-training to generate pseudo-labels and iteratively updates the model by minimizing a weighted loss function to ensure stability and robustness during progressive adaptation in the intermediate domain. The experimental part validates the effectiveness of the method on rotated MNIST, color-shifted MNIST, portrait dataset, and forest cover type dataset, and the results show that it outperforms existing baseline methods. The paper further analyses the impact of the dynamic tuning strategy of the hyperparameter łambda on the performance through ablation experiments, confirming the advantages of progressive domain penetration in mitigating domain bias and enhancing the model generalization capability. The study provides theoretical support and a practical framework for asymptotic domain adaptation and expands its application potential in dynamic environments. Jingzehua Xu, Jinzhu Wei, Shuai Zhang 0015 |
ACM Multimedia | 3 |
| 2025 | Self-Training with Dynamic Weighting for Robust Gradual Domain AdaptationabstractIn this paper, we propose a new method called \textit{Self-Training with Dynamic Weighting} (STDW), which aims to enhance robustness in Gradual Domain Adaptation (GDA) by addressing the challenge of smooth knowledge migration from the source to the target domain. Traditional GDA methods mitigate domain shift through intermediate domains and self-training but often suffer from inefficient knowledge migration or incomplete intermediate data. Our approach introduces a dynamic weighting mechanism that adaptively balances the loss contributions of the source and target domains during training. Specifically, we design an optimization framework governed by a time-varying hyperparameter $\varrho$ (progressing from 0 to 1), which controls the strength of domain-specific learning and ensures stable adaptation. The method leverages self-training to generate pseudo-labels and optimizes a weighted objective function for iterative model updates, maintaining robustness across intermediate domains. Experiments on rotated MNIST, color-shifted MNIST, portrait datasets, and the Cover Type dataset demonstrate that STDW outperforms existing baselines. Ablation studies further validate the critical role of $\varrho$'s dynamic scheduling in achieving progressive adaptation, confirming its effectiveness in reducing domain bias and improving generalization. This work provides both theoretical insights and a practical framework for robust gradual domain adaptation, with potential applications in dynamic real-world scenarios. Yushe Cao, Jinzhu Wei, Jingzehua Xu, Shuai Zhang 0015 |
NeurIPS | 5 |
| 2025 | Multi-Objective-Optimization Multi-Auv Assisted Data Collection Framework for Iout Based on Offline Reinforcement LearningabstractThe Internet of Underwater Things (IoUT) offers significant potential for ocean exploration but encounters challenges due to dynamic underwater environments and severe signal attenuation. Current methods relying on Autonomous Underwater Vehicles (AUVs) based on online reinforcement learning (RL) lead to high computational costs and low data utilization. To address these issues and the constraints of turbulent ocean environments, we propose a multi-AUV assisted data collection framework for IoUT based on multi-agent offline RL. This framework maximizes data rate and the value of information (VoI), minimizes energy consumption, and ensures collision avoidance by utilizing environmental and equipment status data. We introduce a semi-communication decentralized training with decentralized execution (SC-DTDE) paradigm and a multi-agent independent conservative Q -learning algorithm (MAICQL) to effectively tackle the problem. Extensive simulations demonstrate the high applicability, robustness, and data collection efficiency of the proposed framework. Yimian Ding, Xinqi Wang, Jingzehua Xu, Guanwen Xie |
WCNC | 3 |
| 2025 | EFILN: The Electric Field Inversion-Localization Network for High-Precision Underwater PositioningabstractAccurate underwater target localization is essential for underwater exploration. To improve accuracy and efficiency in complex underwater environments, we propose the Electric Field Inversion-Localization Network (EFILN), a deep feedforward neural network that reconstructs position coordinates from underwater electric field signals. By assessing whether the neural network's input-output values satisfy the Coulomb law, the error between the network's inversion solution and the equation's exact solution can be determined. The Adam optimizer was employed first, followed by the L-BFGS optimizer, to progressively improve the output precision of EFILN. A series of noise experiments demonstrated the robustness and practical utility of the proposed method, while small sample data experiments validated its strong small-sample learning (SSL) capabilities. To accelerate relevant research, we have made the codes available as open-source11Codes are available at https://github.com/Xiboxtg/EFILN. Yimian Ding, Jingzehua Xu, Guanwen Xie |
WCNC | 2 |
| 2025 | UPEGSim: An RL-Enabled Simulator for Unmanned Underwater Vehicles Dedicated in the Underwater Pursuit-Evasion GameabstractUnmanned underwater vehicles (UUVs) have been widely used in various ocean applications, such as underwater exploration and data collection. And the underwater pursuit-evasion game (UPEG) is the key to efficient implementation of other tasks, holding significant research value. However, testing the UPEG task in real ocean environment is both costly and risky, and currently, UUV control algorithms that rely on specific environmental models struggle to complete the complicated UPEG task. To address above challenge, we propose UPEGSim, an UUV simulator specifically designed for the UPEG task. Built through Gazebo and robot operating system, UPEGSim provides a reinforcement learning (RL) environment to train UUVs for improving the intelligent performance in the UPEG task. Furthermore, we propose an efficient UPEG training framework (ETFDU), which includes multiagent decentralized training and execution techniques, scene transfer training methods, and offline RL techniques based on decision transformer, to facilitate efficient UUV training. Through training on the UPEG task in UPEGSim, we validate the effectiveness and feasibility of the proposed UPEGSim simulator and the ETFDU training framework. Jingzehua Xu, Guanwen Xie, Xiangwang Hou, Shuai Zhang 0015, Yong Ren 0001, Dusit Niyato |
IEEE Internet Things J. | 1 |
| 2024 | One Process Spatiotemporal Learning of Transformers via Vcls Token for Multivariate Time Series Forecasting
Tao Cai 0003, Haixiang Wu, DeJiao Niu, Xuewen Xia, Jingzehua Xu |
ICANN (6) | 6 |
| 2024 | Multimodal Monocular Dense Depth Estimation with Event-Frame Fusion Using Transformer
Baihui Xiao, Jingzehua Xu, Tianyu Xing, Jingjing Wang 0001, Yong Ren 0001 |
ICANN (2) | 2 |
| 2024 | Robust Navigation for Unmanned Surface Vehicle Utilizing Improved Distributional Soft Actor-Critic
Jingzehua Xu, Ziqi Jia, Tianyu Xing, Jingjing Wang 0001, Yong Ren 0001 |
ICANN (4) | 1 |
| 2024 | AUV Efficient Navigation Relying on Adaptive Proximal Policy Optimization
Jingzehua Xu, Yongming Zeng, Xuanchen Li, Lingru Meng, Haocai Huang, Jingjing Wang 0001, Yong Ren 0001 |
ICONIP (11) | 1 |
| 2024 | Multi-AUV Assisted Seamless Underwater Target Tracking Relying on Deep Learning and Reinforcement LearningabstractSince seamless tracking of the underwater target is crucial for various underwater applications, we propose a fusion algorithm combining deep learning and reinforcement learning for multi-autonomous underwater vehicles (AUVs) to seamlessly track the underwater target. The framework of our proposed fusion algorithm consists of two stages. In the first stage, we propose an underwater target localization method based on convolutional neural network (CNN) that relies on shaft-rate electric fields, in which the data collected by underwater sensors is utilized to train CNN to achieve accurate target localization. In the second stage, we innovatively propose a multi-agent soft actor-critic (MASAC) reinforcement learning algorithm based on centralized training with decentralized execution, in which appropriate reward functions are designed to encourage multiple AUVs to cooperate in seamlessly tracking the target in unknown environments while avoiding obstacles. Simulation results show that the proposed fusion algorithm has excellent performance, while the real-time target localization accuracy is 97.8%, and AUVs can carry out seamlessly cooperative tracking of the target in unknown environment. Jingzehua Xu, Yimian Ding, Guanwen Xie, Ziyuan Wang 0002, Yongming Zeng |
IJCNN | 1 |
| 2024 | UUVSim: Intelligent Modular Simulation Platform for Unmanned Underwater Vehicle LearningabstractUnmanned underwater vehicles (UUVs) face challenges such as high hardware costs, security concerns, a lack of training data in the actual development and debugging. Creating a simulation platform for simulation verification, training, and learning presents a potential solution to address these challenges. However, this area has seen limited prior work, and existing underwater platforms lack accuracy, user-friendliness, and intelligence. Therefore, this paper introduces an intelligent simulation platform “UUVSim” based on the robot operating system and Gazebo. UUVSim modular integrates basic modules such as high-precision simulation scenarios, dynamic models, sensors and controllers, while reserving programming interfaces. In addition, UUVSim provides reinforcement learning environment for UUV intelligent learning, supplemented with scenario transfer training, multi-agent reinforcement learning, offline reinforcement learning techniques to realize efficiently training for complex tasks, multi-robot coordination, and simulation to reality (sim2real) deployment. Further, we validate these technologies through underwater target tracking benchmarks and sim2real experiments, demonstrating the platform’s practicality. Jingzehua Xu, Jun Du 0001, Weishi Mi, Ziyuan Wang 0002, Zonglin Li 0007, Yong Ren 0001 |
IJCNN | 2 |
| 2024 | Vol and Energy-Aware AUV-Assisted Data Collection for Internet of Underwater ThingsabstractIn this study, an autonomous underwater vehicle (AUV) is considered to collect data in Internet of Underwater Things (loUT) networks. The AUV is tasked with timely visits to sensor nodes (SNs) to collect data using a navigation-hover-communication protocol. Considering the AUV's limited energy, the dynamic data upload demand of SNs and the diminishing Value of Information (Vol) during data transmission, efficient path planning for AUV is required. Therefore, we formulate a multi-objective optimization problem and employ the deep deterministic policy gradient (DDPG) algorithm to address it. Our objectives encompass the maximization of the sum data rate, the maximization of the sum Vol, and the minimization of the AUV's energy consumption within specified task time constraints. To mitigate the challenges posed by sparse rewards, we enhance the DDPG algorithm with hindsight experience replay (HER). The simulation results show that the data collection policy trained by our proposed algorithm can converge quickly and has excellent generalization. When the communication range changes, it can still effectively reduce the AUV's energy consumption while ensuring the quality and timeliness of the data collection task. Jingzehua Xu, Ziyuan Wang 0002, Jingjing Wang 0001, Yong Rent |
WCNC | 1 |
| 2024 | Multi-AUV Pursuit-Evasion Game in the Internet of Underwater Things: An Efficient Training Framework via Offline Reinforcement LearningabstractIn this article, we investigate the pursuit-evasion game of multiple autonomous underwater vehicles (AUVs) in a complex ocean environment. The pursuer AUVs need to optimize their trajectories to avoid obstacles and dangerous vortex regions in the environment in order to pursue the escaper AUV. Both the pursuer and escaper can sense each other with limited detection capabilities for further pursuit or escape. As the underwater pursuit-evasion (UPE) game is a high-dimensional NP-hard problem, we innovatively transform it into a finite-horizon Markov game process and propose a decentralized training and decentralized execution efficient training framework based on the offline reinforcement learning. During the training process, we propose multiagent independent soft actor–critic to facilitate policy improvement and generate the offline data set, and propose multiagent independent decision transformer for model training in the UPE game. Extensive simulations demonstrate the scalability and generalization ability of our proposed training framework, which can achieve excellent performance in the UPE games under different conditions and environments with only a few AUVs participating in policy improvement to generate the high-quality offline data set. Jingzehua Xu, Jingjing Wang 0001, Zhu Han 0001, Yong Ren 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Environment- and Energy-Aware AUV-Assisted Data Collection for the Internet of Underwater ThingsabstractConsidering the wide-area distribution and limited transmission power of sensing devices in the Internet of Underwater Things (IoUT), employing autonomous underwater vehicles (AUVs) to collect data is considered a promising solution. While most existing AUV-assisted data collection schemes primarily focus on enhancing data collection throughput and identifying the shortest path, they often overlook the influence of the underwater environment on AUV and the timeliness of data collection. In this article, we design a multi-AUV-assisted data collection system, in which AUVs select their own target devices to collect data according to the data upload urgencies of IoUT devices. Considering the disturbance of turbulent ocean environment and the limited energy of AUV, we propose an environment- and energy-aware AUV-assisted data collection scheme. This scheme aims to conduct path planning for multiple AUVs based on perceived environmental information, including turbulent fields and device statuses. The primary goals are to maximize the sum data collection rate and total data throughput, minimize AUV energy consumption, reduce the average data overflow times. To solve this high-dimensional NP-hard problem, we first model the problem as a Markov decision process, and propose a multiagent independent soft actor–critic to solve it. Extensive simulations validate the effectiveness and adaptability of our approach. Jingzehua Xu, Guanwen Xie, Jingjing Wang 0001, Zhu Han 0001, Yong Ren 0001 |
IEEE Internet Things J. | 2 |