Jiangyu Wang

dblp:223/9955 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive evolutionary inverse reinforcement learning for large-scale interconnected systems
Ding Wang 0001, Jiangyu Wang, Junfei Qiao 0001
Neurocomputing3
2025 Content-Independent Avatar Ownership Detection for Preventing Sockpuppet-Enabled Violations in the Social Metaverse
Jiangyu Wang, Guohao Li 0004, Li Yang 0005, Haixin Ye
UIST1
2025 Evolution-guided Q-learning for tracking control of unknown dynamic systems
Zeqiang Yuan, Ding Wang 0001, Jiangyu Wang, Junfei Qiao 0001
Neurocomputing3
2025 Evolution-Guided Q-Learning With Dual Swarm Intelligence for Model-Free Optimal Control
abstract
In this article, a novel accelerated evolution-guided Q-learning (EGQL) algorithm is introduced to address optimal control problems for unknown nonlinear systems. A novel adaptive evolutionary algorithm is introduced here to enhance problem-solving strategies in two key aspects, which is a concept referred to dual swarm intelligence. First, the novel evolutionary algorithm is incorporated into policy improvement, replacing the gradient descent approach in traditional Q-learning. The method eliminates the need for gradient information in the approximate Q-function, while also enabling more precise policy solutions. Second, the commonly used polynomial model is replaced with a neural network, which is further optimized by the evolutionary algorithm to approximate the critic network. The accuracy of the approximate Q-function is significantly enhanced, improving the precision of value function update. Furthermore, the monotonicity and convergence properties of the algorithm are thoroughly analyzed. Finally, the effectiveness of the EGQL method is validated through two simulation experiments, demonstrating its superiority over traditional Q-learning.
Ding Wang 0001, Zeqiang Yuan, Guohan Tang, Jiangyu Wang, Junfei Qiao 0001
IEEE Trans Autom. Sci. Eng.4
2025 Relaxed Optimal Control With Self-Learning Horizon for Discrete-Time Stochastic Dynamics
abstract
The innovation of optimal learning control methods is profoundly propelled due to the improvement of the learning ability. In this article, we investigate the synthesis of initialization and acceleration for optimal learning control algorithms. This approach contrasts with traditional methods that concentrate solely on either the improvement of initialization or acceleration. Specifically, we establish a novel relaxed policy iteration (PI) algorithm with self-learning horizon for stochastic optimal control. Notably, by suitably utilizing self-learning horizon, we can directly evaluate inadmissible policies to reduce the initialization burden. Meanwhile, the inadmissible policy can be rapidly optimized with few learning iterations. Then, several critical conclusions of relaxed optimal control are established by discussing algorithm convergence and system stability. Furthermore, to provide the convincing application potentials, a class of unconventional problems is effectively solved by the relaxed PI algorithm, including the dynamics with external noises and nonzero equilibrium. Finally, we present a series of nonlinear benchmarks with practical applications to comprehensively evaluate the performance of relaxed PI. The experimental results obtained from these diverse benchmarks uniformly highlight the effectiveness of self-learning horizon mechanism.
Ding Wang 0001, Jiangyu Wang, Ao Liu 0012, Derong Liu 0001, Junfei Qiao 0001
IEEE Trans. Cybern.2
2025 Parallel Multistep Evaluation With Efficient Data Utilization for Safe Neural Critic Control and Its Application to Orbital Maneuver Systems
abstract
Data-driven methods have significantly advanced optimal learning control, but some approaches overlook systematic considerations of data utilization, including safety, efficiency, and error accumulation. To address the neglects in safe neural critic control, this article introduces a parallel multistep evaluation mechanism that combines data from the system interaction with data generated by data-driven models. Based on this evaluation mechanism, we propose a novel parallel multistep Q-learning algorithm that enhances data utilization efficiency and mitigates the error accumulation. Furthermore, we formulate a novel control barrier function (CBF) to ensure safety during learning and control processes, which is capable of dealing with asymmetric constraints and adjusting the constraint strength. In addition, the analysis reveals that multistep information introduced by data-driven models influences the learning performance of actor-critic neural networks (NNs). Finally, parallel multistep Q-learning, which makes use of data in aspects of safety, efficiency, and error bounds, is validated within an orbital maneuver system.
Jiangyu Wang, Ding Wang 0001, Derong Liu 0001, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Intelligent Critic Learning for Data-Driven Output Tracking Control With Lightweight Parallelization
abstract
This article investigates critical challenges in optimal output tracking control, including residual tracking errors, potential instability caused by discount factors, and premature convergence due to inefficient termination criteria. We tackle these issues by developing a data-driven parallelQ-learning algorithm. Specifically, a utility function directly linked to system states is proposed to avoid the instability and error amplification issues in traditional discounted approaches. In addition, the algorithm uses dual lightweight controllers that use convergence properties to enhance learning efficiency. Based on dual controllers, a novel termination criterion is introduced to prevent premature convergence during the training process. Numerical simulations demonstrate that the proposed method eliminates tracking errors, accelerates convergence compared with traditional algorithms, and ensures stable convergence across diverse system dynamics.
Jiangyu Wang, Ding Wang 0001, Junfei Qiao 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Assessing Threats: Security Boundary and Side-Channel Attack Detection in the Metaverse
Ruiyuan Yang, Guohao Li 0004, Li Yang 0005, Jiangyu Wang, Anyuan Sang
ICDF2C (2)5
2024 HEFANet: hierarchical efficient fusion and aggregation segmentation network for enhanced rgb-thermal urban scene parsing
Zhengwen Shen, Zaiyu Pan, Yuchen Weng, Yulian Li, Jiangyu Wang, Jun Wang 0071
Appl. Intell.5
2024 Novel Parallel Formulation for Iterative Reinforcement Learning Control
abstract
Parallelization is widely employed to improve the exploration ability of controllers. However, it is rare to provide a lightweight scheme for reducing homogeneous policies with theoretical guarantees. This article is concerned with a novel parallel scheme for solving optimal control problems. In brief, we design a novel global indicator that inherits the theoretical guarantees of a class of iterative reinforcement learning algorithms. By generating a tentative function, the global indicator can guide and communicate with parallel controllers to accelerate the learning process. Using two typical exploration policies, the novel parallel scheme can rapidly compress the neighborhood of the optimal cost function. Besides, two parallel algorithms based on value iteration and Q-learning are established to improve the data efficiency through different extensions. Finally, two benchmark problems are presented to demonstrate the learning effectiveness of the novel parallel scheme.
Ding Wang 0001, Jiangyu Wang, Lingzhi Hu, Liguo Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Extract Then Adjust: A Two-Stage Approach for Automatic Term Extraction
Jiangyu Wang, Chong Feng 0001
NLPCC (2)1
2023 Event-based online learning control design with eligibility trace for discrete-time unknown nonlinear systems
Ding Wang 0001, Jiangyu Wang, Lingzhi Hu
Eng. Appl. Artif. Intell.2
2023 Dichotomy value iteration with parallel learning design towards discrete-time zero-sum games
Jiangyu Wang, Ding Wang 0001, Xin Li 0055, Junfei Qiao 0001
Neural Networks1
2022 CTFusion: Convolutions Integrate with Transformers for Multi-modal Image Fusion
Zhengwen Shen, Jun Wang 0071, Zaiyu Pan, Jiangyu Wang, Yulian Li
PRCV (1)4