Jun Yang 0028

dblp:181/2799-28 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
18since 2021 · last 2025
0000-0002-9386-5825ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 Episodic Novelty Through Temporal Distance
abstract
Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approaches, which are ineffective in large state spaces, or on similarity-based methods that lack appropriate metrics for state comparison. To address these shortcomings, we propose Episodic Novelty Through Temporal Distance (ETD), a novel approach that introduces temporal distance as a robust metric for state similarity and intrinsic reward computation. By employing contrastive learning, ETD accurately estimates temporal distances and derives intrinsic rewards based on the novelty of states within the current episode. Extensive experiments on various benchmark tasks demonstrate that ETD significantly outperforms state-of-the-art methods, highlighting its effectiveness in enhancing exploration in sparse reward CMDPs.
Yuhua Jiang, Qihan Liu, Yiqin Yang, Xiaoteng Ma, Dianyu Zhong, Hao Hu 0006, Jun Yang 0028, Bin Liang 0001, Bo Xu 0002, Chongjie Zhang, Qianchuan Zhao
ICLR7
2025 DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
abstract
We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm. DSAC models the randomness in both action and rewards, surpassing baseline performances on various continuous control tasks. Unlike standard approaches that solely maximize expected rewards, we propose a unified framework for risk-sensitive learning, one that optimizes the risk-related objective while balancing entropy to encourage exploration. Extensive experiments demonstrate DSAC’s effectiveness in enhancing agent performances for both risk-neutral and risk-sensitive control tasks.
Xiaoteng Ma, Junyao Chen, Jun Yang 0028, Qianchuan Zhao, Zhengyuan Zhou
J. Artif. Intell. Res.4
2025 Celebrating Diversity With Subtask Specialization in Shared Multiagent Reinforcement Learning
abstract
Subtask decomposition offers a promising approach for achieving and comprehending complex cooperative behaviors in multiagent systems. Nonetheless, existing methods often depend on intricate high-level strategies, which can hinder interpretability and learning efficiency. To tackle these challenges, we propose a novel approach that specializes subtasks for subgroups by employing diverse observation representation encoders within information bottlenecks. Moreover, to enhance the efficiency of subtask specialization while promoting sophisticated cooperation, we introduce diversity in both optimization and neural network architectures. These advancements enable our method to achieve state-of-the-art performance and offer interpretable subtask factorization across various scenarios in Google Research Football (GRF).
Chenghao Li 0002, Tonghan Wang 0001, Chengjie Wu, Qianchuan Zhao, Jun Yang 0028, Chongjie Zhang
IEEE Trans. Neural Networks Learn. Syst.5
2025 CVaR-Constrained Policy Optimization for Safe Reinforcement Learning
abstract
Current constrained reinforcement learning (RL) methods guarantee constraint satisfaction only in expectation, which is inadequate for safety-critical decision problems. Since a constraint satisfied in expectation remains a high probability of exceeding the cost threshold, solving constrained RL problems with high probabilities of satisfaction is critical for RL safety. In this work, we consider the safety criterion as a constraint on the conditional value-at-risk (CVaR) of cumulative costs, and propose the CVaR-constrained policy optimization algorithm (CVaR-CPO) to maximize the expected return while ensuring agents pay attention to the upper tail of constraint costs. According to the bound on the CVaR-related performance between two policies, we first reformulate the CVaR-constrained problem in augmented state space using the state extension procedure and the trust-region method. CVaR-CPO then derives the optimal update policy by applying the Lagrangian method to the constrained optimization problem. In addition, CVaR-CPO utilizes the distribution of constraint costs to provide an efficient quantile-based estimation of the CVaR-related value function. We conduct experiments on constrained control tasks to show that the proposed method can produce behaviors that satisfy safety constraints, and achieve comparable performance to most safe RL (SRL) methods.
Shu Leng, Xiaoteng Ma, Qihan Liu, Xueqian Wang 0001, Bin Liang 0001, Yu Liu 0036, Jun Yang 0028
IEEE Trans. Neural Networks Learn. Syst.8
2024 Learning Diverse Risk Preferences in Population-Based Self-Play
abstract
Among the remarkable successes of Reinforcement Learning (RL), self-play algorithms have played a crucial role in solving competitive games. However, current self-play RL methods commonly optimize the agent to maximize the expected win-rates against its current or historical copies, resulting in a limited strategy style and a tendency to get stuck in local optima. To address this limitation, it is important to improve the diversity of policies, allowing the agent to break stalemates and enhance its robustness when facing with different opponents. In this paper, we present a novel perspective to promote diversity by considering that agents could have diverse risk preferences in the face of uncertainty. To achieve this, we introduce a novel reinforcement learning algorithm called Risk-sensitive Proximal Policy Optimization (RPPO), which smoothly interpolates between worst-case and best-case policy learning, enabling policy learning with desired risk preferences. Furthermore, by seamlessly integrating RPPO with population-based self-play, agents in the population optimize dynamic risk-sensitive objectives using experiences gained from playing against diverse opponents. Our empirical results demonstrate that our method achieves comparable or superior performance in competitive games and, importantly, leads to the emergence of diverse behavioral modes. Code is available at https://github.com/Jackory/RPBT.
Yuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li 0002, Yiqin Yang, Jun Yang 0028, Bin Liang 0001, Qianchuan Zhao
AAAI6
2024 Efficient Multi-agent Reinforcement Learning by Planning
abstract
Multi-agent reinforcement learning (MARL) algorithms have accomplished remarkable breakthroughs in solving large-scale decision-making tasks. Nonetheless, most existing MARL algorithms are model-free, limiting sample efficiency and hindering their applicability in more challenging scenarios. In contrast, model-based reinforcement learning (MBRL), particularly algorithms integrating planning, such as MuZero, has demonstrated superhuman performance with limited data in many tasks. Hence, we aim to boost the sample efficiency of MARL by adopting model-based approaches. However, incorporating planning and search methods into multi-agent systems poses significant challenges. The expansive action space of multi-agent systems often necessitates leveraging the nearly-independent property of agents to accelerate learning. To tackle this issue, we propose the MAZero algorithm, which combines a centralized model with Monte Carlo Tree Search (MCTS) for policy search. We design an ingenious network structure to facilitate distributed execution and parameter sharing. To enhance search efficiency in deterministic environments with sizable action spaces, we introduce two novel techniques: Optimistic Search Lambda (OS($\lambda$)) and Advantage-Weighted Policy Optimization (AWPO). Extensive experiments on the SMAC benchmark demonstrate that MAZero outperforms model-free approaches in terms of sample efficiency and provides comparable or better performance than existing model-based methods in terms of both sample and computational efficiency.
Qihan Liu, Jianing Ye, Xiaoteng Ma, Jun Yang 0028, Bin Liang 0001, Chongjie Zhang
ICLR4
2024 Single-Trajectory Distributionally Robust Reinforcement Learning
abstract
To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As a price for robustness gain, DRRL involves optimizing over a set of distributions, which is inherently more challenging than optimizing over a fixed distribution in the non-robust case. Existing DRRL algorithms are either model-based or fail to learn from a single sample trajectory. In this paper, we design a first fully model-free DRRL algorithm, called distributionally robust Q-learning with single trajectory (DRQ). We delicately design a multi-timescale framework to fully utilize each incrementally arriving sample and directly learn the optimal distributionally robust policy without modeling the environment, thus the algorithm can be trained along a single trajectory in a model-free fashion. Despite the algorithm’s complexity, we provide asymptotic convergence guarantees by generalizing classical stochastic approximation tools.Comprehensive experimental results demonstrate the superior robustness and sample complexity of our proposed algorithm, compared to non-robust methods and other robust RL algorithms.
Xiaoteng Ma, Jose H. Blanchet, Jun Yang 0028, Jiheng Zhang, Zhengyuan Zhou
ICML4
2024 More Like Real World Game Challenge for Partially Observable Multi-agent Cooperation
Xueou Feng, Shengqi Shen, Qiyue Yin, Jun Yang 0028
PRCV (4)5
2024 An Asymmetric Game Theoretic Learning Model
Qiyue Yin, Tongtong Yu, Xueou Feng, Jun Yang 0028, Kaiqi Huang
PRCV (3)4
2023 Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery
abstract
Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurring and temporally extended structures in the logged data yields better learning. However, these methods suffer greatly when the primitives have limited representation ability to recover the original policy space, especially in offline settings. In this paper, we give a quantitative characterization of the performance of offline hierarchical learning and highlight the importance of learning lossless primitives. To this end, we propose to use a flow-based structure as the representation for low-level policies. This allows us to represent the behaviors in the dataset faithfully while keeping the expression ability to recover the whole policy space. We show that such lossless primitives can drastically improve the performance of hierarchical policies. The experimental results and extensive ablation studies on the standard D4RL benchmark show that our method has a good representation ability for policies and achieves superior performance in most tasks.
Yiqin Yang, Hao Hu 0006, Siyuan Li 0003, Jun Yang 0028, Qianchuan Zhao, Chongjie Zhang
AAAI5
2023 Uncertainty-Driven Trajectory Truncation for Data Augmentation in Offline Reinforcement Learning
abstract
Equipped with the trained environmental dynamics, model-based offline reinforcement learning (RL) algorithms can often successfully learn good policies from fixed-sized datasets, even some datasets with poor quality. Unfortunately, however, it can not be guaranteed that the generated samples from the trained dynamics model are reliable (e.g., some synthetic samples may lie outside of the support region of the static dataset). To address this issue, we propose Trajectory Truncation with Uncertainty (TATU), which adaptively truncates the synthetic trajectory if the accumulated uncertainty along the trajectory is too large. We theoretically show the performance bound of TATU to justify its benefits. To empirically show the advantages of TATU, we first combine it with two classical model-based offline RL algorithms, MOPO and COMBO. Furthermore, we integrate TATU with several off-the-shelf model-free offline RL algorithms, e.g., BCQ. Experimental results on the D4RL benchmark show that TATU significantly improves their performance, often by a large margin. Code is available here.
Jiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Jun Yang 0028, Le Wan, Xiu Li 0001
ECAI5
2023 Conservative Offline Policy Adaptation in Multi-Agent Games
abstract
Prior research on policy adaptation in multi-agent games has often relied on online interaction with the target agent in training, which can be expensive and impractical in real-world scenarios. Inspired by recent progress in offline reinforcement learn- ing, this paper studies offline policy adaptation, which aims to utilize the target agent’s behavior data to exploit its weakness or enable effective cooperation. We investigate its distinct challenges of distributional shift and risk-free deviation, and propose a novel learning objective, conservative offline adaptation, that optimizes the worst-case performance against any dataset consistent proxy models. We pro- pose an efficient algorithm called Constrained Self-Play (CSP) that incorporates dataset information into regularized policy learning. We prove that CSP learns a near-optimal risk-free offline adaptation policy upon convergence. Empirical results demonstrate that CSP outperforms non-conservative baselines in various environments, including Maze, predator-prey, MuJoCo, and Google Football.
Chengjie Wu, Pingzhong Tang, Jun Yang 0028, Yujing Hu, Tangjie Lv, Changjie Fan, Chongjie Zhang
NeurIPS3
2022 Offline Reinforcement Learning with Value-based Episodic Memory
Xiaoteng Ma, Yiqin Yang, Hao Hu 0006, Jun Yang 0028, Chongjie Zhang, Qianchuan Zhao, Bin Liang 0001, Qihan Liu
ICLR4
2022 Safe Opponent-Exploitation Subgame Refinement
abstract
In zero-sum games, an NE strategy tends to be overly conservative confronted with opponents of limited rationality, because it does not actively exploit their weaknesses. From another perspective, best responding to an estimated opponent model is vulnerable to estimation errors and lacks safety guarantees. Inspired by the recent success of real-time search algorithms in developing superhuman AI, we investigate the dilemma of safety and opponent exploitation and present a novel real-time search framework, called Safe Exploitation Search (SES), which continuously interpolates between the two extremes of online strategy refinement. We provide SES with a theoretically upper-bounded exploitability and a lower-bounded evaluation performance. Additionally, SES enables computationally efficient online adaptation to a possibly updating opponent model, while previous safe exploitation methods have to recompute for the whole game. Empirical results show that SES significantly outperforms NE baselines and previous algorithms while keeping exploitability low at the same time.
Chengjie Wu, Qihan Liu, Yansen Jing, Jun Yang 0028, Pingzhong Tang, Chongjie Zhang
NeurIPS5
2022 Cooperative planning of multi-agent systems based on task-oriented knowledge fusion with graph neural networks
abstract
Cooperative planning is one of the critical problems in the field of multi-agent system gaming. This work focuses on cooperative planning when each agent has only a local observation range and local communication. We propose a novel cooperative planning architecture that combines a graph neural network with a task-oriented knowledge fusion sampling method. Two main contributions of this paper are based on the comparisons with previous work: (1) we realize feasible and dynamic adjacent information fusion using GraphSAGE (i.e., Graph SAmple and aggreGatE), which is the first time this method has been used to deal with the cooperative planning problem, and (2) a task-oriented sampling method is proposed to aggregate the available knowledge from a particular orientation, to obtain an effective and stable training process in our model. Experimental results demonstrate the good performance of our proposed method.
Hanqi Dai, Weining Lu, Jun Yang 0028, Deshan Meng, Yanze Liu, Bin Liang 0001
Frontiers Inf. Technol. Electron. Eng.4
2021 Average-Reward Reinforcement Learning with Trust Region Methods
abstract
Most of reinforcement learning algorithms optimize the discounted criterion which is beneficial to accelerate the convergence and reduce the variance of estimates. Although the discounted criterion is appropriate for certain tasks such as financial related problems, many engineering problems treat future rewards equally and prefer a long-run average criterion. In this paper, we study the reinforcement learning problem with the long-run average criterion. Firstly, we develop a unified trust region theory with discounted and average criteria. With the average criterion, a novel performance bound within the trust region is derived with the Perturbation Analysis (PA) theory. Secondly, we propose a practical algorithm named Average Policy Optimization (APO), which improves the value estimation with a novel technique named Average Value Constraint. To the best of our knowledge, our work is the first one to study the trust region approach with the average criterion and it complements the framework of reinforcement learning beyond the discounted criterion. Finally, experiments are conducted in the continuous control environment MuJoCo. In most tasks, APO performs better than the discounted PPO, which demonstrates the effectiveness of our approach.
Xiaoteng Ma, Xiaohang Tang, Jun Yang 0028, Qianchuan Zhao
IJCAI4
2021 Celebrating Diversity in Shared Multi-Agent Reinforcement Learning
abstract
Recently, deep multi-agent reinforcement learning (MARL) has shown the promise to solve complex cooperative tasks. Its success is partly because of parameter sharing among agents. However, such sharing may lead agents to behave similarly and limit their coordination capacity. In this paper, we aim to introduce diversity in both optimization and representation of shared multi-agent reinforcement learning. Specifically, we propose an information-theoretical regularization to maximize the mutual information between agents' identities and their trajectories, encouraging extensive exploration and diverse individualized behaviors. In representation, we incorporate agent-specific modules in the shared neural network architecture, which are regularized by L1-norm to promote learning sharing among agents while keeping necessary diversity. Empirical results show that our method achieves state-of-the-art performance on Google Research Football and super hard StarCraft II micromanagement tasks.
Chenghao Li 0002, Tonghan Wang 0001, Chengjie Wu, Qianchuan Zhao, Jun Yang 0028, Chongjie Zhang
NeurIPS5
2021 Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning
abstract
Learning from datasets without interaction with environments (Offline Learning) is an essential step to apply Reinforcement Learning (RL) algorithms in real-world scenarios.However, compared with the single-agent counterpart, offline multi-agent RL introduces more agents with the larger state and action space, which is more challenging but attracts little attention. We demonstrate current offline RL algorithms are ineffective in multi-agent systems due to the accumulated extrapolation error. In this paper, we propose a novel offline RL algorithm, named Implicit Constraint Q-learning (ICQ), which effectively alleviates the extrapolation error by only trusting the state-action pairs given in the dataset for value estimation. Moreover, we extend ICQ to multi-agent tasks by decomposing the joint-policy under the implicit constraint. Experimental results demonstrate that the extrapolation error is successfully controlled within a reasonable range and insensitive to the number of agents. We further show that ICQ achieves the state-of-the-art performance in the challenging multi-agent offline tasks (StarCraft II). Our code is public online at https://github.com/YiqinYang/ICQ.
Yiqin Yang, Xiaoteng Ma, Chenghao Li 0002, Zewu Zheng, Gao Huang 0001, Jun Yang 0028, Qianchuan Zhao
NeurIPS7
2020 Conservatism Comparison of State Estimation Error and Residual in Multiple Actuator Faults Detection
abstract
This paper focuses on analyzing and comparing the performance of two robust fault detection (FD) criteria for discrete-time linear parameter varying (LPV) systems with bounded uncertainties, namely the state estimation error-based criterion and the classical residual-based criterion. First, a new FD criterion for the detection of multiple multiplicative actuator faults is proposed by testing consistency between the state estimation errors and the healthy state estimation error sets on-line. Then, a guaranteed FD condition is established based on set-separation of healthy and faulty invariant sets of state estimation error. Moreover, the generalized minimum detectable fault (MDF) for multiple actuator faults is defined and computed in order to characterize the performance of the two FD criteria. Finally, a proof is provided to compare the conservatism of the FD criterion using state estimation errors with the classical one based on residuals. At the end of this paper, a numerical example is used to illustrate the effectiveness of the obtained results.
Bo Min, Junbo Tan, Xueqian Wang 0001, Jun Yang 0028, Bin Liang 0001
SMC4
2019 Modeling and Control of Free-Floating Space Manipulator Using the T-S Fuzzy Descriptor System Approach
abstract
In this paper, a Takagi-Sugeno (T-S) fuzzy descriptor approach for control of a two-link free-floating space manipulator (FFSM) is proposed. The T-S fuzzy descriptor model of the FFSM is first derived from its nonlinear dynamic model, which makes more sense in reality since it avoids the use of joint acceleration measurement and the inversion of inertia matrix. And some nonlinear terms are considered as uncertainties to balance the complexity and accuracy of the model. Then a robust controller based on the Lyapunov stability theory is designed and reformulated as a linear matrix inequality (LMI) optimization problem which can be efficiently solved with the solver SeduMi. Finally, simulation results are carried out with the SimMechanics to demonstrate the effectiveness of the proposed approach.
Jiabao He 0001, Feng Xu 0006, Xueqian Wang 0001, Jun Yang 0028, Bin Liang 0001
SMC4
2008 Redundant design of A CAN BUS Testing and Communication System for space robot arm
abstract
This paper analyzes and designs a Testing and Communication System (TCS), which is used for the simulation of controlling and communicating for the space robot arm. Arming at enhancing the reliability of the operation of the space robot arm, a new Hot Double redundant strategy is adopted during the designing of TCS. The new Hot Double Redundant way is utilized to back up CAN BUS, which is the medium of data transmission and reception. The new network structure based on CAN BUS is validated on a nine nodes experiment system. The simulation results demonstrate the effectiveness of the redundancy strategy, and also the reliability and fault-tolerant ability of the TCS is testified by experiments.
Jun Yang 0028, Tao Zhang 0006, Jingyan Song, Hanxu Sun, Guozhen Shi
ICARCV1