EDBT 2026 Demo / reviewers in the wild / expert
Shangdong Yang
dblp:178/8734
· DBLP profile ↗
30ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAPG2: Multiagent Policy Gradient via Potential Game for Multirobot Task Allocation ProblemsabstractEfficient task allocation among multiple UAVs and autonomous robots is critical in modern IoT scenarios. This is typically modeled as a multi-robot task allocation (MRTA) problem, known to be an NP-hard combinatorial optimization problem. Neural sequential modeling combined with reinforcement learning (RL) optimization has emerged as a promising paradigm for solving this problem, owing to its high efficiency during inference. However, most existing methods assume that each robot is capable of performing only a single type of task. The development of sensing technologies has significantly enhanced the functional diversity of robots, thereby challenging the effectiveness and scalability of traditional methods. This paper considers a variant of the MRTA problem, where each robot is capable of handling multiple tasks, and tasks vary in both their types and required resources. To this end, we present a novel game-theoretic multi-agent RL algorithm called multi-agent policy gradient via potential game (MAPG2). The key components of proposed method consist of three parts. Firstly, we utilize graph-based attention model (GAM) to characterize the representations between tasks. Secondly, we formulate the single-step allocation process as a potential game (PG) to guarantee the consistency and soundness of the reward function design. Lastly, our approach sequentially generates allocation strategies through centralized training and decentralized execution (CTDE) framework. Extensive experiments demonstrate that MAPG2achieves a 10% improvement in task completion rate compared to state-of-the-art baselines, validating its effectiveness and robustness. Shangdong Yang, Hongye Cao, Xingguo Chen, Yansheng Wu, Gongzhi Luo |
IEEE Internet Things J. | 1 |
| 2026 | Bellman error centering
Xingguo Chen, Jinguo Ye, Shangdong Yang |
Neural Networks | 5 |
| 2026 | A unified and efficient training framework for open-ended non-transitive games
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001 |
Neural Networks | 3 |
| 2026 | Model-Based Offline Reinforcement Learning With Adversarial Data AugmentationabstractModel-based offline reinforcement learning (RL) constructs environment models from offline datasets to perform conservative policy optimization. Existing approaches focus on learning state transitions through ensemble models, rolling out conservative estimation to mitigate extrapolation errors. However, the static data makes it challenging to develop a robust policy, and offline agents cannot access the environment to gather new data. To address these challenges, we introduce Model-based Offline Reinforcement learning with AdversariaL data augmentation (MORAL). In MORAL, we replace the fixed horizon rollout by employing adversarial data augmentation to execute alternating sampling with ensemble models to enrich training data. Specifically, this adversarial process dynamically selects ensemble models against policy for biased sampling, mitigating the optimistic estimation of fixed models, thus robustly expanding the training data for policy optimization. Moreover, a differential factor (DF) is integrated into the adversarial process for regularization, ensuring error minimization in extrapolations. This data-augmented optimization adapts to diverse offline tasks without rollout horizon tuning, showing remarkable applicability. Extensive experiments on the D4RL benchmark demonstrate that MORAL outperforms other model-based offline RL methods in terms of policy learning and sample efficiency. Hongye Cao, Jing Huo, Shangdong Yang, Tianpei Yang, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Beyond Mandatory Federations: Balancing Egoism, Utilitarianism and Egalitarianism in Mixed-Motive GamesabstractIn the field of mixed-motive games, extensive multi-agent learning studies have explored the balance between egoism (individual interest), utilitarianism (collective interest), and egalitarianism (fairness). Traditional approaches often rely on manually designed reward functions, social norms, and alliance/federation mechanisms to transition agents from individualistic behaviors toward cooperative strategies. However, these methods typically require all agents to share private local information or to mandatorily participate in federations, which is impractical in real-world applications. To address these issues, this paper proposes a Flexible-Participation Federation (FPF) framework that allows agents to participate in the federation voluntarily. Furthermore, we extend the federation from a global to a Local Multi-Federation (LMF) framework, enabling agents to form multiple localized federations, thereby promoting more efficient and adaptive cooperation. Theoretical evidence demonstrates that the global FPF model, along with the discrepancy between decentralized egoistic policies and federated utilitarian policies, achieves an O(1/T) convergence rate. Agents in the LMF framework also reach consensus within a sublinear gap. Extensive experiments show that agents opting out of federation participation experience a reduction in egoism, and our approach outperforms multiple baselines in terms of both utilitarianism and egalitarianism. Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001 |
AAAI | 3 |
| 2025 | DISTA-Net: Dynamic Closely-Spaced Infrared Small Target UnmixingabstractResolving closely-spaced small targets in dense clusters presents a significant challenge in infrared imaging, as the overlapping signals hinder precise determination of their quantity, sub-pixel positions, and radiation intensities. While deep learning has advanced the field of infrared small target detection, its application to closely-spaced infrared small targets has not yet been explored. This gap exists primarily due to the complexity of separating superimposed characteristics and the lack of an open-source infrastructure. In this work, we propose the Dynamic Iterative Shrinkage Thresholding Network (DISTA-Net), which reconceptualizes traditional sparse reconstruction within a dynamic framework. DISTA-Net adaptively generates convolution weights and thresholding parameters to tailor the reconstruction process in real time. To the best of our knowledge, DISTA-Net is the first deep learning model designed specifically for the unmixing of closely-spaced infrared small targets, achieving superior sub-pixel detection accuracy. Moreover, we have established the first open-source ecosystem to foster further research in this field. This ecosystem comprises three key components: (1) CSIST-100K, a publicly available benchmark dataset; (2) CSO-mAP, a custom evaluation metric for sub-pixel detection; and (3) GrokCSO, an open-source toolkit featuring DISTA-Net and other models. Our code and dataset are available at https://github.com/GrokCV/GrokCSO. Shengdong Han, Shangdong Yang, Yuxuan Li 0004, Xin Zhang 0170, Xiang Li 0041, Jian Yang 0003, Ming-Ming Cheng, Yimian Dai |
ICCV | 2 |
| 2025 | AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving ScenariosabstractBy sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous work focus on vehicle-to-vehicle and vehicle-to-infrastructure collaboration, with limited attention to aerial perspectives provided by UAVs, which uniquely offer dynamic, top-down views to alleviate occlusions and monitor large-scale interactive environments. A major reason for this is the lack of high-quality datasets for aerial-ground collaborative scenarios. To bridge this gap, we present AGC-Drive, the first large-scale real-world dataset for Aerial-Ground Cooperative 3D perception. The data collection platform consists of two vehicles, each equipped with five cameras and one LiDAR sensor, and one UAV carrying a forward-facing camera and a LiDAR sensor, enabling comprehensive multi-view and multi-agent perception. Consisting of approximately 80K LiDAR frames and 360K images, the dataset covers 14 diverse real-world driving scenarios, including urban roundabouts, highway tunnels, and on/off ramps. Notably, 17\% of the data comprises dynamic interaction events, including vehicle cut-ins, cut-outs, and frequent lane changes. AGC-Drive contains 350 scenes, each with approximately 100 frames and fully annotated 3D bounding boxes covering 13 object categories. We provide benchmarks for two 3D perception tasks: vehicle-to-vehicle collaborative perception and vehicle-to-UAV collaborative perception. Additionally, we release an open-source toolkit, including spatiotemporal alignment verification tools, multi-agent visualization systems, and collaborative annotation utilities. The dataset and code are available at https://github.com/PercepX/AGC-Drive. Yunhao Hou, Bochao Zou, Shangdong Yang, Junbao Zhuo, Siheng Chen, Jiansheng Chen 0001, Huimin Ma 0001 |
NeurIPS | 5 |
| 2025 | Efficient Last-Iterate Convergence in Solving Extensive-Form GamesabstractTo establish last-iterate convergence for Counterfactual Regret Minimization (CFR) algorithms in learning a Nash equilibrium (NE) of extensive-form games (EFGs), recent studies reformulate learning an NE of the original EFG as learning the NEs of a sequence of (perturbed) regularized EFGs. Hence, proving last-iterate convergence in solving the original EFG reduces to proving last-iterate convergence in solving (perturbed) regularized EFGs. However, these studies only establish last-iterate convergence for Online Mirror Descent (OMD)-based CFR algorithms instead of Regret Matching (RM)-based CFR algorithms in solving perturbed regularized EFGs, resulting in a poor empirical convergence rate, as RM-based CFR algorithms typically outperform OMD-based CFR algorithms. In addition, as solving multiple perturbed regularized EFGs is required, fine-tuning across multiple perturbed regularized EFGs is infeasible, making parameter-free algorithms highly desirable. This paper show that CFR$^+$, a classical parameter-free RM-based CFR algorithm, achieves last-iterate convergence in learning an NE of perturbed regularized EFGs. This is the first parameter-free last-iterate convergence for RM-based CFR algorithms in perturbed regularized EFGs. Leveraging CFR$^+$ to solve perturbed regularized EFGs, we get Reward Transformation CFR$^+$ (RTCFR$^+$). Importantly, we extend prior work on the parameter-free property of CFR$^+$, enhancing its stability, which is vital for the empirical convergence of RTCFR$^+$. Experiments show that RTCFR$^+$ exhibits a significantly faster empirical convergence rate than existing algorithms that achieve theoretical last-iterate convergence. Interestingly, RTCFR$^+$ show performance no worse than average-iterate convergence CFR algorithms. It is the first last-iterate convergence algorithm to achieve such performance. Our code is available at https://github.com/menglinjian/NeurIPS-2025-RTCFR. Linjian Meng, Tianpei Yang, Youzhi Zhang 0001, Zhenxing Ge, Shangdong Yang, Tianyu Ding, Wenbin Li 0006, Bo An 0001, Yang Gao 0001 |
NeurIPS | 5 |
| 2025 | Last-Iterate Convergence of Smooth Regret Matching$^+$ Variants in Learning Nash EquilibriaabstractRegret Matching$^+$ (RM$^+$) variants are widely used to build superhuman Poker AIs, yet few studies investigate their last-iterate convergence in learning a Nash equilibrium (NE). Although their last-iterate convergence is established for games satisfying the Minty Variational Inequality (MVI), no studies have demonstrated that these algorithms achieve such convergence in the broader class of games satisfying the weak MVI. A key challenge in proving last-iterate convergence for RM$^+$ variants in games satisfying the weak MVI is that even if the game's loss gradient satisfies the weak MVI, RM$^+$ variants operate on a transformed loss feedback which does not satisfy the weak MVI. To provide last-iterate convergence for RM$^+$ variants, we introduce a concise yet novel proof paradigm that involves: (i) transforming an RM$^+$ variant into an Online Mirror Descent (OMD) instance that updates within the original strategy space of the game to recover the weak MVI, and (ii) showing last-iterate convergence by proving the distance between accumulated regrets converges to zero via the recovered weak MVI of the feedback. Inspired by our proof paradigm, we propose Smooth Optimistic Gradient Based RM$^+$ (SOGRM$^+$) and show that it achieves last-iterate and finite-time best-iterate convergence in learning an NE of games satisfying the weak MVI, the weakest condition among all known RM$^+$ variants. Experiments show that SOGRM$^+$ significantly outperforms other algorithms. Our code is available at https://github.com/menglinjian/NeurIPS-2025-SOGRM. Linjian Meng, Youzhi Zhang 0001, Zhenxing Ge, Tianyu Ding, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001 |
NeurIPS | 5 |
| 2025 | Coordinating Multi-Agent Reinforcement Learning via Dual Collaborative Constraints
Shaokang Dong, Shangdong Yang, Yujing Hu, Wenbin Li 0006, Yang Gao 0001 |
Neural Networks | 3 |
| 2025 | Multi-Task Multi-Agent Reinforcement Learning With Interaction and Task RepresentationsabstractMulti-task multi-agent reinforcement learning (MT-MARL) is capable of leveraging useful knowledge across multiple related tasks to improve performance on any single task. While recent studies have tentatively achieved this by learning independent policies on a shared representation space, we pinpoint that further advancements can be realized by explicitly characterizing agent interactions within these multi-agent tasks and identifying task relations for selective reuse. To this end, this article proposes Representing Interactions and Tasks (RIT), a novel MT-MARL algorithm that characterizes both intra-task agent interactions and inter-task task relations. Specifically, for characterizing agent interactions, RIT presents the interactive value decomposition to explicitly take the dependency among agents into policy learning. Theoretical analysis demonstrates that the learned utility value of each agent approximates its Shapley value, thus representing agent interactions. Moreover, we learn task representations based on per-agent local trajectories, which assess task similarities and accordingly identify task relations. As a result, RIT facilitates the effective transfer of interaction knowledge across similar multi-agent tasks. Structurally, RIT develops universal policy structure for scalable multi-task policy learning. We evaluate RIT against multiple state-of-the-art baselines in various cooperative tasks, and its significant performance under both multi-task and zero-shot settings demonstrates its effectiveness. Shaokang Dong, Shangdong Yang, Yujing Hu, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | State Abstraction via Deep Supervised Hash LearningabstractState abstraction is a widely used technique in reinforcement learning (RL) that compresses the state space to accelerate learning algorithms. However, designing an effective abstraction function in large-scale or high-dimensional state space problems remains a significant challenge. In this brief, we present a novel state abstraction method based on deep supervised hash learning (DSH) and provide a theoretical analysis of its near-optimal property. Furthermore, by leveraging the DSH-based representation as the optimization objective, we propose a direct and concise optimization method based on the target value. In addition, we construct an auxiliary learning task for state abstraction that can be combined with various RL algorithms. In particular, we apply the DSH-based state abstraction to both deep Q-learning (DQN) and soft actor-critic (SAC). Extensive experiments are conducted on Atari and several classic control benchmarks to evaluate the effectiveness of the DSH-based state abstraction method, showing that our method surpasses existing state abstraction algorithms in performance. Guang Yang 0066, Jing Huo, Shangdong Yang, Tianyu Ding, Xingguo Chen, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Multi-Agent Sparse Interaction Modeling is an Anomaly Detection ProblemabstractMost real-world multi-agent tasks exhibit the characteristic of sparse interaction, wherein agents interact with each other in a limited number of crucial states while largely acting independently. Effectively modeling the sparse interaction and leveraging the learned interaction structure to instruct agents’ learning processes can enhance the efficiency of multi-agent reinforcement learning algorithms. However, it remains unclear how to identify these specific interactive states solely through trials and errors within current multi-agent tasks. To address this challenge, this paper introduces a novel algorithm called Sparse Interaction as Anomaly (SIA), which innovatively casts the sparse interaction modeling into an anomaly detection problem. The underlying intuition is that interactive states appear rarely in agents’ trajectories and exhibit distinct dynamics compared to other commonplace states. Building upon this insight, SIA first employs variational inference to model the latent dynamics of agents’ trajectories. It then designates states with anomalous dynamics as the elusive interactive states and subsequently instructs agents to explore these states more extensively. This facilitates the emergence of interactive behaviors and promotes the learning of multi-agent policies. Experimental evaluation of SIA across various multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness. Shaokang Dong, Shangdong Yang, Hongye Cao, Yang Gao 0001 |
ICASSP | 3 |
| 2024 | STAR: Spatio-Temporal State Compression for Multi-Agent Tasks with Rich Observations
Yujing Hu, Shangdong Yang, Tangjie Lv, Changjie Fan, Wenbin Li 0012, Chongjie Zhang, Yang Gao 0001 |
IJCAI | 3 |
| 2024 | Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in mixed cooperative and competitive environments
Shaokang Dong, Shangdong Yang, Yang Gao 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Selective policy transfer in multi-agent systems with sparse interactions
Yunkai Zhuang, Yong Liu 0007, Shangdong Yang, Yang Gao 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Egoism, utilitarianism and egalitarianism in multi-agent reinforcement learning
Shaokang Dong, Shangdong Yang, Bo An 0001, Wenbin Li 0006, Yang Gao 0001 |
Neural Networks | 3 |
| 2024 | WToE: Learning When to Explore in Multiagent Reinforcement LearningabstractExisting multiagent exploration works focus on how to explore in the fully cooperative task, which is insufficient in the environment with nonstationarity induced by agent interactions. To tackle this issue, we propose When to Explore (WToE), a simple yet effective variational exploration method to learn WToE under nonstationary environments. WToE employs an interaction-oriented adaptive exploration mechanism to adapt to environmental changes. We first propose a novel graphical model that uses a latent random variable to model the step-level environmental change resulting from interaction effects. Leveraging this graphical model, we employ the supervised variational auto-encoder (VAE) framework to derive a short-term inferred policy from historical trajectories to deal with the nonstationarity. Finally, agents engage in exploration when the short-term inferred policy diverges from the current actor policy. The proposed approach theoretically guarantees the convergence of the Q -value function. In our experiments, we validate our exploration mechanism in grid examples, multiagent particle environments and the battle game of MAgent environments. The results demonstrate the superiority of WToE over multiple baselines and existing exploration methods, such as MAEXQ, NoisyNets, EITI, and PR2. Shaokang Dong, Hangyu Mao, Shangdong Yang, Shengyu Zhu 0001, Wenbin Li 0006, Jianye Hao, Yang Gao 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Modeling Rationality: Toward Better Performance Against Unknown Agents in Sequential GamesabstractOpponent modeling is necessary for autonomous agents to capture the intents of others during strategic interactions. Most previous works assume that they can access enough interaction history to build the model. However, it may not be realistic. To solve this problem, we present a novel rationality-consistent opponent modeling (ROM) method for games with imperfect information. In our approach, a game-theoretical concept of consistence about rationality is proposed to take advantage of the characteristic of imperfect information sequential games that rational behavior at disjoint information sets is correlated through anticipated opponent's behavior. With the correlation between different information sets, agents could infer the opponents' strategies at information sets correlated to observed behavior. To exploit the correlation, ROM attempts to conduct reasoning from the opponent's perspective and rationalize its past behavior. In this way, ROM acquires the ability to better adapt to different opponents and achieves a more accurate opponent model with insufficient observation history, which is verified by experiments in different settings. A heuristic adaptation approach is also applied in ROM, which updates the opponent model in an online manner and significantly reduces the computation cost. We evaluate ROM in both a grid world game and a poker game. Compared with other opponent modeling methods, ROM shows better performance and has more accurate predictions in both games against different types of opponents with limited action interactions. Experimental results also show that ROM's time cost is significantly reduced through heuristic adaptation. Zhenxing Ge, Shangdong Yang, Pinzhuo Tian, Yang Gao 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | Learning Multi-Intersection Traffic Signal Control via Coevolutionary Multi-Agent Reinforcement LearningabstractEffective management of multi-intersection traffic signal control (MTSC) is vital for intelligent transportation systems. Multi-agent reinforcement learning (MARL) has shown promise in achieving MTSC. However, existing MARL-based MTSC algorithms have primarily focused on capturing the spatial relationship between multi-intersection traffic signals but have overlooking the importance of the temporally stable traffic pattern. This pattern refers to the fixed positions and relatively stable traffic flow between intersections over short periods in real-world MTSC scenarios, which indicates that the learned spatial relationships between traffic signals should co-evolve over time. To this end, we propose a novel algorithm calledCoevolutionaryMulti-AgentReinforcementLearning (CoevoMARL). CoevoMARL employs a graph neural network to capture the complex spatial interaction network among traffic signals. Furthermore, we propose a relationship-driven progressive LSTM (RDP-LSTM) that dynamically evolves the learned spatial interaction network over time by leveraging insights from the temporally stable traffic pattern. To accelerate convergence, we also propose the mutual information reward optimization (MIRO) technique, which strengthens the correlation between policy learning and high-performance samples by using a mutual information-based intrinsic reward. Experimental results on both synthetic and realistic datasets demonstrate the superiority of CoevoMARL over existing MTSC algorithms, providing valuable insights into incorporating the temporally stable traffic pattern. Wubing Chen, Shangdong Yang, Wenbin Li 0006, Yujing Hu, Yang Gao 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Learning Explicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning via Polarization Policy GradientabstractCooperative multi-agent policy gradient (MAPG) algorithms have recently attracted wide attention and are regarded as a general scheme for the multi-agent system. Credit assignment plays an important role in MAPG and can induce cooperation among multiple agents. However, most MAPG algorithms cannot achieve good credit assignment because of the game-theoretic pathology known as centralized-decentralized mismatch. To address this issue, this paper presents a novel method, Multi-Agent Polarization Policy Gradient (MAPPG). MAPPG takes a simple but efficient polarization function to transform the optimal consistency of joint and individual actions into easily realized constraints, thus enabling efficient credit assignment in MAPPG. Theoretically, we prove that individual policies of MAPPG can converge to the global optimum. Empirically, we evaluate MAPPG on the well-known matrix game and differential game, and verify that MAPPG can converge to the global optimum for both discrete and continuous action spaces. We also evaluate MAPPG on a set of StarCraft II micromanagement tasks and demonstrate that MAPPG outperforms the state-of-the-art MAPG algorithms. Wubing Chen, Wenbin Li 0006, Shangdong Yang, Yang Gao 0001 |
AAAI | 4 |
| 2023 | Enhancing OOD Generalization in Offline Reinforcement Learning with Energy-Based Policy OptimizationabstractOffline Reinforcement Learning (RL) is an important research domain for real-world applications because it can avert expensive and dangerous online exploration. Offline RL is prone to extrapolation errors caused by the distribution shift between offline datasets and states visited by behavior policy. Existing offline RL methods constrain the policy to offline behavior to prevent extrapolation errors. But these methods limit the generalization potential of agents in Out-Of-Distribution (OOD) regions and cannot effectively evaluate OOD generalization behavior. To improve the generalization of the policy in OOD regions while avoiding extrapolation errors, we propose an Energy-Based Policy Optimization (EBPO) method for OOD generalization. An energy function based on the distribution of offline data is proposed for the evaluation of OOD generalization behavior, instead of relying on model discrepancies to constrain the policy. The way of quantifying exploration behavior in terms of energy values can balance the return and risk. To improve the stability of generalization and solve the problem of sparse reward in complex environment, episodic memory is applied to store successful experiences that can improve sample efficiency. Extensive experiments on the D4RL datasets demonstrate that EBPO outperforms the state-of-the-art methods and achieves robust performance on challenging tasks that require OOD generalization. Hongye Cao, Shangdong Yang, Jing Huo, Xingguo Chen, Yang Gao 0001 |
ECAI | 2 |
| 2023 | Convergence Analysis of Graphical Game-Based Nash Q-Learning using the Interaction Detection Signal of N-Step ReturnabstractThe graphical game provides an effective method for modeling different kinds of sparse interactions in multi-agent reinforcement learning. Most previous work on game abstraction lacks theoretical guarantees of convergence. In this paper, we adopt the ${\mathcal{N}}$-step return signal to detect interactions between agents and build the Markov graphical game based on it. We analyze that the solution of the Markov graphical game is an ϵ-Nash equilibrium which guarantees the convergence of the proposed NSR-G2NashQ algorithm theoretically. Also, we have done experiments in different multi-agent reinforcement learning tasks with both tabular and function approximation solutions. The results show the NSR-G2NashQ algorithm accelerates the convergence of agents to the optimal policy. Yunkai Zhuang, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001 |
ICASSP | 2 |
| 2023 | Modified Retrace for Off-Policy Temporal Difference LearningabstractOff-policy learning is a key to extend reinforcement learning as it allows to learn a target policy from a different behavior policy that generates the data. However, it is well known as “the deadly triad” when combined with bootstrapping and function approximation. Retrace is an efficient and convergent off-policy algorithm with tabular value functions which employs truncated importance sampling ratios. Unfortunately, Retrace is known to be unstable with linear function approximation. In this paper, we propose modified Retrace to correct the off-policy return, derive a new off-policy temporal difference learning algorithm (TD-MRetrace) with linear function approximation, and obtain a convergence guarantee under standard assumptions. Experimental results on counterexamples and control tasks validate the effectiveness of the proposed algorithm compared with traditional algorithms. Xingguo Chen, Xingzhou Ma, Guang Yang 0066, Shangdong Yang, Yang Gao 0001 |
UAI | 5 |
| 2023 | Leveraging transition exploratory bonus for efficient exploration in Hard-Transiting reinforcement learning problems
Shangdong Yang, Shaokang Dong, Xingguo Chen |
Future Gener. Comput. Syst. | 1 |
| 2023 | Online attentive kernel-based temporal difference learning
Xingguo Chen, Guang Yang 0066, Shangdong Yang, Shaokang Dong, Yang Gao 0001 |
Knowl. Based Syst. | 3 |
| 2021 | An Optimal Algorithm for the Stochastic Bandits While Knowing the Near-Optimal Mean RewardabstractThis brief studies a variation of the stochastic multiarmed bandit (MAB) problems, where the agent knows the a priori knowledge named the near-optimal mean reward (NoMR). In common MAB problems, an agent tries to find the optimal arm without knowing the optimal mean reward. However, in more practical applications, the agent can usually get an estimation of the optimal mean reward defined as NoMR. For instance, in an online Web advertising system based on MAB methods, a user's near-optimal average click rate (NoMR) can be roughly estimated from his/her demographic characteristics. As a result, application of the NoMR is efficient at improving the algorithm's performance. First, we formalize the stochastic MAB problem by knowing the NoMR that is in between the suboptimal mean reward and the optimal mean reward. Second, we use the cumulative regret as the performance metric for our problem, and we get that this problem's lower bound of the cumulative regret is Ω(1/∆) , where ∆ is the difference between the suboptimal mean reward and the optimal mean reward. Compared with the conventional MAB problem with the increasing logarithmic lower bound of the regret, our regret lower bound is uniform with the learning step. Third, a novel algorithm, NoMR-BANDIT, is set forth to solve this problem. In NoMR-BANDIT, the NoMR is used to design an efficient exploration strategy. In addition, we analyzed the regret's upper bound in NoMR-BANDIT and concluded that it also has a uniform upper bound of O(1/∆) , which is in the same order as the lower bound. Consequently, NoMR-BANDIT is an optimal algorithm of this problem. To enhance our method's generalization, CASCADE-BANDIT based on NoMR-BANDIT is proposed to solve the problem, where NoMR is less than the suboptimal mean reward. CASCADE-BANDIT has an upper bound of O(∆logn) , where n represents the learning step, and the order of O(∆logn) is the same with that of the conventional MAB methods. Finally, extensive experimental results demonstrated that the established NoMR-BANDIT is more efficient than the compared bandit solutions. After sufficient iterations, NOMR-BANDIT saved 10%-80% more cumulative regret than the state of the art. Shangdong Yang, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | A Contextual Bandit Approach to Personalized Online Recommendation via Sparse Interactions
Hao Wang 0013, Shangdong Yang, Yang Gao 0001 |
PAKDD (2) | 3 |
| 2016 | Efficient Average Reward Reinforcement Learning Using Constant Shifting ValuesabstractThere are two classes of average reward reinforcement learning (RL) algorithms: model-based ones that explicitly maintain MDP models and model-free ones that do not learn such models. Though model-free algorithms are known to be more efficient, they often cannot converge to optimal policies due to the perturbation of parameters. In this paper, a novel model-free algorithm is proposed, which makes use of constant shifting values (CSVs) estimated from prior knowledge. To encourage exploration during the learning process, the algorithm constantly subtracts the CSV from the rewards. A terminating condition is proposed to handle the unboundedness of Q-values caused by such substraction. The convergence of the proposed algorithm is proved under very mild assumptions. Furthermore, linear function approximation is investigated to generalize our method to handle large-scale tasks. Extensive experiments on representative MDPs and the popular game Tetris show that the proposed algorithms significantly outperform the state-of-the-art ones. Shangdong Yang, Yang Gao 0001, Bo An 0001, Hao Wang 0013, Xingguo Chen |
AAAI | 1 |
| 2016 | Incremental Nonnegative Matrix Factorization Based on Matrix Sketching and k-means Clustering
Hao Wang 0013, Shangdong Yang, Yang Gao 0001 |
IDEAL | 3 |