Ying Wen 0001

dblp:41/4203-1 · DBLP profile ↗
← Back
48ranked-venue papers
3as first author
42since 2021 · last 2026
0000-0003-1247-2382ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 3 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Offline Fictitious Self-Play for Competitive Games
abstract
Offline Reinforcement Learning (RL) enables policy improvement from fixed datasets without online interactions, making it highly suitable for real-world applications lacking efficient simulators. Despite its success in the single-agent setting, offline multi-agent RL remains a challenge, especially in competitive games. Firstly, unaware of the game structure, it is impossible to interact with the opponents and conduct a major learning paradigm, self-play, for competitive games. Secondly, real-world datasets cannot cover all the state and action space in the game, resulting in barriers to identifying Nash equilibrium (NE). To address these issues, this paper introduces Off-FSP, the first practical model-free offline RL algorithm for competitive games. We start by simulating interactions with various opponents by adjusting the weights of the fixed dataset with importance sampling. This technique allows us to learn the best responses to different opponents and employ the Offline Self-Play learning framework. To overcome the challenge of partial coverage, we combine the single-agent offline RL method with Fictitious Self-Play (FSP) to approximate NE by constraining the approximate best responses away from out-of-distribution actions. Experiments on matrix games, extensive-form poker, and board games demonstrate that Off-FSP achieves significantly lower exploitability than state-of-the-art baselines. Finally, we validate Off-FSP on a real-world human-robot competitive task, demonstrating its potential for solving complex, hard-to-simulate real-world problems.
Jingxiao Chen, Weiji Xie, Weinan Zhang 0001, Yong Yu 0001, Ying Wen 0001
AAAI5
2026 ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
abstract
Jianghao Lin, Yuanyuan Shi, Xin Peng, Renjie Ding, Hairui Wang, Yuxuan Peng, Bizhe Bai, Weixi Song, Fengshuo Bai, Huacan Chai, Weinan Zhang, Fei Huang, Ying Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jianghao Lin, Renjie Ding, Yuxuan Peng, Bizhe Bai, Weixi Song, Fengshuo Bai, Huacan Chai, Weinan Zhang 0001, Ying Wen 0001
ACL (1)13
2026 DSR: optimization of performance lower bound for hierarchical policy with dynamical skill refinement
Dongxiang Chen, Ying Wen 0001
Frontiers Comput. Sci.2
2026 Computing ex ante equilibrium in heterogeneous zero-sum team game
Naming Liu, Xihuai Wang, Weinan Zhang 0001, Yaodong Yang 0001, Youzhi Zhang 0001, Bo An 0001, Ying Wen 0001
Frontiers Comput. Sci.8
2025 RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
abstract
Evaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into specific behaviors that align with the attacker’s objectives, often bypassing traditional reward-based defenses. Prior methods have primarily focused on reducing cumulative rewards; however, rewards are typically too generic to capture complex safety requirements effectively. As a result, focusing solely on reward reduction can lead to suboptimal attack strategies, particularly in safety-critical scenarios where more precise behavior manipulation is needed. To address these challenges, we propose RAT, a method designed for universal, targeted behavior attacks. RAT trains an intention policy that is explicitly aligned with human preferences, serving as a precise behavioral target for the adversary. Concurrently, an adversary manipulates the victim's policy to follow this target behavior. To enhance the effectiveness of these attacks, RAT dynamically adjusts the state occupancy measure within the replay buffer, allowing for more controlled and effective behavior manipulation. Our empirical results on robotic simulation tasks demonstrate that RAT outperforms existing adversarial attack algorithms in inducing specific behaviors. Additionally, RAT shows promise in improving agent robustness, leading to more resilient policies. We further validate RAT by guiding Decision Transformer agents to adopt behaviors aligned with human preferences in various MuJoCo tasks, demonstrating its effectiveness across diverse tasks.
Fengshuo Bai, Runze Liu 0002, Yali Du 0001, Ying Wen 0001, Yaodong Yang 0001
AAAI4
2025 Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration
abstract
Shao Zhang, Xihuai Wang, Wenhao Zhang, Chaoran Li, Junru Song, Tingyu Li, Lin Qiu, Xuezhi Cao, Xunliang Cai, Wen Yao, Weinan Zhang, Xinbing Wang, Ying Wen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shao Zhang, Xihuai Wang, Junru Song, Xuezhi Cao, Weinan Zhang 0001, Xinbing Wang, Ying Wen 0001
ACL (1)13
2025 Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles
abstract
For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy for Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR.
Jiesong Lian, Yucong Huang, Chengdong Ma, Ying Wen 0001, Long Hu, Yixue Hao
ECAI5
2025 Accelerating Multi‑turn LLM Agent Workflows via Context Templating and Opportunistic Prefill
Hanjing Wang, Ying Wen 0001, Weinan Zhang 0001
ICIC (13)2
2025 PMAT: Optimizing Action Generation Order in Multi-Agent Reinforcement Learning
Muning Wen, Xihuai Wang, Shao Zhang, Yiwei Shi, Minne Li, Minglong Li, Ying Wen 0001
AAMAS8
2025 Unlocking the Potential of Decentralized LLM-based MAS: Privacy Preservation and Monetization in Collective Intelligence
Yingxuan Yang, Qiuying Peng, Jun Wang 0012, Ying Wen 0001, Weinan Zhang 0001
AAMAS4
2025 STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization
abstract
Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning from human feedback. However, due to the high cost of obtaining feedback, PbRL typically relies on a limited set of preference-labeled samples. This data scarcity introduces two key inefficiencies: (1) the reward model overfits to the limited feedback, leading to poor generalization to unseen samples, and (2) the agent exploits the learned reward model, exacerbating overestimation of action values in temporal difference (TD) learning. To address these issues, we propose STAR, an efficient PbRL method that integrates preference margin regularization and policy regularization. Preference margin regularization mitigates overfitting by introducing a bounded margin in reward optimization, preventing excessive bias toward specific feedback. Policy regularization bootstraps a conservative estimate $\widehat{Q}$ from well-supported state-action pairs in the replay memory, reducing overestimation during policy learning. Experimental results show that STAR improves feedback efficiency, achieving 34.8\% higher performance in online settings and 29.7\% in offline settings compared to state-of-the-art methods. Ablation studies confirm that STAR facilitates more robust reward and value function learning. The videos of this project are released at https://sites.google.com/view/pbrl-star.
Fengshuo Bai, Rui Zhao 0001, Hongming Zhang 0003, Sijia Cui, Shao Zhang, Bo Xu 0002, Ying Wen 0001, Yaodong Yang 0001
NeurIPS8
2025 ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
abstract
Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench.
Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004
NeurIPS8
2025 ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning
abstract
Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking—enabling models to monitor, evaluate, and control their reasoning processes for more adaptive and effective problem-solving. However, current single-agent work lacks a specialized design for acquiring meta-thinking, resulting in low efficacy. To address this challenge, we introduce Reinforced Meta-thinking Agents (ReMA), a novel framework that leverages Multi-Agent Reinforcement Learning (MARL) to elicit meta-thinking behaviors, encouraging LLMs to think about thinking. ReMA decouples the reasoning process into two hierarchical agents: a high-level meta-thinking agent responsible for generating strategic oversight and plans, and a low-level reasoning agent for detailed executions. Through iterative reinforcement learning with aligned objectives, these agents explore and learn collaboration, leading to improved generalization and robustness. Empirical results from single-turn experiments demonstrate that ReMA outperforms single-agent RL baselines on complex reasoning tasks, including competitive-level mathematical benchmarks and LLM-as-a-Judge benchmarks. Additionally, we further extend ReMA to multi-turn interaction settings, leveraging turn-level ratio and parameter sharing to improve efficiency. Comprehensive ablation studies further illustrate the evolving dynamics of each distinct agent, providing valuable insights into how the meta-thinking reasoning process enhances the reasoning capabilities of LLMs.
Ziyu Wan, Xiaoyu Wen 0001, Yan Song 0003, Hanjing Wang, Linyi Yang, Mark Schmidt 0001, Jun Wang 0012, Weinan Zhang 0001, Shuyue Hu, Ying Wen 0001
NeurIPS11
2024 DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning
abstract
In this work, we investigate the potential of large language models (LLMs) based agents to automate data science tasks, with the goal of comprehending task requirements, then building and training the best-fit machine learning models. Despite their widespread success, existing LLM agents are hindered by generating unreasonable experiment plans within this scenario. To this end, we present DS-Agent, a novel automatic framework that harnesses LLM agent and case-based reasoning (CBR). In the development stage, DS-Agent follows the CBR framework to structure an automatic iteration pipeline, which can flexibly capitalize on the expert knowledge from Kaggle, and facilitate consistent performance improvement through the feedback mechanism. Moreover, DS-Agent implements a low-resource deployment stage with a simplified CBR paradigm to adapt past successful solutions from the development stage for direct code generation, significantly reducing the demand on foundational capabilities of LLMs. Empirically, DS-Agent with GPT-4 achieves 100% success rate in the development stage, while attaining 36% improvement on average one pass rate across alternative LLMs in the deployment stage. In both stages, DS-Agent achieves the best rank in performance, costing $1.60 and \$0.13 per run with GPT-4, respectively. Our data and code are open-sourced at https://github.com/guosyjlu/DS-Agent.
Siyuan Guo 0001, Cheng Deng 0001, Ying Wen 0001, Hechang Chen, Yi Chang 0001, Jun Wang 0012
ICML3
2024 AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and Training
abstract
Recent works like Tree-of-Thought (ToT) and Reasoning via Planning (RAP) aim to augment the multi-step reasoning capabilities of LLMs by using tree-search algorithms. These methods rely on prompting a pre-trained model to serve as a value function and focus on problems with low search depth. As a result, these methods cannot benefit from in-domain training and only rely on pretraining process — they will not work in domains where the pre-trained LLM does not have enough knowledge to serve as an effective value function or in domains that require long-horizon planning. To address these limitations, we present an AlphaZero-like tree-search learning framework for LLMs (termed TS-LLM), systematically illustrating how tree-search with a learned value function can guide LLM decoding. TS-LLM distinguishes itself in two key ways. (1) Leveraging a learned value function and AlphaZero-like algorithms, our approach can be generally adaptable to a wide range of tasks, language models of any size, and tasks of varying search depths. (2) Our approach can guide LLMs during both inference and training, iteratively improving the LLMs. Empirical results across reasoning, planning, alignment, and decision-making tasks show that TS-LLM outperforms existing approaches and can handle trees with a depth of 64.
Ziyu Wan, Xidong Feng, Muning Wen, Stephen McAleer, Ying Wen 0001, Weinan Zhang 0001, Jun Wang 0012
ICML5
2024 Aligning Individual and Collective Objectives in Multi-Agent Cooperation
abstract
Among the research topics in multi-agent learning, mixed-motive cooperation is one of the most prominent challenges, primarily due to the mismatch between individual and collective goals. The cutting-edge research is focused on incorporating domain knowledge into rewards and introducing additional mechanisms to incentivize cooperation. However, these approaches often face shortcomings such as the effort on manual design and the absence of theoretical groundings. To close this gap, we model the mixed-motive game as a differentiable game for the ease of illuminating the learning dynamics towards cooperation. More detailed, we introduce a novel optimization method named \textbf{\textit{A}}ltruistic \textbf{\textit{G}}radient \textbf{\textit{A}}djustment (\textbf{\textit{AgA}}) that employs gradient adjustments to progressively align individual and collective objectives. Furthermore, we theoretically prove that AgA effectively attracts gradients to stable fixed points of the collective objective while considering individual interests, and we validate these claims with empirical evidence. We evaluate the effectiveness of our algorithm AgA through benchmark environments for testing mixed-motive collaboration with small-scale agents such as the two-player public good game and the sequential social dilemma games, Cleanup and Harvest, as well as our self-developed large-scale environment in the game StarCraft II.
Yang Li 0116, Shao Zhang, Yali Du 0001, Ying Wen 0001, Wei Pan 0004
NeurIPS6
2024 ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination
abstract
Zero-shot coordination (ZSC) is a new cooperative multi-agent reinforcement learning (MARL) challenge that aims to train an ego agent to work with diverse, unseen partners during deployment. The significant difference between the deployment-time partners' distribution and the training partners' distribution determined by the training algorithm makes ZSC a unique out-of-distribution (OOD) generalization challenge. The potential distribution gap between evaluation and deployment-time partners leads to inadequate evaluation, which is exacerbated by the lack of appropriate evaluation metrics. In this paper, we present ZSC-Eval, the first evaluation toolkit and benchmark for ZSC algorithms. ZSC-Eval consists of: 1) Generation of evaluation partner candidates through behavior-preferring rewards to approximate deployment-time partners' distribution; 2) Selection of evaluation partners by Best-Response Diversity (BR-Div); 3) Measurement of generalization performance with various evaluation partners via the Best-Response Proximity (BR-Prox) metric. We use ZSC-Eval to benchmark ZSC algorithms in Overcooked and Google Research Football environments and get novel empirical findings. We also conduct a human experiment of current ZSC algorithms to verify the ZSC-Eval's consistency with human evaluation. ZSC-Eval is now available at https://github.com/sjtu-marl/ZSC-Eval.
Xihuai Wang, Shao Zhang, Jingxiao Chen, Ying Wen 0001, Weinan Zhang 0001
NeurIPS6
2024 Reinforcing LLM Agents via Policy Optimization with Action Decomposition
abstract
Language models as intelligent agents push the boundaries of sequential decision-making agents but struggle with limited knowledge of environmental dynamics and exponentially huge action space. Recent efforts like GLAM and TWOSOME manually constrain the action space to a restricted subset and employ reinforcement learning to align agents' knowledge with specific environments. However, they overlook fine-grained credit assignments for intra-action tokens, which is essential for efficient language agent optimization, and rely on human's prior knowledge to restrict action space. This paper proposes decomposing language agent optimization from the action level to the token level, offering finer supervision for each intra-action token and manageable optimization complexity in environments with unrestricted action spaces. Beginning with the simplification of flattening all actions, we theoretically explore the discrepancies between action-level optimization and this naive token-level optimization. We then derive the Bellman backup with Action Decomposition (BAD) to integrate credit assignments for both intra-action and inter-action tokens, effectively eliminating the discrepancies. Implementing BAD within the PPO algorithm, we introduce Policy Optimization with Action Decomposition (POAD). POAD benefits from a finer-grained credit assignment process and lower optimization complexity, leading to enhanced learning efficiency and generalization abilities in aligning language agents with interactive environments. We validate POAD across diverse testbeds, with results affirming the advantages of our approach and the correctness of our theoretical analysis. The source code can be accessed directly with this link: https://github.com/morning9393/ADRL.
Muning Wen, Ziyu Wan, Jun Wang 0012, Weinan Zhang 0001, Ying Wen 0001
NeurIPS5
2024 TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision
abstract
Several large language model (LLM) agents have been constructed for diverse purposes such as web navigation and online shopping, leveraging the broad knowledge and text comprehension capabilities of LLMs. Many of these works rely on in-context examples to achieve generalization without requiring fine-tuning. However, few have addressed the challenge of selecting and effectively utilizing these examples. Recent approaches have introduced trajectory-level retrieval with task meta-data and the use of trajectories as in-context examples to enhance overall performance in some sequential decision making tasks like computer control. Nevertheless, these methods face issues like plausible examples retrieved without task-specific state transition dynamics and long input with plenty of irrelevant context due to using complete trajectories. In this paper, we propose a novel framework (TRAD) to tackle these problems. TRAD first employs Thought Retrieval for step-level demonstration selection through thought matching, enhancing the quality of demonstrations and reducing irrelevant input noise. Then, Aligned Decision is introduced to complement retrieved demonstration steps with their preceding or subsequent steps, providing tolerance for imperfect thought and offering a balance between more context and less noise. Extensive experiments on ALFWorld and Mind2Web benchmarks demonstrate that TRAD not only surpasses state-of-the-art models but also effectively reduces noise and promotes generalization. Furthermore, TRAD has been deployed in real-world scenarios of a global business insurance company and yields an improved success rate of robotic process automation. Our codes are available at: https://github.com/skyriver-2000/TRAD-Official.
Ruiwen Zhou, Yingxuan Yang, Muning Wen, Ying Wen 0001, Chunling Xi, Yong Yu 0001, Weinan Zhang 0001
SIGIR4
2024 Tackling Cooperative Incompatibility for Zero-Shot Human-AI Coordination
abstract
Securing coordination between AI agent and teammates (human players or AI agents) in contexts involving unfamiliar humans continues to pose a significant challenge in Zero-Shot Coordination. The issue of cooperative incompatibility becomes particularly prominent when an AI agent is unsuccessful in synchronizing with certain previously unknown partners. Traditional algorithms have aimed to collaborate with partners by optimizing fixed objectives within a population, fostering diversity in strategies and behaviors. However, these techniques may lead to learning loss and an inability to cooperate with specific strategies within the population, a phenomenon named cooperative incompatibility in learning. In order to solve cooperative incompatibility in learning and effectively address the problem in the context of ZSC, we introduce the Cooperative Open-ended LEarning (COLE) framework, which formulates open-ended objectives in cooperative games with two players using perspectives of graph theory to evaluate and pinpoint the cooperative capacity of each strategy. We present two practical algorithms, specifically COLESV and COLER, which incorporate insights from game theory and graph theory. We also show that COLE could effectively overcome the cooperative incompatibility from theoretical and empirical analysis. Subsequently, we created an online Overcooked human-AI experiment platform, the COLE platform, which enables easy customization of questionnaires, model weights, and other aspects. Utilizing the COLE platform, we enlist 130 participants for human experiments. Our findings reveal a preference for our approach over state-of-the-art methods using a variety of subjective metrics. Moreover, objective experimental outcomes in the Overcooked game environment indicate that our method surpasses existing ones when coordinating with previously unencountered AI agents and the human proxy model. Our code and demo are publicly available at https://sites.google.com/view/cole-2023.
Yang Li 0116, Shao Zhang, Jichen Sun, Yali Du 0001, Ying Wen 0001, Xinbing Wang, Wei Pan 0004
J. Artif. Intell. Res.6
2024 Cross-Utterance Conditioned VAE for Speech Generation
abstract
Speech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In response, we present the Cross-Utterance Conditioned Variational Autoencoder speech synthesis (CUC-VAE S2) framework to enhance prosody and ensure natural speech generation. This framework leverages the powerful representational capabilities of pre-trained language models and the re-expression abilities of variational autoencoders (VAEs). The core component of the CUC-VAE S2 framework is the cross-utterance CVAE, which extracts acoustic, speaker, and textual features from surrounding sentences to generate context-sensitive prosodic features, more accurately emulating human prosody generation. We further propose two practical algorithms tailored for distinct speech synthesis applications: CUC-VAE TTS for text-to-speech and CUC-VAE SE for speech editing. The CUC-VAE TTS is a direct application of the framework, designed to generate audio with contextual prosody derived from surrounding texts. On the other hand, the CUC-VAE SE algorithm leverages real mel spectrogram sampling conditioned on contextual information, producing audio that closely mirrors real sound and thereby facilitating flexible speech editing based on text such as deletion, insertion, and replacement. Experimental results on the LibriTTS datasets demonstrate that our proposed models significantly enhance speech synthesis and editing, producing more natural and expressive speech.
Yang Li 0116, Guangzhi Sun, Weiqin Zu, Zheng Tian 0002, Ying Wen 0001, Wei Pan 0004, Chao Zhang 0031, Jun Wang 0012, Yang Yang 0001, Fanglei Sun
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Self-Supervised MAFENN for Classifying Low-Labeled Distorted Images Over Mobile Fading Channels
abstract
Image distortion during wireless transmission presents a significant challenge for real-world artificial intelligence (AI) applications. Recent methods have attempted to address this issue by integrating neural networks into the wireless transmission system. However, these approaches often require a large volume of labeled training data, which can be expensive and time-consuming to collect. To address this issue, we propose a novel approach,Self-SupervisedMulti-AgentFeedbackEnabledNeuralNetworks (S2MAFENN). S2MAFENN is designed to improve the efficiency of labeled data in wireless image transmission. It incorporates a Feedbacker agent that emulates the error correction mechanisms observed in primate brains and employs self-supervised contrastive learning to extract representations from unlabeled distorted images independently. From a theoretical perspective, we model the training process of S2MAFENN as a three-player Stackelberg game and provide evidence that S2MAFENN can achieve exponential convergence rates. We then empirically validate our approach by assessing the representations learned through S2MAFENN. We use varied labeled CIFAR10 and CIFAR100 data to simulate real image transmissions over the Rayleigh fading and 5G channels. Our results show that S2MAFENN matches or even surpasses the performance of state-of-the-art self-supervised training methods, even when only 50% of labels are used. Moreover, S2MAFENN yields average accuracy gains of 5.11%, 5.8%, and 4.58% with only 0.1, 0.2, and 0.5 of the labels transmitted over the 5G channel, respectively. For the downstream task of semantic segmentation over the 5G channel, S2MAFENN exhibits significant advancements on the ADE20K dataset. It achieves enhancements of approximately 7% and 8.7% in Mean IoU and DICE metrics, respectively, surpassing the performance of current state-of-the-art methods.
Yang Li 0116, Fanglei Sun, Jingchen Hu, Fan Wu 0006, Kai Li 0022, Ying Wen 0001, Zheng Tian 0002, Yaodong Yang 0001, Jiangcheng Zhu, Jun Wang 0012, Yang Yang 0001
IEEE Trans. Mob. Comput.7
2023 DCAC: Reducing Unnecessary Conservatism in Offline-to-online Reinforcement Learning
abstract
Recent advancements in offline reinforcement learning (RL) have facilitated the training of powerful agents using fixed datasets exclusively. Despite this, the quality of a dataset plays a critical role in determining an agent’s performance, and high-quality datasets are often scarce. This scarcity necessitates the enhancement of agents through subsequent environmental interactions. Particularly, the state-action distribution shift may exert a potentially detrimental effect on well-initialized policies, thus impeding the straightforward application of off-policy RL algorithms to policies trained offline. Predominant offline-to-online RL approaches are typically founded on conservatism, a characteristic that may inadvertently confine the asymptotic performance. In response, we propose a method referred to as Dynamically Constrained Actor-Critic (DCAC), grounded in the mathematical form of dynamically constrained policy optimization. This innovative method enables judicious adjustments to the constraints on policy optimization in accordance with a specified rule, thus stabilizing the initial online learning stage and reducing undue conservatism that restricts asymptotic performance. Through comprehensive experimentation across diverse locomotion tasks, we have ascertained that our method successfully improves the policies trained offline with various datasets via subsequent online environmental interactions. The empirical results substantiate that our method mitigates the harmful effects of distribution shift and consistently attains superior asymptotic performance in comparison to prior works.
Dongxiang Chen, Ying Wen 0001
DAI2
2023 Adaptive Control Strategy for Quadruped Robots in Actuator Degradation Scenarios
abstract
Quadruped robots have strong adaptability to extreme environments but may also experience faults. Once these faults occur, robots must be repaired before returning to the task, reducing their practical feasibility. One prevalent concern among these faults is actuator degradation, stemming from factors like device aging or unexpected operational events. Traditionally, addressing this problem has relied heavily on intricate fault-tolerant design, which demands deep domain expertise from developers and lacks generalizability. Learning-based approaches offer effective ways to mitigate these limitations, but a research gap exists in effectively deploying such methods on real-world quadruped robots. This paper introduces a pioneering teacher-student framework rooted in reinforcement learning, named Actuator Degeneration Adaptation Transformer (Adapt), aimed at addressing this research gap. This framework produces a unified control strategy, enabling the robot to sustain its locomotion and perform tasks despite sudden joint actuator faults, relying exclusively on its internal sensors. Empirical evaluations on the Unitree A1 platform validate the deployability and effectiveness of Adapt on real-world quadruped robots, and affirm the robustness and practicality of our approach.
Hang Lai, Yong Yu 0001, Ying Wen 0001
DAI5
2023 Order Matters: Agent-by-agent Policy Optimization
Xihuai Wang, Zheng Tian 0002, Ziyu Wan, Ying Wen 0001, Jun Wang 0012, Weinan Zhang 0001
ICLR4
2023 Cooperative Open-ended Learning Framework for Zero-Shot Coordination
abstract
Zero-shot coordination in cooperative artificial intelligence (AI) remains a significant challenge, which means effectively coordinating with a wide range of unseen partners. Previous algorithms have attempted to address this challenge by optimizing fixed objectives within a population to improve strategy or behaviour diversity. However, these approaches can result in a loss of learning and an inability to cooperate with certain strategies within the population, known as cooperative incompatibility. To address this issue, we propose the Cooperative Open-ended LEarning (COLE) framework, which constructs open-ended objectives in cooperative games with two players from the perspective of graph theory to assess and identify the cooperative ability of each strategy. We further specify the framework and propose a practical algorithm that leverages knowledge from game theory and graph theory. Furthermore, an analysis of the learning process of the algorithm shows that it can efficiently overcome cooperative incompatibility. The experimental results in the Overcooked game environment demonstrate that our method outperforms current state-of-the-art methods when coordinating with different-level partners. Our demo is available at https://sites.google.com/view/cole-2023.
Yang Li 0116, Shao Zhang, Jichen Sun, Yali Du 0001, Ying Wen 0001, Xinbing Wang, Wei Pan 0004
ICML5
2023 GEAR: A GPU-Centric Experience Replay System for Large Reinforcement Learning Models
abstract
This paper introduces a distributed, GPU-centric experience replay system, GEAR, designed to perform scalable reinforcement learning (RL) with large sequence models (such as transformers). With such models, existing systems such as Reverb face considerable bottlenecks in memory, computation, and communication. GEAR, however, optimizes memory efficiency by enabling the memory resources on GPU servers (including host memory and device memory) to manage trajectory data. Furthermore, it facilitates decentralized GPU devices to expedite various trajectory selection strategies, circumventing computational bottlenecks. GEAR is equipped with GPU kernels capable of collecting trajectories using zero-copy access to host memory, along with remote-directed-memory access over InfiniBand, improving communication efficiency. Cluster experiments have shown that GEAR can achieve performance levels up to 6× greater than Reverb when training state-of-the-art large RL models. GEAR is open-sourced at https:// github.com/bigrl-team/gear.
Hanjing Wang, Man-Kit Sit, Congjie He, Ying Wen 0001, Weinan Zhang 0001, Jun Wang 0012, Yaodong Yang 0001, Luo Mai
ICML4
2023 Large sequence models for sequential decision-making: a survey
Muning Wen, Runji Lin, Hanjing Wang, Yaodong Yang 0001, Ying Wen 0001, Luo Mai, Jun Wang 0012, Haifeng Zhang 0002, Weinan Zhang 0001
Frontiers Comput. Sci.5
2023 MALib: A Parallel Framework for Population-based Multi-agent Reinforcement Learning
abstract
Population-based multi-agent reinforcement learning (PB-MARL) encompasses a range of methods that merge dynamic population selection with multi-agent reinforcement learning algorithms (MARL). While PB-MARL has demonstrated notable achievements in complex multi-agent tasks, its sequential execution is plagued by low computational efficiency due to the diversity in computing patterns and policy combinations. We propose a solution involving a stateless central task dispatcher and stateful workers to handle PB-MARL's subroutines, thereby capitalizing on parallelism across various components for efficient problem-solving. In line with this approach, we introduce MALib, a parallel framework that incorporates a task control model, independent data servers, and an abstraction of MARL training paradigms. The framework has undergone extensive testing and is available under the MIT license (https://github.com/sjtu-marl/malib)
Ming Zhou 0006, Ziyu Wan, Hanjing Wang, Muning Wen, Runzhe Wu, Ying Wen 0001, Yaodong Yang 0001, Yong Yu 0001, Jun Wang 0012, Weinan Zhang 0001
J. Mach. Learn. Res.6
2022 Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech
abstract
Yang Li, Cheng Yu, Guangzhi Sun, Hua Jiang, Fanglei Sun, Weiqin Zu, Ying Wen, Yang Yang, Jun Wang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yang Li 0116, Guangzhi Sun, Fanglei Sun, Weiqin Zu, Ying Wen 0001, Yang Yang 0001, Jun Wang 0012
ACL (1)7
2022 A Game-Theoretic Approach to Multi-agent Trust Region Optimization
Ying Wen 0001, Yaodong Yang 0001, Minne Li, Zheng Tian 0002, Xu Chen 0017, Jun Wang 0012
DAI1
2022 Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 0001, Fanglei Sun, Jun Wang 0012, Yaodong Yang 0001
ICLR4
2022 Greedy when Sure and Conservative when Uncertain about the Opponents
abstract
We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response (conditional additionally on our opponent policy embedding) instead of a finite set of separate best responses against any opponent; 3) computes online a posterior of the current opponent policy embedding, without making the discrete and ineffective decision which type the current opponent belongs to; and 4) selects online between a real-time greedy policy and a fixed conservative policy via an adversarial bandit algorithm, gaining a theoretically better regret than adhering to either. Experimental studies on popular benchmarks demonstrate GSCU’s superiority over the state-of-the-art methods. The code is available online at \url{https://github.com/YeTianJHU/GSCU}.
Haobo Fu, Hongxiang Yu, Weiming Liu 0004, Jiechao Xiong, Ying Wen 0001, Kai Li 0022, Junliang Xing, Qiang Fu 0016, Wei Yang 0032
ICML7
2022 Multi-Agent Reinforcement Learning is a Sequence Modeling Problem
abstract
Large sequence models (SM) such as GPT series and BERT have displayed outstanding performance and generalization capabilities in natural language process, vision and recently reinforcement learning. A natural follow-up question is how to abstract multi-agent decision making also as an sequence modeling problem and benefit from the prosperous development of the SMs. In this paper, we introduce a novel architecture named Multi-Agent Transformer (MAT) that effectively casts cooperative multi-agent reinforcement learning (MARL) into SM problems wherein the objective is to map agents' observation sequences to agents' optimal action sequences. Our goal is to build the bridge between MARL and SMs so that the modeling power of modern sequence models can be unleashed for MARL. Central to our MAT is an encoder-decoder architecture which leverages the multi-agent advantage decomposition theorem to transform the joint policy search problem into a sequential decision making process; this renders only linear time complexity for multi-agent problems and, most importantly, endows MAT with monotonic performance improvement guarantee. Unlike prior arts such as Decision Transformer fit only pre-collected offline data, MAT is trained by online trial and error from the environment in an on-policy fashion. To validate MAT, we conduct extensive experiments on StarCraftII, Multi-Agent MuJoCo, Dexterous Hands Manipulation, and Google Research Football benchmarks. Results demonstrate that MAT achieves superior performance and data efficiency compared to strong baselines including MAPPO and HAPPO. Furthermore, we demonstrate that MAT is an excellent few-short learner on unseen tasks regardless of changes in the number of agents.See our project page at https://sites.google.com/view/multi-agent-transformer.
Muning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 0001, Ying Wen 0001, Jun Wang 0012, Yaodong Yang 0001
NeurIPS5
2022 Investigating the geometric structure of neural activation spaces with convex hull approximations
Yuting Jia, Shao Zhang, Haiwen Wang, Ying Wen 0001, Luoyi Fu, Huan Long, Xinbing Wang, Chenghu Zhou
Neurocomputing4
2022 Multi-Agent Feedback Enabled Neural Networks for Intelligent Communications
abstract
In the intelligent communication field, deep learning (DL) has attracted much attention due to its strong fitting ability and data-driven learning capability. Compared with the typical DL feedforward network structures, an enhancement structure with direct data feedback have been studied and proved to have better performance than the feedfoward networks. However, due to the above simple feedback methods lack sufficient analysis and learning ability on the feedback data, it is inadequate to deal with more complicated nonlinear systems and therefore the performance is limited for further improvement. In this paper, a novel multi-agent feedback enabled neural network (MAFENN) framework is proposed, consisting of three fully cooperative intelligent agents, which make the framework have stronger feedback learning capabilities and more intelligence on feature abstraction, denoising or generation, etc. Furthermore, the MAFENN frame work is theoretically formulated into a three-player Feedback Stackelberg game, and the game is proved to converge to the Feedback Stackelberg equilibrium. The design of MAFENN framework and algorithm are dedicated to enhance the learning capability of the feedfoward DL networks or their variations with the simple data feedback. To verify the MAFENN framework’s feasibility in wireless communications, a multi-agent MAFENN based equalizer (MAFENN-E) is developed for wireless fading channels with inter-symbol interference (ISI). Experimental results show that when the quadrature phase-shift keying (QPSK) modulation scheme is adopted, the SER performance of our proposed method outperforms that of the traditional equalizers by about 2 dB in linear channels. When in nonlinear channels, the SER performance of our proposed method outperforms that of either traditional or DL based equalizers more significantly, which shows the effectiveness and robustness of our proposal in the complex channel environment.
Fanglei Sun, Yang Li 0116, Ying Wen 0001, Jingchen Hu, Jun Wang 0012, Yang Yang 0001, Kai Li 0022
IEEE Trans. Wirel. Commun.3
2021 MAFENN: Multi-Agent Feedback Enabled Neural Network for Wireless Channel Equalization
abstract
Feedback mechanism has been widely used in wireless communication such as channel equalization and resource allocation. In recent years, deep learning (DL) has made great progress in the field of wireless communication. There is now some work that attempts to introduce plain feedback mechanisms into DL algorithm to solve wireless communication problems. However, the improvement of plain feedback DL methods is limited in complex situations due to those methods lack sufficient learning ability on feedback information. In this paper, we propose a Multi-Agent Feedback Enabled Neural Network (MAFENN) equalizer, which consists of a specific learnable feedback agent and two feed-forward agents. Three fully cooperative intelligent agents help the system improve the ability to remove wireless inter-symbol interference (ISI) in receiving ends. We further formulate it into a three-player Stackelberg Game, which helps us to optimize and train this model more efficiently. To verify the feasibility of our proposed MAFENN system and the Stackelberg Game optimization, we conduct a series of experiments to compare the symbol error rate (SER) performance of the MAFENN equalizer and the other methods which utilizes quadrature phase-shift keying (QPSK) modulation scheme. Our performance outperforms that of the other equalizers at different signal-to-noise ratio (SNR) settings for both linear and nonlinear channels.
Yang Li 0116, Fanglei Sun, Weiqin Zu, Wenbin Song, Ying Wen 0001, Jun Wang 0012, Yang Yang 0001, Kai Li 0022, Liantao Wu
GLOBECOM5
2021 Retrospective Thinking based Multi-Agent System for Wireless Video Transmissions
abstract
Benefiting from the breakthrough development of the fifth generation (5G), beyond 5G (B5G) wireless communication networks and Artificial Intelligence (AI) in recent years, the artificial intelligence of things (AIoT) is a new trend in the future. AIoT devices often have high-quality wireless video transmission requirements. However, the propagating signals at millimeter wave suffer from high propagation loss and sensitivity to blockage, resulting in the received video is vulnerable to be interfered. Due to the ability of Deep Learning (DL) to discover and learn good representations, some DL methods have achieved breakthrough performance in video recovery. However, most of these methods cannot exploit information from the higher to lower level to refine themselves. In this paper, we propose a novel retrospective thinking based multi-agent (ReTMA) system to solve the interference problem experienced on wireless channels. Compared with other plain feedback models, we add a retrospective agent on the feedback loop, which makes the entire system have stronger capabilities to learn good representative features. We further formulate it as a Stackelberg game to analyze the dependency relationship between the agents and facilitate the complex training issue of the multiple agents. To verify the feasibility of ReTMA system, we randomly add masks to simulate the severe interference received by the video frames in wireless transmissions. Experimental results show that the performances of similarity index measure (SSIM), peak signal-to-noise ratio (PSNR) and classification accuracy all achieve significant gains compared with those of other plain feedback models at different mask ratios.
Yang Li 0116, Fanglei Sun, Wenbin Song, Ying Wen 0001, Kai Li 0022, Jun Wang 0012, Yang Yang 0001
ICC4
2021 Learning in Nonzero-Sum Stochastic Games with Potentials
abstract
Multi-agent reinforcement learning (MARL) has become effective in tackling discrete cooperative game scenarios. However, MARL has yet to penetrate settings beyond those modelled by team and zero-sum games, confining it to a small subset of multi-agent systems. In this paper, we introduce a new generation of MARL learners that can handle \textit{nonzero-sum} payoff structures and continuous settings. In particular, we study the MARL problem in a class of games known as stochastic potential games (SPGs) with continuous state-action spaces. Unlike cooperative games, in which all agents share a common reward, SPGs are capable of modelling real-world scenarios where agents seek to fulfil their individual goals. We prove theoretically our learning method, $\ourmethod$, enables independent agents to learn Nash equilibrium strategies in \textit{polynomial time}. We demonstrate our framework tackles previously unsolvable tasks such as \textit{Coordination Navigation} and \textit{large selfish routing games} and that it outperforms the state of the art MARL baselines such as MADDPG and COMIX in such scenarios.
David Mguni, Yutong Wu 0005, Yali Du 0001, Yaodong Yang 0001, Ziyi Wang 0004, Minne Li, Ying Wen 0001, Joel Jennings, Jun Wang 0012
ICML7
2021 Modelling Behavioural Diversity for Learning in Open-Ended Games
abstract
Promoting behavioural diversity is critical for solving games with non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). Yet, there is a lack of rigorous treatment for defining diversity and constructing diversity-aware learning dynamics. In this work, we offer a geometric interpretation of behavioural diversity in games and introduce a novel diversity metric based on \emph{determinantal point processes} (DPP). By incorporating the diversity metric into best-response dynamics, we develop \emph{diverse fictitious play} and \emph{diverse policy-space response oracle} for solving normal-form games and open-ended games. We prove the uniqueness of the diverse best response and the convergence of our algorithms on two-player games. Importantly, we show that maximising the DPP-based diversity metric guarantees to enlarge the \emph{gamescape} – convex polytopes spanned by agents’ mixtures of strategies. To validate our diversity-aware solvers, we test on tens of games that show strong non-transitivity. Results suggest that our methods achieve at least the same, and in most games, lower exploitability than PSRO solvers by finding effective and diverse strategies.
Nicolas Perez Nieves, Yaodong Yang 0001, Oliver Slumbers, David Mguni, Ying Wen 0001, Jun Wang 0012
ICML5
2021 Neural Auto-Curricula in Two-Player Zero-Sum Games
abstract
When solving two-player zero-sum games, multi-agent reinforcement learning (MARL) algorithms often create populations of agents where, at each iteration, a new agent is discovered as the best response to a mixture over the opponent population. Within such a process, the update rules of "who to compete with" (i.e., the opponent mixture) and "how to beat them" (i.e., finding best responses) are underpinned by manually developed game theoretical principles such as fictitious play and Double Oracle. In this paper, we introduce a novel framework—Neural Auto-Curricula (NAC)—that leverages meta-gradient descent to automate the discovery of the learning update rule without explicit human design. Specifically, we parameterise the opponent selection module by neural networks and the best-response module by optimisation subroutines, and update their parameters solely via interaction with the game engine, where both players aim to minimise their exploitability. Surprisingly, even without human design, the discovered MARL algorithms achieve competitive or even better performance with the state-of-the-art population-based game solvers (e.g., PSRO) on Games of Skill, differentiable Lotto, non-transitive Mixture Games, Iterated Matching Pennies, and Kuhn Poker. Additionally, we show that NAC is able to generalise from small games to large games, for example training on Kuhn Poker and outperforming PSRO on Leduc Poker. Our work inspires a promising future direction to discover general MARL algorithms solely from data.
Xidong Feng, Oliver Slumbers, Ziyu Wan, Bo Liu 0039, Stephen McAleer, Ying Wen 0001, Jun Wang 0012, Yaodong Yang 0001
NeurIPS6
2021 Towards Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games
abstract
Measuring and promoting policy diversity is critical for solving games with strong non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). With that in mind, maintaining a pool of diverse policies via open-ended learning is an attractive solution, which can generate auto-curricula to avoid being exploited. However, in conventional open-ended learning algorithms, there are no widely accepted definitions for diversity, making it hard to construct and evaluate the diverse policies. In this work, we summarize previous concepts of diversity and work towards offering a unified measure of diversity in multi-agent open-ended learning to include all elements in Markov games, based on both Behavioral Diversity (BD) and Response Diversity (RD). At the trajectory distribution level, we re-define BD in the state-action space as the discrepancies of occupancy measures. For the reward dynamics, we propose RD to characterize diversity through the responses of policies when encountering different opponents. We also show that many current diversity measures fall in one of the categories of BD or RD but not both. With this unified diversity measure, we design the corresponding diversity-promoting objective and population effectivity when seeking the best responses in open-ended learning. We validate our methods in both relatively simple games like matrix game, non-transitive mixture model, and the complex \textit{Google Research Football} environment. The population found by our methods reveals the lowest exploitability, highest population effectivity in matrix game and non-transitive mixture model, as well as the largest goal difference when interacting with opponents of various levels in \textit{Google Research Football}.
Hangtian Jia, Ying Wen 0001, Yujing Hu, Changjie Fan, Zhipeng Hu, Yaodong Yang 0001
NeurIPS3
2020 Multi-Agent Determinantal Q-Learning
abstract
Centralized training with decentralized execution has become an important paradigm in multi-agent learning. Though practical, current methods rely on restrictive assumptions to decompose the centralized value function across agents for execution. In this paper, we eliminate this restriction by proposing multi-agent determinantal Q-learning. Our method is established on Q-DPP, a novel extension of determinantal point process (DPP) to multi-agent setting. Q-DPP promotes agents to acquire diverse behavioral models; this allows a natural factorization of the joint Q-functions with no need for \emph{a priori} structural constraints on the value function or special network architectures. We demonstrate that Q-DPP generalizes major solutions including VDN, QMIX, and QTRAN on decentralizable cooperative tasks. To efficiently draw samples from Q-DPP, we develop a linear-time sampler with theoretical approximation guarantee. Our sampler also benefits exploration by coordinating agents to cover orthogonal directions in the state space during training. We evaluate our algorithm on multiple cooperative benchmarks; its effectiveness has been demonstrated when compared with the state-of-the-art.
Yaodong Yang 0001, Ying Wen 0001, Jun Wang 0012, Kun Shao, David Mguni, Weinan Zhang 0001
ICML2
2020 Modelling Bounded Rationality in Multi-Agent Interactions by Generalized Recursive Reasoning
abstract
Though limited in real-world decision making, most multi-agent reinforcement learning (MARL) models assume perfectly rational agents -- a property hardly met due to individual's cognitive limitation and/or the tractability of the decision problem. In this paper, we introduce generalized recursive reasoning (GR2) as a novel framework to model agents with different \emph{hierarchical} levels of rationality; our framework enables agents to exhibit varying levels of ``thinking'' ability thereby allowing higher-level agents to best respond to various less sophisticated learners. We contribute both theoretically and empirically. On the theory side, we devise the hierarchical framework of GR2 through probabilistic graphical models and prove the existence of a perfect Bayesian equilibrium. Within the GR2, we propose a practical actor-critic solver, and demonstrate its convergent property to a stationary point in two-player games through Lyapunov analysis. On the empirical side, we validate our findings on a variety of MARL benchmarks. Precisely, we first illustrate the hierarchical thinking process on the Keynes Beauty Contest, and then demonstrate significant improvements compared to state-of-the-art opponent modeling baselines on the normal-form games and the cooperative navigation benchmark.
Ying Wen 0001, Yaodong Yang 0001, Jun Wang 0012
IJCAI1
2019 Factorized Q-learning for large-scale multi-agent systems
abstract
Deep Q-learning has achieved significant success in single-agent decision making tasks. However, it is challenging to extend Q-learning to large-scale multi-agent scenarios, due to the explosion of action space resulting from the complex dynamics between the environment and the agents. In this paper, we propose to make the computation of multi-agent Q-learning tractable by treating the Q-function (w.r.t. state and joint-action) as a high-order high-dimensional tensor and then approximate it with factorized pairwise interactions. Furthermore, we utilize a composite deep neural network architecture for computing the factorized Q-function, share the model parameters among all the agents within the same group, and estimate the agents' optimal joint actions through a coordinate descent type algorithm. All these simplifications greatly reduce the model complexity and accelerate the learning process. Extensive experiments on two different multi-agent problems demonstrate the performance gain of our proposed approach in comparison with strong baselines, particularly when there are a large number of agents.
Ming Zhou 0006, Yong Chen 0008, Ying Wen 0001, Yaodong Yang 0001, Yufeng Su, Weinan Zhang 0001, Dell Zhang, Jun Wang 0012
DAI3
2019 Probabilistic Recursive Reasoning for Multi-Agent Reinforcement Learning
Ying Wen 0001, Yaodong Yang 0001, Rui Luo 0001, Jun Wang 0012, Wei Pan 0004
ICLR (Poster)1
2019 A Regularized Opponent Model with Maximum Entropy Objective
abstract
In a single-agent setting, reinforcement learning (RL) tasks can be cast into an inference problem by introducing a binary random variable o, which stands for the "optimality". In this paper, we redefine the binary random variable o in multi-agent setting and formalize multi-agent reinforcement learning (MARL) as probabilistic inference. We derive a variational lower bound of the likelihood of achieving the optimality and name it as Regularized Opponent Model with Maximum Entropy Objective (ROMMEO). From ROMMEO, we present a novel perspective on opponent modeling and show how it can improve the performance of training agents theoretically and empirically in cooperative games. To optimize ROMMEO, we first introduce a tabular Q-iteration method ROMMEO-Q with proof of convergence. We extend the exact algorithm to complex environments by proposing an approximate version, ROMMEO-AC. We evaluate these two algorithms on the challenging iterated matrix game and differential game respectively and show that they can outperform strong MARL baselines.
Zheng Tian 0002, Ying Wen 0001, Zhichen Gong, Faiz Punakkath, Shihao Zou, Jun Wang 0012
IJCAI2
2016 Product-Based Neural Networks for User Response Prediction
abstract
Predicting user responses, such as clicks and conversions, is of great importance and has found its usage inmany Web applications including recommender systems, websearch and online advertising. The data in those applicationsis mostly categorical and contains multiple fields, a typicalrepresentation is to transform it into a high-dimensional sparsebinary feature representation via one-hot encoding. Facing withthe extreme sparsity, traditional models may limit their capacityof mining shallow patterns from the data, i.e. low-order featurecombinations. Deep models like deep neural networks, on theother hand, cannot be directly applied for the high-dimensionalinput because of the huge feature space. In this paper, we proposea Product-based Neural Networks (PNN) with an embeddinglayer to learn a distributed representation of the categorical data, a product layer to capture interactive patterns between interfieldcategories, and further fully connected layers to explorehigh-order feature interactions. Our experimental results on twolarge-scale real-world ad click datasets demonstrate that PNNsconsistently outperform the state-of-the-art models on various metrics.
Yanru Qu, Han Cai, Kan Ren, Weinan Zhang 0001, Yong Yu 0001, Ying Wen 0001, Jun Wang 0012
ICDM6