EDBT 2026 Demo / reviewers in the wild / expert
Muning Wen
dblp:295/0261
· DBLP profile ↗
17ranked-venue papers
3as first author
17since 2021 · last 2025
0009-0000-7868-1262ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Autonomous Goal Detection and Cessation in Reinforcement Learning: A Case Study on Source Term EstimationabstractReinforcement Learning has revolutionized decision-making processes in dynamic environments, yet it often struggles with autonomously detecting and achieving goals without clear feedback signals. For example, in a Source Term Estimation problem, the lack of precise environmental information makes it challenging to provide clear feedback signals and to define and evaluate how the source's location is determined. To address this challenge, the Autonomous Goal Detection and Cessation (AGDC) module was developed, enhancing various RL algorithms by incorporating a self-feedback mechanism for autonomous goal detection and cessation upon task completion. Our method effectively identifies and ceases undefined goals by approximating the agent's belief, significantly enhancing the capabilities of RL algorithms in environments with limited feedback. To validate effectiveness of our approach, we integrated AGDC with deep Q-Network, proximal policy optimization, and deep deterministic policy gradient algorithms, and evaluated its performance on the Source Term Estimation problem. The experimental results showed that AGDC-enhanced RL algorithms significantly outperformed traditional statistical methods such as infotaxis, entrotaxis, and dual control for exploitation and exploration, as well as a non-statistical random action selection method. These improvements were evident in terms of success rate, mean traveled distance, and search time, highlighting AGDC's effectiveness and efficiency in complex, real-world scenarios. Yiwei Shi, Muning Wen, Weinan Zhang 0001, Cunjia Liu, Weiru Liu |
AAAI | 2 |
| 2025 | Diversity-Aware Self-Paced Data Selection for LLM Fine-TuningabstractFine-tuning large language models (LLMs) is challenged by the presence of noisy data and the high computational cost when training on large-scale datasets. While data selection has emerged as a promising approach to reduce training cost and improve data quality, existing methods often rely on static heuristics or manual metrics. These approaches struggle to adapt to the model’s evolving capabilities during training, as its understanding of tasks improves. As the model becomes more powerful, its requirements for data that can enhance performance also change, making it crucial to incorporate this dynamic into the data selection process. Moreover, ensuring data diversity throughout different stages of training is essential for preventing redundancy, reducing overfitting. To address these issues, we propose DSP, a Diversity-Aware Self-Paced data selection framework that evolves with the model. DSP progressively selects training samples based on the model’s own outputs and incorporates a diversity-aware mechanism to enhance generalization and mitigate overfitting. Unlike prior static or rule-based strategies, DSP adaptively adjusts to the model’s internal feedback and training stage. Experiments on two public benchmarks demonstrate that DSP consistently outperforms static and heuristic-based baselines across multiple datasets and backbone models. Our findings highlight the critical role of dynamic, diversity-aware data selection in effective LLM fine-tuning. Yingxuan Yang, Muning Wen, Xiaoyun Mo, Qiuying Peng, Jun Wang 0020, Weinan Zhang 0001 |
ECAI | 3 |
| 2025 | Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement LearningabstractDriven by inherent uncertainty and the sim-to-real gap, robust reinforcement learning (RL) seeks to improve resilience against the complexity and variability in agent-environment sequential interactions. Despite the existence of a large number of RL benchmarks, there is a lack of standardized benchmarks for robust RL. Current robust RL policies often focus on a specific type of uncertainty and are evaluated in distinct, one-off environments. In this work, we introduce Robust-Gymnasium, a unified modular benchmark designed for robust RL that supports a wide variety of disruptions across all key RL components—agents' observed state and reward, agents' actions, and the environment. Offering over sixty diverse task environments spanning control and robotics, safe RL, and multi-agent RL, it provides an open-source and user-friendly tool for the community to assess current methods and foster the development of robust RL algorithms.
In addition, we benchmark existing standard and robust RL algorithms within this framework, uncovering significant deficiencies in each and offering new insights. Shangding Gu, Laixi Shi, Muning Wen, Ming Jin 0002, Eric Mazumdar, Yuejie Chi, Adam Wierman, Costas J. Spanos |
ICLR | 3 |
| 2025 | Robust Function-Calling for On-Device Language Model via Function MaskingabstractLarge language models have demonstrated impressive value in performing as autonomous agents when equipped with external tools and API calls. Nonetheless, effectively harnessing their potential for executing complex tasks crucially relies on enhancements in their function-calling capabilities. This paper identifies a critical gap in existing function-calling models, where performance varies significantly across benchmarks, often due to over-fitting to specific naming conventions. To address such an issue, we introduce Hammer, a novel family of foundation models specifically engineered for on-device function calling. Hammer employs an augmented dataset that enhances models’ sensitivity to irrelevant functions and incorporates function masking techniques to minimize over-fitting. Our empirical evaluations reveal that Hammer not only outperforms larger models but also demonstrates robust generalization across diverse benchmarks, achieving state-of-the-art results. Our open-source contributions include a specialized dataset for irrelevance detection, a tuning framework for enhanced generalization, and the Hammer models, establishing a new standard for function-calling performance. Qiqiang Lin, Muning Wen, Qiuying Peng, Guanyu Nie, Junwei Liao, Xiaoyun Mo, Jiamu Zhou, Yin Zhao, Jun Wang 0012, Weinan Zhang 0001 |
ICLR | 2 |
| 2025 | PMAT: Optimizing Action Generation Order in Multi-Agent Reinforcement Learning
Muning Wen, Xihuai Wang, Shao Zhang, Yiwei Shi, Minne Li, Minglong Li, Ying Wen 0001 |
AAMAS | 2 |
| 2025 | MobileUse: A Hierarchical Reflection-Driven GUI Agent for Autonomous Mobile OperationabstractRecent advances in Multimodal Large Language Models (MLLMs) have enabled the development of mobile agents that can understand visual inputs and follow user instructions, unlocking new possibilities for automating complex tasks on mobile devices. However, applying these models to real-world mobile scenarios remains a significant challenge due to the long-horizon task execution, difficulty in error recovery, and the cold-start problem in unfamiliar environments. To address these challenges, we propose MobileUse, a GUI agent designed for robust and adaptive mobile task execution. To improve resilience in long-horizon tasks and dynamic environments, we introduce a hierarchical reflection architecture that enables the agent to self-monitor, detect, and recover from errors across multiple temporal scales—ranging from individual actions to overall task completion—while maintaining efficiency through a Reflection-on-Demand strategy. To tackle cold-start issues, we further introduce a proactive exploration module, which enriches the agent’s understanding of the environment through self-planned exploration. Evaluations on the AndroidWorld and AndroidLab benchmarks demonstrate that MobileUse establishes new state-of-the-art performance, achieving success rates of 62.9% and 44.2%, respectively. To facilitate real-world applications, we release an out-of-the-box toolkit for automated task execution on physical mobile devices, which is available at https://github.com/MadeAgents/mobile-use. Ning Li 0029, Xiangmou Qu, Jiamu Zhou, Muning Wen, Kounianhua Du, Xingyu Lou, Qiuying Peng, Jun Wang 0012, Weinan Zhang 0001 |
NeurIPS | 4 |
| 2025 | RDHNet: addressing rotational and permutational symmetries in continuous multi-agent systems
Dongzi Wang 0002, Lilan Huang, Muning Wen, Yuanxi Peng, Minglong Li, Teng Li 0011 |
Frontiers Comput. Sci. | 3 |
| 2024 | AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingabstractRecent works like Tree-of-Thought (ToT) and Reasoning via Planning (RAP) aim to augment the multi-step reasoning capabilities of LLMs by using tree-search algorithms. These methods rely on prompting a pre-trained model to serve as a value function and focus on problems with low search depth. As a result, these methods cannot benefit from in-domain training and only rely on pretraining process — they will not work in domains where the pre-trained LLM does not have enough knowledge to serve as an effective value function or in domains that require long-horizon planning. To address these limitations, we present an AlphaZero-like tree-search learning framework for LLMs (termed TS-LLM), systematically illustrating how tree-search with a learned value function can guide LLM decoding. TS-LLM distinguishes itself in two key ways. (1) Leveraging a learned value function and AlphaZero-like algorithms, our approach can be generally adaptable to a wide range of tasks, language models of any size, and tasks of varying search depths. (2) Our approach can guide LLMs during both inference and training, iteratively improving the LLMs. Empirical results across reasoning, planning, alignment, and decision-making tasks show that TS-LLM outperforms existing approaches and can handle trees with a depth of 64. Ziyu Wan, Xidong Feng, Muning Wen, Stephen McAleer, Ying Wen 0001, Weinan Zhang 0001, Jun Wang 0012 |
ICML | 3 |
| 2024 | Reinforcing LLM Agents via Policy Optimization with Action DecompositionabstractLanguage models as intelligent agents push the boundaries of sequential decision-making agents but struggle with limited knowledge of environmental dynamics and exponentially huge action space. Recent efforts like GLAM and TWOSOME manually constrain the action space to a restricted subset and employ reinforcement learning to align agents' knowledge with specific environments. However, they overlook fine-grained credit assignments for intra-action tokens, which is essential for efficient language agent optimization, and rely on human's prior knowledge to restrict action space. This paper proposes decomposing language agent optimization from the action level to the token level, offering finer supervision for each intra-action token and manageable optimization complexity in environments with unrestricted action spaces. Beginning with the simplification of flattening all actions, we theoretically explore the discrepancies between action-level optimization and this naive token-level optimization. We then derive the Bellman backup with Action Decomposition (BAD) to integrate credit assignments for both intra-action and inter-action tokens, effectively eliminating the discrepancies. Implementing BAD within the PPO algorithm, we introduce Policy Optimization with Action Decomposition (POAD). POAD benefits from a finer-grained credit assignment process and lower optimization complexity, leading to enhanced learning efficiency and generalization abilities in aligning language agents with interactive environments. We validate POAD across diverse testbeds, with results affirming the advantages of our approach and the correctness of our theoretical analysis. The source code can be accessed directly with this link: https://github.com/morning9393/ADRL. Muning Wen, Ziyu Wan, Jun Wang 0012, Weinan Zhang 0001, Ying Wen 0001 |
NeurIPS | 1 |
| 2024 | TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned DecisionabstractSeveral large language model (LLM) agents have been constructed for diverse purposes such as web navigation and online shopping, leveraging the broad knowledge and text comprehension capabilities of LLMs. Many of these works rely on in-context examples to achieve generalization without requiring fine-tuning. However, few have addressed the challenge of selecting and effectively utilizing these examples. Recent approaches have introduced trajectory-level retrieval with task meta-data and the use of trajectories as in-context examples to enhance overall performance in some sequential decision making tasks like computer control. Nevertheless, these methods face issues like plausible examples retrieved without task-specific state transition dynamics and long input with plenty of irrelevant context due to using complete trajectories. In this paper, we propose a novel framework (TRAD) to tackle these problems. TRAD first employs Thought Retrieval for step-level demonstration selection through thought matching, enhancing the quality of demonstrations and reducing irrelevant input noise. Then, Aligned Decision is introduced to complement retrieved demonstration steps with their preceding or subsequent steps, providing tolerance for imperfect thought and offering a balance between more context and less noise. Extensive experiments on ALFWorld and Mind2Web benchmarks demonstrate that TRAD not only surpasses state-of-the-art models but also effectively reduces noise and promotes generalization. Furthermore, TRAD has been deployed in real-world scenarios of a global business insurance company and yields an improved success rate of robotic process automation. Our codes are available at: https://github.com/skyriver-2000/TRAD-Official. Ruiwen Zhou, Yingxuan Yang, Muning Wen, Ying Wen 0001, Chunling Xi, Yong Yu 0001, Weinan Zhang 0001 |
SIGIR | 3 |
| 2024 | RoMAT: Role-based multi-agent transformer for generalizable heterogeneous cooperation
Dongzi Wang 0002, Fangwei Zhong, Minglong Li, Muning Wen, Yuanxi Peng, Teng Li 0011, Yaodong Yang 0001 |
Neural Networks | 4 |
| 2024 | Safe Multiagent Learning With Soft Constrained Policy Optimization in Real Robot ControlabstractDue to a lack of safety considerations, a wide range of multiagent reinforcement learning (MARL) applications are limited in real-world environments. Thus, ensuring MARL safety is essential and urgent in the domain. However, merely a few studies consider the safe MARL problem, and the investigation of real-world applications using safe MARL algorithms still needs to be improved. To fill this gap, we provide a framework with soft constrained policy optimization, in which we develop practical algorithms to address the problem in a cooperative game setting. First, the problem formulation of safe MARL is introduced. Second, the safe policy optimization of safe MARL algorithms based on soft constrained optimization is analyzed, and we further propose a safe learning framework for safe MARL. The framework can be plugged into MARL algorithms without manually fine-tuning safety bounds. Third, we investigate the sim-to-real problems, and conduct simulation and real-world experiments to evaluate the effectiveness of our algorithms. Finally, the comprehensive experimental results indicate that our method has significant benefits regarding the balance between reward and safety performance and outperforms several strong baselines. Shangding Gu, Dianye Huang, Muning Wen, Guang Chen 0001, Alois C. Knoll |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Large sequence models for sequential decision-making: a survey
Muning Wen, Runji Lin, Hanjing Wang, Yaodong Yang 0001, Ying Wen 0001, Luo Mai, Jun Wang 0012, Haifeng Zhang 0002, Weinan Zhang 0001 |
Frontiers Comput. Sci. | 1 |
| 2023 | MALib: A Parallel Framework for Population-based Multi-agent Reinforcement LearningabstractPopulation-based multi-agent reinforcement learning (PB-MARL) encompasses a range of methods that merge dynamic population selection with multi-agent reinforcement learning algorithms (MARL). While PB-MARL has demonstrated notable achievements in complex multi-agent tasks, its sequential execution is plagued by low computational efficiency due to the diversity in computing patterns and policy combinations. We propose a solution involving a stateless central task dispatcher and stateful workers to handle PB-MARL's subroutines, thereby capitalizing on parallelism across various components for efficient problem-solving. In line with this approach, we introduce MALib, a parallel framework that incorporates a task control model, independent data servers, and an abstraction of MARL training paradigms. The framework has undergone extensive testing and is available under the MIT license (https://github.com/sjtu-marl/malib) Ming Zhou 0006, Ziyu Wan, Hanjing Wang, Muning Wen, Runzhe Wu, Ying Wen 0001, Yaodong Yang 0001, Yong Yu 0001, Jun Wang 0012, Weinan Zhang 0001 |
J. Mach. Learn. Res. | 4 |
| 2022 | Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 0001, Fanglei Sun, Jun Wang 0012, Yaodong Yang 0001 |
ICLR | 3 |
| 2022 | Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemabstractLarge sequence models (SM) such as GPT series and BERT have displayed outstanding performance and generalization capabilities in natural language process, vision and recently reinforcement learning. A natural follow-up question is how to abstract multi-agent decision making also as an sequence modeling problem and benefit from the prosperous development of the SMs. In this paper, we introduce a novel architecture named Multi-Agent Transformer (MAT) that effectively casts cooperative multi-agent reinforcement learning (MARL) into SM problems wherein the objective is to map agents' observation sequences to agents' optimal action sequences. Our goal is to build the bridge between MARL and SMs so that the modeling power of modern sequence models can be unleashed for MARL. Central to our MAT is an encoder-decoder architecture which leverages the multi-agent advantage decomposition theorem to transform the joint policy search problem into a sequential decision making process; this renders only linear time complexity for multi-agent problems and, most importantly, endows MAT with monotonic performance improvement guarantee. Unlike prior arts such as Decision Transformer fit only pre-collected offline data, MAT is trained by online trial and error from the environment in an on-policy fashion. To validate MAT, we conduct extensive experiments on StarCraftII, Multi-Agent MuJoCo, Dexterous Hands Manipulation, and Google Research Football benchmarks. Results demonstrate that MAT achieves superior performance and data efficiency compared to strong baselines including MAPPO and HAPPO. Furthermore, we demonstrate that MAT is an excellent few-short learner on unseen tasks regardless of changes in the number of agents.See our project page at https://sites.google.com/view/multi-agent-transformer. Muning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang 0001, Ying Wen 0001, Jun Wang 0012, Yaodong Yang 0001 |
NeurIPS | 1 |
| 2021 | Settling the Variance of Multi-Agent Policy GradientsabstractPolicy gradient (PG) methods are popular reinforcement learning (RL) methods where a baseline is often applied to reduce the variance of gradient estimates. In multi-agent RL (MARL), although the PG theorem can be naturally extended, the effectiveness of multi-agent PG (MAPG) methods degrades as the variance of gradient estimates increases rapidly with the number of agents. In this paper, we offer a rigorous analysis of MAPG methods by, firstly, quantifying the contributions of the number of agents and agents' explorations to the variance of MAPG estimators. Based on this analysis, we derive the optimal baseline (OB) that achieves the minimal variance. In comparison to the OB, we measure the excess variance of existing MARL algorithms such as vanilla MAPG and COMA. Considering using deep neural networks, we also propose a surrogate version of OB, which can be seamlessly plugged into any existing PG methods in MARL. On benchmarks of Multi-Agent MuJoCo and StarCraft challenges, our OB technique effectively stabilises training and improves the performance of multi-agent PPO and COMA algorithms by a significant margin. Code is released at \url{https://github.com/morning9393/Optimal-Baseline-for-Multi-agent-Policy-Gradients}. Jakub Grudzien Kuba, Muning Wen, Linghui Meng 0001, Shangding Gu, Haifeng Zhang 0002, David Mguni, Jun Wang 0012, Yaodong Yang 0001 |
NeurIPS | 2 |