EDBT 2026 Demo / reviewers in the wild / expert
Wenhao Li 0001
dblp:11/444-1
· DBLP profile ↗
28ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0003-2985-1098ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On-policy Reinforcement Fine-tuning with Offline reward for Multi-step Embodied PlanningabstractEmbodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and verbal goals.While recent vision-language models (VLMs) excel at static perception tasks, they struggle in interactive environments.Reinforcement learning (RL) offers a natural way to address this limitation, yet online RL approaches suffer from costly interaction and sparse rewards in embodied settings.This paper introduces ORBIT, an On-policy Reinforcement finetuning (RFT) framework with offline rewards for EmBodIed Task Planning, that preserves the generalization benefits of RFT while addressing the challenges of costly interaction and sparse rewards, supported by solid theoretical guarantees.Our approach is evaluated on EmbodiedBench, a recent benchmark for interactive embodied tasks, covering both in-domain and out-of-domain scenarios.Experimental results show that ORBIT achieves SOTA performance on EB-ALFRED, outperforming all closed-source and online-RL-based methods, while being substantially more efficient in training speed and computational cost, remaining robust to sub-optimal expert trajectories, and exhibiting strong generalization to unseen environments.We released all code and data at https://github.com/mail-taii/Reinforced- Reasoning-for Chloe Gu, Guanbo Wang, Wenhao Li 0001, Bo Jin 0003 |
ACL (1) | 6 |
| 2026 | Learning Roles With Emergent Social Value OrientationsabstractSocial dilemmas can be considered situations where individual rationality leads to collective irrationality. The multi-agent reinforcement learning community has leveraged ideas from social science, such as social value orientations (SVO), to solve social dilemmas in complex cooperative tasks. In this paper, we first introduce the typical "division of labor or roles" mechanism in human society, and provide a promising solution for intertemporal social dilemmas (ISD) with SVOs. A novel learning framework, called Learning Roles with Emergent SVOs (RESVO), is proposed to transform the learning of roles into the social value orientation emergence, which is symmetrically solved by endowing agents with altruism to share rewards with other agents. An SVO-based role embedding space is then constructed by individual conditioning policies on roles with a novel rank regularizer and mutual information maximizer. Experiments show that RESVO achieves a stable division of labor and cooperation in ISDs with different complexity. Wenhao Li 0001, Xiangfeng Wang 0001, Bo Jin 0003, Hongyuan Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Multi-Agent Credit Assignment with Pretrained Language ModelsabstractThe difficulty of appropriately assigning credit is particularly heightened in cooperative MARL with sparse reward, due to the concurrent time and structural scales involved. Automatic subgoal generation (ASG) has recently emerged as a viable MARL approach inspired by utilizing subgoals in intrinsically motivated reinforcement learning. However, end-to-end learning of complex task planning from sparse rewards without prior knowledge, undoubtedly requires massive training samples. Moreover, the diversity-promoting nature of existing ASG methods can lead to the "over-representation" of subgoals, generating numerous spurious subgoals of limited relevance to the actual task reward and thus decreasing the sample efficiency of the algorithm. To address this problem and inspired by the disentangled representation learning, we propose a novel "disentangled" decision-making method, Semantically Aligned task decomposition in MARL (SAMA), that prompts pretrained language models with chain-of-thought that can suggest potential goals, provide suitable goal decomposition and subgoal allocation as well as self-reflection-based replanning. Additionally, SAMA incorporates language-grounded MARL to train each agent’s subgoal-conditioned policy. SAMA demonstrates considerable advantages in sample efficiency compared to state-of-the-art ASG methods, as evidenced by its performance on two challenging sparse-reward tasks, Overcooked and MiniRTS. The code is available at \url{https://anonymous.4open.science/r/SAMA/.} Wenhao Li 0001, Baoxiang Wang 0001, Xiangfeng Wang 0001, Hao Shen 0002, Bo Jin 0003, Hongyuan Zha |
AISTATS | 1 |
| 2025 | Reward Translation via Reward Machine in Semi-Alignable MDPsabstractAddressing reward design complexities in deep reinforcement learning is facilitated by knowledge transfer across different domains. To this end, we define reward translation to describe the cross-domain reward transfer problem. However, current methods struggle with non-pairable and non-time-alignable incompatible MDPs. This paper presents an adaptable reward translation framework neural reward translation featuring semi-alignable MDPs, which allows efficient reward translation under relaxed constraints while handling the intricacies of incompatible MDPs. Given the inherent difficulty of directly mapping semi-alignable MDPs and transferring rewards, we introduce an indirect mapping method through reward machines, created using limited human input or LLM-based automated learning. Graph-matching techniques establish links between reward machines from distinct environments, thus enabling cross-domain reward translation within semi-alignable MDP settings. This broadens the applicability of DRL across multiple domains. Experiments substantiate our approach’s effectiveness in tasks under environments with semi-alignable MDPs. Yun Hua, Wenhao Li 0001, Bo Jin 0003, Baoxiang Wang 0001, Hongyuan Zha, Xiangfeng Wang 0002 |
ICML | 3 |
| 2025 | Dynamic Conservative Degree Allocation for Offline Multi-Agent Reinforcement Learning
Yun Hua, Junjie Sheng, Wenhao Li 0001, Bo Jin 0003, Xiangfeng Wang 0001 |
AAMAS | 4 |
| 2025 | Negotiated Reasoning: On Provably Addressing Relative Over-Generalization
Junjie Sheng, Wenhao Li 0001, Bo Jin 0003, Hongyuan Zha, Jun Wang 0006, Xiangfeng Wang 0001 |
AAMAS | 2 |
| 2025 | SkyRover: A Modular Simulator for Cross-Domain PathfindingabstractUnmanned Aerial Vehicles (UAVs) and Automated Guided Vehicles (AGVs) increasingly collaborate in logistics, surveillance, inspection tasks and etc. However, existing simulators often focus on a single domain, limiting cross-domain study. This paper presents the SkyRover, a modular simulator for UAV-AGV multi-agent pathfinding (MAPF). SkyRover supports realistic agent dynamics, configurable 3D environments, and convenient APIs for external solvers and learning methods. By unifying ground and aerial operations, it facilitates cross-domain algorithm design, testing, and benchmarking. Experiments highlight SkyRover’s capacity for efficient pathfinding and high-fidelity simulations in UAV-AGV coordination. We believe the SkyRover fills a key gap in MAPF research. Project is available at https://sites.google.com/view/mapf3d/home. Wenhui Ma, Wenhao Li 0001, Bo Jin 0003, Changhong Lu, Xiangfeng Wang 0001 |
IJCAI | 2 |
| 2025 | Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM AgentsabstractLarge Language Models (LLMs) are increasingly deployed as autonomous agents in multi-agent systems, and promising coordination has been demonstrated in handling complex tasks under predefined roles and scripted workflows.
However, significant challenges remain in open-ended environments, where agents are inherently self-interested and explicit coordination guidelines are absent.
In such scenarios, misaligned incentives frequently lead to social dilemmas and inefficient collective outcomes.
Inspired by how human societies tackle similar coordination challenges—through temporary collaborations like employment or subcontracting—a cooperative workflow \textbf{Shapley-Coop} is proposed.
This workflow enables self-interested Large Language Model (LLM) agents to engage in emergent collaboration by using a fair credit allocation mechanism to ensure each agent’s contributions are appropriately recognized and rewarded.
Shapley-Coop introduces structured negotiation protocols and Shapley-inspired reasoning to estimate agents’ marginal contributions, thereby enabling effective task-time coordination and equitable post-task outcome redistribution.
This results in effective coordination that fosters collaboration while preserving agent autonomy, through a rational pricing mechanism that encourages cooperative behavior.
Evaluated in two multi-agent games and a software engineering simulation, Shapley-Coop consistently enhances LLM agent collaboration and facilitates equitable outcome redistribution, accurately reflecting individual contributions during the task execution process. Yun Hua, Shiqin Wang, Wenhao Li 0001, Xiangfeng Wang 0002 |
NeurIPS | 4 |
| 2025 | LOPT: Learning Optimal Pigovian Tax in Sequential Social DilemmasabstractMulti-agent reinforcement learning (MARL) has emerged as a powerful framework for modeling autonomous agents that independently optimize their individual objectives. However, in mixed-motive MARL environments, rational self-interested behaviors often lead to collectively suboptimal outcomes situations commonly referred to as social dilemmas.
A key challenge in addressing social dilemmas lies in accurately quantifying and representing them in a numerical form that captures how self-interested agent behaviors impact social welfare.
To address this challenge, \textit{externalities} in the economic concept is adopted and extended to denote the unaccounted-for impact of one agent's actions on others, as a means to rigorously quantify social dilemmas.
Based on this measurement, a novel method, \textbf{L}earning \textbf{O}ptimal \textbf{P}igovian \textbf{T}ax (\textbf{LOPT}) is proposed. Inspired by Pigovian taxes, which are designed to internalize externalities by imposing cost on negative societal impacts, LOPT employs an auxiliary tax agent that learns an optimal Pigovian tax policy to reshape individual rewards aligned with social welfare, thereby promoting agent coordination and mitigating social dilemmas. We support LOPT with theoretical analysis and validate it on standard MARL benchmarks, including Escape Room and Cleanup. Results show that by effectively internalizing externalities that quantify social dilemmas, LOPT aligns individual objectives with collective goals, significantly improving social welfare over state-of-the-art baselines. Yun Hua, Wenhao Li 0001, Bo Jin 0003, Xiangfeng Wang 0001, Hongyuan Zha |
NeurIPS | 3 |
| 2025 | Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache RetrievalabstractReducing the key-value (KV) cache burden in Large Language Models (LLMs) significantly accelerates inference. Dynamically selecting critical KV caches during decoding helps maintain performance. Existing methods use random linear hashing to identify important tokens, but this approach is inefficient due to the orthogonal distribution of queries and keys within two narrow cones in LLMs. We introduce Spotlight Attention, a novel method that employs non-linear hashing functions to optimize the embedding distribution of queries and keys, enhancing coding efficiency and robustness. We also developed a lightweight, stable training framework using a Bradley-Terry ranking-based loss, enabling optimization of the non-linear hashing module on GPUs with 16GB memory in 8 hours. Experimental results show that Spotlight Attention drastically improves retrieval precision while shortening the length of the hash code at least 5$\times$ compared to traditional linear hashing. Finally, we exploit the computational advantages of bitwise operations by implementing specialized CUDA kernels, achieving hashing retrieval for 512K tokens in under 100$\mu$s on a single A100 GPU, with end-to-end throughput up to 3$\times$ higher than vanilla decoding. Wenhao Li 0001, Yuxin Zhang 0002, Gen Luo, Haiyuan Wan, Ziyang Gong, Fei Chao 0001, Rongrong Ji |
NeurIPS | 1 |
| 2025 | Interpretable Hybrid-Rule Temporal Point Processes
Yunyang Cao, Juekai Lin, Hongye Wang, Wenhao Li 0001, Bo Jin 0003 |
ECML/PKDD (3) | 4 |
| 2024 | Efficient Planning with Latent DiffusionabstractTemporal abstraction and efficient planning pose significant challenges in offline reinforcement learning, mainly when dealing with domains that involve temporally extended tasks and delayed sparse rewards. Existing methods typically plan in the raw action space and can be inefficient and inflexible. Latent action spaces offer a more flexible approach, capturing only possible actions within the behavior policy support and decoupling the temporal structure between planning and modeling. However, current latent-action-based methods are limited to discrete spaces and require expensive planning steps. This paper presents a unified framework for continuous latent action space representation learning and planning by leveraging latent, score-based diffusion models. We establish the theoretical equivalence between planning in the latent action space and energy-guided sampling with a pretrained diffusion model and introduce a novel sequence-level exact sampling method. Our proposed method, $\texttt{LatentDiffuser}$, demonstrates competitive performance on low-dimensional locomotion control tasks and surpasses existing methods in higher-dimensional tasks. Wenhao Li 0001 |
ICLR | 1 |
| 2024 | Carbon Market Simulation with Adaptive Mechanism Design
Wenhao Li 0001, Hongyuan Zha, Baoxiang Wang 0001 |
IJCAI | 2 |
| 2024 | Optimizing Efficiency and Effectiveness in Sequential Prompt Strategy for SAM Using Reinforcement Learning
Chuyun Shen, Wenhao Li 0001, Xiangfeng Wang 0001, Bo Jin 0003, Haibin Cai |
MICCAI (8) | 3 |
| 2024 | Complementary information mutual learning for multimodality medical image segmentation
Chuyun Shen, Wenhao Li 0001, Haoqing Chen, Xiaoling Wang 0004, Fengping Zhu, Xiangfeng Wang 0001, Bo Jin 0003 |
Neural Networks | 2 |
| 2023 | Temporally-Extended Prompts Optimization for SAM in Interactive Medical Image SegmentationabstractThe Segmentation Anything Model (SAM) has recently emerged as a foundation model for addressing image segmentation. Owing to the intrinsic complexity of medical images and the high annotation cost, the medical image segmentation (MIS) community has been encouraged to investigate SAM’s zero-shot capabilities to facilitate automatic annotation. Inspired by the extraordinary accomplishments of the interactive medical image segmentation (IMIS) paradigm, this paper focuses on assessing the potential of SAM’s zero-shot capabilities within the IMIS paradigm to amplify its benefits in the MIS domain. Regrettably, we observe that SAM’s vulnerability to prompt forms (e.g., points, bounding boxes) becomes notably pronounced in IMIS. This leads us to develop a mechanism that adaptively offers suitable prompt forms for human experts. We refer to the mechanism above as temporally-extended prompts optimization (TEPO) and model it as a Markov decision process, solvable through reinforcement learning. Numerical experiments on the standardized benchmark Brats2020 demonstrate that the learned TEPO agent can further enhance SAM’s zero-shot capability in the MIS context. Chuyun Shen, Wenhao Li 0001, Ya Zhang 0002, Yanfeng Wang 0001, Xiangfeng Wang 0001 |
BIBM | 2 |
| 2023 | Hierarchical Diffusion for Offline Decision MakingabstractOffline reinforcement learning typically introduces a hierarchical structure to solve the long-horizon problem so as to address its thorny issue of variance accumulation. Problems of deadly triad, limited data and reward sparsity, however, still remain, rendering the design of effective, hierarchical offline RL algorithms for general-purpose policy learning a formidable challenge. In this paper, we first formulate the problem of offline long-horizon decision-$\mathbf{M}$ak$\mathbf{I}$ng from the perspective of conditional generative modeling by incorporating goals into the control-as-inference graphic models. A $\mathbf{H}$ierarchical trajectory-level $\mathbf{D}$iffusion probabilistic model is then proposed with classifier-free guidance. HDMI employs a cascade framework that utilizes the reward-conditional goal diffuser for the subgoal discovery and the goal-conditional trajectory diffuser for generating the corresponding action sequence of subgoals. Planning-based subgoal extraction and transformer-based diffusion are employed to deal with the sub-optimal data pollution and long-range subgoal dependencies in the goal diffusion. Numerical experiments verify the advantages of HDMI on long-horizon decision-making compared to SOTA offline RL methods and conditional generative models. Wenhao Li 0001, Xiangfeng Wang 0001, Bo Jin 0003, Hongyuan Zha |
ICML | 1 |
| 2023 | Information Design in Multi-Agent Reinforcement LearningabstractReinforcement learning (RL) is inspired by the way human infants and animals learn from the environment. The setting is somewhat idealized because, in actual tasks, other agents in the environment have their own goals and behave adaptively to the ego agent. To thrive in those environments, the agent needs to influence other agents so their actions become more helpful and less harmful. Research in computational economics distills two ways to influence others directly: by providing tangible goods (mechanism design) and by providing information (information design). This work investigates information design problems for a group of RL agents. The main challenges are two-fold. One is the information provided will immediately affect the transition of the agent trajectories, which introduces additional non-stationarity. The other is the information can be ignored, so the sender must provide information that the receiver is willing to respect. We formulate the Markov signaling game, and develop the notions of signaling gradient and the extended obedience constraints that address these challenges. Our algorithm is efficient on various mixed-motive tasks and provides further insights into computational economics. Our code is publicly available at https://github.com/YueLin301/InformationDesignMARL. Yue Lin 0006, Wenhao Li 0001, Hongyuan Zha, Baoxiang Wang 0001 |
NeurIPS | 2 |
| 2023 | F2A2: Flexible Fully-decentralized Approximate Actor-critic for Cooperative Multi-agent Reinforcement LearningabstractTraditional centralized multi-agent reinforcement learning (MARL) algorithms are sometimes unpractical in complicated applications due to non-interactivity between agents, the curse of dimensionality, and computation complexity. Hence, several decentralized MARL algorithms are motivated. However, existing decentralized methods only handle the fully cooperative setting where massive information needs to be transmitted in training. The block coordinate gradient descent scheme they used for successive independent actor and critic steps can simplify the calculation, but it causes serious bias. This paper proposes a flexible fully decentralized actor-critic MARL framework, which can combine most of the actor-critic methods and handle large-scale general cooperative multi-agent settings. A primal-dual hybrid gradient descent type algorithm framework is designed to learn individual agents separately for decentralization. From the perspective of each agent, policy improvement and value evaluation are jointly optimized, which can stabilize multi-agent policy learning. Furthermore, the proposed framework can achieve scalability and stability for the large-scale environment. This framework also reduces information transmission by the parameter sharing mechanism and novel modeling-other-agents methods based on theory-of-mind and online supervised learning. Sufficient experiments in cooperative Multi-agent Particle Environment and StarCraft II show that the proposed decentralized MARL instantiation algorithms perform competitively against conventional centralized and decentralized methods. Wenhao Li 0001, Bo Jin 0003, Xiangfeng Wang 0001, Junchi Yan, Hongyuan Zha |
J. Mach. Learn. Res. | 1 |
| 2023 | Interactive medical image segmentation with self-adaptive confidence calibrationabstractInteractive medical image segmentation based on human-in-the-loop machine learning is a novel paradigm that draws on human expert knowledge to assist medical image segmentation. However, existing methods often fall into what we call interactive misunderstanding, the essence of which is the dilemma in trading off short- and long-term interaction information. To better use the interaction information at various timescales, we propose an interactive segmentation framework, called interactive MEdical image segmentation with self-adaptive Confidence CAlibration (MECCA), which combines action-based confidence learning and multi-agent reinforcement learning. A novel confidence network is learned by predicting the alignment level of the action with short-term interaction information. A confidence-based reward-shaping mechanism is then proposed to explicitly incorporate confidence in the policy gradient calculation, thus directly correcting the model’s interactive misunderstanding. MECCA also enables user-friendly interactions by reducing the interaction intensity and difficulty via label generation and interaction guidance, respectively. Numerical experiments on different segmentation tasks show that MECCA can significantly improve short- and long-term interaction information utilization efficiency with remarkably fewer labeled samples. The demo video is available at https://bit.ly/mecca-demo-video . Chuyun Shen, Wenhao Li 0001, Qisen Xu, Bo Jin 0003, Haibin Cai, Fengping Zhu, Xiangfeng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | Dealing with Non-Stationarity in MARL via Trust-Region Decomposition
Wenhao Li 0001, Xiangfeng Wang 0001, Bo Jin 0003, Junjie Sheng, Hongyuan Zha |
ICLR | 1 |
| 2022 | Multi-Agent Path Finding with Prioritized Communication LearningabstractMulti-agent pathfinding (MAPF) has been widely used to solve large-scale real-world problems, e.g., automation warehouses. The learning-based, fully decentralized framework has been introduced to alleviate real-time problems and simultaneously pursue optimal planning policy. However, existing methods might generate significantly more vertex conflicts (or collisions), which lead to a low success rate or more makespan. In this paper, we propose a PrIoritized COmmunication learning method (PICO), which incorporates the implicit planning priorities into the communication topology within the decentralized multi-agent reinforcement learning framework. Assembling with the classic coupled planners, the implicit priority learning module can be utilized to form the dynamic communication topology, which also builds an effective collision-avoiding mechanism. PICO performs significantly better in large-scale MAPF tasks in success rates and collision rates than state-of-the-art learning-based planners. Wenhao Li 0001, Bo Jin 0003, Wenzhe Tan, Hongyuan Zha, Xiangfeng Wang 0001 |
ICRA | 1 |
| 2022 | VMAgent: A Practical Virtual Machine Scheduling PlatformabstractVirtual machine (VM) scheduling is one of the critical tasks in cloud computing. Many works have attempted to incorporate machine learning, especially reinforcement learning, to empower VM scheduling procedures. Although improved results are shown in several demo simulators, the performances in real-world scenarios are still underexploited. In this paper, we design a practical VM scheduling platform, i.e., VMAgent, to assist researchers in developing their methods on the VM scheduling problem. VMAgent consists of three components: simulator, scheduler, and visualizer. The simulator abstracts three general realistic scheduling scenarios (fading, recovering, and expansion) based on Huawei Cloud’s scheduling data, which is the core of our platform. Flexible configurations are further provided to make the simulator compatible with practical cloud computing architecture (i.e., Multi Non-Uniform Memory Access) and scenarios. Researchers then need to instantiate the scheduler to interact with the simulator, which is also pre-built in various types (e.g., heuristic, machine learning, and operations research) of scheduling algorithms to speed up the algorithm design. The visualizer, as an auxiliary component of the simulator and scheduler, facilitates researchers to conduct an in-depth analysis of the scheduling procedure and comprehensively compare different scheduling algorithms. We believe that VMAgent would shed light on the AI for the VM scheduling community, and the demo video is presented in https://bit.ly/vmagent-demo-video. Junjie Sheng, Shengliang Cai, Haochuan Cui, Wenhao Li 0001, Yun Hua, Bo Jin 0003, Yiqiu Hu, Hongyuan Zha, Xiangfeng Wang 0001 |
IJCAI | 4 |
| 2022 | Learning structured communication for multi-agent reinforcement learning
Junjie Sheng, Xiangfeng Wang 0001, Bo Jin 0003, Junchi Yan, Wenhao Li 0001, Tsung-Hui Chang, Jun Wang 0006, Hongyuan Zha |
Auton. Agents Multi Agent Syst. | 5 |
| 2022 | Structured Cooperative Reinforcement Learning With Time-Varying Composite Action SpaceabstractIn recent years, reinforcement learning has achieved excellent results in low-dimensional static action spaces such as games and simple robotics. However, the action space is usually composite, composed of multiple sub-action with different functions, and time-varying for practical tasks. The existing sub-actions might be temporarily invalid due to the external environment, while unseen sub-actions can be added to the current system. To solve the robustness and transferability problems in time-varying composite action spaces, we propose a structured cooperative reinforcement learning algorithm based on the centralized critic and decentralized actor framework, called SCORE. We model the single-agent problem with composite action space as a fully cooperative partially observable stochastic game and further employ a graph attention network to capture the dependencies between heterogeneous sub-actions. To promote tighter cooperation between the decomposed heterogeneous agents, SCORE introduces a hierarchical variational autoencoder, which maps the heterogeneous sub-action space into a common latent action space. We also incorporate an implicit credit assignment structure into the SCORE to overcome the multi-agent credit assignment problem in the fully cooperative partially observable stochastic game. Performance experiments on the proof-of-concept task and precision agriculture task show that SCORE has significant advantages in robustness and transferability for time-varying composite action space. Wenhao Li 0001, Xiangfeng Wang 0001, Bo Jin 0003, Dijun Luo, Hongyuan Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemabstractIn spite of the success of existing meta reinforcement learning methods, they still have difficulty in learning a meta policy effectively for RL problems with sparse reward. In this respect, we develop a novel meta reinforcement learning framework called Hyper-Meta RL(HMRL), for sparse reward RL problems. It is consisted with three modules including the cross-environment meta state embedding module which constructs a common meta state space to adapt to different environments; the meta state based environment-specific meta reward shaping which effectively extends the original sparse reward trajectory by cross-environmental knowledge complementarity and as a consequence the meta policy achieves better generalization and efficiency with the shaped meta reward. Experiments with sparse-reward environments show the superiority of HMRL on both transferability and policy learning efficiency. Yun Hua, Xiangfeng Wang 0001, Bo Jin 0003, Wenhao Li 0001, Junchi Yan, Hongyuan Zha |
KDD | 4 |
| 2021 | Distributed and Parallel ADMM for Structured Nonconvex Optimization ProblemabstractThe nonconvex optimization problems have recently attracted significant attention. However, both efficient algorithm and solid theory are still very limited. The difficulty is even pronounced for structured large-scale problems in many real-world applications. This article proposes an application-driven algorithmic framework for structured nonconvex optimization problems with distributed and parallel techniques, which jointly handles the high dimensionality of model parameters and distributed training data. The theoretical convergence of our algorithm is established under moderate assumptions. We apply the proposed method to popular multitask applications, including a multitask reinforcement learning problem. The promising performance demonstrates our framework is effective and efficient. Xiangfeng Wang 0001, Junchi Yan, Bo Jin 0003, Wenhao Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2020 | Iteratively-Refined Interactive 3D Medical Image Segmentation With Multi-Agent Reinforcement LearningabstractExisting automatic 3D image segmentation methods usually fail to meet the clinic use. Many studies have explored an interactive strategy to improve the image segmentation performance by iteratively incorporating user hints. However, the dynamic process for successive interactions is largely ignored. We here propose to model the dynamic process of iterative interactive image segmentation as a Markov decision process (MDP) and solve it with reinforcement learning (RL). Unfortunately, it is intractable to use single-agent RL for voxel-wise prediction due to the large exploration space. To reduce the exploration space to a tractable size, we treat each voxel as an agent with a shared voxel-level behavior strategy so that it can be solved with multi-agent reinforcement learning. An additional advantage of this multi-agent model is to capture the dependency among voxels for segmentation task. Meanwhile, to enrich the information of previous segmentations, we reserve the prediction uncertainty in the state space of MDP and derive an adjustment action space leading to a more precise and finer segmentation. In addition, to improve the efficiency of exploration, we design a relative cross-entropy gain-based reward to update the policy in a constrained direction. Experimental results on various medical datasets have shown that our method significantly outperforms existing state-of-the-art methods, with the advantage of less interactions and a faster convergence. Xuan Liao, Wenhao Li 0001, Qisen Xu, Xiangfeng Wang 0001, Bo Jin 0003, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002 |
CVPR | 2 |