Minglong Li

dblp:184/0900 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CABTO: Context-Aware Behavior Tree Grounding for Robot Manipulation
Yishuai Cai, Xinglin Chen, Yunxin Mao, Minglong Li
AAAI5
2026 FD-TE Diagnosis: Enhancing Microservice Fault Diagnosis With Frequency Domain Features and Centrality-Aware Time Encoding
abstract
Fault diagnosis in microservice systems requires high availability, driving research towards multimodal learning that leverages heterogeneous monitoring data, including logs, metrics, and traces. These data inherently contain time series and data streams across various modalities. However, existing frameworks fail to exploit temporal dependencies in these dynamic streams effectively. First, existing methods rely too much on original time-domain metrics, which makes it difficult for them to capture periodic patterns and sudden events. Second, most language-bound integrations treat timestamps only as sequential labels or numerical inputs, which ignores the rich contextual information within timestamps. As a result, existing models struggle to comprehend the temporal sequences and data characteristics associated with faults in multimodal data. In this work, we introduce FD-TE Diagnosis, a framework for microservice fault diagnosis that addresses these challenges by applying frequency domain feature analysis and time encoding. To enhance the detection of abnormal events, we integrate a frequency-domain approach that utilizes the Fast Fourier Transform (FFT) to extract robust features from metric data. Next, we encode the timestamps of all events using a dedicated time encoding layer. These temporal representations are then incorporated to strengthen event embedding for fault inference. Experimental results demonstrate that our method enhances the sensitivity of multimodal models in interpreting temporal features, resulting in improved diagnostic outcomes.
Yangfan Li 0001, Haotian Wang 0001, Minglong Li, Fengxiao Tang, Wenjing Yang 0002
IEEE Trans. Reliab.4
2025 MRBTP: Efficient Multi-Robot Behavior Tree Planning and Collaboration
abstract
Multi-robot task planning and collaboration are critical challenges in robotics. While Behavior Trees (BTs) have been established as a popular control architecture and are plannable for a single robot, the development of effective multi-robot BT planning algorithms remains challenging due to the complexity of coordinating diverse action spaces. We propose the Multi-Robot Behavior Tree Planning (MRBTP) algorithm, with theoretical guarantees of both soundness and completeness. MRBTP features cross-tree expansion to coordinate heterogeneous actions across different BTs to achieve the team's goal. For homogeneous actions, we retain backup structures among BTs to ensure robustness and prevent redundant execution through intention sharing. While MRBTP is capable of generating BTs for both homogeneous and heterogeneous robot teams, its efficiency can be further improved. We then propose an optional plugin for MRBTP when Large Language Models (LLMs) are available to reason goal-related actions for each robot. These relevant actions can be pre-planned to form long-horizon subtrees, significantly enhancing the planning speed and collaboration efficiency of MRBTP. We evaluate our algorithm in warehouse management and everyday service scenarios. Results demonstrate MRBTP's robustness and execution efficiency under varying settings, as well as the ability of the pre-trained LLM to generate effective task-specific subtrees for MRBTP.
Yishuai Cai, Xinglin Chen, Zhongxuan Cai, Yunxin Mao, Minglong Li, Wenjing Yang 0002, Ji Wang 0001
AAAI5
2025 Automated Exposure Mapping for Networked Interference
abstract
By characterizing interactions and influences across individuals, networked interference aims to estimate cross-individual treatment effects. For each individual, one of the central components of existing approaches is to manually design an exposure mapping from their neighboring covariates (including their own ones) to different exposure conditions. However, handcraft neighboring structures defined by such manual schemes struggle to capture the complex and flexible structures exhibited by real-world social networks. To bridge this gap, we propose an Automated Exposure Mapping Network (AEMNet) by capturing networked interference conditions automatically with Graph Neural Networks (GNNs) and achieving mapping with deep embedded clustering. The learned representations between individuals in the graph structure reveal patterns and structures hidden behind data, facilitating application on large-scale, relationally complex networked data. We conducted extensive experiments demonstrating that our approach outperforms the baselines in both quality and flexibility, underscoring its ability to better characterize the interference relationships.
Yunxin Mao, Haotian Wang 0001, Yishuai Cai, Minglong Li, Ji Wang 0001, Wenjing Yang 0002
ICASSP4
2025 HBTP: Heuristic Behavior Tree Planning with Large Language Model Reasoning
abstract
Behavior Trees (BTs) are increasingly becoming a popular control structure in robotics due to their modularity, reactivity, and robustness. In terms of BT generation methods, BT planning shows promise for generating reliable BTs. However, the scalability of BT planning is often constrained by prolonged planning times in complex scenarios, largely due to a lack of domain knowledge. In contrast, pre-trained Large Language Models (LLMs) have demonstrated task reasoning capabilities across various domains, though the correctness and safety of their planning remain uncertain. This paper proposes integrating BT planning with LLM reasoning, introducing Heuristic Behavior Tree Planning (HBTP)-a reliable and efficient framework for BT generation. The key idea in HBTP is to leverage LLMs for task-specific reasoning to generate a heuristic path, which BT planning can then follow to expand efficiently. We first introduce the heuristic BT expansion process, along with two heuristic variants designed for optimal planning and satisficing planning, respectively. Then, we propose methods to address the inaccuracies of LLM reasoning, including action space pruning and reflective feedback, to further enhance both reasoning accuracy and planning efficiency. Experiments demonstrate the theoretical bounds of HBTP, and results from four datasets confirm its practical effectiveness in everyday service robot applications.
Yishuai Cai, Xinglin Chen, Yunxin Mao, Minglong Li, Shaowu Yang, Wenjing Yang 0002, Ji Wang 0001
ICRA4
2025 FutureNet-LoF: Joint Trajectory Prediction and Lane Occupancy Field Prediction with Future Context Encoding
abstract
Most prior motion prediction endeavors in autonomous driving have inadequately encoded future scenarios, leading to predictions that may fail to accurately capture the diverse movements of agents (e.g., vehicles or pedestrians). To address this, we propose FutureNet, which explicitly integrates initially predicted trajectories into the future scenario and further encodes these future contexts to enhance subsequent forecasting. Additionally, most previous motion forecasting works have focused on predicting independent futures for each agent. However, safe and smooth autonomous driving requires accurately predicting the diverse future behaviors of numerous surrounding agents jointly in complex dynamic environments. Given that all agents occupy certain potential travel spaces and possess lane driving priority, we propose Lane Occupancy Field (LOF), a new representation with lane semantics for motion forecasting in autonomous driving. LOF can simultaneously capture the joint probability distribution of all road participants' future spatial-temporal positions. Due to the high compatibility between lane occupancy field prediction and trajectory prediction, we propose a novel network for joint prediction of these two tasks. Our approach ranks 1st on two large-scale motion forecasting benchmarks: Argoverse 1 and Argoverse 2, while it is also the champion method of the CVPR 2024 Argoverse 2 motion forecasting challenge.
Mingkun Wang, Xiaoguang Ren, Ruochun Jin, Minglong Li, Xiaochuan Zhang, Changqian Yu, Wenjing Yang 0002
ICRA4
2025 EffBT: An Efficient Behavior Tree Reactive Synthesis and Execution Framework
abstract
Behavior Trees (BTs), originated from the control of Non-Player-Characters (NPCs), have been widely embraced in robotics and software engineering communities due to their modularity, reactivity, and other beneficial characteristics. It is highly desirable to synthesize BTs automatically. The consequent challenges are to ensure the generated BTs semantically correct, well-structured, and efficiently executable. To address these challenges, in this paper, we present a novel reactive synthesis method for BTs, namely EffBT, to generate correct and efficient controllers from formal specifications in GR(1) automatically. The idea is to construct BTs soundly from the intermediate strategies derived during the algorithm of GR(1) realizability check. Additionally, we introduce pruning strategies and use of Parallel nodes to improve BT execution, while none of the priors explored before. We prove the soundness of the EffBT method, and the experimental results demonstrate its effectiveness in various scenarios and datasets.
Ziji Wu, Peishan Huang, Shanghua Wen, Minglong Li, Ji Wang 0001
ICSE5
2025 PMAT: Optimizing Action Generation Order in Multi-Agent Reinforcement Learning
Muning Wen, Xihuai Wang, Shao Zhang, Yiwei Shi, Minne Li, Minglong Li, Ying Wen 0001
AAMAS7
2025 BTPG: A Platform and Benchmark for Behavior Tree Planning in Everyday Service Robots
abstract
Behavior Trees (BTs) are a widely used control architecture in robotics, renowned for their robustness and safety, which are especially crucial for everyday service robots. Recently, several methods have been proposed to automatically plan BTs to accomplish specific tasks. However, existing research in BT planning lacks two main aspects: (1) the absence of a standard platform for modeling and planning BTs, along with testing benchmarks; and (2) insufficient metrics for a comprehensive evaluation of BT planning algorithms. In this paper, we propose Behavior Tree Planning Gym (BTPG), the first platform and benchmark for BT planning in everyday service robots. In BTPG, behavior nodes are represented by predicate logic, and objects are categorized to better define the predicate domains and action models. The BT planning problem is then formulated in the STRIPS style. We support four environments and three simulators with different action models, which cover most of the needs of everyday service activities. We design a dataset generator for each environment and test three state-of-the-art BT planning algorithms, as well as one proposed by us, using various common metrics. In addition, we design three advanced metrics, planning progress, region distance, and execution robustness, to gain deeper insights into these BT planning algorithms. With a standard test benchmark, we hope BTPG can inspire and accelerate progress in the field of BT planning. Our codes are available at https://github.com/DIDS-EI/BTPG.
Xinglin Chen, Yishuai Cai, Minglong Li, Yunxin Mao, Wenjing Yang 0002, Ji Wang 0001
IJCAI3
2025 Learning Adaptive Reserve Price in Display Advertising
abstract
Real-Time Bidding (RTB) is a trading mechanism that allocates advertising (ad) requests through online auctions. Participants in these auctions typically include an ad exchange (AdX) and several demand-side platforms (DSPs). When an RTB auction begins, the AdX first establishes the reserve price set by publishers as the starting bid, after which the DSPs bid to compete for potential ad impressions. The reserve price strategy is crucial to the ad revenue of publishers; however, due to the strategic and dynamic bidding behavior of DSPs, optimizing the reserve price presents a significant challenge. In this work, we report a novel adaptive reserve price strategy based on reinforcement learning (RL). In our scheme, value bucket identification is leveraged to estimate the intrinsic values of ad inventories. Following this estimation, specialized reward functions are utilized to generate informative reward signals for RL models. Furthermore, we study the issue of risk management on the publisher side and develop a risk-aware instantiation to model risk tendency, considering both empirical expert knowledge and real-time trading conditions. Extensive experiments using real-world datasets collected from operational environments have demonstrated the effectiveness of the proposed method.
Kun Hu 0009, Lixia Wu, Yongjun Dai, Minfang Lu, Yuting Qiang, Minglong Li
KDD (1)7
2025 RDHNet: addressing rotational and permutational symmetries in continuous multi-agent systems
Dongzi Wang 0002, Lilan Huang, Muning Wen, Yuanxi Peng, Minglong Li, Teng Li 0011
Frontiers Comput. Sci.5
2024 Task Allocation in Heterogeneous Multi-Robot Systems Based on Preference-Driven Hedonic Game
abstract
Multiple preferences between robots and tasks have been largely overlooked in previous research on Multi-Robot Task Allocation (MRTA) problems. In this paper, we propose a preference-driven approach based on hedonic game to address the task allocation problem of muti-robot systems in emergency rescue scenarios. We present a distributed framework considering various preferences between robots and tasks to determine the division of coalitions in such problems and evaluate the scalability and adaptability of our algorithm through relevant experiments. Furthermore, considering the strict communication limitations in emergency rescue scenarios, we have verified that our algorithm can efficiently converge to a Nash-stable coalition partition even in conditions of insufficient communication distance.
Liwang Zhang, Minglong Li, Wenjing Yang 0002, Shaowu Yang
ICRA2
2024 Integrating Intent Understanding and Optimal Behavior Planning for Behavior Tree Generation from Human Instructions
Xinglin Chen, Yishuai Cai, Yunxin Mao, Minglong Li, Wenjing Yang 0002, Ji Wang 0001
IJCAI4
2024 Coalition Formation Game Approach for Task Allocation in Heterogeneous Multi-Robot Systems under Resource Constraints
abstract
This paper studies a case of the multi-robot task allocation (MRTA) problem, where each unmanned aerial vehicle (UAV) is endowed with multiple but limited resources. Completing each task necessitates UAVs to combine different resources through coalition formation, which will incur various costs including flight cost, execution cost, and cooperation cost. To minimize the total cost while maximizing both task completion rate and resource utilization rate, we model the MRTA problem of the UAVs as a leader-follower coalition formation game. In this game, leader UAVs coordinate follower UAVs to fulfill task resource requisites. Meanwhile, follower UAVs select suitable coalitions to join based on the altruistic preference. Theoretical analysis confirms the existence of a Nash stable partition in the coalition formation game. To achieve this stable partition, we propose a coalition formation algorithm. Simulation experiments validate that the proposed algorithm outperforms existing methods for the MRTA problem under resource constraints in terms of both task completion rate and resource utilization rate.
Liwang Zhang, Minglong Li, Wenjing Yang 0002, Shaowu Yang
IROS3
2024 RoMAT: Role-based multi-agent transformer for generalizable heterogeneous cooperation
Dongzi Wang 0002, Fangwei Zhong, Minglong Li, Muning Wen, Yuanxi Peng, Teng Li 0011, Yaodong Yang 0001
Neural Networks3
2024 XNV: Explainable Network Verification
abstract
Network verification has recently made strides, focusing on the satisfiability of configurations and policies or the performance and versatility of their methods. However, they generally ignore explainability, which is the ability to explain why a network violates or satisfies a certain forwarding policy. In this paper, we propose an explainable network verification framework XNV, which uses a novel interpretable fault analysis method to construct an effective explainable network verifier using knowledge graph (KG). XNV provides appropriate explanations to help operators understand the verification results, improving the transparency and trustworthiness of the verification system. First, XNV uses the KG as an intermediate representation of the configuration semantic level, storing the configuration semantics and routing protocol states. Then, XNV constructs human-logical fault trees for policies and implements root-cause analysis of policy violations based on KG queries and minimum cut set matching. Experiments and case evaluations show that our system provides good interpretability while balancing performance, accelerated understanding, and handling of misconfigurations.
Fuliang Li, Minglong Li, Yunhang Pu, Xingwei Wang 0001, Jiannong Cao 0001
IEEE/ACM Trans. Netw.2
2023 Memory-based Exploration-value Evaluation Model for Visual Navigation
abstract
We propose a hierarchical visual navigation solution, called Memory-based Exploration-value Evaluation Model (MEEM), to improve the agent's navigation performance. MEEM employs a hierarchical policy to tackle the challenge of sparse rewards, holds an episodic memory to store the historical information of the agent, and applies an Exploration-value Evaluation Model to calculate an exploration-value for action planning at each location in the observable area. We experimentally verify MEEM by navigation performance comparison on two datasets including the grid-map dataset and the 3D scenes Gibson dataset, where our approach achieves state-of-the-art performance on both. Specifically, the overall success rate of MEEM is 95% on the grid-map dataset while the best competitor reaches 68% only. As for the Gibson dataset, the success rate of ours and the best competitor SemExp are 69.8% and 54.4%, respectively. Ablation analysis on the tile-map dataset indicates that all three components of MEEM have positive effects.
Yongquan Feng, Minglong Li, Ruochun Jin, Shaowu Yang, Wenjing Yang 0002
ICRA3
2023 Task2Morph: Differentiable Task-Inspired Framework for Contact-Aware Robot Design
abstract
Optimizing the morphologies and the controllers that adapt to various tasks is a critical issue in the field of robot design, aka. embodied intelligence. Previous works typically model it as a joint optimization problem and use search-based methods to find the optimal solution in the morphology space. However, they ignore the implicit knowledge of task-to-morphology mapping which can directly inspire robot design. For example, flipping heavier boxes tends to require more muscular robot arms. This paper proposes a novel and general differentiable task-inspired framework for contact-aware robot design called Task2Morph. We abstract task features highly related to task performance and use them to build a task-to-morphology mapping. Further, we embed the mapping into a differentiable robot design process, where the gradient information is leveraged for both the mapping learning and the whole optimization. The experiments are conducted on three scenarios, and the results validate that Task2Morph outperforms DiffHand, which lacks a task-inspired morphology module, in terms of efficiency and effectiveness.
Yishuai Cai, Shaowu Yang, Minglong Li, Xinglin Chen, Yunxin Mao, Xiaodong Yi 0002, Wenjing Yang 0002
IROS3
2023 Evolving Physical Instinct for Morphology and Control Co-Adaption
abstract
The capability of a robot to perform tasks depends not only on precise motion control, but also on a well-suited body morphology. Adapting both morphology and control of robots to improve their task performance has been a widely studied and long-standing issue. While the bio-inspired bi-level optimization framework has gained popularity in recent years, it suffers from high computation complexity due to the time-consuming and inefficient learning process for each morphology. In fact, in nature, besides the adaptive morphology and the intelligent brain, animals also possess an important gift, which is physical instinct. These instincts allow animals to respond quickly to their surroundings in the neonatal period, facilitating skills acquisition. Inspired by this, we propose an evolvable instinct controller to enhance the morphology-control co-adaption. The instinct controller suggests rough motion inclinations, which require minimal domain knowledge and entail less sophisticated design. Its purpose is to assist the main controller in learning fine-grained and robust control efficiently. We implemented this idea in the context of legged locomotion and designed the instinct controller using phase-based FSMs. We propose the instinct-based co-adaption algorithm and construct GPU parallel simulation experiments on different morphology prototypes. The results indicate that combining the co-adaption process with instinct evolution leads to the development of superior morphologies and robust controllers compared with the conventional co-adaption approach, with minimal additional time cost.
Xinglin Chen, Minglong Li, Yishuai Cai, Zhuoer Wen, Zhongxuan Cai, Wenjing Yang 0002
IROS3
2023 Formal Verification Based Synthesis for Behavior Trees
Weijiang Hong, Zhenbang Chen 0001, Minglong Li, Peishan Huang, Ji Wang 0001
SETTA3
2021 BT Expansion: a Sound and Complete Algorithm for Behavior Planning of Intelligent Robots with Behavior Trees
abstract
Behavior Trees (BTs) have attracted much attention in the robotics field in recent years, which generalize existing control architectures and bring unique advantages for building robot systems. Automated synthesis of BTs can reduce human workload and build behavior models for complex tasks beyond the ability of human design, but theoretical studies are almost missing in existing methods because it is difficult to conduct formal analysis with the classic BT representations. As a result, they may fail in tasks that are actually solvable. This paper proposes BT expansion, an automated planning approach to building intelligent robot behaviors with BTs, and proves the soundness and completeness through the state-space formulation of BTs. The advantages of blended reactive planning and acting are formally discussed through the region of attraction of BTs, by which robots with BT expansion are robust to any resolvable external disturbances. Experiments with a mobile manipulator and test sets are simulated to validate the effectiveness and efficiency, where the proposed algorithm surpasses the baseline by virtue of its soundness and completeness. To the best of our knowledge, it is the first time to leverage the state-space formulation to synthesize BTs with a complete theoretical basis.
Zhongxuan Cai, Minglong Li, Wanrong Huang, Wenjing Yang 0002
AAAI2
2021 Dec-SGTS: Decentralized Sub-Goal Tree Search for Multi-Agent Coordination
abstract
Multi-agent coordination tends to benefit from efficient communication, where cooperation often happens based on exchanging information about what the agents intend to do, i.e. intention sharing. It becomes a key problem to model the intention by some proper abstraction. Currently, it is either too coarse such as final goals or too fined as primitive steps, which is inefficient due to the lack of modularity and semantics. In this paper, we design a novel multi-agent coordination protocol based on subgoal intentions, defined as the probability distribution over feasible subgoal sequences. The subgoal intentions encode macro-action behaviors with modularity so as to facilitate joint decision making at higher abstraction. Built over the proposed protocol, we present Dec-SGTS (Decentralized Sub-Goal Tree Search) to solve decentralized online multi-agent planning hierarchically and efficiently. Each agent runs Dec-SGTS asynchronously by iteratively performing three phases including local sub-goal tree search, local subgoal intention update and global subgoal intention sharing. We conduct the experiments on courier dispatching problem, and the results show that Dec-SGTS achieves much better reward while enjoying a significant reduction of planning time and communication cost compared with Dec-MCTS (Decentralized Monte Carlo Tree Search).
Minglong Li, Zhongxuan Cai, Wenjing Yang 0002, Lixia Wu, Ji Wang 0001
AAAI1
2021 Role-based attention in deep reinforcement learning for games
abstract
Abstract Reinforcement learning method that learns while interacting with the environment, relies heavily on the concept of state as the input to the policy and value function. In the task, the view of agent contains a lot of information, and it is difficult for the agent to learn to ignore the irrelevant information and focus on the key information. Inspired by recent work in attention models for computer vision, we present a role‐based attention model for reinforcement learning. The proposed model uses convolutional neural networks to generate soft attention maps, adding crucial role information in the task, forcing the agent to focus on important features and distinguish task‐related information. To validate the performance in complex problems, the proposed approach is evaluated in a challenging scenario, Football Academy in Google Research Football Environment, a newly released reinforcement learning environment with physics‐based three‐dimensional simulator. The experimental results demonstrate that agents using role‐based attention mechanism can perform better in football games.
Dong Yang 0010, Wenjing Yang 0002, Minglong Li, Qiong Yang
Comput. Animat. Virtual Worlds3
2020 Energy Minimum Regularization in Continual Learning
abstract
How to give agents the ability of continuous learning like human and animals is still a challenge. In the regularized continual learning method OWM, the constraint of the model on the energy compression of the learned task is ignored, which results in the poor performance of the method on the dataset with a large number of learning tasks. In this paper, we propose an energy minimization regularization(EMR) method to constrain the energy of learned tasks, providing enough learning space for the following tasks that are not learned, and increasing the capacity of the model to the number of learning tasks. A large number of experiments show that our method can effectively increase the capacity of the model and reduce the sensitivity of the model to the number of tasks and the size of the network.
Xiaobin Li 0006, Lianlei Shan, Minglong Li, Weiqiang Wang 0001
ICPR3
2020 Global-Local Attention Network for Semantic Segmentation in Aerial Images
abstract
Errors in semantic segmentation could be classified into two types: the large area misclassification and inaccurate local boundaries. Previously attention-based methods typically capture rich global contextual information, which benefits the large area classification but cannot address the local errors of boundaries. In this paper, we propose a Global-Local Attention Network (GLANet) which can simultaneously consider the global context and local details. Specifically, our GLANet consists of two branches: (1) the global attention branch and (2) local attention branch. Furthermore, three different modules are embedded in GLANet for respectively modelling the semantic interdependencies in spatial, channel and boundary dimension. Lastly, we merge the outputs of different branches to enhance the feature representation further, resulting in more precise segmentation. Overall, the proposed method achieves the competitive segmentation accuracy on two public aerial image datasets, bringing significant improvements over the existing baselines.
Minglong Li, Lianlei Shan, Xiaobin Li 0006, Dengji Zhou, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001
ICPR1
2020 UHRSNet: A Semantic Segmentation Network Specifically for Ultra-High-Resolution Images
abstract
Semantic segmentation is a basic task in computer vision, but only limited attention has been devoted to the ultra-high-resolution (UHR) image segmentation. Since UHR images occupy too much memory, they cannot be directly put into GPU for training. Previous methods are cropping images to small patches or downsampling the whole images. Cropping and downsampling cause the loss of contexts and details, which is essential for segmentation accuracy. To solve this problem, we improve and simplify the local and global feature fusion method in previous works. Local features are extracted from patches and global features are from downsampled images. Meanwhile, we propose one new fusion called local feature fusion for the first time, which can make patches get information from surrounding patches. We call the network with these two fusions ultra-high-resolution segmentation network (UHRSNet). These two fusions can effectively and efficiently solve the problem caused by cropping and downsampling. Experiments show a remarkable improvement on Deepglobe dataset [1].
Lianlei Shan, Minglong Li, Xiaobin Li 0006, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001, Weiqiang Wang 0001
ICPR2
2019 Parallel Gym Gazebo: a Scalable Parallel Robot Deep Reinforcement Learning Platform
abstract
Deep reinforcement learning is making advances in robotics with the platforms of realistic environment simulation. However, as shown in this paper, the realistic simulation introduces vast time cost which is the bottleneck of the learning procedure. To solve this problem generally, we propose a parallel reinforcement learning platform which follows the master-slave principle and integrates learning programs with multiple distributedrobot simulators. The platform is intrinsically scalable and requires no modification to existing serially designed learning environments or algorithms. Experimental results demonstrate that our platform significantly accelerates the learning progress of robots, in direct proportion to the parallel scale. The parallelism also brings richer exploration and sampling, enhancing the performance of deep reinforcement learning algorithms compared with existing serial platforms.
Zhongxuan Cai, Minglong Li, Wenjing Yang 0002
ICTAI3
2019 Integrating Decision Sharing with Prediction in Decentralized Planning for Multi-Agent Coordination under Uncertainty
abstract
The performance of decentralized multi-agent systems tends to benefit from information sharing and its effective utilization. However, too much or unnecessary sharing may hinder the performance due to the delay, instability and additional overhead of communications. Aiming to a satisfiable coordination performance, one would prefer the cost of communications as less as possible. In this paper, we propose an approach for improving the sharing utilization by integrating information sharing with prediction in decentralized planning. We present a novel planning algorithm by combining decision sharing and prediction based on decentralized Monte Carlo Tree Search called Dec-MCTS-SP. Each agent grows a search tree guided by the rewards calculated by the joint actions, which can not only be sampled from the shared probability distributions over action sequences, but also be predicted by a sufficiently-accurate and computationally-cheap heuristics-based method. Besides, several policies including sparse and discounted UCT and DIY-bonus are leveraged for performance improvement. We have implemented Dec-MCTS-SP in the case study on multi-agent information gathering under threat and uncertainty, which is formulated as Decentralized Partially Observable Markov Decision Process (Dec-POMDP). The factored belief vectors are integrated into Dec-MCTS-SP to handle the uncertainty. Comparing with the random, auction-based algorithm and Dec-MCTS, the evaluation shows that Dec-MCTS-SP can reduce communication cost significantly while still achieving a surprisingly higher coordination performance.
Minglong Li, Wenjing Yang 0002, Zhongxuan Cai, Shaowu Yang, Ji Wang 0001
IJCAI1
2016 ALLIANCE-ROS: A Software Architecture on ROS for Fault-Tolerant Cooperative Multi-robot Systems
Minglong Li, Zhongxuan Cai, Xiaodong Yi 0002, Xuejun Yang
PRICAI1