VLDB 2026 Research / reviewers in the wild / expert
Lanxiao Huang
dblp:255/6012
· DBLP profile ↗
13ranked-venue papers
2as first author
11since 2021 · last 2025
0009-0005-6366-4781ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration TestingabstractLarge Language Models (LLMs) have been explored for automating or enhancing penetration testing tasks, but their effectiveness and reliability across diverse attack phases remain open questions.This study presents a comprehensive evaluation of multiple LLM-based agents, ranging from singular to modular designs, across realistic penetration testing scenarios, analyzing their empirical performance and recurring failure patterns.We further investigate the impact of core functional capabilities on agent success, operationalized through five targeted augmentations: Global Context Memory (GCM), Inter-Agent Messaging (IAM), Context-Conditioned Invocation (CCI), Adaptive Planning (AP), and Real-Time Monitoring (RTM).These interventions respectively support the capabilities of Context Coherence & Retention, Inter-Component Coordination & State Management, Tool Usage Accuracy & Selective Execution, Multi-Step Strategic Planning & Error Detection & Recovery, and Real-Time Dynamic Responsiveness.Our findings reveal that while some architectures natively exhibit select properties, targeted augmentations significantly enhance modular agent performance-particularly in complex, multi-step, and real-time penetration testing scenarios. Lanxiao Huang, Daksh Dave, Tyler Cody, Peter A. Beling, Ming Jin 0002 |
EMNLP | 1 |
| 2025 | Mini Honor of Kings: A Lightweight Environment for Multiagent Reinforcement LearningabstractGames are widely used as research environments for multiagent reinforcement learning (MARL), but they pose three significant challenges: limited customization, high computational demands, and oversimplification. To address these issues, we introduce the first publicly available map editor for the popular mobile gameHonor of Kingsand design a lightweight environment,Mini Honor of Kings(Mini HoK), for researchers to conduct experiments. Mini HoK is highly efficient, allowing experiments to be run on personal PCs or laptops while still presenting sufficient challenges for existing MARL algorithms. We have tested our environment on common MARL algorithms and demonstrated that these algorithms have yet to surpass the performance of rule based policies, indicating that current MARL methods are not able to solve this environment. This facilitates the dissemination and advancement of MARL methods within the research community. In addition, we hope that more researchers will leverage theHonor of Kingsmap editor to develop innovative and scientifically valuable new maps. Lin Liu 0016, Jian Zhao 0018, Zhengtao Cao, Youpeng Zhao 0001, Zhenbin Ye, Zhaofeng He 0001, Houqiang Li, Xia Lin, Lanxiao Huang |
IEEE Trans. Games | 12 |
| 2024 | Multi-agent Multi-game Entity Transformer: Towards Generalist Models in MARLabstractBuilding large-scale generalist pre-trained models for many tasks is becoming an emerging and potential direction in reinforcement learning (RL).Research such as Gato and Multi-Game Decision Transformer have displayed outstanding performance and generalization capabilities on many games and domains.However, there exists a research blank about developing highly capable and generalist models in multi-agent RL (MARL), which can substantially accelerate progress toward general AI.To fill this gap, we propose Multi-Agent multi-Game ENtity TrAnsformer (MA-GENTA) from the entity perspective as orthogonal research to previous time-sequential modeling.Specifically, to deal with different state/observation spaces in different games, we analogize games as languages by aligning one single game to one single language, thus training different "tokenizers" and a shared transformer for various games.The feature inputs are split according to different entities and tokenized in the same continuous space.Then, two types of transformer-based models are proposed as permutationinvariant architectures to deal with various numbers of entities and capture the attention of different entities.MAGENTA is trained on Xianhan Zeng, Liang Wang 0015, Zhengjie Liang, Yiming Gao 0007, Feiyu Liu, Siqin Li, Xianliang Wang, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Longtao Zheng, Zinovi Rabinovich, Bo An 0001 |
DAI | 12 |
| 2024 | Enhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human GainabstractExisting game AI research mainly focuses on enhancing agents' abilities to win games, but this does not inherently make humans have a better experience when collaborating with these agents. For example, agents may dominate the collaboration and exhibit unintended or detrimental behaviors, leading to poor experiences for their human partners. In other words, most game AI agents are modeled in a "self-centered" manner. In this paper, we propose a "human-centered" modeling scheme for collaborative agents that aims to enhance the experience of humans. Specifically, we model the experience of humans as the goals they expect to achieve during the task. We expect that agents should learn to enhance the extent to which humans achieve these goals while maintaining agents' original abilities (e.g., winning games). To achieve this, we propose the Reinforcement Learning from Human Gain (RLHG) approach. The RLHG approach introduces a "baseline", which corresponds to the extent to which humans primitively achieve their goals, and encourages agents to learn behaviors that can effectively enhance humans in achieving their goals better. We evaluate the RLHG agent in the popular Multi-player Online Battle Arena (MOBA) game, Honor of Kings, by conducting real-world human-agent tests. Both objective performance and subjective preference results show that the RLHG agent provides participants better gaming experience. Yiming Gao 0007, Feiyu Liu, Liang Wang 0015, Dehua Zheng, Zhenjie Lian, Siqin Li, Xianliang Wang, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Wei Liu 0005 |
ICLR | 14 |
| 2023 | Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication Perspective
Yiming Gao 0007, Feiyu Liu, Liang Wang 0015, Zhenjie Lian, Siqin Li, Xianliang Wang, Xianhan Zeng, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Wei Liu 0005 |
ICLR | 13 |
| 2023 | Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning BenchmarksabstractThe advancement of Offline Reinforcement Learning (RL) and Offline Multi-Agent Reinforcement Learning (MARL) critically depends on the availability of high-quality, pre-collected offline datasets that represent real-world complexities and practical applications. However, existing datasets often fall short in their simplicity and lack of realism. To address this gap, we propose Hokoff, a comprehensive set of pre-collected datasets that covers both offline RL and offline MARL, accompanied by a robust framework, to facilitate further research. This data is derived from Honor of Kings, a recognized Multiplayer Online Battle Arena (MOBA) game known for its intricate nature, closely resembling real-life situations. Utilizing this framework, we benchmark a variety of offline RL and offline MARL algorithms. We also introduce a novel baseline algorithm tailored for the inherent hierarchical action space of the game. We reveal the incompetency of current offline RL approaches in handling task complexity, generalization and multi-task learning. Yun Qu 0002, Jianzhun Shao, Yuhang Jiang 0001, Zhenbin Ye, Lin Lai, Hongyang Qin, Minwen Deng, Juchao Zhuo, Deheng Ye, Qiang Fu 0016, Yang Guang, Wei Yang 0032, Lanxiao Huang, Xiangyang Ji |
NeurIPS | 17 |
| 2022 | MCS: An In-battle Commentary System for MOBA GamesabstractThis paper introduces a generative system for in-battle real-time commentary in mobile MOBA games. Event commentary is important for battles in MOBA games, which is applicable to a wide range of scenarios like live streaming, e-sports commentary and combat information analysis. The system takes real-time match statistics and events as input, and an effective transform method is designed to convert match statistics and utterances into consistent encoding space. This paper presents the general framework and implementation details of the proposed system, and provides experimental results on large-scale real-world match data. Xiaofeng Qi, Zhongping Liang, Jigang Liu, Yuanxin Wei, Lanxiao Huang |
COLING | 9 |
| 2022 | Exposing Surveillance Detection Routes via Reinforcement Learning, Attack Graphs, and Cyber TerrainabstractReinforcement learning (RL) operating on attack graphs leveraging cyber terrain principles are used to develop reward and state associated with determination of surveillance detection routes (SDR). This work extends previous efforts on developing RL methods for path analysis within enterprise networks. This work focuses on building SDR where the routes focus on exploring the network services while trying to evade risk. RL is utilized to support the development of these routes by building a reward mechanism that would help in realization of these paths. The RL algorithm is modified to have a novel warm-up phase which decides in the initial exploration which areas of the network are safe to explore based on the rewards and penalty scale factor. Lanxiao Huang, Tyler Cody, Christopher Redino, Abdul Rahman, Akshay Kakkar, Deepak Kushwaha, Cheng Wang 0040, Ryan Clark, Daniel Radke, Peter A. Beling, Edward Bowen |
ICMLA | 1 |
| 2022 | Honor of Kings Arena: an Environment for Generalization in Competitive Reinforcement LearningabstractThis paper introduces Honor of Kings Arena, a reinforcement learning (RL) environment based on the Honor of Kings, one of the world’s most popular games at present. Compared to other environments studied in most previous work, ours presents new generalization challenges for competitive reinforcement learning. It is a multi-agent problem with one agent competing against its opponent; and it requires the generalization ability as it has diverse targets to control and diverse opponents to compete with. We describe the observation, action, and reward specifications for the Honor of Kings domain and provide an open-source Python-based interface for communicating with the game engine. We provide twenty target heroes with a variety of tasks in Honor of Kings Arena and present initial baseline results for RL-based methods with feasible computing resources. Finally, we showcase the generalization challenges imposed by Honor of Kings Arena and possible remedies to the challenges. All of the software, including the environment-class, are publicly available. Hua Wei 0001, Jingxiao Chen, Xiyang Ji, Hongyang Qin, Minwen Deng, Siqin Li, Liang Wang 0015, Weinan Zhang 0001, Yong Yu 0001, Lanxiao Huang, Deheng Ye, Qiang Fu 0016, Wei Yang 0032 |
NeurIPS | 11 |
| 2022 | Supervised Learning Achieves Human-Level Performance in MOBA Games: A Case Study of Honor of KingsabstractWe present JueWu-SL, the first supervised-learning-based artificial intelligence (AI) program that achieves human-level performance in playing multiplayer online battle arena (MOBA) games. Unlike prior attempts, we integrate the macro-strategy and the micromanagement of MOBA-game-playing into neural networks in a supervised and end-to-end manner. Tested on Honor of Kings, the most popular MOBA at present, our AI performs competitively at the level of High King players in standard 5v5 games. Deheng Ye, Peilin Zhao, Fuhao Qiu, Bo Yuan 0008, Mingfei Sun 0001, Siqin Li, Zhenjie Lian, Bei Shi, Liang Wang 0015, Tengfei Shi, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang |
IEEE Trans. Neural Networks Learn. Syst. | 18 |
| 2021 | Learning Diverse Policies in MOBA Games via Macro-GoalsabstractRecently, many researchers have made successful progress in building the AI systems for MOBA-game-playing with deep reinforcement learning, such as on Dota 2 and Honor of Kings. Even though these AI systems have achieved or even exceeded human-level performance, they still suffer from the lack of policy diversity. In this paper, we propose a novel Macro-Goals Guided framework, called MGG, to learn diverse policies in MOBA games. MGG abstracts strategies as macro-goals from human demonstrations and trains a Meta-Controller to predict these macro-goals. To enhance policy diversity, MGG samples macro-goals from the Meta-Controller prediction and guides the training process towards these goals. Experimental results on the typical MOBA game Honor of Kings demonstrate that MGG can execute diverse policies in different matches and lineups, and also outperform the state-of-the-art methods over 102 heroes. Yiming Gao 0007, Bei Shi, Xueying Du, Liang Wang 0015, Guangwei Chen, Zhenjie Lian, Fuhao Qiu, Guoan Han, Deheng Ye, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang |
NeurIPS | 13 |
| 2020 | Mastering Complex Control in MOBA Games with Deep Reinforcement LearningabstractWe study the reinforcement learning problem of complex action control in the Multi-player Online Battle Arena (MOBA) 1v1 games. This problem involves far more complicated state and action spaces than those of traditional 1v1 games, such as Go and Atari series, which makes it very difficult to search any policies with human-level performance. In this paper, we present a deep reinforcement learning framework to tackle this problem from the perspectives of both system and algorithm. Our system is of low coupling and high scalability, which enables efficient explorations at large scale. Our algorithm includes several novel strategies, including control dependency decoupling, action mask, target attention, and dual-clip PPO, with which our proposed actor-critic network can be effectively trained in our system. Tested on the MOBA game Honor of Kings, the trained AI agents can defeat top professional human players in full 1v1 games. Deheng Ye, Mingfei Sun 0001, Bei Shi, Peilin Zhao, Hongsheng Yu, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin, Tengfei Shi, Liang Wang 0015, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang |
AAAI | 18 |
| 2020 | Towards Playing Full MOBA Games with Deep Reinforcement LearningabstractMOBA games, e.g., Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc. Developing AI for playing MOBA games has raised much attention accordingly. However, existing work falls short in handling the raw game complexity caused by the explosion of agent combinations, i.e., lineups, when expanding the hero pool in case that OpenAI's Dota AI limits the play to a pool of only 17 heroes. As a result, full MOBA games without restrictions are far from being mastered by any existing AI system. In this paper, we propose a MOBA AI learning paradigm that methodologically enables playing full MOBA games with deep reinforcement learning. Specifically, we develop a combination of novel and existing learning techniques, including off-policy adaption, multi-head value estimation, curriculum self-play learning, policy distillation, and Monte-Carlo tree-search, in training and playing a large pool of heroes, meanwhile addressing the scalability issue skillfully. Tested on Honor of Kings, a popular MOBA game, we show how to build superhuman AI agents that can defeat top esports players. The superiority of our AI is demonstrated by the first large-scale performance test of MOBA AI agent in the literature. Deheng Ye, Bo Yuan 0008, Fuhao Qiu, Hongsheng Yu, Yinyuting Yin, Bei Shi, Liang Wang 0015, Tengfei Shi, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Wei Liu 0005 |
NeurIPS | 17 |