EDBT 2026 Demo / reviewers in the wild / expert
Kun Shao
dblp:44/1667
· DBLP profile ↗
35ranked-venue papers
4as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AgentSwift: Efficient LLM Agent Design via Value-Guided Hierarchical SearchabstractLarge language model (LLM) agents have demonstrated strong capabilities across diverse domains, yet automated agent design remains a significant challenge. Current automated agent design approaches are often constrained by limited search spaces that primarily optimize workflows but fail to integrate crucial human-designed components like memory, planning, and tool use. Furthermore, these methods are hampered by high evaluation costs, as evaluating even a single new agent on a benchmark can require tens of dollars. The difficulty of this exploration is further exacerbated by inefficient search strategies that struggle to navigate the large design space effectively, making the discovery of novel agents a slow and resource-intensive process. To address these challenges, we propose AgentSwift, a novel framework for automated agent design. We formalize a hierarchical search space that jointly models agentic workflow and composable functional components. This structure moves beyond optimizing workflows alone by co-optimizing functional components, which enables the discovery of more complex and effective agent architectures. To make exploration within this expansive space feasible, we mitigate high evaluation costs by training a value model on a high-quality dataset, generated via a novel strategy combining combinatorial coverage and balanced Bayesian sampling for low-cost evaluation. Guiding the entire process is a hierarchical Monte Carlo Tree Search (MCTS) strategy, which is informed by uncertainty to efficiently navigate the search space. Evaluated across a comprehensive set of seven benchmarks spanning embodied, math, web, tool, and game domains, AgentSwift discovers agents that achieve an average performance gain of 8.34\% over both existing automated agent search methods and manually designed agents. Moreover, our framework exhibits steeper and more stable search trajectories. By enabling the efficient, automated composition of workflow with functional components, AgentSwift provides a scalable methodology to explore complex agent designs. Our framework serves as a launchpad for researchers to rapidly prototype and discover powerful agent architectures without the impediment of prohibitive evaluation costs. Yu Li 0022, Lehui Li, Qingmin Liao, Jianye Hao, Kun Shao, Fengli Xu |
AAAI | 6 |
| 2026 | Adaptive Theory of Mind for LLM-based Multi-Agent CoordinationabstractTheory of Mind (ToM) refers to the ability to reason about others’ mental states, and higher-order ToM involves considering that others also possess their own ToM. Equipping large language model (LLM)-driven agents with ToM has long been considered to improve their coordination in multiagent collaborative tasks. However, we find that misaligned ToM orders—mismatches in the depth of ToM reasoning between agents—can lead to insufficient or excessive reasoning about others, thereby impairing their coordination. To address this issue, we design an adaptive ToM (A-ToM) agent, which can align in ToM orders with its partner. Based on prior interactions, the agent estimates the partner’s likely ToM order and leverages this estimation to predict the partner’s action, thereby facilitating behavioral coordination. We conduct empirical evaluations on four multi-agent coordination tasks: a repeated matrix game, two grid navigation tasks and an Overcooked task. The results validate our findings on ToM alignment and demonstrate the effectiveness of our AToM agent. Furthermore, we discuss the generalizability of our A-ToM to non-LLM-based agents, as well as what would diminish the importance of ToM alignment. Chunjiang Mu, Ya Zeng, Qiaosheng Zhang 0002, Kun Shao, Chen Chu, Danyang Jia, Zhen Wang 0004, Shuyue Hu |
AAAI | 4 |
| 2026 | ResMAS: Resilience Optimization in LLM-based Multi-agent SystemsabstractLarge Language Model-based Multi-Agent Systems (LLM-based MAS), where multiple LLM agents collaborate to solve complex tasks, have shown impressive performance in many areas. However, MAS are typically distributed across different devices or environments, making them vulnerable to perturbations such as agent failures. While existing works have studied the adversarial attacks and corresponding defense strategies, they mainly focus on reactively detecting and mitigating attacks after they occur rather than proactively designing inherently resilient systems. In this work, we study the resilience of LLM-based MAS under perturbations and find that both the communication topology and prompt design significantly influence system resilience. Motivated by these findings, we propose ResMAS: a two-stage framework for enhancing MAS resilience. First, we train a reward model to predict the MAS’s resilience, based on which we train a topology generator to automatically design resilient topology for specific tasks through reinforcement learning. Second, we introduce a topology-aware prompt optimization method that refines each agent’s prompt based on its connections and interactions with other agents. Extensive experiments across a range of tasks show that our approach substantially improves MAS resilience under various constraints. Moreover, our framework demonstrates strong generalization ability to new tasks and models, highlighting its potential for building resilient MASs. Zhilun Zhou, Jiahe Liu, Qingyu Shao, Kun Shao, Depeng Jin, Fengli Xu |
AAAI | 6 |
| 2026 | AFE-Master: Enhancing LLM-Driven Autonomous Feature Engineering with Domain-Specific Language Parsing and Guided Local SearchabstractAutonomous Feature Engineering (AFE) is critical for improving predictive performance on tabular data by relieving humans from manual feature crafting. However, traditional AFE lacks the semantic guidance needed to fully exploit domain knowledge. Although large language models (LLMs) can, in principle, emulate experts, existing approaches typically operate in an open code space that directly generates and rewrites entire features; without a compositional structural representation and invariant constraints, edits are coarse and non-local, making it hard to distill interpretable features with high information content and rich hierarchical structure. Hebin Liang, Jianye Hao, Jinyi Liu 0002, Yi Ma 0005, Zilin Cao, Kun Shao, Zhaocheng Du, Fei Ni 0001, Yifu Yuan, Yan Zheng 0002 |
WWW | 7 |
| 2026 | Mask-Aware Kernel Learning for Action RecognitionabstractAction recognition aims to identify an action from video frames. The actions are usually surrounded by irrelevant backgrounds. The action/background information is diverse in different video frames, which hinders learning the implicit action patterns. In this work, we propose a Mask-aware Kernel Model (MKM), which ensures implicit action pattern learning by integrating kernel learning with proper cluster relations. The MKM provides novel cluster-aware kernels to enhance the action representation for frame patches. The MKM is deployed on a temporal Vision Transformer, and introduces a kernel clustering learner, kernel masking filter, and a kernel attention selector.First, to learn temporal features, the temporal Vision Transformer uses temporal correlation to ensure the action features for kernel learning.Second, to analyze the action kernels for frame patches, we design a kernel clustering learner module. This module learns cluster relations with patch- wise convolutions to describe the common action among patches. The cluster relations are learned in each frame, which ensures cluster-aware kernel learning with input frame adaptivity.Third, to analyze the action kernels with spatial adaptivity, we design a kernel masking filter module. This module introduces a location mask by analyzing the region patterns with spatial convolution. The patch-level mask ensures the kernel learning with region-aware selection.Fourth, after learning multiple channel features by convolution with multiple kernels, we design a kernel attention selector module. This module excites kernel-aware features by learning channel- wise attention with channel- wise convolutions, which ensures the kernel learning with channel- wise selection for effective action representation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on Something-Something V1 & V2, Kinetics-400, UAV-human, and Diving 48 datasets. Kewei Wu, Chongjia Zhu, Zhao Xie, Kun Shao, Dan Guo 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Learning Precise Affordances From Egocentric Videos for Robotic Manipulation
Gen Li 0008, Nikolaos Tsagkas, Jifei Song, Ruaridh Mon-Williams, Sethu Vijayakumar, Kun Shao, Laura Sevilla-Lara |
ICCV | 6 |
| 2025 | Spa-Bench: a comprehensive Benchmark for Smartphone Agent EvaluationabstractSmartphone agents are increasingly important for helping users control devices efficiently, with (Multimodal) Large Language Model (MLLM)-based approaches emerging as key contenders. Fairly comparing these agents is essential but challenging, requiring a varied task scope, the integration of agents with different implementations, and a generalisable evaluation pipeline to assess their strengths and weaknesses. In this paper, we present SPA-Bench, a comprehensive SmartPhone Agent Benchmark designed to evaluate (M)LLM-based agents in an interactive environment that simulates real-world conditions. SPA-Bench offers three key contributions: (1) A diverse set of tasks covering system and third-party apps in both English and Chinese, focusing on features commonly used in daily routines; (2) A plug-and-play framework enabling real-time agent interaction with Android devices, integrating over ten agents with the flexibility to add more; (3) A novel evaluation pipeline that automatically assesses agent performance across multiple dimensions, encompassing seven metrics related to task completion and resource consumption. Our extensive experiments across tasks and agents reveal challenges like interpreting mobile user interfaces, action grounding, memory retention, and execution costs. We propose future research directions to ease these difficulties, moving closer to real-world smartphone agent applications. Jingxuan Chen, Derek Yuen, Yuhao Yang 0008, Gongwei Chen, Li Yixing, Xurui Zhou, Weiwen Liu, Shuai Wang 0020, Kaiwen Zhou 0001, Rui Shao 0001, Liqiang Nie, Yasheng Wang, Jianye Hao, Jun Wang 0012, Kun Shao |
ICLR | 17 |
| 2025 | Lightweight Neural App ControlabstractThis paper introduces a novel mobile phone control architecture, Lightweight Multi-modal App Control (LiMAC), for efficient interactions and control across various Android apps. LiMAC takes as input a textual goal and a sequence of past mobile observations, such as screenshots and corresponding UI trees, to generate precise actions. To address the computational constraints inherent to smartphones, we introduce a small Action Transformer (AcT) integrated with a fine-tuned vision-language model (VLM) for real-time decision-making and task execution. We evaluate LiMAC on two open-source mobile control datasets, demonstrating the superior performance of our small-form-factor approach against fine-tuned versions of open-source VLMs, such as Florence2 and Qwen2-VL. It also significantly outperforms prompt engineering baselines utilising closed-source foundation models like GPT-4o. More specifically, LiMAC increases the overall action accuracy by up to 19% compared to fine-tuned VLMs, and up to 42% compared to prompt-engineering baselines. Filippos Christianos, Georgios Papoudakis, Thomas Coste, Jianye Hao, Jun Wang 0012, Kun Shao |
ICLR | 6 |
| 2025 | DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control AgentabstractOn-device control agents, especially on mobile devices, are responsible for operating mobile devices to fulfill users' requests, enabling seamless and intuitive interactions. Integrating Multimodal Large Language Models (MLLMs) into these agents enhances their ability to understand and execute complex commands, thereby improving user experience. However, fine-tuning MLLMs for on-device control presents significant challenges due to limited data availability and inefficient online training processes. This paper introduces DistRL, a novel framework designed to enhance the efficiency of online RL fine-tuning for mobile device control agents. DistRL employs centralized training and decentralized data acquisition to ensure efficient fine-tuning in the context of dynamic online interactions. Additionally, the framework is backed by our tailor-made RL algorithm, which effectively balances exploration with the prioritized utilization of collected data to ensure stable and robust training. Our experiments show that, on average, DistRL delivers a 3$\times$ improvement in training efficiency and enables training data collection 2.4$\times$ faster than the leading synchronous multi-machine methods. Notably, after training, DistRL achieves a 20\% relative improvement in success rate compared to state-of-the-art methods on general Android tasks from an open benchmark, significantly outperforming existing approaches while maintaining the same training time. These results validate DistRL as a scalable and efficient solution, offering substantial improvements in both training efficiency and agent performance for real-world, in-the-wild device control tasks. Taiyi Wang, Jianheng Liu, Jianye Hao, Jun Wang 0012, Kun Shao |
ICLR | 6 |
| 2025 | Taming Multi-Agent Reinforcement Learning with Estimator Variance Reduction
Taher Jafferjee, Juliusz Krysztof Ziomek, Tianpei Yang, Zipeng Dai, Matthew E. Taylor, Kun Shao, Jun Wang 0012, David Mguni |
AAMAS | 7 |
| 2025 | Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Video Temporal GroundingabstractVideo Temporal Grounding (TG) aims to temporally locate video segments matching a natural language description (a query) in a long video. While Vision-Language Models (VLMs) are effective at holistic semantic matching, they often struggle with fine-grained temporal
localisation. Recently, Group Relative Policy Optimisation (GRPO) reformulates the inference process as a reinforcement learning task, enabling fine-grained grounding and achieving strong in-domain performance. However, GRPO relies on labelled data, making it unsuitable in unlabelled domains. Moreover, because videos are large and expensive to store and process, performing full-scale adaptation introduces prohibitive latency and computational overhead, making it impractical for real-time deployment. To overcome both problems, we introduce a Data-Efficient Unlabelled Cross-domain Temporal Grounding method, from which a model is first trained on a labelled source domain, then adapted to a target domain using only a small number of {\em unlabelled videos from the target domain}. This approach eliminates the need for target annotation and keeps both computational and storage overhead low enough to run in real time. Specifically, we introduce \textbf{U}ncertainty-quantified \textbf{R}ollout \textbf{P}olicy \textbf{A}daptation (\textbf{URPA}) for cross-domain knowledge transfer in learning video temporal grounding without target labels. URPA generates multiple candidate predictions using GRPO rollouts, averages them to form a pseudo label, and estimates confidence from the variance across these rollouts. This confidence then weights the training rewards, guiding the model to focus on reliable supervision. Experiments on three datasets across six cross-domain settings show that URPA generalises well using only a few unlabelled target videos. Codes are given in supplemental materials. Zixu Cheng, Shaogang Gong, Isabel Guan, Jianye Hao, Jun Wang 0012, Kun Shao |
NeurIPS | 7 |
| 2025 | ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM ReasoningabstractEvaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench. Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004 |
NeurIPS | 9 |
| 2025 | Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App ControlabstractReinforcement learning (RL) using foundation models for policy approximations in multi-turn tasks remains challenging. We identify two main limitations related to sparse reward settings and policy gradient updates, based on which we formulate a key insight: updates from positive samples with high returns typically do not require policy regularisation, whereas updates from negative samples, reflecting undesirable behaviour, can harm model performance. This paper introduces Succeed or Learn Slowly (SoLS), a novel off-policy RL algorithm evaluated on mobile app control tasks. SoLS improves sample efficiency when fine-tuning foundation models for user interface navigation via a modified off-policy actor-critic approach, applying direct policy updates for positive samples and conservative, regularised updates for negative ones to prevent model degradation. We augment SoLS with Successful Transition Replay (STR), which prioritises learning from successful interactions, further improving sample efficiency. We evaluate SoLS on the AndroidWorld benchmark, where it significantly outperforms existing methods (at least 17\% relative increase), including prompt-engineering and RL approaches, while requiring substantially fewer computational resources than GPT-4o-based methods with 5-60x faster inference. Georgios Papoudakis, Thomas Coste, Jianye Hao, Jun Wang 0012, Kun Shao |
NeurIPS | 5 |
| 2024 | Distilling Morphology-Conditioned Hypernetworks for Efficient Universal Morphology ControlabstractLearning a universal policy across different robot morphologies can significantly improve learning efficiency and enable zero-shot generalization to unseen morphologies. However, learning a highly performant universal policy requires sophisticated architectures like transformers (TF) that have larger memory and computational cost than simpler multi-layer perceptrons (MLP). To achieve both good performance like TF and high efficiency like MLP at inference time, we propose HyperDistill, which consists of: (1) A morphology-conditioned hypernetwork (HN) that generates robot-wise MLP policies, and (2) A policy distillation approach that is essential for successful training. We show that on UNIMAL, a benchmark with hundreds of diverse morphologies, HyperDistill performs as well as a universal TF teacher policy on both training and unseen test robots, but reduces model size by 6-14 times, and computational cost by 67-160 times in different environments. Our analysis attributes the efficiency advantage of HyperDistill at inference time to knowledge decoupling, i.e., the ability to decouple inter-task and intra-task knowledge, a general principle that could also be applied to improve inference efficiency in other domains. The code is publicly available at https://github.com/MasterXiong/Universal-Morphology-Control. Zheng Xiong, Risto Vuorio, Jacob Beck, Matthieu Zimmer, Kun Shao, Shimon Whiteson |
ICML | 5 |
| 2024 | Cooperative Multiagent Transfer Learning With Coalition Pattern DecompositionabstractKnowledge transfer in cooperative multi-agent reinforcement learning (MARL) has drawn increasing attention in recent years. Unlike generalizing policies in single-agent tasks, it is more important to consider coordination knowledge than individual knowledge in multi-agent transfer learning. However, most of the existing methods only focus on knowledge transfer of the individual agent policy, which leads to coordination bias and finally affects the final performance in cooperative MARL. In this paper, we propose a level-adaptive MARL framework called “LA-QTransformer”, to realize the knowledge transfer on the coordination level via efficiently decomposing the agent coordination into multi-level coalition patterns for different agents. Compatible with centralized training with decentralized execution (CTDE) regime, LA-QTransformer utilizes the Level- Adaptive Transformer to generate suitable coalition patterns and then realizes the credit assignment for each agent. Besides, to deal with unexpected changes in the number of agents in the coordination transfer phase, we design a policy network called “Population invariant agent with Transformer (PIT)” to adapt dynamic observation and action space. We evaluate the LAQTransformer and PIT in the StarCraft II micro-management benchmark by comparing them with several state-of-the-art MARL baselines. The experimental results demonstrate the superiority of LA-QTransformer and PIT and verify the feasibility of coordination knowledge transfer. Tianze Zhou, Fubiao Zhang, Kun Shao, Zipeng Dai, Kai Li 0022, Wenhan Huang, Weixun Wang, Bin Wang 0034, Dong Li 0016, Wulong Liu, Jianye Hao |
IEEE Trans. Games | 3 |
| 2023 | Learning to Shape Rewards Using a Game of Two PartnersabstractReward shaping (RS) is a powerful method in reinforcement learning (RL) for overcoming the problem of sparse or uninformative rewards. However, RS typically relies on manually engineered shaping-reward functions whose construc- tion is time-consuming and error-prone. It also requires domain knowledge which runs contrary to the goal of autonomous learning. We introduce Reinforcement Learning Optimising Shaping Algorithm (ROSA), an automated reward shaping framework in which the shaping-reward function is constructed in a Markov game between two agents. A reward-shaping agent (Shaper) uses switching controls to determine which states to add shaping rewards for more efficient learning while the other agent (Controller) learns the optimal policy for the task using these shaped rewards. We prove that ROSA, which adopts existing RL algorithms, learns to construct a shaping-reward function that is beneficial to the task thus ensuring efficient convergence to high performance policies. We demonstrate ROSA’s properties in three didactic experiments and show its superior performance against state-of-the-art RS algorithms in challenging sparse reward environments. David Mguni, Taher Jafferjee, Nicolas Perez Nieves, Wenbin Song, Feifei Tong, Matthew E. Taylor, Tianpei Yang, Zipeng Dai, Jiangcheng Zhu, Kun Shao, Jun Wang 0012, Yaodong Yang 0001 |
AAAI | 12 |
| 2023 | Traj-MAE: Masked Autoencoders for Trajectory PredictionabstractTrajectory prediction has been a crucial task in building a reliable autonomous driving system by anticipating possible dangers. One key issue is to generate consistent trajectory predictions without colliding. To overcome the challenge, we propose an efficient masked autoencoder for trajectory prediction (Traj-MAE) that better represents the complicated behaviors of agents in the driving environment. Specifically, our Traj-MAE employs diverse masking strategies to pre-train the trajectory encoder and map encoder, allowing for the capture of social and temporal information among agents while leveraging the effect of environment from multiple granularities. To address the catastrophic forgetting problem that arises when pre-training the network with multiple masking strategies, we introduce a continual pre-training framework, which can help Traj-MAE learn valuable and diverse information from various strategies efficiently. Our experimental results in both multi-agent and single-agent settings demonstrate that Traj-MAE achieves competitive results with state-of-the-art methods and significantly outperforms our baseline model. Project page: https://jiazewang.com/projects/trajmae.html. Hao Chen 0193, Kun Shao, Furui Liu, Jianye Hao, Chenyong Guan, Guangyong Chen, Pheng-Ann Heng |
ICCV | 3 |
| 2023 | Timing is Everything: Learning to Act Selectively with Costly Actions and Budgetary Constraints
David Mguni, Aivar Sootla, Juliusz Krysztof Ziomek, Oliver Slumbers, Zipeng Dai, Kun Shao, Jun Wang 0012 |
ICLR | 6 |
| 2023 | ChessGPT: Bridging Policy Learning and Language ModelingabstractWhen solving decision-making tasks, humans typically depend on information from two key sources: (1) Historical policy data, which provides interaction replay from the environment, and (2) Analytical insights in natural language form, exposing the invaluable thought process or strategic considerations. Despite this, the majority of preceding research focuses on only one source: they either use historical replay exclusively to directly learn policy or value functions, or engaged in language model training utilizing mere language corpus. In this paper, we argue that a powerful autonomous agent should cover both sources. Thus, we propose ChessGPT, a GPT model bridging policy learning and language modeling by integrating data from these two sources in Chess games. Specifically, we build a large-scale game and language dataset related to chess. Leveraging the dataset, we showcase two model examples ChessCLIP and ChessGPT, integrating policy learning and language modeling. Finally, we propose a full evaluation framework for evaluating language model's chess ability. Experimental results validate our model and dataset's effectiveness. We open source our code, model, and dataset at https://github.com/waterhorse1/ChessGPT. Xidong Feng, Yicheng Luo, Hongrui Tang, Mengyue Yang, Kun Shao, David Mguni, Yali Du 0001, Jun Wang 0012 |
NeurIPS | 6 |
| 2022 | Efficient Dual-Process Cognitive Recommender Balancing Accuracy and Diversity
Yixu Gao, Kun Shao, Zhijian Duan 0001, Zhongyu Wei, Dong Li 0016, Bin Wang 0034, Mengchen Zhao, Jianye Hao |
DASFAA (3) | 2 |
| 2022 | Heterogeneous Graph Neural Network-Based Imitation Learning for Gate Sizing AccelerationabstractGate Sizing is an important step in logic synthesis, where the cells are resized to optimize metrics such as area, timing, power, leakage, etc. In this work, we consider the gate sizing problem for leakage power optimization with timing constraints. Lagrangian Relaxation is a widely employed optimization method for gate sizing problems. We accelerate Lagrangian Relaxation-based algorithms by narrowing down the range of cells to resize. In particular, we formulate a heterogeneous directed graph to represent the timing graph, propose a heterogeneous graph neural network as the encoder, and train in the way of imitation learning to mimic the selection behavior of each iteration in Lagrangian Relaxation. This network is used to predict the set of cells that need to be changed during the optimization process of Lagrangian Relaxation. Experiments show that our accelerated gate sizer could achieve comparable performance to the baseline with an average of 22.5% runtime reduction. Xinyi Zhou 0010, Junjie Ye 0002, Chak-Wa Pui, Kun Shao, Guangliang Zhang, Bin Wang 0034, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
ICCAD | 4 |
| 2022 | Promoting Quality and Diversity in Population-based Reinforcement Learning via Hierarchical Trajectory Space ExplorationabstractQuality Diversity (QD) algorithms in population-based reinforcement learning aim to optimize agents' returns and diversity among the population simultaneously. It is conducive to solving exploration problems in reinforcement learning and potentially getting multiple good and diverse strategies. However, previous methods typically define behavioral embedding in action space or outcome space, which neglect trajectory characteristics during the execution process. In this paper, we introduce a trajectory embedding model trained by Variational Autoencoder with similarity constraint to characterize trajectory features. Based on that, we propose a hierarchical trajectory-space exploration (HTSE) framework using Determinantal Point Processes (DPP) to generate high-quality and diverse solutions in the selection and mutation process. The experimental results show that our HTSE method effectively completes several simulated tasks, outperforming other Quality-Diversity Reinforcement Learning algorithms. Jiayu Miao, Tianze Zhou, Kun Shao, Ming Zhou 0006, Weinan Zhang 0001, Jianye Hao, Yong Yu 0001, Jun Wang 0012 |
ICRA | 3 |
| 2022 | Multiagent Q-learning with Sub-Team CoordinationabstractIn many real-world cooperative multiagent reinforcement learning (MARL) tasks, teams of agents can rehearse together before deployment, but then communication constraints may force individual agents to execute independently when deployed. Centralized training and decentralized execution (CTDE) is increasingly popular in recent years, focusing mainly on this setting. In the value-based MARL branch, credit assignment mechanism is typically used to factorize the team reward into each individual’s reward — individual-global-max (IGM) is a condition on the factorization ensuring that agents’ action choices coincide with team’s optimal joint action. However, current architectures fail to consider local coordination within sub-teams that should be exploited for more effective factorization, leading to faster learning. We propose a novel value factorization framework, called multiagent Q-learning with sub-team coordination (QSCAN), to flexibly represent sub-team coordination while honoring the IGM condition. QSCAN encompasses the full spectrum of sub-team coordination according to sub-team size, ranging from the monotonic value function class to the entire IGM function class, with familiar methods such as QMIX and QPLEX located at the respective extremes of the spectrum. Experimental results show that QSCAN’s performance dominates state-of-the-art methods in matrix games, predator-prey tasks, the Switch challenge in MA-Gym. Additionally, QSCAN achieves comparable performances to those methods in a selection of StarCraft II micro-management tasks. Wenhan Huang, Kai Li 0022, Kun Shao, Tianze Zhou, Matthew E. Taylor, Jun Luo 0009, Dongge Wang 0001, Hangyu Mao, Jianye Hao, Jun Wang 0012, Xiaotie Deng |
NeurIPS | 3 |
| 2022 | Postoperative MPA-AUC0-12h Prediction for Kidney Transplant Recipients based on Few-shot LearningabstractMycophenolic acid (MPA) is a commonly used immunosuppressive drug.The anti-immune rejection effect of mycophenolic acid is closely related to its exposure level in the body.In clinical practice, mycophenolate acid drug exposure level is usually reflected by monitoring the area under the drug-time curve MPA-AUC0-12h.Calculating the MPA-AUC0-12h requires numerous blood sampling time points.Not only does the medical staff have more work, but patients suffer more as well.Limited sampling strategies (LSS) are generally used to reduce the number of time points.Nevertheless, this method involves complicated calculations and the predictive accuracy is very low for small sample data.A new method of predicting the MPA-AUC0-12h value is proposed based on the selection of SHAP features with an improved neural network for small sample data.The experimental results show that the average prediction errors of the MPA-AUC0-12h values on different data sets by our method are better than that of the baseline models. Qiao Pan, Kun Shao |
SEKE | 4 |
| 2022 | The triggers that open the NLP model backdoors are hidden in the adversarial samples
Kun Shao, Junan Yang, Xiaoshuai Li, Hui Liu 0032 |
Comput. Secur. | 1 |
| 2021 | BDDR: An Effective Defense Against Textual Backdoor Attacks
Kun Shao, Junan Yang, Yang Ai, Hui Liu 0032 |
Comput. Secur. | 1 |
| 2020 | Multi-Agent Determinantal Q-LearningabstractCentralized training with decentralized execution has become an important paradigm in multi-agent learning. Though practical, current methods rely on restrictive assumptions to decompose the centralized value function across agents for execution. In this paper, we eliminate this restriction by proposing multi-agent determinantal Q-learning. Our method is established on Q-DPP, a novel extension of determinantal point process (DPP) to multi-agent setting. Q-DPP promotes agents to acquire diverse behavioral models; this allows a natural factorization of the joint Q-functions with no need for \emph{a priori} structural constraints on the value function or special network architectures. We demonstrate that Q-DPP generalizes major solutions including VDN, QMIX, and QTRAN on decentralizable cooperative tasks. To efficiently draw samples from Q-DPP, we develop a linear-time sampler with theoretical approximation guarantee. Our sampler also benefits exploration by coordinating agents to cover orthogonal directions in the state space during training. We evaluate our algorithm on multiple cooperative benchmarks; its effectiveness has been demonstrated when compared with the state-of-the-art. Yaodong Yang 0001, Ying Wen 0001, Jun Wang 0012, Kun Shao, David Mguni, Weinan Zhang 0001 |
ICML | 5 |
| 2020 | Cooperative Multi-Agent Deep Reinforcement Learning with Counterfactual RewardabstractIn partially observable fully cooperative games, agents generally tend to maximize global rewards with joint actions, so it is difficult for each agent to deduce their own contribution. To address this credit assignment problem, we propose a multi-agent reinforcement learning algorithm with counterfactual reward mechanism, which is termed as CoRe algorithm. CoRe computes the global reward difference in condition that the agent does not take its actual action but takes other actions, while other agents fix their actual actions. This approach can determine each agent's contribution for the global reward. We evaluate CoRe in a simplified Pig Chase game with a decentralised Deep Q Network (DQN) framework. The proposed method helps agents learn end-to-end collaborative behaviors. Compared with other DQN variants with global reward, CoRe significantly improves learning efficiency and achieves better results. In addition, CoRe shows excellent performances in various size game environments. Kun Shao, Yuanheng Zhu, Zhentao Tang, Dongbin Zhao |
IJCNN | 1 |
| 2020 | An improved CapsNet applied to recognition of 3D vertebral images
Hao Wang 0008, Kun Shao, Xing Huo |
Appl. Intell. | 2 |
| 2019 | A Cooperative Particle Swarm Optimization Algorithm Based on Greedy Disturbance
Xing Huo, Jieqing Tan, Kun Shao |
PRCV (3) | 5 |
| 2019 | Deep sparse representation-based mid-level visual elements discovery in fine-grained classification
Le Lv, Dongbin Zhao, Kun Shao |
Soft Comput. | 3 |
| 2018 | Visual Navigation with Actor-Critic Deep Reinforcement LearningabstractVisual navigation in complex environments is crucial for intelligent agents. In this paper, we propose an efficient deep reinforcement learning (DRL) method to tackle visual navigation tasks. We present the synchronous advantage actor-critic (A2C) with generalized advantage estimator (GAE) algorithm. The A2C enables agents to learn from multiple processes, which significantly reduces the training time. The GAE used to estimate the advantage function improves the policy gradient estimates. We focus on visual navigation tasks in ViZDoom, and train agents in two health gathering scenarios. The experimental results show this method successfully teaches our agents to navigate in these scenarios. The A2C with GAE agent reaches the highest score in the first task, and a competitive score in the second task. In addition, this agent has better average scores and lower variances in both tasks. Kun Shao, Dongbin Zhao, Yuanheng Zhu |
IJCNN | 1 |
| 2018 | Small sample image recognition using improved Convolutional Neural Network
Kun Shao |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Practical and accurate pinpointing of configuration errors using static analysisabstractSoftware misconfigurations are responsible for a substantial part of today's system failures, causing about one-quarter of all customer-reported issues. Identifying their root causes can be costly in terms of time and human resources. We present an approach to automatically pinpoint such defects without error reproduction. It uses static analysis to infer the correlation degree between each configuration option and program sites affected by an exception. The only run-time information required by our approach is the stack trace of a failure. This is an essential advantage compared to existing approaches which require to reproduce errors or to provide testing oracles. We evaluate our approach on 29 errors from 4 configurable software programs, namely JChord, Randoop, Hadoop, and Hbase. Our approach can successfully diagnose 27 out of 29 errors. For 20 errors, the failure-inducing configuration option is ranked first. Zhen Dong 0004, Artur Andrzejak 0001, Kun Shao |
ICSME | 3 |
| 2001 | A Parliamentary Architecture of multi-agent systemabstractThe paper presents an architecture of MAS called Parliamentary Architecture. It employs the concepts of management role and councilman role to perform management tasks and to speak for the agent community. The mechanism of electing a public servant role can significantly enhance the capability of self-organization. In order to achieve interoperability of a legacy system, the architecture allows for two types of communication languages: global communicative language and local communicative language. It briefly introduces a case to illustrate the architecture. Zhiyonq Sun, Zongtian Liu, Kun Shao |
SMC | 3 |