VLDB 2026 Research / reviewers in the wild / expert
Zhaopeng Meng
dblp:67/8175 · also Zhao-Peng Meng
· DBLP profile ↗
45ranked-venue papers
0as first author
26since 2021 · last 2025
0000-0001-6019-5952ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 9 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Qibo: A Large Language Model for traditional Chinese medicine
Yongzhe Jia, Xin Wang 0030, Heyi Zhang, Zhaopeng Meng, Pengwei Zhuang, Jianguo Wei |
Expert Syst. Appl. | 5 |
| 2025 | Nearest neighbor regression for evolutionary dynamic multiobjective optimization
Youpeng Deng, Haobo Gao, Yan Zheng 0002, Zhaopeng Meng, Yueyang Hua, Qiangguo Jin, Leilei Cao |
Inf. Sci. | 4 |
| 2024 | ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
Jianye Hao, Yi Ma 0005, Jinyi Liu 0002, Yan Zheng 0002, Zhaopeng Meng |
IJCAI | 6 |
| 2024 | Learning with noisy labels using collaborative sample selection and contrastive semi-supervised learning
Xiaohe Wu, Chao Xu 0003, Yanli Ji, Wangmeng Zuo, Yiwen Guo, Zhaopeng Meng |
Knowl. Based Syst. | 7 |
| 2024 | Corrigendum to "Learning with Noisy Labels Using Collaborative Sample Selection and Contrastive Semi-Supervised Learning" [Knowledge-Based Systems 296 (2024) 111860]
Xiaohe Wu, Chao Xu 0003, Yanli Ji, Wangmeng Zuo, Yiwen Guo, Zhaopeng Meng |
Knowl. Based Syst. | 7 |
| 2024 | Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent DomainabstractDeep reinforcement learning (DRL) and deep multiagent reinforcement learning (MARL) have achieved significant success across a wide range of domains, including game artificial intelligence (AI), autonomous vehicles, and robotics. However, DRL and deep MARL agents are widely known to be sample inefficient that millions of interactions are usually needed even for relatively simple problem settings, thus preventing the wide application and deployment in real-industry scenarios. One bottleneck challenge behind is the well-known exploration problem, i.e., how efficiently exploring the environment and collecting informative experiences that could benefit policy learning toward the optimal ones. This problem becomes more challenging in complex environments with sparse rewards, noisy distractions, long horizons, and nonstationary co-learners. In this article, we conduct a comprehensive survey on existing exploration methods for both single-agent RL and multiagent RL. We start the survey by identifying several key challenges to efficient exploration. Then, we provide a systematic survey of existing approaches by classifying them into two major categories: uncertainty-oriented exploration and intrinsic motivation-oriented exploration. Beyond the above two main branches, we also include other notable exploration methods with different ideas and techniques. In addition to algorithmic analysis, we provide a comprehensive and unified empirical comparison of different exploration methods for DRL on a set of commonly used benchmarks. According to our algorithmic and empirical investigation, we finally summarize the open problems of exploration in DRL and deep MARL and point out a few future directions. Jianye Hao, Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu 0002, Zhaopeng Meng, Peng Liu 0008, Zhen Wang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Node-Wise Domain Adaptation Based on Transferable Attention for Recognizing Road Rage via EEGabstractRoad rage is a social problem that deserves attention, but few research has been done so far. In this paper, based on the biological topology of multi-channel electroencephalogram (EEG) signals, we propose a model which combines transferable attention (TA) and regularized graph neural network (RGNN). First, topology-aware information aggregation is performed on EEG signals, and complex relationships between channels are dynamically learned. Then, the transferability of each channel is quantified based on the results of the node-wise domain classifier, which is embedded into the emotion classifier as attention score. Importantly, we recruited 10 subjects and collected their EEG signals in pleasure and rage states in simulated driving conditions. We verify the effectiveness of our method on this dataset and compare it with other methods. The results indicate that our method is simple and efficient, with 85.63% accuracy in cross-subject experiments. It can be used to identify road rage. Xueqi Gao, Chao Xu 0003, Yihang Song, Jing Hu 0007, Zhaopeng Meng |
ICASSP | 6 |
| 2023 | ERL-Re$^2$: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy Representation
Jianye Hao, Pengyi Li 0001, Hongyao Tang, Yan Zheng 0002, Xian Fu, Zhaopeng Meng |
ICLR | 6 |
| 2023 | Reining Generalization in Offline Reinforcement Learning via Representation DistinctionabstractOffline Reinforcement Learning (RL) aims to address the challenge of distribution shift between the dataset and the learned policy, where the value of out-of-distribution (OOD) data may be erroneously estimated due to overgeneralization. It has been observed that a considerable portion of the benefits derived from the conservative terms designed by existing offline RL approaches originates from their impact on the learned representation. This observation prompts us to scrutinize the learning dynamics of offline RL, formalize the process of generalization, and delve into the prevalent overgeneralization issue in offline RL. We then investigate the potential to rein the generalization from the representation perspective to enhance offline RL. Finally, we present Representation Distinction (RD), an innovative plug-in method for improving offline RL algorithm performance by explicitly differentiating between the representations of in-sample and OOD state-action pairs generated by the learning policy. Considering scenarios in which the learning policy mirrors the behavioral policy and similar samples may be erroneously distinguished, we suggest a dynamic adjustment mechanism for RD based on an OOD data generator to prevent data representation collapse and further enhance policy performance. We demonstrate the efficacy of our approach by applying RD to specially-designed backbone algorithms and widely-used offline RL algorithms. The proposed RD method significantly improves their performance across various continuous control tasks on D4RL datasets, surpassing several state-of-the-art offline RL algorithms. Yi Ma 0005, Hongyao Tang, Dong Li 0016, Zhaopeng Meng |
NeurIPS | 4 |
| 2023 | On better detecting and leveraging noisy samples for learning with severe label noise
Xiaohe Wu, Chao Xu 0003, Wangmeng Zuo, Zhaopeng Meng |
Pattern Recognit. | 5 |
| 2023 | Delayed rectification of discriminative correlation filters for visual tracking
Chao Xu 0003, Wangmeng Zuo, Zhaopeng Meng |
Vis. Comput. | 5 |
| 2022 | What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorabstractWe study Policy-extended Value Function Approximator (PeVFA) in Reinforcement Learning (RL), which extends conventional value function approximator (VFA) to take as input not only the state (and action) but also an explicit policy representation. Such an extension enables PeVFA to preserve values of multiple policies at the same time and brings an appealing characteristic, i.e., value generalization among policies. We formally analyze the value generalization under Generalized Policy Iteration (GPI). From theoretical and empirical lens, we show that generalized value estimates offered by PeVFA may have lower initial approximation error to true values of successive policies, which is expected to improve consecutive value approximation during GPI. Based on above clues, we introduce a new form of GPI with PeVFA which leverages the value generalization along policy improvement path. Moreover, we propose a representation learning framework for RL policy, providing several approaches to learn effective policy embeddings from policy network parameters or state-action pairs. In our experiments, we evaluate the efficacy of value generalization offered by PeVFA and policy representation learning in several OpenAI Gym continuous control tasks. For a representative instance of algorithm implementation, Proximal Policy Optimization (PPO) re-implemented under the paradigm of GPI with PeVFA achieves about 40% performance improvement on its vanilla counterpart in most environments. Hongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen 0077, Daniel Graves, Dong Li 0016, Changmin Yu, Hangyu Mao, Wulong Liu, Yaodong Yang 0002, Wenyuan Tao |
AAAI | 2 |
| 2022 | HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation
Hongyao Tang, Yan Zheng 0002, Jianye Hao, Pengyi Li 0001, Zhen Wang 0004, Zhaopeng Meng |
ICLR | 7 |
| 2022 | PAnDR: Fast Adaptation to New Environments from Offline Experiences via Decoupling Policy and Environment RepresentationsabstractDeep Reinforcement Learning (DRL) has been a promising solution to many complex decision-making problems. Nevertheless, the notorious weakness in generalization among environments prevent widespread application of DRL agents in real-world scenarios. Although advances have been made recently, most prior works assume sufficient online interaction on training environments, which can be costly in practical cases. To this end, we focus on an offline-training-online-adaptation setting, in which the agent first learns from offline experiences collected in environments with different dynamics and then performs online policy adaptation in environments with new dynamics. In this paper, we propose Policy Adaptation with Decoupled Representations (PAnDR) for fast policy adaptation. In offline training phase, the environment representation and policy representation are learned through contrastive learning and policy recovery, respectively. The representations are further refined by mutual information optimization to make them more decoupled and complete. With learned representations, a Policy-Dynamics Value Function (PDVF) network is trained to approximate the values for different combinations of policies and environments from offline experiences. In online adaptation phase, with the environment context inferred from few experiences collected in new environments, the policy is optimized by gradient ascent with respect to the PDVF. Our experiments show that PAnDR outperforms existing algorithms in several representative policy adaptation problems. Tong Sang, Hongyao Tang, Yi Ma 0005, Jianye Hao, Yan Zheng 0002, Zhaopeng Meng, Zhen Wang 0004 |
IJCAI | 6 |
| 2022 | Semi-supervised Histological Image Segmentation via Hierarchical Consistency Enforcement
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leyi Wei, Zhenyu Fang, Zhaopeng Meng, Ran Su |
MICCAI (2) | 7 |
| 2022 | Multi-view Stereo Network with Attention Thin Volume
Zihang Wan, Chao Xu 0003, Jing Hu 0007, Zhaopeng Meng, Jitai Chen |
PRICAI (3) | 5 |
| 2021 | Addressing Action Oscillations through Learning Policy InertiaabstractDeep reinforcement learning (DRL) algorithms have been demonstrated to be effective on a wide range of challenging decision making and control tasks. However, these methods typically suffer from severe action oscillations in particular in discrete action setting, which means that agents select different actions within consecutive steps even though states only slightly differ. This issue is often neglected since we usually evaluate the quality of a policy using cumulative rewards only. Action oscillation strongly affects the user experience and even causes serious potential security menace especially in real-world domains with the main concern of safety, such as autonomous driving. In this paper, we introduce Policy Inertia Controller (PIC) which serves as a generic plug-in framework to off-the-shelf DRL algorithms, to enable adaptive balance between the optimality and smoothness in a formal way. We propose Nested Policy Iteration as a general training algorithm for PIC-augmented policy which ensures monotonically non-decreasing updates.Further, we derive a practical DRL algorithm, namely Nested Soft Actor-Critic. Experiments on a collection of autonomous driving tasks and several Atari games suggest that our approach demonstrates substantial oscillation reduction than a range of commonly adopted baselines with almost no performance degradation. Chen Chen 0077, Hongyao Tang, Jianye Hao, Wulong Liu, Zhaopeng Meng |
AAAI | 5 |
| 2021 | Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionabstractValue function is the central notion of Reinforcement Learning (RL). Value estimation, especially with function approximation, can be challenging since it involves the stochasticity of environmental dynamics and reward signals that can be sparse and delayed in some cases. A typical model-free RL algorithm usually estimates the values of a policy by Temporal Difference (TD) or Monte Carlo (MC) algorithms directly from rewards, without explicitly taking dynamics into consideration. In this paper, we propose Value Decomposition with Future Prediction (VDFP), providing an explicit two-step understanding of the value estimation process: 1) first foresee the latent future, 2) and then evaluate it. We analytically decompose the value function into a latent future dynamics part and a policy-independent trajectory return part, inducing a way to model latent dynamics and returns separately in value estimation. Further, we derive a practical deep RL algorithm, consisting of a convolutional model to learn compact trajectory representation from past experiences, a conditional variational auto-encoder to predict the latent future dynamics and a convex return model that evaluates trajectory representation. In experiments, we empirically demonstrate the effectiveness of our approach for both off-policy and on-policy RL in several OpenAI Gym continuous control tasks as well as a few challenging variants with delayed reward. Hongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen 0003, Chen Chen 0077, Yaodong Yang 0002, Luo Zhang 0002, Wulong Liu, Jianye Hao |
AAAI | 2 |
| 2021 | Uncertainty-Aware Low-Rank Q-Matrix Estimation for Deep Reinforcement Learning
Tong Sang, Hongyao Tang, Jianye Hao, Yan Zheng 0002, Zhaopeng Meng |
DAI | 5 |
| 2021 | L-DPSNet: Deep Photometric Stereo Network via Local Diffuse Reflectance Maxima
Kanghui Zeng, Chao Xu 0003, Jing Hu 0007, Yushi Li, Zhaopeng Meng |
ICONIP (5) | 5 |
| 2021 | A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsabstractThe Dynamic Pickup and Delivery Problem (DPDP) is an essential problem in the logistics domain, which is NP-hard. The objective is to dynamically schedule vehicles among multiple sites to serve the online generated orders such that the overall transportation cost could be minimized. The critical challenge of DPDP is the orders are not known a priori, i.e., the orders are dynamically generated in real-time. To address this problem, existing methods partition the overall DPDP into fixed-size sub-problems by caching online generated orders and solve each sub-problem, or on this basis to utilize the predicted future orders to optimize each sub-problem further. However, the solution quality and efficiency of these methods are unsatisfactory, especially when the problem scale is very large. In this paper, we propose a novel hierarchical optimization framework to better solve large-scale DPDPs. Specifically, we design an upper-level agent to dynamically partition the DPDP into a series of sub-problems with different scales to optimize vehicles routes towards globally better solutions. Besides, a lower-level agent is designed to efficiently solve each sub-problem by incorporating the strengths of classical operational research-based methods with reinforcement learning-based policies. To verify the effectiveness of the proposed framework, real historical data is collected from the order dispatching system of Huawei Supply Chain Business Unit and used to build a functional simulator. Extensive offline simulation and online testing conducted on the industrial order dispatching system justify the superior performance of our framework over existing baselines. Yi Ma 0005, Xiaotian Hao, Jianye Hao, Xialiang Tong, Mingxuan Yuan, Jie Tang 0001, Zhaopeng Meng |
NeurIPS | 10 |
| 2021 | An Efficient Transfer Learning Framework for Multiagent Reinforcement LearningabstractTransfer Learning has shown great potential to enhance single-agent Reinforcement Learning (RL) efficiency. Similarly, Multiagent RL (MARL) can also be accelerated if agents can share knowledge with each other. However, it remains a problem of how an agent should learn from other agents. In this paper, we propose a novel Multiagent Policy Transfer Framework (MAPTF) to improve MARL efficiency. MAPTF learns which agent's policy is the best to reuse for each agent and when to terminate it by modeling multiagent policy transfer as the option learning problem. Furthermore, in practice, the option module can only collect all agent's local experiences for update due to the partial observability of the environment. While in this setting, each agent's experience may be inconsistent with each other, which may cause the inaccuracy and oscillation of the option-value's estimation. Therefore, we propose a novel option learning algorithm, the successor representation option learning to solve it by decoupling the environment dynamics from rewards and learning the option-value under each agent's preference. MAPTF can be easily combined with existing deep RL and MARL approaches, and experimental results show it significantly boosts the performance of existing methods in both discrete and continuous state spaces. Tianpei Yang, Weixun Wang, Hongyao Tang, Jianye Hao, Zhaopeng Meng, Hangyu Mao, Dong Li 0016, Wulong Liu, Yujing Hu, Changjie Fan, Chengwei Zhang 0001 |
NeurIPS | 5 |
| 2021 | Efficient policy detecting and reusing for non-stationarity in Markov games
Yan Zheng 0002, Jianye Hao, Zongzhang Zhang, Zhaopeng Meng, Tianpei Yang, Yanran Li, Changjie Fan |
Auton. Agents Multi Agent Syst. | 4 |
| 2021 | Robust resistance to noise and outliers: Screened Poisson Surface Reconstruction using adaptive kernel density estimation
Chao Xu 0003, Jing Hu 0007, Zhaopeng Meng |
Comput. Graph. | 4 |
| 2021 | Domain adaptation based self-correction model for COVID-19 infection segmentation in CT images
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Leyi Wei, Ran Su |
Expert Syst. Appl. | 4 |
| 2021 | Free-form tumor synthesis in computed tomography images via richer generative adversarial network
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Ran Su |
Knowl. Based Syst. | 4 |
| 2020 | Continuous Multiagent Control Using Collective Behavior Entropy for Large-Scale Home Energy ManagementabstractWith the increasing popularity of electric vehicles, distributed energy generation and storage facilities in smart grid systems, an efficient Demand-Side Management (DSM) is urgent for energy savings and peak loads reduction. Traditional DSM works focusing on optimizing the energy activities for a single household can not scale up to large-scale home energy management problems. Multi-agent Deep Reinforcement Learning (MA-DRL) shows a potential way to solve the problem of scalability, where modern homes interact together to reduce energy consumers consumption while striking a balance between energy cost and peak loads reduction. However, it is difficult to solve such an environment with the non-stationarity, and existing MA-DRL approaches cannot effectively give incentives for expected group behavior. In this paper, we propose a collective MA-DRL algorithm with continuous action space to provide fine-grained control on a large scale microgrid. To mitigate the non-stationarity of the microgrid environment, a novel predictive model is proposed to measure the collective market behavior. Besides, a collective behavior entropy is introduced to reduce the high peak loads incurred by the collective behaviors of all householders in the smart grid. Empirical results show that our approach significantly outperforms the state-of-the-art methods regarding power cost reduction and daily peak loads optimization. Yan Zheng 0002, Jianye Hao, Zhaopeng Meng, Yang Liu 0003 |
AAAI | 4 |
| 2020 | Generating Behavior-Diverse Game AIs with Evolutionary Multi-Objective Deep Reinforcement LearningabstractGenerating diverse behaviors for game artificial intelligence (Game AI) has been long recognized as a challenging task in the game industry. Designing a Game AI with a satisfying behavioral characteristic (style) heavily depends on the domain knowledge and is hard to achieve manually. Deep reinforcement learning sheds light on advancing the automatic Game AI design. However, most of them focus on creating a superhuman Game AI, ignoring the importance of behavioral diversity in games. To bridge the gap, we introduce a new framework, named EMOGI, which can automatically generate desirable styles with almost no domain knowledge. More importantly, EMOGI succeeds in creating a range of diverse styles, providing behavior-diverse Game AIs. Evaluations on the Atari and real commercial games indicate that, compared to existing algorithms, EMOGI performs better in generating diverse behaviors and significantly improves the efficiency of Game AI design. Ruimin Shen, Yan Zheng 0002, Jianye Hao, Zhaopeng Meng, Changjie Fan, Yang Liu 0003 |
IJCAI | 4 |
| 2020 | Efficient Deep Reinforcement Learning via Adaptive Policy TransferabstractTransfer learning has shown great potential to accelerate Reinforcement Learning (RL) by leveraging prior knowledge from past learned policies of relevant tasks. Existing approaches either transfer previous knowledge by explicitly computing similarities between tasks or select appropriate source policies to provide guided explorations. However, how to directly optimize the target policy by alternatively utilizing knowledge from appropriate source policies without explicitly measuring the similarities is currently missing. In this paper, we propose a novel Policy Transfer Framework (PTF) by taking advantage of this idea. PTF learns when and which source policy is the best to reuse for the target policy and when to terminate it by modeling multi-policy transfer as an option learning problem. PTF can be easily combined with existing DRL methods and experimental results show it significantly accelerates RL and surpasses state-of-the-art policy transfer methods in terms of learning efficiency and final performance in both discrete and continuous action spaces. Tianpei Yang, Jianye Hao, Zhaopeng Meng, Zongzhang Zhang, Yujing Hu, Changjie Fan, Weixun Wang, Wulong Liu, Zhaodong Wang, Jiajie Peng |
IJCAI | 3 |
| 2020 | Fuzzy dynamic timetable scheduling for public transit
Qinghua Hu, Zhaopeng Meng, Anca L. Ralescu |
Fuzzy Sets Syst. | 3 |
| 2020 | Efficient Multiagent Policy Optimization Based on Weighted Estimators in Stochastic Cooperative Environments
Yan Zheng 0002, Jianye Hao, Zongzhang Zhang, Zhaopeng Meng, Xiaotian Hao |
J. Comput. Sci. Technol. | 4 |
| 2020 | Construction of Retinal Vessel Segmentation Models Based on Convolutional Neural Network
Qiangguo Jin, Zhaopeng Meng, Ran Su |
Neural Process. Lett. | 3 |
| 2019 | Towards Efficient Detection and Optimal Response against Sophisticated OpponentsabstractMultiagent algorithms often aim to accurately predict the behaviors of other agents and find a best response accordingly. Previous works usually assume an opponent uses a stationary strategy or randomly switches among several stationary ones. However, an opponent may exhibit more sophisticated behaviors by adopting more advanced reasoning strategies, e.g., using a Bayesian reasoning strategy. This paper proposes a novel approach called Bayes-ToMoP which can efficiently detect the strategy of opponents using either stationary or higher-level reasoning strategies. Bayes-ToMoP also supports the detection of previously unseen policies and learning a best-response policy accordingly. We provide a theoretical guarantee of the optimality on detecting the opponent's strategies. We also propose a deep version of Bayes-ToMoP by extending Bayes-ToMoP with DRL techniques. Experimental results show both Bayes-ToMoP and deep Bayes-ToMoP outperform the state-of-the-art approaches when faced with different types of opponents in two-agent competitive games. Tianpei Yang, Jianye Hao, Zhaopeng Meng, Chongjie Zhang, Yan Zheng 0002, Ze Zheng |
IJCAI | 3 |
| 2019 | Wuji: Automatic Online Combat Game Testing Using Evolutionary Deep Reinforcement LearningabstractGame testing has been long recognized as a notoriously challenging task, which mainly relies on manual playing and scripting based testing in game industry. Even until recently, automated game testing still remains to be largely untouched niche. A key challenge is that game testing often requires to play the game as a sequential decision process. A bug may only be triggered until completing certain difficult intermediate tasks, which requires a certain level of intelligence. The recent success of deep reinforcement learning (DRL) sheds light on advancing automated game testing, without human competitive intelligent support. However, the existing DRLs mostly focus on winning the game rather than game testing. To bridge the gap, in this paper, we first perform an in-depth analysis of 1349 real bugs from four real-world commercial game products. Based on this, we propose four oracles to support automated game testing, and further propose Wuji, an on-the-fly game testing framework, which leverages evolutionary algorithms, DRL and multi-objective optimization to perform automatic game testing. Wuji balances between winning the game and exploring the space of the game. Winning the game allows the agent to make progress in the game, while space exploration increases the possibility of discovering bugs. We conduct a large-scale evaluation on a simple game and two popular commercial games. The results demonstrate the effectiveness of Wuji in exploring space and detecting bugs. Moreover, Wuji found 3 previously unknown bugs, which have been confirmed by the developers, in the commercial games. Yan Zheng 0002, Changjie Fan, Xiaofei Xie, Ting Su 0001, Lei Ma 0003, Jianye Hao, Zhaopeng Meng, Yang Liu 0003, Ruimin Shen |
ASE | 7 |
| 2019 | LoopFix: an approach to automatic repair of buggy loops
Weichao Wang, Zhaopeng Meng, Shuang Liu 0007, Jianye Hao |
J. Syst. Softw. | 2 |
| 2019 | DUNet: A deformable network for retinal vessel segmentation
Qiangguo Jin, Zhaopeng Meng, Tuan D. Pham, Leyi Wei, Ran Su |
Knowl. Based Syst. | 2 |
| 2018 | A Deep Bayesian Policy Reuse Approach Against Non-Stationary AgentsabstractIn multiagent domains, coping with non-stationary agents that change behaviors from time to time is a challenging problem, where an agent is usually required to be able to quickly detect the other agent's policy during online interaction, and then adapt its own policy accordingly. This paper studies efficient policy detecting and reusing techniques when playing against non-stationary agents in Markov games. We propose a new deep BPR+ algorithm by extending the recent BPR+ algorithm with a neural network as the value-function approximator. To detect policy accurately, we propose the \textit{rectified belief model} taking advantage of the \textit{opponent model} to infer the other agent's policy from reward signals and its behaviors. Instead of directly storing individual policies as BPR+, we introduce \textit{distilled policy network} that serves as the policy library in BPR+, using policy distillation to achieve efficient online policy learning and reuse. Deep BPR+ inherits all the advantages of BPR+ and empirically shows better performance in terms of detection accuracy, cumulative rewards and speed of convergence compared to existing algorithms in complex Markov games with raw visual inputs. Yan Zheng 0002, Zhaopeng Meng, Jianye Hao, Zongzhang Zhang, Tianpei Yang, Changjie Fan |
NeurIPS | 2 |
| 2018 | Weighted Double Deep Multiagent Reinforcement Learning in Stochastic Cooperative Environments
Yan Zheng 0002, Zhaopeng Meng, Jianye Hao, Zongzhang Zhang |
PRICAI | 2 |
| 2016 | Accelerating Norm Emergence Through Hierarchical Heuristic LearningabstractSocial norms serve as an important mechanism to regulate the behaviours of agents and to facilitate coordination among them in multiagent systems. One important research question is how a norm can rapidly emerge through repeated local interaction within agent societies under different environments when their coordination space becomes large. To address this problem, we propose a hierarchically heuristic learning strategy (HHLS) under the hierarchical social learning framework. Subordinate agents report their information to their supervisors, while supervisors can generate instructions (rules and suggestions) based on the information collected from their subordinates. Subordinate agents heuristically update their strategies based on both their own experience and the instructions from their supervisors. Extensive experiment evaluations show that HHLS can support the emergence of desirable social norms more efficiently and can be applicable in a much wider range of multiagent interaction scenarios compared with previous work. The influence of key related factors (e.g., different topologies, population, neighbourhood and action space size, cluster size) are also investigated and new insights are obtained as well. Tianpei Yang, Zhaopeng Meng, Jianye Hao, Sandip Sen, Chao Yu 0004 |
ECAI | 2 |
| 2016 | Affective experience modeling based on interactive synergetic dependence in big data
Chao Xu 0003, Zhiyong Feng 0002, Zhaopeng Meng |
Future Gener. Comput. Syst. | 3 |
| 2015 | An Interactive Radial Visualization of Geoscience Observation DataabstractGeoscience observation data refers to the datasets consisting of time series of multiple parameters generated from the sensors at fixed locations. Although a few works have attempted to visualize features of these data, none of them views these data as a specific type and attempts to show the overview in all the space, time and attribute aspects. It is important for domain experts to select interested subsets from huge amounts of observation data according to the high level patterns shown in the overview. We present a novel approach to visualizing geoscience observation data in a compact radial view. Our solution consists of three visual elements. A map showing the spatial aspect is in the center of the visualization, while temporal and attribute aspects are seamlessly combined with the spatial information. Our approach is equipped with interactive mechanisms for highlighting the selected features, adjusting the display range, as well as interactively generating a fisheye view. We demonstrate the effectiveness and usability of our approach with a usability experiment. Eye tracking records and user feedbacks obtained in the experiment prove the effectiveness of our approach. Jie Li 0006, Zhaopeng Meng, Mao Lin Huang, Kang Zhang 0001 |
VINCI | 2 |
| 2015 | A Short-Text Oriented Clustering Method for Hot Topics ExtractionabstractA major challenge in document clustering is the extremely high dimensionality as well as the sparsity of the sample matrix. In this paper, we propose a new short-text oriented analysis approach to cluster short text automatically and extract the hot topics from each cluster. Different from the previous studies focused on long text, our analysis approach mainly focused on short-text cases. The approach consists of three stages: Firstly, generate feature vector for each sample so as to obtain the whole high-dimensional Vector Space Model; Secondly, use Singular Value Decomposition to achieve the dimensions reduction; Lastly, apply cosine similarity and k-means method to cluster samples on the low-dimensional matrix and extract the hot topics for each cluster. The experimental results show that our analysis approach can deal with the short-text samples and find out the hot topics efficiently and effectively. Yan Zheng 0002, Zhaopeng Meng, Chao Xu 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2015 | Robust visual tracking via online multiple instance learning with Fisher information
Chao Xu 0003, Wenyuan Tao, Zhaopeng Meng, Zhiyong Feng 0002 |
Pattern Recognit. | 3 |
| 2015 | Layered modeling and generation of Pollock's drip style
Yan Zheng 0002, Xuecheng Nie, Zhaopeng Meng, Wei Feng 0005, Kang Zhang 0001 |
Vis. Comput. | 3 |
| 2010 | A New Closeness Metric for Social Networks Based on the k Shortest Paths
Chun Shang, Yuexian Hou, Zhaopeng Meng |
ISNN (2) | 4 |