VLDB 2026 Research / reviewers in the wild / expert
Wei Yang 0032
dblp:03/1094-32
· DBLP profile ↗
34ranked-venue papers
0as first author
31since 2021 · last 2025
0000-0002-6488-2546ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive ModelingabstractWe present CHARM, a novel parametric representation and generative framework for anime hairstyle modeling. While traditional hair modeling methods focus on realistic hair using strand-based or volumetric representations, anime hairstyle exhibits highly stylized, piecewise-structured geometry that challenges existing techniques. Existing works often rely on dense mesh modeling or hand-crafted spline curves, making them inefficient for editing and unsuitable for scalable learning. CHARM introduces a compact, invertible control-point-based parameterization, where a sequence of control points represents each hair card, and each point is encoded with only five geometric parameters. This efficient and accurate representation supports both artist-friendly design and learning-based generation. Built upon this representation, CHARM introduces an autoregressive generative framework that effectively generates anime hairstyles from input images or point clouds. By interpreting anime hairstyles as a sequential “hair language”, our autoregressive transformer captures both local geometry and global hairstyle topology, resulting in high-fidelity anime hairstyle creation. To facilitate both training and evaluation of anime hairstyle generation, we construct AnimeHair, a large-scale dataset of 37K high-quality anime hairstyles with separated hair cards and processed mesh data. Extensive experiments demonstrate state-of-the-art performance of CHARM in both reconstruction accuracy and generation quality, offering an expressive and scalable solution for anime hairstyle modeling. Project page: https://hyzcluster.github.io/charm Yanning Zhou 0003, Wang Zhao 0001, Jingwen Ye, Yushi Bai, Kaiwen Xiao, Yong-Jin Liu 0001, Zhongqian Sun, Wei Yang 0032 |
SIGGRAPH Asia | 9 |
| 2024 | Multi-agent Multi-game Entity Transformer: Towards Generalist Models in MARLabstractBuilding large-scale generalist pre-trained models for many tasks is becoming an emerging and potential direction in reinforcement learning (RL).Research such as Gato and Multi-Game Decision Transformer have displayed outstanding performance and generalization capabilities on many games and domains.However, there exists a research blank about developing highly capable and generalist models in multi-agent RL (MARL), which can substantially accelerate progress toward general AI.To fill this gap, we propose Multi-Agent multi-Game ENtity TrAnsformer (MA-GENTA) from the entity perspective as orthogonal research to previous time-sequential modeling.Specifically, to deal with different state/observation spaces in different games, we analogize games as languages by aligning one single game to one single language, thus training different "tokenizers" and a shared transformer for various games.The feature inputs are split according to different entities and tokenized in the same continuous space.Then, two types of transformer-based models are proposed as permutationinvariant architectures to deal with various numbers of entities and capture the attention of different entities.MAGENTA is trained on Xianhan Zeng, Liang Wang 0015, Zhengjie Liang, Yiming Gao 0007, Feiyu Liu, Siqin Li, Xianliang Wang, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Longtao Zheng, Zinovi Rabinovich, Bo An 0001 |
DAI | 11 |
| 2024 | Enhancing Human Experience in Human-Agent Collaboration: A Human-Centered Modeling Approach Based on Positive Human GainabstractExisting game AI research mainly focuses on enhancing agents' abilities to win games, but this does not inherently make humans have a better experience when collaborating with these agents. For example, agents may dominate the collaboration and exhibit unintended or detrimental behaviors, leading to poor experiences for their human partners. In other words, most game AI agents are modeled in a "self-centered" manner. In this paper, we propose a "human-centered" modeling scheme for collaborative agents that aims to enhance the experience of humans. Specifically, we model the experience of humans as the goals they expect to achieve during the task. We expect that agents should learn to enhance the extent to which humans achieve these goals while maintaining agents' original abilities (e.g., winning games). To achieve this, we propose the Reinforcement Learning from Human Gain (RLHG) approach. The RLHG approach introduces a "baseline", which corresponds to the extent to which humans primitively achieve their goals, and encourages agents to learn behaviors that can effectively enhance humans in achieving their goals better. We evaluate the RLHG agent in the popular Multi-player Online Battle Arena (MOBA) game, Honor of Kings, by conducting real-world human-agent tests. Both objective performance and subjective preference results show that the RLHG agent provides participants better gaming experience. Yiming Gao 0007, Feiyu Liu, Liang Wang 0015, Dehua Zheng, Zhenjie Lian, Siqin Li, Xianliang Wang, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Wei Liu 0005 |
ICLR | 13 |
| 2024 | Dynamics-Adaptive Continual Reinforcement Learning via Progressive ContextualizationabstractA key challenge of continual reinforcement learning (CRL) in dynamic environments is to promptly adapt the reinforcement learning (RL) agent's behavior as the environment changes over its lifetime while minimizing the catastrophic forgetting of the learned information. To address this challenge, in this article, we propose DaCoRL, that is, dynamics-adaptive continual RL. DaCoRL learns a context-conditioned policy using progressive contextualization, which incrementally clusters a stream of stationary tasks in the dynamic environment into a series of contexts and opts for an expandable multihead neural network to approximate the policy. Specifically, we define a set of tasks with similar dynamics as an environmental context and formalize context inference as a procedure of online Bayesian infinite Gaussian mixture clustering on environment features, resorting to online Bayesian inference to infer the posterior distribution over contexts. Under the assumption of a Chinese restaurant process (CRP) prior, this technique can accurately classify the current task as a previously seen context or instantiate a new context as needed without relying on any external indicator to signal environmental changes in advance. Furthermore, we employ an expandable multihead neural network whose output layer is synchronously expanded with the newly instantiated context and a knowledge distillation regularization term for retaining the performance on learned tasks. As a general framework that can be coupled with various deep RL algorithms, DaCoRL features consistent superiority over existing methods in terms of stability, overall performance, and generalization ability, as verified by extensive experiments on several robot navigation and MuJoCo locomotion tasks. Tiantian Zhang 0002, Zichuan Lin, Deheng Ye, Qiang Fu 0016, Wei Yang 0032, Xueqian Wang 0001, Bin Liang 0001, Bo Yuan 0003, Xiu Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement LearningabstractWhole-slide images (WSI) in computational pathology have high resolution with gigapixel size, but are generally with sparse regions of interest, which leads to weak diagnostic relevance and data inefficiency for each area in the slide. Most of the existing methods rely on a multiple instance learning framework that requires densely sampling local patches at high magnification. The limitation is evident in the application stage as the heavy computation for extracting patch-level features is inevitable. In this paper, we develop RLogist, a benchmarking deep reinforcement learning (DRL) method for fast observation strategy on WSIs. Imitating the diagnostic logic of human pathologists, our RL agent learns how to find regions of observation value and obtain representative features across multiple resolution levels, without having to analyze each part of the WSI at the high magnification. We benchmark our method on two whole-slide level classification tasks, including detection of metastases in WSIs of lymph node sections, and subtyping of lung cancer. Experimental results demonstrate that RLogist achieves competitive classification performance compared to typical multiple instance learning algorithms, while having a significantly short observation path. In addition, the observation path given by RLogist provides good decision-making interpretability, and its ability of reading path navigation can potentially be used by pathologists for educational/assistive purposes. Our code is available at: https://github.com/tencent-ailab/RLogist. Boxuan Zhao, Jun Zhang 0018, Deheng Ye, Jian Cao 0001, Xiao Han 0011, Qiang Fu 0016, Wei Yang 0032 |
AAAI | 7 |
| 2023 | ASM: Adaptive Skinning Model for High-Quality 3D Face ModelingabstractThe research fields of parametric face model and 3D face reconstruction have been extensively studied. However, a critical question remains unanswered: how to tailor the face model for specific reconstruction settings. We argue that reconstruction with multi-view uncalibrated images demands a new model with stronger capacity. Our study shifts attention from data-dependent 3D Morphable Models (3DMM) to an understudied human-designed skinning model. We propose Adaptive Skinning Model (ASM), which redefines the skinning model with more compact and fully tunable parameters. With extensive experiments, we demonstrate that ASM achieves significantly improved capacity than 3DMM, with the additional advantage of model size and easy implementation for new topology. We achieve state-of-the-art performance with ASM for multi-view reconstruction on the Florence MICC Coop benchmark. Our quantitative analysis demonstrates the importance of a high-capacity model for fully exploiting abundant information from multi-view input in reconstruction. Furthermore, our model with physical-semantic parameters can be directly utilized for real-world applications, such as in-game avatar creation. As a result, our work opens up new research direction for parametric face model and facilitates future research on multi-view reconstruction. Hong Shang, Tianyang Shi, Xinghan Chen, Jingkai Zhou, Zhongqian Sun, Wei Yang 0032 |
ICCV | 7 |
| 2023 | Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication Perspective
Yiming Gao 0007, Feiyu Liu, Liang Wang 0015, Zhenjie Lian, Siqin Li, Xianliang Wang, Xianhan Zeng, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Wei Liu 0005 |
ICLR | 12 |
| 2023 | Quality-Similar Diversity via Population Based Reinforcement Learning
Jian Yao 0008, Haobo Fu, Chao Qian 0001, Yaodong Yang 0001, Qiang Fu 0016, Wei Yang 0032 |
ICLR | 8 |
| 2023 | Opponent-Limited Online Search for Imperfect Information GamesabstractIn recent years, online search has been playing an increasingly important role in imperfect information games (IIGs). Previous online search is known as common-knowledge subgame solving, which has to consider all the states in a common-knowledge closure. This is only computationally tolerable for medium size games, such as poker. To handle larger games, order-1 Knowledge-Limited Subgame Solving (1-KLSS) only considers the states in a knowledge-limited closure, which results in a much smaller subgame. However, 1-KLSS is unsafe. In this paper, we first extend 1-KLSS to Safe-1-KLSS and prove its safeness. To make Safe-1-KLSS applicable to even larger games, we propose Opponent-Limited Subgame Solving (OLSS) to limit how the opponent reaches a subgame and how it acts in the subgame. Limiting the opponent's strategy dramatically reduces the subgame size and improves the efficiency of subgame solving while still preserving some safety in the limit. Experiments in medium size poker show that Safe-1-KLSS and OLSS are orders of magnitude faster than previous common-knowledge subgame solving. Also, OLSS significantly improves the online performance in a two-player Mahjong game, whose game size prohibits the use of previous common-knowledge subgame-solving methods. Weiming Liu 0004, Haobo Fu, Qiang Fu 0016, Wei Yang 0032 |
ICML | 4 |
| 2023 | Dynamic Low-Rank Instance Adaptation for Universal Neural Image Compression
Yue Lv, Jinxi Xiang, Jun Zhang 0018, Wenming Yang, Xiao Han 0011, Wei Yang 0032 |
ACM Multimedia | 6 |
| 2023 | Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration InformationabstractThe state-of-the-art neural video codecs have outperformed the most sophisticated traditional codecs in terms of rate-distortion (RD) performance in certain cases. However, utilizing them for practical applications is still challenging for two major reasons. 1) Cross-platform computational errors resulting from floating point operations can lead to inaccurate decoding of the bitstream. 2) The high computational complexity of the encoding and decoding process poses a challenge in achieving real-time performance. In this paper, we propose a real-time cross-platform neural video codec, which is capable of efficiently decoding (25FPS) of 720P video bitstream from other encoding platforms on a consumer-grade GPU (e.g., NVIDIA RTX 2080). First, to solve the problem of inconsistency of codec caused by the uncertainty of floating point calculations across platforms, we design a calibration transmitting system to guarantee the consistent quantization of entropy parameters between the encoding and decoding stages. The parameters that may have transboundary quantization between encoding and decoding are identified in the encoding stage, and their coordinates will be delivered by auxiliary transmitted bitstream. By doing so, these inconsistent parameters can be processed properly in the decoding stage. Furthermore, to reduce the bitrate of the auxiliary bitstream, we rectify the distribution of entropy parameters using a piecewise Gaussian constraint. Second, to match the computational limitations on the decoding side for real-time video codec, we design a lightweight model. A series of efficiency techniques, such as model pruning, motion downsampling, and arithmetic coding skipping, enable our model to achieve 25 FPS decoding speed on NVIDIA RTX 2080 GPU. Experimental results demonstrate that our model can achieve real-time decoding of 720P videos while encoding on another platform. Furthermore, the real-time model brings up to a maximum of 24.2% BD-rate improvement from the perspective of PSNR with the anchor H.265 (medium). Kuan Tian, Yonghang Guan, Jinxi Xiang, Jun Zhang 0018, Xiao Han 0011, Wei Yang 0032 |
ACM Multimedia | 6 |
| 2023 | A Robust and Opponent-Aware League Training Method for StarCraft IIabstractIt is extremely difficult to train a superhuman Artificial Intelligence (AI) for games of similar size to StarCraft II. AlphaStar is the first AI that beat human professionals in the full game of StarCraft II, using a league training framework that is inspired by a game-theoretic approach. In this paper, we improve AlphaStar's league training in two significant aspects. We train goal-conditioned exploiters, whose abilities of spotting weaknesses in the main agent and the entire league are greatly improved compared to the unconditioned exploiters in AlphaStar. In addition, we endow the agents in the league with the new ability of opponent modeling, which makes the agent more responsive to the opponent's real-time strategy. Based on these improvements, we train a better and superhuman AI with orders of magnitude less resources than AlphaStar (see Table 1 for a full comparison). Considering the iconic role of StarCraft II in game AI research, we believe our method and results on StarCraft II provide valuable design principles on how one would utilize the general league training framework for obtaining a least-exploitable strategy in various, large-scale, real-world games. Ruozi Huang, Xipeng Wu, Hongsheng Yu, Zhong Fan, Haobo Fu, Qiang Fu 0016, Wei Yang 0032 |
NeurIPS | 7 |
| 2023 | Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning BenchmarksabstractThe advancement of Offline Reinforcement Learning (RL) and Offline Multi-Agent Reinforcement Learning (MARL) critically depends on the availability of high-quality, pre-collected offline datasets that represent real-world complexities and practical applications. However, existing datasets often fall short in their simplicity and lack of realism. To address this gap, we propose Hokoff, a comprehensive set of pre-collected datasets that covers both offline RL and offline MARL, accompanied by a robust framework, to facilitate further research. This data is derived from Honor of Kings, a recognized Multiplayer Online Battle Arena (MOBA) game known for its intricate nature, closely resembling real-life situations. Utilizing this framework, we benchmark a variety of offline RL and offline MARL algorithms. We also introduce a novel baseline algorithm tailored for the inherent hierarchical action space of the game. We reveal the incompetency of current offline RL approaches in handling task complexity, generalization and multi-task learning. Yun Qu 0002, Jianzhun Shao, Yuhang Jiang 0001, Zhenbin Ye, Lin Lai, Hongyang Qin, Minwen Deng, Juchao Zhuo, Deheng Ye, Qiang Fu 0016, Yang Guang, Wei Yang 0032, Lanxiao Huang, Xiangyang Ji |
NeurIPS | 16 |
| 2023 | Policy Space Diversity for Non-Transitive GamesabstractPolicy-Space Response Oracles (PSRO) is an influential algorithm framework for approximating a Nash Equilibrium (NE) in multi-agent non-transitive games. Many previous studies have been trying to promote policy diversity in PSRO. A major weakness with existing diversity metrics is that a more diverse (according to their diversity metrics) population does not necessarily mean (as we proved in the paper) a better approximation to a NE. To alleviate this problem, we propose a new diversity metric, the improvement of which guarantees a better approximation to a NE. Meanwhile, we develop a practical and well-justified method to optimize our diversity metric using only state-action samples. By incorporating our diversity regularization into the best response solving of PSRO, we obtain a new PSRO variant, \textit{Policy Space Diversity} PSRO (PSD-PSRO). We present the convergence property of PSD-PSRO. Empirically, extensive experiments on single-state games, Leduc, and Goofspiel demonstrate that PSD-PSRO is more effective in producing significantly less exploitable policies than state-of-the-art PSRO variants. Jian Yao 0008, Weiming Liu 0004, Haobo Fu, Yaodong Yang 0001, Stephen McAleer, Qiang Fu 0016, Wei Yang 0032 |
NeurIPS | 7 |
| 2023 | RetCCL: Clustering-guided contrastive learning for whole-slide image retrieval
Yuexi Du, Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
Medical Image Anal. | 7 |
| 2023 | A generalizable and robust deep learning algorithm for mitosis detection in multicenter breast histopathological images
Jun Zhang 0018, Sen Yang 0006, Jingxi Xiang, Feng Luo 0003, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
Medical Image Anal. | 8 |
| 2023 | Merging nucleus datasets by correlation-based cross-training
Jun Zhang 0018, Sen Yang 0006, Junzhou Huang, Wei Yang 0032, Xiao Han 0011 |
Medical Image Anal. | 6 |
| 2022 | Node-aligned Graph Convolutional Network for Whole-slide Image Representation and ClassificationabstractThe large-scale whole-slide images (WSIs) facilitate the learning-based computational pathology methods. However, the gigapixel size of WSIs makes it hard to train a conventional model directly. Current approaches typically adopt multiple-instance learning (MIL) to tackle this problem. Among them, MIL combined with graph convolutional network (GCN) is a significant branch, where the sampled patches are regarded as the graph nodes to further discover their correlations. However, it is difficult to build correspondence across patches from different WSIs. Therefore, most methods have to perform non-ordered node pooling to generate the bag-level representation. Direct non-ordered pooling will lose much structural and contextual information, such as patch distribution and heterogeneous patterns, which is critical for WSI representation. In this paper, we propose a hierarchical global-to-local clustering strategy to build a Node-Aligned GCN (NAGCN) to represent WSI with rich local structural information as well as global distribution. We first deploy a global clustering operation based on the instance features in the dataset to build the correspondence across different WSIs. Then, we perform a local clustering-based sampling strategy to select typical instances belonging to each cluster within the WSI. Finally, we employ the graph convolution to obtain the representation. Since our graph construction strategy ensures the alignment among different WSIs, WSI-level representation can be easily generated and used for the subsequent classification. The experiment results on two cancer subtype classification datasets demonstrate our method achieves better performance compared with the state-of-the-art methods. Yonghang Guan, Jun Zhang 0018, Kuan Tian, Sen Yang 0006, Pei Dong, Jinxi Xiang, Wei Yang 0032, Junzhou Huang, Yuyao Zhang 0005, Xiao Han 0011 |
CVPR | 7 |
| 2022 | Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game
Haobo Fu, Weiming Liu 0004, Kai Li 0022, Junliang Xing, Bin Li 0025, Qiang Fu 0016, Wei Yang 0032 |
ICLR | 11 |
| 2022 | Greedy when Sure and Conservative when Uncertain about the OpponentsabstractWe develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response (conditional additionally on our opponent policy embedding) instead of a finite set of separate best responses against any opponent; 3) computes online a posterior of the current opponent policy embedding, without making the discrete and ineffective decision which type the current opponent belongs to; and 4) selects online between a real-time greedy policy and a fixed conservative policy via an adversarial bandit algorithm, gaining a theoretically better regret than adhering to either. Experimental studies on popular benchmarks demonstrate GSCU’s superiority over the state-of-the-art methods. The code is available online at \url{https://github.com/YeTianJHU/GSCU}. Haobo Fu, Hongxiang Yu, Weiming Liu 0004, Jiechao Xiong, Ying Wen 0001, Kai Li 0022, Junliang Xing, Qiang Fu 0016, Wei Yang 0032 |
ICML | 11 |
| 2022 | JueWu-MC: Playing Minecraft with Sample-efficient Hierarchical Reinforcement LearningabstractLearning rational behaviors in open-world games like Minecraft remains to be challenging for Reinforcement Learning (RL) research due to the compound challenge of partial observability, high-dimensional visual perception and delayed reward. To address this, we propose JueWu-MC, a sample-efficient hierarchical RL approach equipped with representation learning and imitation learning to deal with perception and exploration. Specifically, our approach includes two levels of hierarchy, where the high-level controller learns a policy to control over options and the low-level workers learn to solve each sub-task. To boost the learning of sub-tasks, we propose a combination of techniques including 1) action-aware representation learning which captures underlying relations between action and representation, 2) discriminator-based self-imitation learning for efficient exploration, and 3) ensemble behavior cloning with consistency filtering for policy robustness. Extensive experiments show that JueWu-MC significantly improves sample efficiency and outperforms a set of baselines by a large margin. Notably, we won the championship of the NeurIPS MineRL 2021 research competition and achieved the highest performance score ever. Zichuan Lin, Junyou Li, Jianing Shi, Deheng Ye, Qiang Fu 0016, Wei Yang 0032 |
IJCAI | 6 |
| 2022 | Honor of Kings Arena: an Environment for Generalization in Competitive Reinforcement LearningabstractThis paper introduces Honor of Kings Arena, a reinforcement learning (RL) environment based on the Honor of Kings, one of the world’s most popular games at present. Compared to other environments studied in most previous work, ours presents new generalization challenges for competitive reinforcement learning. It is a multi-agent problem with one agent competing against its opponent; and it requires the generalization ability as it has diverse targets to control and diverse opponents to compete with. We describe the observation, action, and reward specifications for the Honor of Kings domain and provide an open-source Python-based interface for communicating with the game engine. We provide twenty target heroes with a variety of tasks in Honor of Kings Arena and present initial baseline results for RL-based methods with feasible computing resources. Finally, we showcase the generalization challenges imposed by Honor of Kings Arena and possible remedies to the challenges. All of the software, including the environment-class, are publicly available. Hua Wei 0001, Jingxiao Chen, Xiyang Ji, Hongyang Qin, Minwen Deng, Siqin Li, Liang Wang 0015, Weinan Zhang 0001, Yong Yu 0001, Lanxiao Huang, Deheng Ye, Qiang Fu 0016, Wei Yang 0032 |
NeurIPS | 14 |
| 2022 | SCL-WC: Cross-Slide Contrastive Learning for Weakly-Supervised Whole-Slide Image ClassificationabstractWeakly-supervised whole-slide image (WSI) classification (WSWC) is a challenging task where a large number of unlabeled patches (instances) exist within each WSI (bag) while only a slide label is given. Despite recent progress for the multiple instance learning (MIL)-based WSI analysis, the major limitation is that it usually focuses on the easy-to-distinguish diagnosis-positive regions while ignoring positives that occupy a small ratio in the entire WSI. To obtain more discriminative features, we propose a novel weakly-supervised classification method based on cross-slide contrastive learning (called SCL-WC), which depends on task-agnostic self-supervised feature pre-extraction and task-specific weakly-supervised feature refinement and aggregation for WSI-level prediction. To enable both intra-WSI and inter-WSI information interaction, we propose a positive-negative-aware module (PNM) and a weakly-supervised cross-slide contrastive learning (WSCL) module, respectively. The WSCL aims to pull WSIs with the same disease types closer and push different WSIs away. The PNM aims to facilitate the separation of tumor-like patches and normal ones within each WSI. Extensive experiments demonstrate state-of-the-art performance of our method in three different classification tasks (e.g., over 2% of AUC in Camelyon16, 5% of F1 score in BRACS, and 3% of AUC in DiagSet). Our method also shows superior flexibility and scalability in weakly-supervised localization and semi-supervised classification experiments (e.g., first place in the BRIGHT challenge). Our code will be available at https://github.com/Xiyue-Wang/SCL-WC. Jinxi Xiang, Jun Zhang 0018, Sen Yang 0006, Zhongyi Yang, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
NeurIPS | 8 |
| 2022 | Transformer-based unsupervised contrastive learning for histopathological image classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
Medical Image Anal. | 6 |
| 2022 | Knowledge-Based Representation Learning for Nucleus Instance Classification From Histopathological ImagesabstractThe classification of nuclei in H&E-stained histopathological images is a fundamental step in the quantitative analysis of digital pathology. Most existing methods employ multi-class classification on the detected nucleus instances, while the annotation scale greatly limits their performance. Moreover, they often downplay the contextual information surrounding nucleus instances that is critical for classification. To explicitly provide contextual information to the classification model, we design a new structured input consisting of a content-rich image patch and a target instance mask. The image patch provides rich contextual information, while the target instance mask indicates the location of the instance to be classified and emphasizes its shape. Benefiting from our structured input format, we propose Structured Triplet for representation learning, a triplet learning framework on unlabelled nucleus instances with customized positive and negative sampling strategies. We pre-train a feature extraction model based on this framework with a large-scale unlabeled dataset, making it possible to train an effective classification model with limited annotated data. We also add two auxiliary branches, namely the attribute learning branch and the conventional self-supervised learning branch, to further improve its performance. As part of this work, we will release a new dataset of H&E-stained pathology images with nucleus instance masks, containing 20,187 patches of size 1024 ×1024 , where each patch comes from a different whole-slide image. The model pre-trained on this dataset with our framework significantly reduces the burden of extensive labeling. We show a substantial improvement in nucleus classification accuracy compared with the state-of-the-art methods. Jun Zhang 0018, Sen Yang 0006, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Supervised Learning Achieves Human-Level Performance in MOBA Games: A Case Study of Honor of KingsabstractWe present JueWu-SL, the first supervised-learning-based artificial intelligence (AI) program that achieves human-level performance in playing multiplayer online battle arena (MOBA) games. Unlike prior attempts, we integrate the macro-strategy and the micromanagement of MOBA-game-playing into neural networks in a supervised and end-to-end manner. Tested on Honor of Kings, the most popular MOBA at present, our AI performs competitively at the level of High King players in standard 5v5 games. Deheng Ye, Peilin Zhao, Fuhao Qiu, Bo Yuan 0008, Mingfei Sun 0001, Siqin Li, Zhenjie Lian, Bei Shi, Liang Wang 0015, Tengfei Shi, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang |
IEEE Trans. Neural Networks Learn. Syst. | 17 |
| 2021 | Boosting Offline Reinforcement Learning with Residual Generative ModelingabstractOffline reinforcement learning (RL) tries to learn the near-optimal policy with recorded offline experience without online exploration.Current offline RL research includes: 1) generative modeling, i.e., approximating a policy using fixed data; and 2) learning the state-action value function. While most research focuses on the state-action function part through reducing the bootstrapping error in value function approximation induced by the distribution shift of training data, the effects of error propagation in generative modeling have been neglected. In this paper, we analyze the error in generative modeling. We propose AQL (action-conditioned Q-learning), a residual generative model to reduce policy approximation error for offline RL. We show that our method can learn more accurate policy approximations in different benchmark datasets. In addition, we show that the proposed offline RL method can learn more competitive AI agents in complex control tasks under the multiplayer online battle arena (MOBA) game, Honor of Kings. Hua Wei 0001, Deheng Ye, Bo Yuan 0008, Qiang Fu 0016, Wei Yang 0032, Zhenhui Li |
IJCAI | 7 |
| 2021 | MapGo: Model-Assisted Policy Optimization for Goal-Oriented TasksabstractIn Goal-oriented Reinforcement learning, relabeling the raw goals in past experience to provide agents with hindsight ability is a major solution to the reward sparsity problem. In this paper, to enhance the diversity of relabeled goals, we develop FGI (Foresight Goal Inference), a new relabeling strategy that relabels the goals by looking into the future with a learned dynamics model. Besides, to improve sample efficiency, we propose to use the dynamics model to generate simulated trajectories for policy training. By integrating these two improvements, we introduce the MapGo framework (Model-Assisted Policy optimization for Goal-oriented tasks). In our experiments, we first show the effectiveness of the FGI strategy compared with the hindsight one, and then show that the MapGo framework achieves higher sample efficiency when compared to model-free baselines on a set of complicated tasks. Menghui Zhu, Minghuan Liu, Jian Shen 0003, Weinan Zhang 0001, Deheng Ye, Yong Yu 0001, Qiang Fu 0016, Wei Yang 0032 |
IJCAI | 10 |
| 2021 | TransPath: Transformer-Based Self-supervised Learning for Histopathological Image Classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Junzhou Huang, Wei Yang 0032, Xiao Han 0011 |
MICCAI (8) | 7 |
| 2021 | Learning Diverse Policies in MOBA Games via Macro-GoalsabstractRecently, many researchers have made successful progress in building the AI systems for MOBA-game-playing with deep reinforcement learning, such as on Dota 2 and Honor of Kings. Even though these AI systems have achieved or even exceeded human-level performance, they still suffer from the lack of policy diversity. In this paper, we propose a novel Macro-Goals Guided framework, called MGG, to learn diverse policies in MOBA games. MGG abstracts strategies as macro-goals from human demonstrations and trains a Meta-Controller to predict these macro-goals. To enhance policy diversity, MGG samples macro-goals from the Meta-Controller prediction and guides the training process towards these goals. Experimental results on the typical MOBA game Honor of Kings demonstrate that MGG can execute diverse policies in different matches and lineups, and also outperform the state-of-the-art methods over 102 heroes. Yiming Gao 0007, Bei Shi, Xueying Du, Liang Wang 0015, Guangwei Chen, Zhenjie Lian, Fuhao Qiu, Guoan Han, Deheng Ye, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang |
NeurIPS | 12 |
| 2021 | Which Heroes to Pick? Learning to Draft in MOBA Games With Neural Networks and Tree SearchabstractHero drafting is essential in multiplayer online battle arena (MOBA) game playing as it builds the team of each side and directly affects the match outcome. State-of-the-art drafting methods fail to consider: 1) drafting efficiency when the hero pool is expanded; 2) the multiround nature of a MOBA 5v5 match series, i.e., two teams play best-of-$N$and the same hero is only allowed to be drafted once throughout the series. In this article, we formulate the drafting process as a multiround combinatorial game and propose a novel drafting algorithm based on neural networks and Monte Carlo tree search, named JueWuDraft. Specifically, we design a long-term value estimation mechanism to handle the best-of-$N$drafting case. TakingHonor of Kings, one of the most popular MOBA games at present, as a running case, we demonstrate the practicality and effectiveness of JueWuDraft when compared to state-of-the-art drafting methods. Menghui Zhu, Deheng Ye, Weinan Zhang 0001, Qiang Fu 0016, Wei Yang 0032 |
IEEE Trans. Games | 6 |
| 2020 | Mastering Complex Control in MOBA Games with Deep Reinforcement LearningabstractWe study the reinforcement learning problem of complex action control in the Multi-player Online Battle Arena (MOBA) 1v1 games. This problem involves far more complicated state and action spaces than those of traditional 1v1 games, such as Go and Atari series, which makes it very difficult to search any policies with human-level performance. In this paper, we present a deep reinforcement learning framework to tackle this problem from the perspectives of both system and algorithm. Our system is of low coupling and high scalability, which enables efficient explorations at large scale. Our algorithm includes several novel strategies, including control dependency decoupling, action mask, target attention, and dual-clip PPO, with which our proposed actor-critic network can be effectively trained in our system. Tested on the MOBA game Honor of Kings, the trained AI agents can defeat top professional human players in full 1v1 games. Deheng Ye, Mingfei Sun 0001, Bei Shi, Peilin Zhao, Hongsheng Yu, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin, Tengfei Shi, Liang Wang 0015, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang |
AAAI | 17 |
| 2020 | Towards Playing Full MOBA Games with Deep Reinforcement LearningabstractMOBA games, e.g., Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc. Developing AI for playing MOBA games has raised much attention accordingly. However, existing work falls short in handling the raw game complexity caused by the explosion of agent combinations, i.e., lineups, when expanding the hero pool in case that OpenAI's Dota AI limits the play to a pool of only 17 heroes. As a result, full MOBA games without restrictions are far from being mastered by any existing AI system. In this paper, we propose a MOBA AI learning paradigm that methodologically enables playing full MOBA games with deep reinforcement learning. Specifically, we develop a combination of novel and existing learning techniques, including off-policy adaption, multi-head value estimation, curriculum self-play learning, policy distillation, and Monte-Carlo tree-search, in training and playing a large pool of heroes, meanwhile addressing the scalability issue skillfully. Tested on Honor of Kings, a popular MOBA game, we show how to build superhuman AI agents that can defeat top esports players. The superiority of our AI is demonstrated by the first large-scale performance test of MOBA AI agent in the literature. Deheng Ye, Bo Yuan 0008, Fuhao Qiu, Hongsheng Yu, Yinyuting Yin, Bei Shi, Liang Wang 0015, Tengfei Shi, Qiang Fu 0016, Wei Yang 0032, Lanxiao Huang, Wei Liu 0005 |
NeurIPS | 16 |
| 2019 | Leveraging Other Datasets for Medical Imaging Classification: Evaluation of Transfer, Multi-task and Semi-supervised Learning
Hong Shang, Zhongqian Sun, Wei Yang 0032, Xinghui Fu, Jia Chang, Junzhou Huang |
MICCAI (5) | 3 |