Zhipeng Hu

dblp:95/8843 · DBLP profile ↗
← Back
66ranked-venue papers
13as first author
63since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 4 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite Feedback
abstract
Large language models (LLMs) have shown promise in providing scalable mental health support, while evaluating their counseling capability remains crucial to ensure both efficacy and safety. Existing evaluations are limited by the static assessment that focuses on knowledge tests, the single perspective that centers on user experience, and the open-loop framework that lacks actionable feedback. To address these issues, we propose Ψ-Arena, an interactive framework for comprehensive assessment and optimization of LLM-based counselors, featuring three key characteristics: (1) Realistic arena interactions that simulate real-world counseling through multi-stage dialogues with psychologically profiled NPC clients; (2) Tripartite evaluation that integrates assessments from the client, supervisor, and counselor perspectives; (3) Closed-loop optimization that iteratively improves LLM counselors using diagnostic feedback. Experiments across eight state-of-the-art LLMs show significant performance variations in different real-world scenarios and evaluation perspectives. Moreover, reflection-based optimization results in up to a 141% improvement in counseling performance. We hope Ψ-Arena provides a foundational resource for advancing reliable and human-aligned LLM applications in mental healthcare.
Shijing Zhu, Zhuang Chen 0002, Guanqun Bi, Binghang Li, Yaxi Deng, Dazhen Wan, Libiao Peng, Xiyao Xiao, Tangjie Lv, Zhipeng Hu, Minlie Huang
AAAI11
2026 DCFANet: Merging dynamic context clustering mamba and context-to-focus attention for medical image segmentation
Xiaoyan Kui, Zhipeng Hu, Zexin Ji, Shen Jiang, Qianmu Xiao, Ziwei Zou, Qinsong Li, Yang Li 0111, Beiji Zou 0001, Liming Chen 0002
Neurocomputing2
2026 TFKAN: Time-frequency KAN for long-term time series forecasting
Xiaoyan Kui, Canwei Liu, Qinsong Li, Zhipeng Hu, Yangyang Shi, Weixin Si, Beiji Zou 0001
Neurocomputing4
2026 RMViM-Net: Residual multi-path vision mamba with graph interaction attention for medical image segmentation
Shen Jiang, Xiaoyan Kui, Xingzhuo Bao, Qinsong Li, Zhipeng Hu, Beiji Zou 0001
Knowl. Based Syst.5
2026 Essential Proteins Prediction Using Features Synergy Model and GO Pure Centrality
abstract
Essential proteins are a crucial component of living organisms, and their absence will lead to cell death or reproductive arrest. Discovering these proteins can propel advancements in synthetic biology and facilitate the development of novel antibiotics and therapies for various diseases. However, current computational methods suffer from two major drawbacks that hinder their discovery rate: one is the significant noise in protein-protein interaction (PPI) data, and the other is the inadequate consideration of feature relationships. To enhance identification capabilities, this study proposes a novel essential protein prediction method, Feature Synergy Method (FSM), which leverages a features synergy model and GO pure centrality. The FSM is described as follows:Firstly, based on the principle of co-expression, gene expression data are integrated with the original PPI network to construct a pure PPI network (PPIN). Subsequently, GO annotation data are employed to calculate GO_sim weights for the interactions within the original PPI network, forming a GS_PIN. The PPIN and GS_PIN are then fused to establish the GS_PPIN, which helps mitigate the impact of noise in PPI data. Secondly, a new centrality measure, GO pure centrality (GPC), is designed based on this GO similarity-weighted pure PPI network. Thirdly, an evolutionary conservation score (ECS) is extracted from subcellular localization and orthologous proteins data. Fourthly, after analyzing the relationship between GPC and ECS, a novel fusion model, the features synergy model, is developed to integrate GPC and ECS, ultimately leading to the proposal of the new essential protein prediction method, FSM. To validate the performance of FSM, six computational methods (PeC, WDC, ION, NCCO, E_POC, and JDC) and six centrality measures (NC, IC, EC, SC, CC, and DC) were evaluated on three distinct yeast datasets. The results demonstrate that FSM achieves a higher essential protein identification rate. Similarly, GPC identifies more essential proteins compared to the six centrality-based approaches (NC, IC, EC, SC, CC, and DC).
Xinlong Luo 0002, Gaoshi Li, Zhipeng Hu, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu
IEEE Trans. Comput. Biol. Bioinform.3
2026 DPBL: Denoised Player Behavior Representation Learning
abstract
The video game industry has emerged as a significant economic force, driving extensive research on optimizing the gaming environment and improving gaming experiences. Among these endeavors, player behavior representation learning has become a critical way to model valuable player properties and is beneficial for a wide range of downstream tasks. However, some common factors, such as login rewards and daily tasks, can trigger similar behaviors among different players, which are informative and noisy for learning high-quality player behavior representations. Existing methods ignore the low signal-to-noise ratio in player behavior data and waste too much modeling capacity on less informative behaviors, resulting in their learned representations being noisy. In this paper, we propose a novel model for Denoised Player Behavior representation Learning, namely DPBL, which consists of two key modules. The first module extracts various player behavior patterns and isolates them from less informative noise. The second module utilizes the extracted patterns to refine the embedding of each behavior and eliminates noise. To optimize DPBL, two contrastive learning strategies are proposed to identify the noise that should be eliminated and to learn distinguishable representations, respectively. With the above design, DPBL is capable of mitigating the impact of noise in the data and learning high-quality representations that effectively capture player characteristics. We conducted extensive experiments on two real-world datasets, and DPBL outperforms all baselines on various downstream tasks with an improvement of 1.4% ∼ 18.1%. The results also show that DPBL achieves an improvement of 5.6% ∼ 23.0% in the denoising experiments, which proves that DPBL is more robust to noisy behaviors. Code is available at https://github.com/LwbXc/DPBL.
Wenbin Li 0012, Di Yao 0001, Zijie Xu 0006, Chang Gong 0001, Quanliang Jing, Runze Wu 0001, Haining Tan, Zhipeng Hu, Tangjie Lv, Changjie Fan, Jingping Bi
IEEE Trans. Games9
2025 LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation
abstract
Diffusion models have exhibited substantial success in text-to-image generation. However, they often encounter challenges when dealing with complex and dense prompts involving multiple objects, attribute binding, and long descriptions. In this paper, we propose a novel framework called LLM4GEN, which enhances the semantic understanding of text-to-image diffusion models by leveraging the representation of Large Language Models (LLMs). It can be seamlessly incorporated into various diffusion models as a plug-and-play component. A specially designed Cross-Adapter Module (CAM) integrates the original text features of text-to-image models with LLM features, thereby enhancing text-to-image generation. Additionally, to facilitate and correct entity-attribute relationships in text prompts, we develop an entity-guided regularization loss to further improve generation performance. We also introduce DensePrompts, which contains 7,000 dense prompts to provide a comprehensive evaluation for the text-to-image generation task. Experiments indicate that LLM4GEN significantly improves the semantic alignment of SD1.5 and SDXL, demonstrating increases of 9.69% and 12.90% in color on T2I-CompBench, respectively. Moreover, it surpasses existing models in terms of sample quality, image-text alignment, and human evaluation.
Mushui Liu, Jun Dan, Zeng Zhao, Zhipeng Hu, Bai Liu 0002, Changjie Fan
AAAI7
2025 Storynizor: Consistent Story Generation via Inter-Frame Synchronized and Shuffled ID Injection
abstract
Recent advances in text-to-image diffusion models have spurred significant interest in continuous story image generation. In this paper, we introduce Storynizor, a model capable of generating coherent stories with strong inter-frame character consistency, effective foreground-background separation, and diverse pose variation. The core innovation of Storynizor lies in its key modules: ID-Synchronizer and ID-Injector. The ID-Synchronizer employs an auto-mask self-attention module and a mask perceptual loss across inter-frame images to improve the consistency of character generation, vividly representing their postures and backgrounds. The ID-Injector utilize a Shuffling Reference Strategy (SRS) to integrate ID features into specific locations, enhancing ID-based consistent character generation. Additionally, to facilitate the training of Storynizor, we have curated a novel dataset called StoryDB comprising 100, 000 images. This dataset contains single and multiple-character sets in diverse environments, layouts, and gestures with detailed descriptions. Experimental results indicate that Storynizor demonstrates superior coherent story generation with high-fidelity character consistency, flexible postures, and vivid backgrounds compared to other character-specific methods.
Wenting Xu, Chaoyi Zhao, Keqiang Sun, Qinfeng Jin, Xiaoda Yang, Zeng Zhao, Changjie Fan, Zhipeng Hu
AAAI9
2025 DialogDraw: Image Generation and Editing System Based on Multi-Turn Dialogue
abstract
In recent years, diffusion modeling has shown great potential for image generation and editing. Beyond single-model approaches, various drawing workflows now exist to handle diverse drawing tasks. However, few solutions effectively identify user intentions through dialogue and progressively complete drawings. We introduce DialogDraw, which facilitates image generation and editing through continuous dialogue interaction. DialogDraw enables users to create and refine drawings using natural language and integrates with numerous open-source drawing workflows and models. The system accurately recognizes intentions and extracts user inputs via parameterization, adapts to various drawing function parameters, and provides an intuitive interaction mode. It effectively executes user instructions, supports dozens of image generation and editing methods, and offers robust scalability. Moreover, we employ SFT and RLHF to iterate the Intention Recognition and Parameter Extraction Model (IRPEM). To evaluate DialogDraw's functionality, we propose DrawnConvos, a dataset rich in drawing functions and command dialogue data collected from the open-source community. Our evaluation demonstrates that DialogDraw excels in command compliance, identifying and adapting to user drawing intentions, thereby proving the effectiveness of our method.
Zeng Zhao, Bai Liu 0002, Changjie Fan, Zhipeng Hu
AAAI6
2025 CharacterBench: Benchmarking Character Customization of Large Language Models
abstract
Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs’ character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters’ responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark’s potential to optimize LLMs’ character customization.
Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Pei Ke, Zhuang Chen 0002, Xiyao Xiao, Libiao Peng, Kuntian Tang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang
AAAI14
2025 LRD-3DSAM: A Novel SAM-Based 3D Medical Segmentation Network for Modeling Long-Range Dependencies
abstract
Although the Segment Anything Model (SAM) performs well in natural image segmentation, its application to 3D medical imaging remains challenging due to limited modeling of complex spatial structures and long-range slice dependencies, as well as high sensitivity to input prompt, which affects segmentation stability and clinical reliability. To address these issues, we propose LRD-3DSAM, a novel network tailored for 3D medical image segmentation. It incorporates an Adaptive Spatial Adapter to enhance 3D spatial context modeling and effectively capture cross-slice dependencies. On top of this, we introduce a Tumor Prototype Prompt Encoder, which aggregates tumor features across samples using structured representations from the adapter. This yields globally consistent tumor prototypes, improving semantic stability, inter-slice consistency, and robustness to annotation noise. Experiments on multiple public 3D tumor segmentation datasets show that LRD-3DSAM outperforms SAM and its variants in segmentation accuracy, inter-slice consistency, and cross-dataset generalization. On the pancreas tumor dataset, it achieves a 1.2% higher Dice score and 1.66% improvement in NSD, setting a new state-of-the-art.
Zhipeng Hu, Wei Wang 0229, Xin Wang 0078
BIBM2
2025 EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting
abstract
Character customization, or ’face crafting,’ is a vital feature in role-playing games (RPGs), enhancing player engagement by enabling the creation of personalized avatars. Existing automated methods often struggle with generalizability across diverse game engines due to their reliance on the intermediate constraints of specific image domain and typically support only one type of input, either text or image. To overcome these challenges, we introduce EasyCraft, an innovative end-to-end feedforward framework that automates character crafting by uniquely supporting both text and image inputs. Our approach employs a translator capable of converting facial images of any style into crafting parameters. We first establish a unified feature distribution in the translator’s image encoder through self-supervised learning on a large-scale dataset, enabling photos of any style to be embedded into a unified feature representation.Subsequently, we map this unified feature distribution to crafting parameters specific to a game engine, a process that can be easily adapted to most game engines and thus enhances EasyCraft’s generalizability. By integrating text-to-image techniques with our translator, EasyCraft also facilitates precise, text-based character crafting. EasyCraft’s ability to integrate diverse inputs significantly enhances the versatility and accuracy of avatar creation. Extensive experiments on two RPG games demonstrate the effectiveness of our method, achieving state-of-the-art results and facilitating adaptability across various avatar engines.
Suzhen Wang 0001, Wei Zhang 0219, Minda Zhao, Lincheng Li, Zhipeng Hu, Xin Yu 0002
CVPR7
2025 Crisp: Cognitive Restructuring of Negative Thoughts through Multi-turn Supportive Dialogues
abstract
Jinfeng Zhou, Yuxuan Chen, Jianing Yin, Yongkang Huang, Yihan Shi, Xikun Zhang, Libiao Peng, Rongsheng Zhang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jinfeng Zhou, Yongkang Huang, Yihan Shi, Xikun Zhang 0008, Libiao Peng, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang
EMNLP10
2025 Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
abstract
Within the domain of Massively Multiplayer Online (MMO) economy research, Agent-Based Modeling (ABM) has emerged as a robust tool for analyzing game economics, evolving from rule-based agents to decision-making agents enhanced by reinforcement learning. Nevertheless, existing works encounter significant challenges when attempting to emulate human-like economic activities among agents, particularly regarding agent reliability, sociability, and interpretability.In this study, we take a preliminary step in introducing a novel approach using Large Language Models (LLMs) in MMO economy simulation. Leveraging LLMs' role-playing proficiency, generative capacity, and reasoning aptitude, we design LLM-driven agents with human-like decision-making and adaptability. These agents are equipped with the abilities of role-playing, perception, memory, and reasoning, addressing the aforementioned challenges effectively. Simulation experiments focusing on in-game economic activities demonstrate that LLM-empowered agents can promote emergent phenomena like role specialization and price fluctuations in line with market rules.
Bihan Xu, Runze Wu 0001, Zhenya Huang, Zhipeng Hu, Kai Wang 0064, Haoyu Liu 0002, Tangjie Lv, Changjie Fan, Xin T. Tong, Jiangze Han
KDD (2)6
2025 MNMO: discover driver genes from a multi-omics data based-multi-layer network
abstract
MOTIVATION: Cancer as a public health problem is driven by genomic variations in "cancer driver" genes. The identification of driver genes is critical for the discovery of key biomarkers and the development of personalized therapy. RESULTS: We propose a prediction method MNMO: a multi-layer network model based on multi-omics data. MNMO firstly constructs a dynamically adjusted four-layer network composed of miRNAs and three kinds of genes with different features. Then three kinds of scores, i.e. control capacity, mutation score, and network score, are devised and calculated by harmonic mean to produce the integrated gene score. Experiments were performed on three kinds of real cancer data to compare the identification performance of method MNMO with that of six state-of-the-art ones. The results indicate that method MNMO presents the best identification performance under most circumstances. The genes prioritized by method MNMO not only have a better match to the benchmark ones than those identified by the other methods, but also are all associated with the development and progression of cancers. In addition, some extended versions of method MNMO can further achieve better performance on most evaluation metrics for some specific datasets. They may be more conducive to identifying tissue-specific genes, which has been verified through a number of experiments. AVAILABILITY AND IMPLEMENTATION: The source code and the R package "MNMO" are available at https://github.com/Zheng-D/MNMO. The dataset and code are archived at https://doi.org/10.5281/zenodo.14969986.
Zheng Deng, Jingli Wu, Xiaorong Chen, Gaoshi Li, Jiafei Liu 0001, Zhipeng Hu, Rongyuan Li, Wansu Deng
Bioinform.6
2025 Improving cancer driver genes identifying based on graph embedding hypergraph and hierarchical synergy dominance model
Zhipeng Hu, Xiaoyan Kui, Canwei Liu, Zanbo Sun, Shen Jiang, Kai Zhu 0009, Beiji Zou 0001
Expert Syst. Appl.1
2025 Gl-MambaNet: A global-local hybrid Mamba network for medical image segmentation
Xiaoyan Kui, Shen Jiang, Qinsong Li, Yifei Peng, Zhipeng Hu, Beiji Zou 0001
Neurocomputing5
2025 Knowledge enhanced graph contrastive learning for match outcome prediction
Junji Jiang, Likang Wu, Zhipeng Hu, Runze Wu 0001, Hongke Zhao
Inf. Process. Manag.3
2025 Predicting Driver Genes From Multi-Omics Data Using Hierarchical Multi-Feature Synergy Model
abstract
Cancer is an extremely complex disease, whose occurrence and development are influenced by a multitude of factors, among which the abnormal activity of cancer driver genes plays a crucial role in the pathological process. Identifying these genes allows researchers to understand pathogenic mechanisms and biological functions of cancer, facilitating the development of targeted therapies. Current methods for identifying driver genes often ignore the synergism among genes and the importance of features, thereby affecting identification accuracy. In this paper, we propose a cancer driver genes identification method called HMFS, which is based on the hierarchical multi-feature synergy model. Firstly, a hypergraph is constructed using Node2vec and K-means algorithm. By analyzing the topological feature and mutual exclusion degree of genes in each hyperedge, the Mutation Aggregation Coefficient is extracted. Then, based on the functional expression mechanism of genes, differential expression analysis is performed using miRNA and mRNA expression data. Finally, by analyzing the importance among features, the Hierarchical Multi-Feature Synergy is proposed for features fusion. In this paper, experiments are conducted on three real cancer datasets. Compared with seven representative methods, HMFS has the best performance on all evaluation indicators.
Zhipeng Hu, Xiaoyan Kui, Canwei Liu, Shen Jiang, Ziwei Zou, Beiji Zou 0001
IEEE Trans. Comput. Biol. Bioinform.1
2025 Using Multi-Feature Weak Consensus Model to Discover Essential Proteins
abstract
Essential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a phenomenon among features is found, called weak consensus, then a weak consensus model to fuse these features is proposed in this paper. After analyzing the relationship between a protein and its neighbors in protein-protein interaction networks, a new centrality, namely neighborhood aggregation centrality(NAC) is developed in this paper. Then, a Max-Min strategy is used to integrate NAC with Pearson correlation coefficient and Jaccard similarity coefficient based on gene expression data to obtain local importance score. In addition, orthologous feature score is used to measure proteins conservation. Finally, by using the weak consensus model to fuse orthologous feature score with local importance score, a new method WOL is proposed in this paper. Then experiments are performed on S.cerevisiae data. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate.
Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Jiafei Liu 0001, Jingli Wu, Wei Peng 0004, Xiaoshu Zhu
IEEE Trans. Comput. Biol. Bioinform.1
2025 TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
abstract
Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to find and hence limits their usage. In this work, we propose TalkCLIP, a framework that can generate talking heads where the expressions are specified by natural language, hence allowing for specifying expressions more conveniently. To model the mapping from text to expressions, we first construct a text-video paired talking head dataset where each video has diverse text descriptions that depict both coarse-grained emotions and fine-grained facial movements. Leveraging the proposed dataset, we introduce a CLIP-based style encoder that projects natural language-based descriptions to the representations of expressions. TalkCLIP can even infer expressions for descriptions unseen during training. TalkCLIP can also use text to modulate expression intensity and edit expressions. Extensive experiments demonstrate that TalkCLIP achieves the advanced capability of generating photo-realistic talking heads with vivid facial expressions guided by text descriptions.
Yifeng Ma 0001, Suzhen Wang 0001, Yu Ding 0001, Tangjie Lv, Changjie Fan, Zhipeng Hu, Zhidong Deng, Xin Yu 0002
IEEE Trans. Multim.7
2025 ICE: Interactive 3D Game Character Facial Editing via Dialogue
abstract
Most recent popular Role-Playing Games (RPGs) allow players to create in-game characters with hundreds of adjustable parameters, including bone positions and various makeup options. Although text-driven auto-customization systems have been developed to simplify the complex process of adjusting these intricate character parameters, they are limited by their single-round generation and lack the capability for further editing and fine-tuning. In this paper, we propose an Interactive Character Editing framework (ICE) to achieve a multi-round dialogue-based refinement process. In a nutshell, our ICE offers a more user-friendly way to enable players to convey creative ideas iteratively while ensuring that created characters align with the expectations of players. Specifically, we propose an Instruction Parsing Module (IPM) that utilizes large language models (LLMs) to parse multi-round dialogues into clear editing instruction prompts in each round. To reliably and swiftly modify character control parameters at a fine-grained level, we propose a Semantic-guided Low-dimension Parameter Solver (SLPS) that edits character control parameters according to prompts in a zero-shot manner. Our SLPS first localizes the character control parameters related to the fine-grained modification, and then optimizes the corresponding parameters in a low-dimension space to avoid unrealistic results. Extensive experimental results demonstrate the effectiveness of our proposed ICE for in-game character creation and the superior editing performance of ICE. Code:https://github.com/NeteaseFuxi/ICE-Interactive-3D-Game-Character.
Haoqian Wu, Minda Zhao, Zhipeng Hu, Changjie Fan, Lincheng Li, Rui Zhao 0019, Xin Yu 0002
IEEE Trans. Multim.3
2024 Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured Representations
abstract
Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured representations, i.e., representations of objects, attributes, and relations. The models cannot make a distinction between "An astronaut rides a horse" and "A horse rides an astronaut". This is because they fail to fully leverage structured knowledge when learning multi-modal representations. In this paper, we present an end-to-end framework Structure-CLIP, which integrates Scene Graph Knowledge (SGK) to enhance multi-modal structured representations. Firstly, we use scene graphs to guide the construction of semantic negative examples, which results in an increased emphasis on learning structured representations. Moreover, a Knowledge-Enhance Encoder (KEE) is proposed to leverage SGK as input to further enhance structured representations. To verify the effectiveness of the proposed framework, we pre-train our model with the aforementioned approaches and conduct experiments on downstream tasks. Experimental results demonstrate that Structure-CLIP achieves state-of-the-art (SOTA) performance on VG-Attribution and VG-Relation datasets, with 12.5% and 4.1% ahead of the multi-modal SOTA model respectively. Meanwhile, the results on MSCOCO indicate that Structure-CLIP significantly enhances the structured representations while maintaining the ability of general representations. Our code is available at https://github.com/zjukg/Structure-CLIP.
Jiji Tang, Zhuo Chen 0007, Zeng Zhao, Tangjie Lv, Zhipeng Hu, Wen Zhang 0015
AAAI10
2024 EnMatch: Matchmaking for Better Player Engagement via Neural Combinatorial Optimization
abstract
Matchmaking is a core task in e-sports and online games, as it contributes to player engagement and further influences the game's lifecycle. Previous methods focus on creating fair games at all times. They divide players into different tiers based on skill levels and only select players from the same tier for each game. Though this strategy can ensure fair matchmaking, it is not always good for player engagement. In this paper, we propose a novel Engagement-oriented Matchmaking (EnMatch) framework to ensure fair games and simultaneously enhance player engagement. Two main issues need to be addressed. First, it is unclear how to measure the impact of different team compositions and confrontations on player engagement during the game considering the variety of player characteristics. Second, such a detailed consideration on every single player during matchmaking will result in an NP-hard combinatorial optimization problem with non-linear objectives. In light of these challenges, we turn to real-world data analysis to reveal engagement-related factors. The resulting insights guide the development of engagement modeling, enabling the estimation of quantified engagement before a match is completed. To handle the combinatorial optimization problem, we formulate the problem into a reinforcement learning framework, in which a neural combinatorial optimization problem is built and solved. The performance of EnMatch is finally demonstrated through the comparison with other state-of-the-art methods based on several real-world datasets and online deployments on two games.
Kai Wang 0064, Haoyu Liu 0002, Zhipeng Hu, Xiaochuan Feng, Minghao Zhao 0002, Runze Wu 0001, Tangjie Lv, Changjie Fan
AAAI3
2024 Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
abstract
Diffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or off-line processing, greatly reducing their efficiency. In this paper, we propose a novel and efficient image editing method for Text-to-Image (T2I) diffusion models, termed Instant Diffusion Editing (InstDiffEdit). In particular, InstDiffEdit aims to employ the cross-modal attention ability of existing diffusion models to achieve instant mask guidance during the diffusion steps. To reduce the noise of attention maps and realize the full automatics, we equip InstDiffEdit with a training-free refinement scheme to adaptively aggregate the attention distributions for the automatic yet accurate mask generation. Meanwhile, to supplement the existing evaluations of DIE, we propose a new benchmark called Editing-Mask to examine the mask accuracy and local editing ability of existing methods. To validate InstDiffEdit, we also conduct extensive experiments on ImageNet and Imagen, and compare it with a bunch of the SOTA methods. The experimental results show that InstDiffEdit not only outperforms the SOTA methods in both image quality and editing results, but also has a much faster inference speed, i.e., +5 to +6 times. Our code available at https://anonymous.4open.science/r/InstDiffEdit-C306
Siyu Zou, Jiji Tang, Yiyi Zhou, Chaoyi Zhao, Zhipeng Hu, Xiaoshuai Sun
AAAI7
2024 EfficientDreamer: High-Fidelity and Stable 3D Creation via Orthogonal-view Diffusion Priors
abstract
While image diffusion models have made significant progress in text-driven 3D content creation, they often fail to accurately capture the intended meaning of text prompts, especially for view information. This limitation leads to the Janus problem, where multi-faced 3D models are generated under the guidance of such diffusion models. In this paper, we propose a robust high-quality 3D content generation pipeline by exploiting orthogonal-view image guidance. First, we introduce a novel 2D diffusion model that generates an image consisting of four orthogonal-view sub-images based on the given text prompt. Then, the 3D content is created using this diffusion model. Notably, the generated orthogonal-view image provides strong geometric structure priors and thus improves 3D consistency. As a result, it effectively resolves the Janus problem and significantly enhances the quality of 3D content creation. Additionally, we present a 3D synthesis fusion network that can further improve the details of the generated 3D contents. Both quantitative and qualitative evaluations demonstrate that our method surpasses previous text-to-3D techniques. Project page: https://efficientdreamer.github.io.
Zhipeng Hu, Minda Zhao, Chaoyi Zhao, Lincheng Li, Zeng Zhao, Changjie Fan, Xiaowei Zhou 0001, Xin Yu 0002
CVPR1
2024 Towards a Simultaneous and Granular Identity-Expression Control in Personalized Face Generation
abstract
In human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted por-trait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts towards personalized face generation. To this end, we propose a novel multi-modal face generation frame-work, capable of simultaneous identity-expression control and more fine-grained expression synthesis. Our expression control is so sophisticated that it can be specialized by the fine-grained emotional vocabulary. We devise a novel dif-fusion model that can undertake the task of simultaneously face swapping and reenactment. Due to the entanglement of identity and expression, separately and precisely control-ling them within one framework is a nontrivial task, thus has not been explored yet. To overcome this, we propose sev-eral innovative designs in the conditional diffusion model, including balancing identity and expression encoder, improved midpoint sampling, and explicitly background con-ditioning. Extensive experiments have demonstrated the controllability and scalability of the proposed framework, in comparison with state-of-the-art text-to-image, face swap-ping, and face reenactment methods.
Renshuai Liu, Wei Zhang 0219, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001
CVPR4
2024 Text-Guided 3D Face Synthesis - From Generation to Editing
abstract
Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct gen-eration, ignoring the editing, restricting them from synthe-sizing customized 3D faces through iterative adjustments. In this paper, we propose a unified text-guided framework from face generation to editing. In the generation stage, we propose a geometry-texture decoupled generation to miti-gate the loss of geometric details caused by coupling. Be-sides, decoupling enables us to utilize the generated geom-etry as a condition for texture generation, yielding highly geometry-texture aligned results. We further employ a fine-tuned texture diffusion model to enhance texture quality in both RGB and YUV space. In the editing stage, we first em-ploy a pre-trained diffusion model to update facial geometry or texture based on the texts. To enable sequential editing, we introduce a UV domain consistency preservation reg-ularization, preventing unintentional changes to irrelevant facial attributes. Besides, we propose a self-guided consis-tency weight strategy to improve editing efficacy while pre-serving consistency. Through comprehensive experiments, we showcase our method's superiority in face synthesis. Project page: https://faceg2e.github.io/.
Yunjie Wu, Yapeng Meng, Zhipeng Hu, Lincheng Li, Haoqian Wu, Kun Zhou 0001, Weiwei Xu 0003, Xin Yu 0002
CVPR3
2024 CPT-VR: Improving Surface Rendering via Closest Point Transform with View-Reflection Appearance
Zhipeng Hu, Yongqiang Zhang 0003, Chen Liu 0028, Lincheng Li, Sida Peng, Xiaowei Zhou 0001, Changjie Fan, Xin Yu 0002
ECCV (73)1
2024 AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
abstract
Aligning agent behaviors with diverse human preferences remains a challenging problem in reinforcement learning (RL), owing to the inherent abstractness and mutability of human preferences. To address these issues, we propose AlignDiff, a novel framework that leverages RLHF to quantify human preferences, covering abstractness, and utilizes them to guide diffusion planning for zero-shot behavior customizing, covering mutability. AlignDiff can accurately match user-customized behaviors and efficiently switch from one to another. To build the framework, we first establish the multi-perspective human feedback datasets, which contain comparisons for the attributes of diverse behaviors, and then train an attribute strength model to predict quantified relative strengths. After relabeling behavioral datasets with relative strengths, we proceed to train an attribute-conditioned diffusion model, which serves as a planner with the attribute strength model as a director for preference aligning at the inference phase. We evaluate AlignDiff on various locomotion tasks and demonstrate its superior performance on preference matching, switching, and covering compared to other baselines. Its capability of completing unseen downstream tasks under human instructions also showcases the promising potential for human-AI collaboration. More visualization videos are released on https://aligndiff.github.io/.
Zibin Dong, Yifu Yuan, Jianye Hao, Fei Ni 0001, Yao Mu 0001, Yan Zheng 0002, Yujing Hu, Tangjie Lv, Changjie Fan, Zhipeng Hu
ICLR10
2024 Stylized Offline Reinforcement Learning: Extracting Diverse High-Quality Behaviors from Heterogeneous Datasets
abstract
Previous literature on policy diversity in reinforcement learning (RL) either focuses on the online setting or ignores the policy performance. In contrast, offline RL, which aims to learn high-quality policies from batched data, has yet to fully leverage the intrinsic diversity of the offline dataset. Addressing this dichotomy and aiming to balance quality and diversity poses a significant challenge to extant methodologies. This paper introduces a novel approach, termed Stylized Offline RL (SORL), which is designed to extract high-performing, stylistically diverse policies from a dataset characterized by distinct behavioral patterns. Drawing inspiration from the venerable Expectation-Maximization (EM) algorithm, SORL innovatively alternates between policy learning and trajectory clustering, a mechanism that promotes policy diversification. To further augment policy performance, we introduce advantage-weighted style learning into the SORL framework. Experimental evaluations across multiple environments demonstrate the significant superiority of SORL over previous methods in extracting high-quality policies with diverse behaviors. A case in point is that SORL successfully learns strong policies with markedly distinct playing patterns from a real-world human dataset of a popular basketball video game "Dunk City Dynasty."
Yihuan Mao, Chengjie Wu, Hao Hu 0006, Ji Jiang, Tianze Zhou, Tangjie Lv, Changjie Fan, Zhipeng Hu, Yi Wu 0013, Yujing Hu, Chongjie Zhang
ICLR9
2024 MGMatch: Fast Matchmaking with Nonlinear Objective and Constraints via Multimodal Deep Graph Learning
abstract
As a core problem of online games, matchmaking is to assign players into multiple teams to maximize their gaming experience. With the rapid development of game industry, it is increasingly difficulty to explicitly model players' experiences as linear functions. Instead, it is often modeled in a data-driven way by training a neural network. Meanwhile, complex rules must be satisfied to ensure the robustness of matchmaking, which are often described using logical operators. Therefore, matchmaking in practical scenarios is a challenging combinatorial optimization problem with nonlinear objective, linear constraints and logical constraints, which receives much less attention in previous research. In this paper, we propose a novel deep learning method for high-quality matchmaking in real-time. We first cast the problem as standard mixed-integer programming (MIP) by linearizing ReLU networks and logical constraints. Then, based on supervised learning, we design and train a multi-modal graph learning architecture to predict optimal solutions end-to-end from instance data, and solve a surrogate problem to efficiently obtain feasible solutions. Evaluation results on real industry datasets show that our method can deliver near-optimal solutions within 100ms.
Yu Sun 0051, Kai Wang 0064, Zhipeng Hu, Runze Wu 0001, Yaoxin Wu, Wen Song 0004, Tangjie Lv, Changjie Fan
KDD3
2024 XRL-Bench: A Benchmark for Evaluating and Comparing Explainable Reinforcement Learning Techniques
abstract
Reinforcement Learning (RL) has demonstrated substantial potential across diverse fields, yet understanding its decision-making process, especially in real-world scenarios where rationality and safety are paramount, is an ongoing challenge. This paper delves in to Explainable RL (XRL), a subfield of Explainable AI (XAI) aimed at unravelling the complexities of RL models. Our focus rests on state-explaining techniques, a crucial subset within XRL methods, as they reveal the underlying factors influencing an agent's actions at any given time. Despite their significant role, the lack of a unified evaluation framework hinders assessment of their accuracy and effectiveness. To address this, we introduce XRL-Bench, a unified standardized benchmark tailored for the evaluation and comparison of XRL methods, encompassing three main modules: standard RL environments, explainers based on state importance, and standard evaluators. XRL-Bench supports both tabular and image data for state explanation. We also propose TabularSHAP, an innovative and competitive XRL method. We demonstrate the practical utility of TabularSHAP in real-world online gaming services and offer an open-source benchmark platform for the straightforward implementation and evaluation of XRL methods. Our contributions facilitate the continued progression of XRL technology.
Zhipeng Hu, Runze Wu 0001, Xingchen Fang, Ji Jiang, Tianze Zhou, Yujing Hu, Haoyu Liu 0002, Tangjie Lyu, Changjie Fan
KDD2
2024 Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval Method
abstract
In recent years, Vision-Language Pre-training (VLP) models have demonstrated rich prior knowledge for multimodal alignment, prompting investigations into their application in Specific Domain Image-Text Retrieval(SDITR) such as Text-Image Person Re-identification (TIReID) and Remote Sensing Image-Text Retrieval (RSITR). Due to the unique data characteristics in specific scenarios, the primary challenge is to leverage discriminative fine-grained local information for improved mapping of images and text into a shared space. Current approaches interact with all multimodal local features for alignment, implicitly focusing on discriminative local information to distinguish data differences, which may bring noise and uncertainty. Furthermore, their VLP feature extractors like CLIP often focus on instance-level representations, potentially reducing the discriminability of fine-grained local features. To alleviate these issues, we propose an Explicit Key Local information Selection and Reconstruction Framework (EKLSR), which explicitly selects key local information to enhance feature representation. Specifically, we introduce a Key Local information Selection and Fusion (KLSF) that utilizes hidden knowledge from the VLP model to select interpretably and fuse key local information. Secondly, we employ Key Local segment Reconstruction (KLR) based on multimodal interaction to reconstruct the key local segments of images (text), significantly enriching their discriminative information and enhancing both inter-modal and intra-modal interaction alignment. To demonstrate the effectiveness of our approach, we conducted experiments on five datasets across TIReID and RSITR. Notably, our EKLSR model achieves state-of-the-art performance on two RSITR datasets.
Yu Liao, Rui Yang 0038, Jianwei Tao, Bai Liu 0002, Zhipeng Hu, Shuang Wang 0001, Zeng Zhao
ACM Multimedia6
2024 FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation Model
abstract
Video-driven 3D facial animation transfer aims to drive avatars to reproduce the expressions of actors. Existing methods have achieved remarkable results by constraining both geometric and perceptual consistency. However, geometric constraints (like those designed on facial landmarks) are insufficient to capture subtle emotions, while expression features trained on classification tasks lack fine granularity for complex emotions. To address this, we propose \textbf{FreeAvatar}, a robust facial animation transfer method that relies solely on our learned expression representation. Specifically, FreeAvatar consists of two main components: the expression foundation model and the facial animation transfer model. In the first component, we initially construct a facial feature space through a face reconstruction task and then optimize the expression feature space by exploring the similarities among different expressions. Benefiting from training on the amounts of unlabeled facial images and re-collected expression comparison dataset, our model adapts freely and effectively to any in-the-wild input facial images. In the facial animation transfer component, we propose a novel Expression-driven Multi-avatar Animator, which first maps expressive semantics to the facial control parameters of 3D avatars and then imposes perceptual constraints between the input and output images to maintain expression consistency. To make the entire process differentiable, we employ a trained neural renderer to translate rig parameters into corresponding images. Furthermore, unlike previous methods that require separate decoders for each avatar, we propose a dynamic identity injection module that allows for the joint training of multiple avatars within a single network.
Wei Zhang 0219, Chen Liu 0028, Rudong An, Lincheng Li, Yu Ding 0001, Changjie Fan, Zhipeng Hu, Xin Yu 0002
SIGGRAPH Asia8
2024 Hybrid CtrlFormer: Learning Adaptive Search Space Partition for Hybrid Action Control via Transformer-based Monte Carlo Tree Search
abstract
Hybrid action control tasks are common in the real world, which require controlling some discrete and continuous actions simultaneously. To solve these tasks, existing Deep Reinforcement learning (DRL) methods either directly build a separate policy for each type of action or simplify the hybrid action space into a discrete or continuous action control problem. However, these methods neglect the challenge of exploration resulting from the complexity of the hybrid action space. Thus, it is necessary to design more sample efficient algorithms. To this end, we propose a novel Hybrid Control Transformer (Hybrid CtrlFormer), to achieve better exploration and exploitation for the hybrid action control problems. The core idea is: 1) we construct a hybrid action space tree with the discrete actions at the higher level and the continuous parameter space at the lower level. Each parameter space is split into multiple subregions. 2) To simplify the exploration space, a Transformer-based Monte-Carlo tree search method is designed to efficiently evaluate and partition the hybrid action space into good and bad subregions along the tree. Our method achieves state-of-the-art performance and sample efficiency in a variety of environments with discrete-continuous action space.
Jiashun Liu, Xiaotian Hao, Jianye Hao, Yan Zheng 0002, Yujing Hu, Changjie Fan, Tangjie Lv, Zhipeng Hu
UAI8
2024 MarkerNet: A divide-and-conquer solution to motion capture solving from raw markers
abstract
Abstract Marker‐based optical motion capture (MoCap) aims to localize 3D human motions from a sequence of input raw markers. It is widely used to produce physical movements for virtual characters in various games such as the role‐playing game, the fighting game, and the action‐adventure game. However, the conventional MoCap cleaning and solving process is extremely labor‐intensive, time‐consuming, and usually the most costly part of game animation production. Thus, there is a high demand for automated algorithms to replace costly manual operations and achieve accurate MoCap cleaning and solving in the game industry. In this article, we design a divide‐and‐conquer‐based MoCap solving network, dubbed MarkerNet, to estimate human skeleton motions from sequential raw markers effectively. In a nutshell, our key idea is to decompose the task of direct solving of global motion from all markers into first modeling sub‐motions of local parts from the corresponding marker subsets and then aggregating sub‐motions into a global one. In this manner, our model can effectively capture local motion patterns w.r.t. different marker subsets, thus producing more accurate results compared to the existing methods. Extensive experiments on both real and synthetic data verify the effectiveness of the proposed method.
Zhipeng Hu, Jilin Tang, Lincheng Li, Xin Yu 0002, Jiajun Bu
Comput. Animat. Virtual Worlds1
2024 Promoting human-AI interaction makes a better adoption of deep reinforcement learning: a real-world application in game industry
Zhipeng Hu, Haoyu Liu 0002, Lizi Wang, Runze Wu 0001, Yujing Hu, Tangjie Lyu, Changjie Fan
Multim. Tools Appl.1
2024 Learning a compact embedding for fine-grained few-shot static gesture recognition
Zhipeng Hu, Wei Zhang 0219, Yu Ding 0001, Tangjie Lv, Changjie Fan
Multim. Tools Appl.1
2024 StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads
abstract
Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse speaking styles in the final videos. To address this challenge, we propose a one-shot style-controllable talking face generation method that can obtain speaking styles from reference speaking videos and drive the one-shot portrait to speak with the reference speaking styles and another piece of audio. Our method aims to synthesize the style-controllable coefficients of a 3D Morphable Model (3DMM), including facial expressions and head movements, in a unified framework. Specifically, the proposed framework first leverages a style encoder to extract the desired speaking styles from the reference videos and transform them into style codes. Then, the framework uses a style-aware decoder to synthesize the coefficients of 3DMM from the audio input and style codes. During decoding, our framework adopts a two-branch architecture, which generates the stylized facial expression coefficients and stylized head movement coefficients, respectively. After obtaining the coefficients of 3DMM, an image renderer renders the expression coefficients into a specific person's talking-head video. Extensive experiments demonstrate that our method generates visually authentic talking head videos with diverse speaking styles from only one portrait image and an audio clip.
Suzhen Wang 0001, Yifeng Ma 0001, Yu Ding 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Zhidong Deng, Xin Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Facial Action Unit Detection and Intensity Estimation From Self-Supervised Representation
abstract
As a fine-grained and local expression behavior measurement, facial action unit (FAU) analysis (e.g., detection and intensity estimation) has been documented for its time-consuming, labor-intensive, and error-prone annotation. Thus a long-standing challenge of FAU analysis arises from the data scarcity of manual annotations, limiting the generalization ability of trained models to a large extent. Amounts of previous works have made efforts to alleviate this issue via semi/weakly supervised methods and extra auxiliary information. However, these methods still require domain knowledge and have not yet avoided the high dependency on data annotation. This article introduces a robust facial representation model MAE-Face for AU analysis. Using masked autoencoding as the self-supervised pre-training approach, MAE-Face first learns a high-capacity model from a feasible collection of face images without additional data annotations. Then after being fine-tuned on AU datasets, MAE-Face exhibits convincing performance for both AU detection and AU intensity estimation, achieving a new state-of-the-art on nearly all the evaluation results. Further investigation shows that MAE-Face achieves decent performance even when fine-tuned on only 1% of the AU training set, strongly proving its robustness and generalization performance. The pre-trained model is available at our GitHub repository.
Rudong An, Wei Zhang 0219, Yu Ding 0001, Zeng Zhao, Tangjie Lv, Changjie Fan, Zhipeng Hu
IEEE Trans. Affect. Comput.9
2024 Identification of Cancer Driver Genes based on Dynamic Incentive Model
abstract
Cancer is a complex genomic mutation disease, and identifying cancer driver genes promotes the development of targeted drugs and personalized therapies. The current computational method takes less consideration of the relationship among features and the effect of noise in protein-protein interaction(PPI) data, resulting in a low recognition rate. In this paper, we propose a cancer driver genes identification method based on dynamic incentive model, DIM. This method firstly constructs a hypergraph to reduce the impact of false positive data in PPI. Then, the importance of genes in each hyperedge in hypergraph is considered from three perspectives, network and functional score(NFS) is proposed. By analyzing the relation among features, the dynamic incentive model is proposed to fuse NFS, the differential expression score of mRNA and the differential expression score of miRNA. DIM is compared with some classical methods on breast cancer, lung cancer, prostate cancer, and pan-cancer datasets. The results show that DIM has the best performance on statistical evaluation indicators, functional consistency and the partial area under the ROC curve, and has good cross-cancer capability.
Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu, Jingli Wu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2024 VESPA: A General System for Vision-Based Extrasensory Perception Anticheating in Online FPS Games
abstract
Cheating is widespread in online games, particularly in competitive games like First-Person Shooter (FPS) games. One of the most common types of cheating is Extrasensory Perception (ESP), which involves illicitly obtaining visual information to gain an unfair advantage over normal players. To protect the gaming experience of legitimate players and the interests of game companies, there is an urgent need for anti-cheating applications. In this paper, we propose a general system for ESP anti-cheating in online FPS games, considering the business characteristics and industrial applications. We present a vision-based anti-cheating framework that incorporates both supervised and unsupervised solutions for comprehensive cheating detection. Based on this framework, we design and deploy a dual-audit human-in-the-loop system for industrial gaming anti-cheating applications. We evaluate our proposed framework from multiple online and offline perspectives and demonstrate its practical significance with superior performance.
Jiaheng Qi, Zhipeng Hu, Runze Wu 0001, Tangjie Lv, Changjie Fan
IEEE Trans. Games3
2024 PU-Detector: A PU Learning-based Framework for Real Money Trading Detection in MMORPG
abstract
Massive multiplayer online role-playing games (MMORPG) have been becoming one of the most popular and exciting online games. In recent years, a cheating phenomenon called real money trading (RMT) has arisen and damaged the fantasy world in many ways. RMT is the sale of in-game items, currency, or even characters to earn real money, breaking the balance of the game economy ecosystem and damaging the game experience. Therefore, some studies have emerged to address the problem of RMT detection. However, they cannot well handle the label uncertainty problem in practice, where there are only labeled RMT samples (positive samples) and unlabeled samples, which could either be RMT samples or normal transactions (negative samples). Meanwhile, the trading relationship between RMTers is modeled in a simple way, leading to some normal transactions being falsely classified as RMT. In this article, we propose PU-Detector, a novel framework based on PU learning (learning from positive and unlabeled data) for RMT detection, considering the fact that there are only labeled RMT samples and other unlabeled transactions. We first automatically estimate the likelihood of one transaction being RMT by developing an improved PU learning method and proposing an assessment rule. Sequentially, we use the estimated likelihood as edge weight to construct a trading graph to learn trader representation. Then, with the trader representations and basic trading features, we detect RMT samples by the improved PU learning method. PU-Detector is evaluated on a large-scale real world dataset consisting of 33,809,956 transaction logs generated by 43,217 unique players. Compared with other approaches, it achieves the state-of-the-art performance and demonstrates its advantages in detecting underlying RMT samples.
Yilin Wang 0014, Sha Zhao, Runze Wu 0001, Yuhong Xu, Jianrong Tao, Tangjie Lv, Shijian Li, Zhipeng Hu, Gang Pan 0001
ACM Trans. Knowl. Discov. Data9
2024 Calligraphy Font Generation via Explicitly Modeling Location-Aware Glyph Component Deformations
abstract
Automatic font generation is a challenging and time-consuming task, particularly in languages that consist of large amounts of characters with complicated structures. Typical component-wise font generation methods decompose the source character into components and search for them from the reference glyph set as candidate components. These candidate components are then utilized to learn the local styles of the target glyph. However, these methods overlook that the same component at different locations may have different profiles. When the candidate components locate differently from their corresponding components in the target glyph, the style of a generated glyph will look inconsistent. It is observed that for arbitrary components at two specific locations, the deformation patterns are similar. Driven by this, we present a location-aware component-deformable font generation method. Specifically, we search for candidate components and their corresponding deformative component pairs from the reference glyph set. Each deformative component pair can accurately depict how to deform the candidate component to the desired profile in the target glyph. Hence, we introduce a location-dependent deformation module to perform component warping. In this way, we significantly improve the component deformation ability. Lastly, we integrate deformed components into target glyphs while enforcing their styles to be consistent with the reference ones. Extensive experiments demonstrate that our method produces target-font consistent glyphs and outperforms the state-of-the-art on both seen and unseen fonts.
Minda Zhao, Xingqun Qi, Zhipeng Hu, Lincheng Li, Yongqiang Zhang 0003, Zi Huang, Xin Yu 0002
IEEE Trans. Multim.3
2023 Generating Coherent Narratives by Learning Dynamic and Discrete Entity States with a Contrastive Framework
abstract
Despite advances in generating fluent texts, existing pretraining models tend to attach incoherent event sequences to involved entities when generating narratives such as stories and news. We conjecture that such issues result from representing entities as static embeddings of superficial words, while neglecting to model their ever-changing states, i.e., the information they carry, as the text unfolds. Therefore, we extend the Transformer model to dynamically conduct entity state updates and sentence realization for narrative generation. We propose a contrastive framework to learn the state representations in a discrete space, and insert additional attention layers into the decoder to better exploit these states. Experiments on two narrative datasets show that our model can generate more coherent and diverse narratives than strong baselines with the guidance of meaningful entity states.
Jian Guan 0002, Zhipeng Hu, Minlie Huang
AAAI4
2023 StyleTalk: One-Shot Talking Head Generation with Controllable Speaking Styles
abstract
Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate diverse speaking styles in the final talking head videos. To tackle this problem, we propose a one-shot style-controllable talking face generation framework. In a nutshell, we aim to attain a speaking style from an arbitrary reference speaking video and then drive the one-shot portrait to speak with the reference speaking style and another piece of audio. Specifically, we first develop a style encoder to extract dynamic facial motion patterns of a style reference video and then encode them into a style code. Afterward, we introduce a style-controllable decoder to synthesize stylized facial animations from the speech content and style code. In order to integrate the reference speaking style into generated videos, we design a style-aware adaptive transformer, which enables the encoded style code to adjust the weights of the feed-forward layers accordingly. Thanks to the style-aware adaptation mechanism, the reference speaking style can be better embedded into synthesized videos during decoding. Extensive experiments demonstrate that our method is capable of generating talking head videos with diverse speaking styles from only one portrait image and an audio clip while achieving authentic visual effects. Project Page: https://github.com/FuxiVirtualHuman/styletalk.
Yifeng Ma 0001, Suzhen Wang 0001, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001, Zhidong Deng, Xin Yu 0002
AAAI3
2023 DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution Video
abstract
For few-shot learning, it is still a critical challenge to realize photo-realistic face visually dubbing on high-resolution videos. Previous works fail to generate high-fidelity dubbing results. To address the above problem, this paper proposes a Deformation Inpainting Network (DINet) for high-resolution face visually dubbing. Different from previous works relying on multiple up-sample layers to directly generate pixels from latent embeddings, DINet performs spatial deformation on feature maps of reference images to better preserve high-frequency textural details. Specifically, DINet consists of one deformation part and one inpainting part. In the first part, five reference facial images adaptively perform spatial deformation to create deformed feature maps encoding mouth shapes at each frame, in order to align with input driving audio and also the head poses of input source images. In the second part, to produce face visually dubbing, a feature decoder is responsible for adaptively incorporating mouth movements from the deformed feature maps and other attributes (i.e., head pose and upper facial expression) from the source feature maps together. Finally, DINet achieves face visually dubbing with rich textural details. We conduct qualitative and quantitative comparisons to validate our DINet on high-resolution videos. The experimental results show that our method outperforms state-of-the-art works.
Zhipeng Hu, Wenjin Deng, Changjie Fan, Tangjie Lv, Yu Ding 0001
AAAI2
2023 Essential proteins identification based on weak consensus model and neighborhood aggregation centrality
abstract
Essential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a weak consensus model to fuse features is proposed in this paper. Then, this paper uses the weak consensus model to fuse protein-protein interaction network, gene expression data, and orthologous data, thus proposing a new method, WOL. Then experiments are performed on one S.cerevisiae dataset. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate.
Zhipeng Hu, Gaoshi Li, Jingli Wu, Xinlong Luo 0002, Jiafei Liu 0001, Wei Peng 0004, Xiaoshu Zhu
BIBM1
2023 Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation
abstract
Recent popular Role-Playing Games (RPGs) saw the great success of character auto-creation systems. The bone-drivenface model controlled by continuous parameters (like the position of bones) and discrete parameters (like the hairstyles) makes it possible for users to personalize and customize in-game characters. Previous in-game character auto-creation systems are mostly image-driven, where facial parameters are optimized so that the rendered character looks similar to the reference face photo. This paper proposes a novel text-to-parameter translation method (T2P) to achieve zero-shot text-driven game character auto-creation. With our method, users can create a vivid in-game character with arbitrary text description without using any reference photo or editing hundreds of parameters manually. In our method, taking the power of large-scale pre-trained multi-modal CLIP and neural rendering, T2P searches both continuous facial parameters and discrete facial parameters in a unified framework. Due to the discontinuous parameter representation, previous methods have difficulty in effectively learning discrete facial parameters. T2p, to our best knowledge, is the first method that can handle the optimization of both discrete and continuous parameters. Experimental results show that T2P can generate high-quality and vivid game characters with given text prompts. T2P outperforms other SOTA text-to-3D generation methods on both objective evaluations and subjective evaluations.
Rui Zhao 0019, Wei Li 0224, Zhipeng Hu, Lincheng Li, Zhengxia Zou, Zhenwei Shi 0001, Changjie Fan
CVPR3
2023 NeFII: Inverse Rendering for Reflectance Decomposition with Near-Field Indirect Illumination
abstract
Inverse rendering methods aim to estimate geometry, materials and illumination from multi-view RGB images. In order to achieve better decomposition, recent approaches attempt to model indirect illuminations reflected from different materials via Spherical Gaussians (SG), which, however, tends to blur the high-frequency reflection details. In this paper, we propose an end-to-end inverse rendering pipeline that decomposes materials and illumination from multi-view images, while considering near-field indirect illumination. In a nutshell, we introduce the Monte Carlo sampling based path tracing and cache the indirect illumination as neural radiance, enabling a physics-faithful and easy-to-optimize inverse rendering method. To enhance efficiency and practicality, we leverage SG to represent the smooth environment illuminations and apply importance sampling techniques. To supervise indirect illuminations from unobserved directions, we develop a novel radiance consistency constraint between implicit neural radiance and path tracing results of unobserved rays along with the joint optimization of materials and illuminations, thus significantly improving the decomposition performance. Extensive experiments demonstrate that our method outperforms the state-of-the-art on multiple synthetic and real datasets, especially in terms of inter-reflection decomposition.
Haoqian Wu, Zhipeng Hu, Lincheng Li, Yongqiang Zhang 0003, Changjie Fan, Xin Yu 0002
CVPR2
2023 Towards Unbiased Volume Rendering of Neural Implicit Surfaces with Geometry Priors
abstract
Learning surface by neural implicit rendering has been a promising way for multi-view reconstruction in recent years. Existing neural surface reconstruction methods, such as NeuS [24] and VolSDF [32], can produce reliable meshes from multi-view posed images. Although they build a bridge between volume rendering and Signed Distance Function (SDF), the accuracy is still limited. In this paper, we argue that this limited accuracy is due to the bias of their volume rendering strategies, especially when the viewing direction is close to be tangent to the surface. We revise and provide an additional condition for the unbiased volume rendering. Following this analysis, we propose a new rendering method by scaling the SDF field with the angle between the viewing direction and the surface normal vector. Experiments on simulated data indicate that our rendering method reduces the bias of SDF-based volume rendering. Moreover, there still exists non-negligible bias when the learnable standard deviation of SDF is large at early stage, which means that it is hard to supervise the rendered depth with depth priors. Alternatively we supervise zero-level set with surface points obtained from a pre-trained Multi-View Stereo network. We evaluate our method on the DTU dataset and show that it outperforms the state-of-the-arts neural implicit surface methods without mask supervision.
Yongqiang Zhang 0003, Zhipeng Hu, Haoqian Wu, Minda Zhao, Lincheng Li, Zhengxia Zou, Changjie Fan
CVPR2
2023 I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning
abstract
Image Captioning is a traditional vision-and-language task that aims to generate the language description of an image. Recent studies focus on scaling up the model size and the number of training data, which significantly increase the cost of model training. Different to these heavy-cost models, we introduce a lightweight image captioning framework (I-Tuning), which contains a small number of trainable parameters. We design a novel I-Tuning cross-attention module to connect the non-trainable pre-trained language decoder GPT2 and vision encoder CLIP-ViT. Since most parameters are not required to be updated during training, our framework is lightweight and fast. Experimental results conducted on three image captioning benchmarks reveal that our frame-work achieves comparable or better performance than the large-scale baseline systems. But our models contain up to 10 times fewer trainable parameters and require much fewer data for training compared with state-of-the-art baselines.
Zhipeng Hu, Yadong Xi, Jing Ma 0004
ICASSP2
2023 Tailoring Language Generation Models under Total Variation Distance
Haozhe Ji, Pei Ke, Zhipeng Hu, Minlie Huang
ICLR3
2023 Deep learning applications in games: a survey from a data perspective
Zhipeng Hu, Yu Ding 0001, Runze Wu 0001, Lincheng Li, Yujing Hu, Kai Wang 0064, Yongqiang Zhang 0003, Ji Jiang, Yadong Xi, Jiashu Pu, Wei Zhang 0219, Suzhen Wang 0001, Ke Chen 0005, Tianze Zhou, Jiarui Chen, Tangjie Lv, Changjie Fan
Appl. Intell.1
2023 Explainable AI for Cheating Detection and Churn Prediction in Online Games
abstract
Online gaming is a multibillion dollar industry that entertains a large, global population. Empowering online games with AI has made a great success, however, ignores the explainability of black-box model makes AI less responsible and hinders its further development. In this article, we introduce and discuss the audience and the concept of XAI (eXplainable AI) in online games. We propose a GXAI workflow, which combines the strong expressiveness of multiview data sources and the clear transparency of multiview black-box models. We present four specific classifiers and explainers in the character portrait view, the behavior sequence view, the client image view, and the social graph view. Experiments conducted on real-world datasets for game cheating detection and player churn prediction show the accuracy of classification and the rationality of explanation. We also discover and present numerous interesting and valuable findings from the individual, local, and global explanations. We implement and deploy three practical applications, including evidence and reason generation, model debugging and testing, and model compression and comparison in NetEase Games and have received quite positive reviews from user studies. More future work is in progress since this is the first work that introduces XAI in online games.
Jianrong Tao, Runze Wu 0001, Tangjie Lyu, Changjie Fan, Zhipeng Hu, Sha Zhao, Gang Pan 0001
IEEE Trans. Games8
2022 LaMemo: Language Modeling with Look-Ahead Memory
abstract
Haozhe Ji, Rongsheng Zhang, Zhenyu Yang, Zhipeng Hu, Minlie Huang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Haozhe Ji, Zhipeng Hu, Minlie Huang
NAACL-HLT4
2021 Automatic Translation of Music-to-Dance for In-Game Characters
abstract
Music-to-dance translation is an emerging and powerful feature in recent role-playing games. Previous works of this topic consider music-to-dance as a supervised motion generation problem based on time-series data. However, these methods require a large amount of training data pairs and may suffer from the degradation of movements. This paper provides a new solution to this task where we re-formulate the translation as a piece-wise dance phrase retrieval problem based on the choreography theory. With such a design, players are allowed to optionally edit the dance movements on top of our generation while other regression-based methods ignore such user interactivity. Considering that the dance motion capture is expensive that requires the assistance of professional dancers, we train our method under a semi-supervised learning fashion with a large unlabeled music dataset (20x than our labeled one) and also introduce self-supervised pre-training to improve the training stability and generalization performance. Experimental results suggest that our method not only generalizes well over various styles of music but also succeeds in choreography for game players. Our project including the large-scale dataset and supplemental materials is available at https://github.com/FuxiCV/music-to-dance.
Yinglin Duan, Tianyang Shi, Zhipeng Hu, Zhengxia Zou, Changjie Fan, Yi Yuan 0002
IJCAI3
2021 A 57.2-Gb/s PAM4 Driver for a Segmented Silicon-Photonics Mach-Zehnder Modulator with Extinction Ratio >9-dB in 45-nm RF-SOI CMOS Technology
abstract
A 57.2-Gb/s four-level pulse-amplitude modulation (PAM4) driver for silicon photonic Mach-Zehnder modulator (MZM) is presented. The driver is designed in a 45-nm RF-SOI CMOS technology and consists of a pre-driver and an outputdriver. The pre-driver is made up with scaled, cascaded single- ended CMOS inverter transimpedance amplifiers with inductive and resistive feedback and interstage series inductive peaking. The output-driver adapts the structure of series-stacked cascoded CMOS inverter to overcome the breakdown voltage (BV) limit of transistors in advance technologies while getting high differential voltage swing >4 Vpp for MZM. The MZM is configured as an optical digital-analog converter (ODAC) and consists of a 1.1 mm LSB segment and a 1.9 mm MSB segment. The post-layout simulation results show that 5.6-Vpp differential swing, 9.1 dB extinction ratio (ER) and 95.2% ratio of level mismatch (RLM) at 28.6-GBaud (57.2-Gb/s) with a total power consumption of 750 mW are realized. Furthermore, a larger ER can be obtained by increasing the supply voltage. This work achieves an optimized tradeoff between power consumption, data rate, output swing, chip area and extinction ratio.
Min Tan 0004, Dezhi Xing, Sizhu Shao, Zhipeng Hu, Junbo Feng
ISCAS5
2021 Globally Optimized Matchmaking in Online Games
abstract
As one of the core components of online games, matchmaking is the process of arranging multiple players into matches, where the quality of matchmaking systems directly determines player satisfaction and further affects the life cycle of game products. With the number of candidate players increases, the number of possible match combinations grows exponentially, which makes the current implementation for multiplayer matchmaking can only obtain locally optimal arrangement in an inefficient fashion. In this paper, we focus on the globally optimized matchmaking problem, in which the objective is to decide an optimal matching sequence for the queuing players. To tackle this challenging problem, we propose a novel data-driven matchmaking framework, called GloMatch, based on machine learning principles. Through transforming the matchmaking problem into a sequential decision problem, we solve it with the help of an effective policy-based deep reinforcement learning algorithm. Quantitative experiments on simulation and online game environments demonstrate the effectiveness of the presented framework.
Kai Wang 0064, Zhipeng Hu, Runze Wu 0001, Linxia Gong, Jianrong Tao, Changjie Fan, Peng Cui 0001
KDD4
2021 Towards Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games
abstract
Measuring and promoting policy diversity is critical for solving games with strong non-transitive dynamics where strategic cycles exist, and there is no consistent winner (e.g., Rock-Paper-Scissors). With that in mind, maintaining a pool of diverse policies via open-ended learning is an attractive solution, which can generate auto-curricula to avoid being exploited. However, in conventional open-ended learning algorithms, there are no widely accepted definitions for diversity, making it hard to construct and evaluate the diverse policies. In this work, we summarize previous concepts of diversity and work towards offering a unified measure of diversity in multi-agent open-ended learning to include all elements in Markov games, based on both Behavioral Diversity (BD) and Response Diversity (RD). At the trajectory distribution level, we re-define BD in the state-action space as the discrepancies of occupancy measures. For the reward dynamics, we propose RD to characterize diversity through the responses of policies when encountering different opponents. We also show that many current diversity measures fall in one of the categories of BD or RD but not both. With this unified diversity measure, we design the corresponding diversity-promoting objective and population effectivity when seeking the best responses in open-ended learning. We validate our methods in both relatively simple games like matrix game, non-transitive mixture model, and the complex \textit{Google Research Football} environment. The population found by our methods reveals the lowest exploitability, highest population effectivity in matrix game and non-transitive mixture model, as well as the largest goal difference when interacting with opponents of various levels in \textit{Google Research Football}.
Hangtian Jia, Ying Wen 0001, Yujing Hu, Changjie Fan, Zhipeng Hu, Yaodong Yang 0001
NeurIPS7
2021 GLIB: towards automated test oracle for graphically-rich applications
abstract
Graphically-rich applications such as games are ubiquitous with attractive visual effects of Graphical User Interface (GUI) that offers a bridge between software applications and end-users. However, various types of graphical glitches may arise from such GUI complexity and have become one of the main component of software compatibility issues. Our study on bug reports from game development teams in NetEase Inc. indicates that graphical glitches frequently occur during the GUI rendering and severely degrade the quality of graphically-rich applications such as video games. Existing automated testing techniques for such applications focus mainly on generating various GUI test sequences and check whether the test sequences can cause crashes. These techniques require constant human attention to captures non-crashing bugs such as bugs causing graphical glitches. In this paper, we present the first step in automating the test oracle for detecting non-crashing bugs in graphically-rich applications. Specifically, we propose GLIB based on a code-based data augmentation technique to detect game GUI glitches. We perform an evaluation of GLIB on 20 real-world game apps (with bug reports available) and the result shows that GLIB can achieve 100% precision and 99.5% recall in detecting non-crashing bugs such as game GUI glitches. Practical application of GLIB on another 14 real-world games (without bug reports) further demonstrates that GLIB can effectively uncover GUI glitches, with 48 of 53 bugs reported by GLIB having been confirmed and fixed so far.
Ke Chen 0005, Yufei Li 0001, Changjie Fan, Zhipeng Hu, Wei Yang 0013
ESEC/SIGSOFT FSE5
2021 Asyncflow: A visual programming tool for game artificial intelligence
abstract
Visual programming tools are widely applied in the game industry to assist game designers in developing game artificial intelligence (game AI) and gameplay. However, testing multiple game engines is a time-consuming operation, which degrades development efficiency. To provide an asynchronous platform for game designers, this paper introduces Asyncflow, an open-source visual programming solution. It consists of a flowchart maker for game logic explanation and a runtime framework integrating an asynchronous mechanism based on an event-driven architecture. Asyncflow supports multiple programming languages and can be easily embedded in various game engines to run flowcharts created by game designers.
Zhipeng Hu, Changjie Fan, Qiwei Zheng, Bai Liu 0002
Vis. Informatics1
2013 Selective Eigenbackground for Background Modeling and Subtraction in Crowded Scenes
abstract
Background subtraction is a fundamental preprocessing step in many surveillance video analysis tasks. In spite of significant efforts, however, background subtraction in crowded scenes remains challenging, especially, when a large number of foreground objects move slowly or just keep still. To address the problem, this paper proposes a selective eigenbackground method for background modeling and subtraction in crowded scenes. The contributions of our method are three-fold: First, instead of training eigenbackgrounds using the original video frames that may contain more or less foregrounds, a virtual frame construction algorithm is utilized to assemble clean background pixels from different original frames so as to construct some virtual frames as the training and update samples. This can significantly improve the purity of the trained eigenbackgrounds. Second, for a crowded scene with diversified environmental conditions (e.g., illuminations), it is difficult to use only one eigenbackground model to deal with all these variations, even using some online update strategies. Thus given several models trained offline, we utilize peak signal-to-noise ratio to adaptively choose the optimal one to initialize the online eigenbackground model. Third, to tackle the problem that not all pixels can obtain the optimal results when the reconstruction is performed at once for the whole frame, our method selects the best eigenbackground for each pixel to obtain an improved quality of the reconstructed background image. Extensive experiments on the TRECVID-SED dataset and the Road video dataset show that our method outperforms several state-of-the-art methods remarkably.
Yonghong Tian 0001, Yaowei Wang 0001, Zhipeng Hu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2011 Selective eigenbackgrounds method for background subtraction in crowed scenes
abstract
In this paper, a selective eigenbackgrounds method is proposed for background subtraction in crowded scenes. In order to train and update the eigenbackground model with frames containing few objects (i.e. clean frames), virtual frames are constructed based on a frame selection map. Then, the eigenbackground that best depicts background is selected for each pixel based on an eigenbackground selection map. Experimental results show the performance of the proposed method is better than those of some state-of-the-art methods in crowded scenes.
Zhipeng Hu, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001
ICIP1
2010 ESUR: A system for Events detection in SURveillance video
abstract
In this paper, we present our eSur (Event detection system on SURveillance video) system, which is derived from TRECVID'09 surveillance tasks. Currently, eSur attempts to detect two categories of events: 1) single-actor events (i.e., PersonRuns and ElevatorNoEntry) irrespective of any interaction between individuals, and 2) pair-activity events (i.e., PeopleMeet, PeopleSplitUp, and Embrace) involves more than one individual. eSur consists of three major stages, i.e., preprocessing, event classification, and post-processing. The preprocessing involves view classification, background subtraction, head-shoulder detection, human body detection and object tracking. Event classification fuses One-vs.-All SVM and rule-based classifiers to identify single-actor and pair-activity events in an ensemble way. To reduce false alarms, we introduce prior knowledge into the post-processing, and in particular, we apply a so-called event merging process over TRECVID dataset. Extensive experiments have been performed over TRECVid'08 and '09 ED data corpus involving in total 144 hours surveillance video of London Gatwick airport. According to the TRECVid-ED formal evaluation, our prototype has yielded fairly promising results over TRECVid'09 dataset, with top Act.DCR of 1.023, 1.025, 1.02, and 0.334 for PeopleMeet, PeopleSplitUp, Embrace, and ElevatorNoEntry, respectively.
Yaowei Wang 0001, Yonghong Tian 0001, Ling-Yu Duan, Zhipeng Hu, Guochen Jia
ICIP4