VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0219
dblp:10/4661-219
· DBLP profile ↗
25ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-5907-7342ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAVA: Decoding Art With Visual Analytics Through Feature Modeling and Multi-Agent CollaborationabstractFigurative art, as a culturally embedded medium, encodes narrative, symbolic, and emotional meanings that reflect artistic choices and historical realities. The recent availability of large-scale digital collections of figurative artworks creates opportunities for computational analysis, but existing methods mostly focus on classification or style detection, lacking structured modeling of high-level figurative elements and integration of cultural context. We present DAVA, a visual analytics system that supports interdisciplinary exploration of figurative art. First, we model paintings across three structural levels: facial expressions (micro), posture features (meso), and object co-occurrence (macro). Second, we employ a vision-language model to discover latent patterns from these features and present them through novel visualization designs. Third, we introduce domain-informed AI agents that simulate interdisciplinary research teams to interpret artworks in cultural and historical context. To evaluate DAVA, we first conducted a quantitative evaluation demonstrating the accuracy and consistency of the multi-agent interpretation mechanism. Case studies and expert interviews then confirmed the system's utility and support for semantically and historically informed exploration of figurative art. Wei Zhang 0219, Hengru Liu, Peiyi Jiang, Xianfeng Peng, Zhenqian Xu, Yifang Wang 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | EasyCraft: A Robust and Efficient Framework for Automatic Avatar CraftingabstractCharacter customization, or ’face crafting,’ is a vital feature in role-playing games (RPGs), enhancing player engagement by enabling the creation of personalized avatars. Existing automated methods often struggle with generalizability across diverse game engines due to their reliance on the intermediate constraints of specific image domain and typically support only one type of input, either text or image. To overcome these challenges, we introduce EasyCraft, an innovative end-to-end feedforward framework that automates character crafting by uniquely supporting both text and image inputs. Our approach employs a translator capable of converting facial images of any style into crafting parameters. We first establish a unified feature distribution in the translator’s image encoder through self-supervised learning on a large-scale dataset, enabling photos of any style to be embedded into a unified feature representation.Subsequently, we map this unified feature distribution to crafting parameters specific to a game engine, a process that can be easily adapted to most game engines and thus enhances EasyCraft’s generalizability. By integrating text-to-image techniques with our translator, EasyCraft also facilitates precise, text-based character crafting. EasyCraft’s ability to integrate diverse inputs significantly enhances the versatility and accuracy of avatar creation. Extensive experiments on two RPG games demonstrate the effectiveness of our method, achieving state-of-the-art results and facilitating adaptability across various avatar engines. Suzhen Wang 0001, Wei Zhang 0219, Minda Zhao, Lincheng Li, Zhipeng Hu, Xin Yu 0002 |
CVPR | 3 |
| 2025 | CultiVerse: Towards Cross-Cultural Understanding for Paintings with Large Language ModelabstractUnderstanding cultural heritage through technology faces challenges in connecting with diverse audiences, especially when interpreting art across cultures. In this work, we present CultiVerse, a visual analytics system that leverages Large Language Models (LLMs) to support cross-cultural appreciation of Traditional Chinese Paintings (TCPs). CultiVerse operates within a mixed-initiative framework and guides users through three stages: extracting cultural context, aligning cross-cultural symbols, and extrapolating meaning in the viewer's cultural frame. By combining an interactive interface with LLM-powered analysis, the system enables deeper engagement with symbolic meanings and encourages serendipitous cross-cultural discoveries. Our approach bridges AI interpretation and human insight to foster mutual understanding in a multicultural setting. A curated TCP dataset supports exploration, while empirical evaluations confirm that CultiVerse enhances user understanding, interpretation accuracy, and cultural empathy. Wei Zhang 0219, Kamkwai Wong, Biying Xu, Yiwen Ren, Yuhuai Li, Yingchaojie Feng, Minfeng Zhu 0001, Wei Chen 0001 |
ACM Multimedia | 1 |
| 2025 | CausalPrism: A visual analytics approach for subgroup-based causal heterogeneity exploration
Xingyu Liu 0003, Jiehui Zhou, Xumeng Wang, Kamkwai Wong, Wei Zhang 0219, Juntian Zhang, Minfeng Zhu 0001, Wei Chen 0001 |
Comput. Graph. | 5 |
| 2025 | InkSpirit: An expert knowledge-driven approach for enhancing the visual logic of traditional Chinese painting text-to-image generation
Xiangsheng Zeng, Runqiao Xia, Yongbo Jiang, Yingchaojie Feng, Wei Zhang 0219, Wei Chen 0001 |
Comput. Graph. | 8 |
| 2025 | JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language ModelsabstractThe proliferation of large language models (LLMs) has underscored concerns regarding their security vulnerabilities, notably against jailbreak attacks, where adversaries design jailbreak prompts to circumvent safety mechanisms for potential misuse. Addressing these concerns necessitates a comprehensive analysis of jailbreak prompts to evaluate LLMs' defensive capabilities and identify potential weaknesses. However, the complexity of evaluating jailbreak performance and understanding prompt characteristics makes this analysis laborious. We collaborate with domain experts to characterize problems and propose an LLM-assisted framework to streamline the analysis process. It provides automatic jailbreak assessment to facilitate performance evaluation and support analysis of components and keywords in prompts. Based on the framework, we design JailbreakLens, a visual analysis system that enables users to explore the jailbreak performance against the target model, conduct multi-level analysis of prompt characteristics, and refine prompt instances to verify findings. Through a case study, technical evaluations, and expert interviews, we demonstrate our system's effectiveness in helping users evaluate model security and identify model weaknesses. Yingchaojie Feng, Zhizhang (David) Chen, Zhining Kang, Wei Zhang 0219, Minfeng Zhu 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Towards a Simultaneous and Granular Identity-Expression Control in Personalized Face GenerationabstractIn human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted por-trait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts towards personalized face generation. To this end, we propose a novel multi-modal face generation frame-work, capable of simultaneous identity-expression control and more fine-grained expression synthesis. Our expression control is so sophisticated that it can be specialized by the fine-grained emotional vocabulary. We devise a novel dif-fusion model that can undertake the task of simultaneously face swapping and reenactment. Due to the entanglement of identity and expression, separately and precisely control-ling them within one framework is a nontrivial task, thus has not been explored yet. To overcome this, we propose sev-eral innovative designs in the conditional diffusion model, including balancing identity and expression encoder, improved midpoint sampling, and explicitly background con-ditioning. Extensive experiments have demonstrated the controllability and scalability of the proposed framework, in comparison with state-of-the-art text-to-image, face swap-ping, and face reenactment methods. Renshuai Liu, Wei Zhang 0219, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding 0001 |
CVPR | 3 |
| 2024 | Norface: Improving Facial Expression Analysis by Identity Normalization
Hanwei Liu, Rudong An, Wei Zhang 0219, Yujing Hu, Yu Ding 0001 |
ECCV (55) | 5 |
| 2024 | FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation ModelabstractVideo-driven 3D facial animation transfer aims to drive avatars to reproduce the expressions of actors. Existing methods have achieved remarkable results by constraining both geometric and perceptual consistency. However, geometric constraints (like those designed on facial landmarks) are insufficient to capture subtle emotions, while expression features trained on classification tasks lack fine granularity for complex emotions. To address this, we propose \textbf{FreeAvatar}, a robust facial animation transfer method that relies solely on our learned expression representation. Specifically, FreeAvatar consists of two main components: the expression foundation model and the facial animation transfer model. In the first component, we initially construct a facial feature space through a face reconstruction task and then optimize the expression feature space by exploring the similarities among different expressions. Benefiting from training on the amounts of unlabeled facial images and re-collected expression comparison dataset, our model adapts freely and effectively to any in-the-wild input facial images. In the facial animation transfer component, we propose a novel Expression-driven Multi-avatar Animator, which first maps expressive semantics to the facial control parameters of 3D avatars and then imposes perceptual constraints between the input and output images to maintain expression consistency. To make the entire process differentiable, we employ a trained neural renderer to translate rig parameters into corresponding images. Furthermore, unlike previous methods that require separate decoders for each avatar, we propose a dynamic identity injection module that allows for the joint training of multiple avatars within a single network. Wei Zhang 0219, Chen Liu 0028, Rudong An, Lincheng Li, Yu Ding 0001, Changjie Fan, Zhipeng Hu, Xin Yu 0002 |
SIGGRAPH Asia | 2 |
| 2024 | Learning facial expression-aware global-to-local representation for robust action unit detection
Rudong An, Aobo Jin, Wei Chen 0157, Wei Zhang 0219, Hao Zeng 0001, Zhigang Deng 0001, Yu Ding 0001 |
Appl. Intell. | 4 |
| 2024 | Computational Approaches for Traditional Chinese Painting: From the "Six Principles of Painting" Perspective
Wei Zhang 0219, Jianwei Zhang 0015, Kamkwai Wong, Yi-Fang Wang, Yingchaojie Feng, Lu-Wei Wang, Wei Chen 0001 |
J. Comput. Sci. Technol. | 1 |
| 2024 | Learning a compact embedding for fine-grained few-shot static gesture recognition
Zhipeng Hu, Wei Zhang 0219, Yu Ding 0001, Tangjie Lv, Changjie Fan |
Multim. Tools Appl. | 4 |
| 2024 | Facial Action Unit Detection and Intensity Estimation From Self-Supervised RepresentationabstractAs a fine-grained and local expression behavior measurement, facial action unit (FAU) analysis (e.g., detection and intensity estimation) has been documented for its time-consuming, labor-intensive, and error-prone annotation. Thus a long-standing challenge of FAU analysis arises from the data scarcity of manual annotations, limiting the generalization ability of trained models to a large extent. Amounts of previous works have made efforts to alleviate this issue via semi/weakly supervised methods and extra auxiliary information. However, these methods still require domain knowledge and have not yet avoided the high dependency on data annotation. This article introduces a robust facial representation model MAE-Face for AU analysis. Using masked autoencoding as the self-supervised pre-training approach, MAE-Face first learns a high-capacity model from a feasible collection of face images without additional data annotations. Then after being fine-tuned on AU datasets, MAE-Face exhibits convincing performance for both AU detection and AU intensity estimation, achieving a new state-of-the-art on nearly all the evaluation results. Further investigation shows that MAE-Face achieves decent performance even when fine-tuned on only 1% of the AU training set, strongly proving its robustness and generalization performance. The pre-trained model is available at our GitHub repository. Rudong An, Wei Zhang 0219, Yu Ding 0001, Zeng Zhao, Tangjie Lv, Changjie Fan, Zhipeng Hu |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Detecting Facial Action Units From Global-Local Fine-Grained ExpressionsabstractSince Facial Action Unit (AU) annotations require domain expertise, common AU datasets only contain a limited number of subjects. As a result, a crucial challenge for AU detection is addressing identity overfitting. We find that AUs and facial expressions are highly associated, and existing facial expression datasets often contain a large number of identities. In this paper, we aim to utilize the expression datasets without AU labels to facilitate AU detection. Specifically, we develop a novel AU detection framework aided by the Global-Local facial Expressions Embedding, dubbed GLEE-Net. Our GLEE-Net consists of three branches to extract identity-independent expression features for AU detection. We introduce a global branch for modeling the overall facial expression while eliminating the impacts of identities. We also design a local branch focusing on specific local face regions. The combined output of global and local branches is firstly pre-trained on an expression dataset as an identity-independent expression embedding, and then finetuned on AU datasets. Therefore, we significantly alleviate the issue of limited identities. Furthermore, we introduce a 3D global branch that extracts expression coefficients through 3D face reconstruction to consolidate 2D AU descriptions. Finally, a Transformer-based multi-label classifier is employed to fuse all the representations for AU detection. Extensive experiments demonstrate that our method significantly outperforms the state-of-the-art on the widely-used DISFA, BP4D and BP4D+ datasets. Wei Zhang 0219, Lincheng Li, Yu Ding 0001, Wei Chen 0157, Zhigang Deng 0001, Xin Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | ScrollTimes: Tracing the Provenance of Paintings as a Window Into HistoryabstractThe study of cultural artifact provenance, tracing ownership and preservation, holds significant importance in archaeology and art history. Modern technology has advanced this field, yet challenges persist, including recognizing evidence from diverse sources, integrating sociocultural context, and enhancing interactive automation for comprehensive provenance analysis. In collaboration with art historians, we examined the handscroll, a traditional Chinese painting form that provides a rich source of historical data and a unique opportunity to explore history through cultural artifacts. We present a three-tiered methodology encompassing artifact, contextual, and provenance levels, designed to create a "Biography" for handscroll. Our approach incorporates the application of image processing techniques and language models to extract, validate, and augment elements within handscroll using various cultural heritage databases. To facilitate efficient analysis of non-contiguous extracted elements, we have developed a distinctive layout. Additionally, we introduce ScrollTimes, a visual analysis system tailored to support the three-tiered analysis of handscroll, allowing art historians to interactively create biographies tailored to their interests. Validated through case studies and expert interviews, our approach offers a window into history, fostering a holistic understanding of handscroll provenance and historical significance. Wei Zhang 0219, Kamkwai Wong, Yitian Chen 0004, Ailing Jia, Luwei Wang, Jianwei Zhang 0015, Lechao Cheng, Huamin Qu, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | FlowFace: Semantic Flow-Guided Shape-Aware Face SwappingabstractIn this work, we propose a semantic flow-guided two-stage framework for shape-aware face swapping, namely FlowFace. Unlike most previous methods that focus on transferring the source inner facial features but neglect facial contours, our FlowFace can transfer both of them to a target face, thus leading to more realistic face swapping. Concretely, our FlowFace consists of a face reshaping network and a face swapping network. The face reshaping network addresses the shape outline differences between the source and target faces. It first estimates a semantic flow (i.e. face shape differences) between the source and the target face, and then explicitly warps the target face shape with the estimated semantic flow. After reshaping, the face swapping network generates inner facial features that exhibit the identity of the source face. We employ a pre-trained face masked autoencoder (MAE) to extract facial features from both the source face and the target face. In contrast to previous methods that use identity embedding to preserve identity information, the features extracted by our encoder can better capture facial appearances and identity information. Then, we develop a cross-attention fusion module to adaptively fuse inner facial features from the source face with the target facial attributes, thus leading to better identity preservation. Extensive quantitative and qualitative experiments on in-the-wild faces demonstrate that our FlowFace outperforms the state-of-the-art significantly. Hao Zeng 0001, Wei Zhang 0219, Changjie Fan, Tangjie Lv, Suzhen Wang 0001, Lincheng Li, Yu Ding 0001, Xin Yu 0002 |
AAAI | 2 |
| 2023 | NFTVis: Visual Analysis of NFT PerformanceabstractA non-fungible token (NFT) is a data unit stored on the blockchain. Nowadays, more and more investors and collectors (NFT traders), who participate in transactions of NFTs, have an urgent need to assess the performance of NFTs. However, there are two challenges for NFT traders when analyzing the performance of NFT. First, the current rarity models have flaws and are sometimes not convincing. In addition, NFT performance is dependent on multiple factors, such as images (high-dimensional data), history transactions (network), and market evolution (time series). It is difficult to take comprehensive consideration and analyze NFT performance efficiently. To address these challenges, we propose NFTVis, a visual analysis system that facilitates assessing individual NFT performance. A new NFT rarity model is proposed to quantify NFTs with images. Four well-coordinated views are designed to represent the various factors affecting the performance of the NFT. Finally, we evaluate the usefulness and effectiveness of our system using two case studies and user studies. Fan Yan, Xumeng Wang, Ketian Mao, Wei Zhang 0219, Wei Chen 0001 |
PacificVis | 4 |
| 2023 | Deep learning applications in games: a survey from a data perspective
Zhipeng Hu, Yu Ding 0001, Runze Wu 0001, Lincheng Li, Yujing Hu, Kai Wang 0064, Yongqiang Zhang 0003, Ji Jiang, Yadong Xi, Jiashu Pu, Wei Zhang 0219, Suzhen Wang 0001, Ke Chen 0005, Tianze Zhou, Jiarui Chen, Tangjie Lv, Changjie Fan |
Appl. Intell. | 15 |
| 2023 | Face identity and expression consistency for game character face swapping
Hao Zeng 0001, Wei Zhang 0219, Lincheng Li, Yu Ding 0001 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Visual Reasoning for Uncertainty in Spatio-Temporal Events of Historical FiguresabstractThe development of digitized humanity information provides a new perspective on data-oriented studies of history. Many previous studies have ignored uncertainty in the exploration of historical figures and events, which has limited the capability of researchers to capture complex processes associated with historical phenomena. We propose a visual reasoning system to support visual reasoning of uncertainty associated with spatio-temporal events of historical figures based on data from the China Biographical Database Project. We build a knowledge graph of entities extracted from a historical database to capture uncertainty generated by missing data and error. The proposed system uses an overview of chronology, a map view, and an interpersonal relation matrix to describe and analyse heterogeneous information of events. The system also includes uncertainty visualization to identify uncertain events with missing or imprecise spatio-temporal information. Results from case studies and expert evaluations suggest that the visual reasoning system is able to quantify and reduce uncertainty generated by the data. Wei Zhang 0219, Siwei Tan, Siming Chen 0001, Linghao Meng, Tian-Ye Zhang, Rongchen Zhu, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | CohortVA: A Visual Analytic System for Interactive Exploration of Cohorts based on Historical DataabstractIn history research, cohort analysis seeks to identify social structures and figure mobilities by studying the group-based behavior of historical figures. Prior works mainly employ automatic data mining approaches, lacking effective visual explanation. In this paper, we present CohortVA, an interactive visual analytic approach that enables historians to incorporate expertise and insight into the iterative exploration process. The kernel of CohortVA is a novel identification model that generates candidate cohorts and constructs cohort features by means of pre-built knowledge graphs constructed from large-scale history databases. We propose a set of coordinated views to illustrate identified cohorts and features coupled with historical events and figure profiles. Two case studies and interviews with historians demonstrate that CohortVA can greatly enhance the capabilities of cohort identifications, figure authentications, and hypothesis generation. Wei Zhang 0219, Jason K. Wong, Xumeng Wang, Youcheng Gong, Rongchen Zhu, Siwei Tan, Huamin Qu, Siming Chen 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | Paste You Into Game: Towards Expression and Identity Consistency Face SwappingabstractCustomizing game characters for individual players has been a long-standing attractive feature in the game industry. However, traditional solutions like manual editing within a game engine are always time-consuming and unsatisfying. Our work proposes a novel automatic face swapping method for arbitrary users and game characters, addressing three challenges including style gap between human and game faces, identity preservation, and expression consistency. A game face dataset is collected to handle the cross-style gap; an identity compound embedding is proposed to ease the bias existing in the commonly-used ID identifiers and it provides a more robust identity representation; a novel expression embedding loss is proposed to enforce the expression consistency between the swapped and target faces and it achieves better expression consistency than the previous methods, especially when the expression is very subtle. The visualized results, as well as the qualitative and quantitative comparisons, reveal the significance and effectiveness of our proposed solutions. Hao Zeng 0001, Wei Zhang 0219, Lincheng Li, Yu Ding 0001 |
CoG | 2 |
| 2022 | Semantic-Rich Facial Emotional Expression RecognitionabstractThe ability to perceive human facial emotions is an essential feature of various multi-modal applications, especially in the intelligent human-computer interaction (HCI) area. In recent decades, considerable efforts have been put into researching automatic facial emotion recognition (FER). However, most of the existing FER methods only focus on either basic emotions such as the seven/eight categories (e.g.,happiness, angerandsurprise) or abstract dimensions (valence, arousal, etc.), while neglecting the fruitful nature of emotion statements. In real-world scenarios, there is definitely a larger vocabulary for describing human's inner feelings as well as their reflection on facial expressions. In this work, we propose to address the semantic richness issue in the FER problem, with an emphasis on the granularity of the emotion concepts. Particularly, we take inspiration from former psycho-linguistic research, which conducted a prototypicality rating study and chose 135 emotion names from hundreds of English emotion terms. Based on the 135 emotion categories, we investigate the corresponding facial expressions by collecting a large-scale 135-class FER image dataset and propose a consequent facial emotion recognition framework. To demonstrate the accessibility of prompting FER research to a fine-grained level, we conduct extensive evaluations on the dataset credibility and the accompanying baseline classification model. The qualitative and quantitative results prove that the problem is meaningful and our solution is effective. To the best of our knowledge, this is the first work aimed at exploiting such a large semantic space for emotion representation in the FER problem. Changjie Fan, Wei Zhang 0219, Yu Ding 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2021 | Learning a Facial Expression Embedding Disentangled From IdentityabstractThe facial expression analysis requires a compact and identity-ignored expression representation. In this paper, we model the expression as the deviation from the identity by a subtraction operation, extracting a continuous and identity-invariant expression embedding. We propose a Deviation Learning Network (DLN) with a pseudo-siamese structure to extract the deviation feature vector. To reduce the optimization difficulty caused by additional fully connection layers, DLN directly provides high-order polynomial to nonlinearly project the high-dimensional feature to a low-dimensional manifold. Taking label noise into account, we add a crowd layer to DLN for robust embedding extraction. Also, to achieve a more compact representation, we use hierarchical annotation for data augmentation. We evaluate our facial expression embedding on the FEC validation set. The quantitative results prove that we achieve the state-of-the-art, both in terms of fine-grained and identity-invariant property. We further conduct extensive experiments to show that our expression embedding is of high quality for expression recognition, image retrieval, and face manipulation. Wei Zhang 0219, Xianpeng Ji, Yu Ding 0001, Changjie Fan |
CVPR | 1 |
| 2021 | Visual storytelling of Song Ci and the poets in the social-cultural context of Song dynastyabstractSong Ci is treasured in traditional Chinese culture, which indicates social and cultural evolution in ancient times. Despite the efforts by historians and litterateurs in investigating the characteristics of Song Ci, it is still unclear how to effectively distribute and promote Song Ci in the public sphere. The complexity and abstraction of Song Ci hamper the general public from closely reading, analyzing, and appreciating these excellent works. By means of a set of visual analysis methods, e.g. the spatio-temporal visualization, we exploit visual storytelling to explicitly present the latent and abstractive features of Song Ci. We apply straightway visual charts and lighten the burden of understanding the stories, in order to achieve an effective public distribution. The effectiveness and aesthetics of our work are demonstrated by a user study of three participants with different backgrounds. The result reveals that our story is effective in the distribution, understanding, and promotion of Song Ci. Wei Zhang 0219, Rusheng Pan, Wei Chen 0001 |
Vis. Informatics | 1 |