Yangfu Zhu

dblp:263/5827 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0003-2880-9517ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Diffusion-based Augmentation for Recommendation
abstract
Recommendation systems often rely on implicit feedback, where only positive user-item interactions can be observed. Negative sampling is therefore crucial to provide proper negative training signals. However, existing methods tend to mislabel potentially positive but unobserved items as negatives and lack precise control over negative sample selection. We aim to address these by generating controllable negative samples, rather than sampling from the existing item pool. In this context, we propose Adaptive Diffusion-based Augmentation for Recommendation (ADAR), a novel and model-agnostic module that leverages diffusion to synthesize informative negatives. Inspired by the progressive corruption process in diffusion, ADAR simulates a continuous transition from positive to negative, allowing for fine-grained control over sample hardness. To mine suitable negative samples, we theoretically identify the transition point at which a positive sample turns negative and derive a score-aware function to adaptively determine the optimal sampling timestep. By identifying this transition point, ADAR generates challenging negative samples that effectively refine the model's decision boundary. Experiments confirm that ADAR is broadly compatible and boosts the performance of existing recommendation models substantially, including collaborative filtering and sequential recommendation, without architectural modifications.
Fanghui Sun, Yan Zou, Yangfu Zhu, Xiatian Zhu
AAAI4
2026 MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
abstract
Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning.Yet, how much visual information truly contributes to reasoning remains unclear.Existing benchmarks report strong overall performance but seldom isolate the role of the image modality, leaving open whether VLMs genuinely leverage visual understanding or merely depend on linguistic priors.To address this, we present MathSight, a university-level multimodal mathematical reasoning benchmark designed to disentangle and quantify the effect of visual input.Each problem includes multiple visual variants-original, hand-drawn, photocaptured-and a text-only condition for controlled comparison.Experiments on state-ofthe-art VLMs reveal a consistent trend: the contribution of visual information diminishes with increasing problem difficulty.Remarkably, Qwen3-VL without any image input surpasses both its multimodal variants and GPT-5, underscoring the need for benchmarks like MathSight to advance genuine vision-grounded reasoning in future models.The project page is available at https
Yuandong Wang 0002, Yao Cui, Zhen Yang 0034, Yangfu Zhu, Zhenzhou Shao
ACL (1)5
2026 Debiased Multimodal Personality Understanding through Dual Causal Intervention
abstract
Multimodal personality understanding plays a critical role in human-centered artificial intelligence. Previous work mainly focus on learning rich multimodal representations for video personality understanding. However, they often suffer from potential harm caused by subject bias (e.g., observable age and unobservable mental states), as subjects originate from diverse demographic backgrounds. Learning such spurious associations between multimodal features and traits may lead to unfair personality understanding. In this work, we construct a Structural Causal Model (SCM) to analyze the impact of these biases from a causal perspective, and propose a novel Dual Causal Adjustment Network (DCAN) to mitigate the interference of subject attributes on personality understanding. Specifically, we design a Back-door Adjustment Causal Learning (BACL) module to block spurious correlations from observable demographic factors via a prototype-based confounder dictionary, and subsequently apply a Front-door Adjustment Causal Learning (FACL) module to address latent and unobservable biases through a learned mediator dictionary intervention, thereby achieving causal disentanglement of representations for deconfounded reasoning. Importantly, we construct a Demographic-annotated Multimodal Student Personality (DMSP) dataset to support the analysis and discussion of fairness-related factors. Extensive experiments on the benchmark dataset CFI-V2 and our DMSP dataset demonstrate that DCAN consistently improves prediction accuracy, reaching 92.11% and 92.90%, respectively. Meanwhile, the improvements in the fairness metrics of equal opportunity and demographic parity are 6.57% and 7.97% on CFI-V2, and 15.38% and 20.06% on the DMSP dataset. Our code and DMSP dataset are available at https://github.com/Sabrina-han/DCAN
Yangfu Zhu, Zitong Han, Nianwen Ning, Yuandong Wang 0002, Hang Feng, Zhenzhou Shao
SIGIR1
2025 Sketch-Based Poetry Retrieval with Unsupervised Vision-and-Language Pre-training
Yangfu Zhu, Bin Wu 0001
DASFAA (3)3
2025 Learning Causally Disentangled Representations for Fair Personality Detection
abstract
Personality detection aims to identify the personality traits implied in social posts. Existing methods mainly focus on learning the mapping between user-generated posts and personality trait labels but inevitably suffer from potential harm caused by individual bias, as these posts are written by authors from different backgrounds. Learning such spurious associations between posts and traits may lead to the formation of stereotypes, ultimately restricting the detection of personality in different kind of individual. To tackle the issue, we first investigate individual bias in personality detection from the causality perspective. We propose an Interventional Personality Detection Network (IPDN) to learn implicit confounders in user-generated posts and exploit the true causal effect to train the detection model. Specifically, our IPDN disentangled the causal and biased features behind user-generated posts, and then the biased features are accumulatively clustered as confounder prototypes as the training iterations increase. In parallel, the reconstruction network is reused to approximate backdoor adjustment on raw posts, ensuring that traits see each confounder equally before detection. Extensive experiments conducted on three real-world datasets demonstrate that our IPDN outperforms state-of-the-art methods in personality detection.
Yangfu Zhu, Di Liu 0032, Bin Wu 0001
IJCAI1
2025 RAR: A Refined-Attitude Reasoning Framework on LLMs for Zero Shot Stance Detection
abstract
In the context of the rapid evolution of social media, zero-shot stance detection has become an essential task in Natural Language Processing, particularly in effectively identifying attitudes toward unseen targets in texts using large language models. However, existing methods have not fully taken advantage of Large Language Models (LLMs) capabilities in complex reasoning and contextual understanding, often leading to cumbersome approaches. This paper introduces a new framework named Refined-Attitude Reasoning (RAR), which combines example reasoning and attitude refinement techniques to significantly improve the performance of LLMs on stance detection tasks. Specifically, RAR first constructs a case library using information from the training dataset, retrieving examples similar to new input contexts by calculating a mixed weighted representation of text-target pairs. Second, it refines the expression of attitudes through backward reasoning from the cases and extracts broader attitude words to enhance the accuracy of stance identification in texts. Finally, during the generation phase, it combines optimized attitude words with examples for forward stance reasoning and implements a checking mechanism to ensure the accuracy and rationality of the generated results. Experiments show that our RAR framework achieves an overall F1 score exceeding that of other baseline methods on the zero-shot stance detection dataset VAST, demonstrating the importance of fine-grained attitude refinement and in-context learning in enhancing model performance.
Yiming Lee, Yangfu Zhu, Xinming Chen, Bin Wu 0001
IJCNN2
2025 Cause and Effect: Video Social Relationship Recognition from Causal Perspective
abstract
Video social relation recognition is a fundamental task in video understanding, which is dedicated to the construction of multi-modal knowledge graphs. Previous work mainly focuses on multi-modal fusion and the construction of special character graphs. However, they often treat the global frame sequence equally, ignoring the influence of key frame sequence on relation recognition. Specifically, the key frame sequence that significantly reflect character relationships in a video tends to be sparse and short. At the same time, the key frames have not only temporal but also strong causal relationship. Therefore, we propose a novel Video Local Causal Frame (VLCF) model to explore the causal relationship between frames. Inspired by Granger causality theory, we estimate inter-frame causal relationships by comparing the predicted result frames with and without masking the premise frame. We then construct global connections between video frames. Multiple local causal frame sequences and global frame sequences are extracted to capture the key information and global information in the video. Extensive experiments conducted on the ViSR dataset and the MovieGraphs dataset demonstrate that the proposed model achieves state-of-the-art performance.
Yangfu Zhu, Haorui Wang, Guangyao Su, Bin Wu 0001
ACM Multimedia4
2025 A diachronic language model for long-time span classical Chinese
Yangfu Zhu, Yuanxing Xu, Bin Wu 0001
Inf. Process. Manag.3
2025 A general debiasing framework with counterfactual reasoning for multimodal public speaking anxiety detection
Yangfu Zhu, Bin Wu 0001, Chunping Zheng, Jiachen Tan, Zihua Xiong
Neural Networks2
2024 Data Augmented Graph Neural Networks for Personality Detection
abstract
Personality detection is a fundamental task for user psychology research. One of the biggest challenges in personality detection lies in the quantitative limitation of labeled data collected by completing the personality questionnaire, which is very time-consuming and labor-intensive. Most of the existing works are mainly devoted to learning the rich representations of posts based on labeled data. However, they still suffer from the inherent weakness of the amount limitation of labels, which potentially restricts the capability of the model to deal with unseen data. In this paper, we construct a heterogeneous personality graph for each labeled and unlabeled user and develop a novel psycholinguistic augmented graph neural network to detect personality in a semi-supervised manner, namely Semi-PerGCN. Specifically, our model first explores a supervised Personality Graph Neural Network (PGNN) to refine labeled user representation on the heterogeneous graph. For the remaining massive unlabeled users, we utilize the empirical psychological knowledge of the Linguistic Inquiry and Word Count (LIWC) lexicon for multi-view graph augmentation and perform unsupervised graph consistent constraints on the parameters shared PGNN. During the learning process of finite labeled users, noise-invariant learning on a large scale of unlabeled users is combined to enhance the generalization ability. Extensive experiments on three real-world datasets, Youtube, PAN2015, and MyPersonality demonstrate the effectiveness of our Semi-PerGCN in personality detection, especially in scenarios with limited labeled users.
Yangfu Zhu, Yue Xia, Bin Wu 0001
AAAI1
2024 Knowledge-enhanced Multi-Granularity Interaction Network for Political Perspective Detection
abstract
Nowadays, with political ideologies becoming increasingly polarized, detecting political perspectives has become more and more crucial. Previous studies mainly focused on reasoning with real-world entities as background knowledge, while they fail to effectively model hierarchical relationships in the text structure and leave out the document-level knowledge such as topics in a news article. To overcome these limitations, we propose KMGN, a novel Knowledge-enhanced Multi-Granularity Interaction Network for political perspective prediction which consists of two key components: (1) a knowledge infusion method for learning both the topic representations and political entities of a news article (2) a multi-granularity interaction network to learn semantic representations for words, sentences, and document and integrate the infused knowledge across different granularities. Extensive experiments shows that our proposed approach is effective in predicting the political stance in a news article. An ablation study demonstrates the contribution of each component in our model.
Xinming Chen, Yuanxing Xu, Bin Wu 0001, Yangfu Zhu
IJCNN5
2024 Prompt-Enhanced Prototype Framework for Few-shot Event Detection
abstract
Few-shot event detection (ED) aims at identifying and typing event mentions from text with limited annotations. Most existing methods for few-shot ED use event ontology and related knowledge to construct prototypes and fail to fully leverage the rich knowledge of pre-trained language models (PLMs) which could help improve the representation of prototypes. Motivated by this, we propose an prompt-enhanced prototype framework which combines prototype and prompt for few-shot ED. Considering the scarcity of labeled data, we also introduce contrastive learning to enrich prototypes. Specifically, we use heuristic rules to align FrameNet with annotated data to get corresponding prompts for each event and convert them into prompt prototype. We then leverage contrastive learning to aggregate event mentions into prototypes and maintain these prototypes for few-shot ED. Furthermore, We explore diverse prompt formats for representing prompt prototypes and introduce a more comprehensive lexical prompt which improves the performance of few-shot ED. We conduct extensive experiments on the MAVEN corpus to reveal the effectiveness of the proposed framework compared to state-of-the-art methods.
Xinming Chen, Yangfu Zhu, Bin Wu 0001
IJCNN3
2024 Enhancing multimodal depression detection with intra- and inter-sample contrastive learning
Yangfu Zhu, Bin Wu 0001
Inf. Sci.3
2024 HG-PerCon: Cross-view contrastive learning for personality prediction
Yangfu Zhu, Bin Wu 0001
Neural Networks2
2024 A cross-temporal contrastive disentangled model for ancient Chinese understanding
Yangfu Zhu, Ting Bai 0004, Bin Wu 0001
Neural Networks2
2024 Knowledge-Guided Transformer for Joint Theme and Emotion Classification of Chinese Classical Poetry
abstract
The classifications of the theme and emotion are essential for understanding and organizing Chinese classical poetry. Existing works often overlook the rich semantic knowledge derived from poem annotations, which contain crucial insights into themes and emotions and are instrumental in semantic understanding. Additionally, the complex interdependence and diversity of themes and emotions within poems are frequently disregarded. Hence, this paper introduces a Poetry Knowledge-augmented Joint Model (Poka) specifically designed for the multi-label classification of themes and emotions in Chinese classical poetry. Specifically, we first employ an automated approach to construct two semantic knowledge graphs for theme and emotion. These graphs facilitate a deeper understanding of the poems by bridging the semantic gap between the obscure ancient words and their modern Chinese counterparts. Representations related to themes and emotions are then acquired through a knowledge-guided mask-transformer. Moreover, Poka leverages the inherent correlations between themes and emotions by adopting a joint classification strategy with shared training parameters. Extensive experiments demonstrate that our model achieves state-of-the-art performance on both theme and emotion classifications, especially on tail labels.
Linmei Hu, Yangfu Zhu, Bin Wu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 PCENet: Psychological Clues Exploration Network for Multimodal Personality Assessment
abstract
Multimodal personality assessment aims to identify and express human personality traits in videos. Existing methods primarily focus on multimodal fusion while ignoring the inherent psychological clues essential for this interdisciplinary task. Modality clues: personality traits are stable over time due to their genetic and environmental origins, resulting in stable personality traits in the multimodal data. Trait clues: multiple traits often co-occur with non-negligible correlations, which can collectively aid trait identification. To simultaneously capture the above psychological clues, we propose a novel Psychological Clues Exploration Network (PCENet) for multimodal personality assessment, which is a human-like judgment paradigm with more generalization capability. Specifically, we first devise a multimodal hierarchical disentanglement, which clearly aligns stable representations among different modalities and separates the mutability of each modality. Subsequently, a Transformer-backbone decoder equipped with modality-to-trait attention is exploited to adaptively generate a tailored representation for each trait with the guidance of trait semantics. The trait semantics are obtained by exploiting trait correlations through self-attention. Extensive experiments on the First Impression V2 dataset demonstrate that our PCENet outperforms the state-of-the-art methods for multimodal personality assessment.
Yangfu Zhu, Bin Wu 0001
CIKM1
2023 Self-adaptive Prompt-tuning for Event Extraction in Ancient Chinese Literature
abstract
Extracting different types of war events from ancient Chinese literature is significant, as war is an important factor in driving the development of Chinese history. The existing trend of event extraction models utilizes template-based generative approaches, which do not take into account the brevity and obscurity of ancient Chinese, as well as the diversity of templates for similar event types. In this paper, we propose a novel Knowledge Graph-based generative event extraction framework with a self-Adaptive Prompt (KGAP) for ancient Chinese war. Specifically, we construct a self-adaptive prompt, which considers its unique trigger words for different types of wars and is designed to solve the problem of the similarity in events. Moreover, we construct a semantic knowledge graph of ancient literature, assisting the pre-trained language model to better understand the ancient Chinese text. Since there is no public dataset for the ancient Chinese event extraction task, we provide an event extraction dataset and conduct experiments on it. Experimental results show that our model is more state-of-the-art than both the classification-based and generative-based methods for event extraction in ancient Chinese literature.
Yangfu Zhu, Bin Wu 0001
IJCNN3
2023 Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long Videos
abstract
Social Relation Recognition is an important part of Video Understanding, providing insights into the information that videos convey. Most previous works mainly focused on graph generation for characters, instead of edges which are more suitable for relation modelling. Furthermore, previous methods tend to recognize social relations for single frames or short video clips within their receptive fields, neglecting the importance of continuous reasoning throughout the entire video. To tackle these challenges, we propose a novel Shifted GCN-GAT and Cumulative-Transformer framework, named SGCAT-CT. The overall architecture consists of an SGCAT module for shifted graph operations on novel relation graphs and a CT module for temporal processing with memory. SGCAT-CT conducts continuous recognition of social relations and memorizes information from as early as the beginning of a long video. Experiments conducted on several video datasets demonstrate encouraging performance on long videos. Our code will be released at https://github.com/HarryWgCN/SGCAT-CT.
Haorui Wang, Yibo Hu 0005, Yangfu Zhu, Jinsheng Qi, Bin Wu 0001
ACM Multimedia3
2023 A Social Topic Diffusion Model Based on Rumor, Anti-Rumor, and Motivation-Rumor
abstract
The spread of online rumor poses challenges to social peace and public order. Traditional research on rumor diffusion commences from the rumor itself, without considering the symbiosis and confrontation of anti-rumor and motivation-rumor. This study proposed a diffusion method for online rumor based on three messages: rumor, anti-rumor, and motivation-rumor. First, considering the ability of representation learning to learn unsupervised features, we decided to use representation learning method to the diversity and complexity of the content and structure feature space. In particular, we designed a new representation method—Rumor2vec—for the potential structural feature of the rumor diffusion network. Second, considering the mutual promotion and suppression of the three messages, we constructed a new network topology using the cooperative and competitive relationships based on the evolutionary game theory. Finally, considering the ability of graph convolutional network (GCN) to convolute non-Euclidean structure data such as social network, and in view of the time effectiveness of topic evolution, this study proposed a dynamic and game-GCN (evolutionary game theory GCN)-based rumor diffusion model. Experiments show that the model can not only predict the group behavior under the rumor topic but also accurately reflect the cooperation and competition among multiple messages.
Xuemei Mou, Yangfu Zhu, Qian Li 0009, Yunpeng Xiao 0001
IEEE Trans. Comput. Soc. Syst.3
2022 Contrastive Graph Transformer Network for Personality Detection
abstract
Personality detection is to identify the personality traits underlying social media posts. Most of the existing work is mainly devoted to learning the representations of posts based on labeled data. Yet the ground-truth personality traits are collected through time-consuming questionnaires. Thus, one of the biggest limitations lies in the lack of training data for this data-hungry task. In addition, the correlations among traits should be considered since they are important psychological cues that could help collectively identify the traits. In this paper, we construct a fully-connected post graph for each user and develop a novel Contrastive Graph Transformer Network model (CGTN) which distills potential labels of the graphs based on both labeled and unlabeled data. Specifically, our model first explores a self-supervised Graph Neural Network (GNN) to learn the post embeddings. We design two types of post graph augmentations to incorporate different priors based on psycholinguistic knowledge of Linguistic Inquiry and Word Count (LIWC) and post semantics. Then, upon the post embeddings of the graph, a Transformer-based decoder equipped with post-to-trait attention is exploited to generate traits sequentially. Experiments on two standard datasets demonstrate that our CGTN outperforms the state-of-the-art methods for personality detection.
Yangfu Zhu, Linmei Hu, Xinkai Ge, Wanrong Peng, Bin Wu 0001
IJCAI1
2022 PerKG: A Personality Knowledge Graph for Personality Analysis
abstract
With the blossoming of online social networks (OSN), personality analysis based on OSN texts has gained much research attention in recent years. The previous methods mainly focus on human-designed features extracted through psychological dictionaries or semantic features extracted through language models. However, the shallow statistics features can not fully convey the personality information and the language models can not capture enough psychological background knowledge. Besides, the lack of large labeled datasets has been a serious obstacle impending further research. To tackle these problems, we propose a personality analysis model, namely PerKG, which combines personality knowledge graph and heterogeneous graph representation learning to exploit external knowledge from psycholinguistics and learn the group-level information to predict users’ personalities accurately. Specifically, we construct a personality knowledge graph based on existing psycholinguistics knowledge. And then, for each user, we align the user information with the knowledge graph to obtain the personality heterogeneous graph. Finally, the personality vector of each entity node is learned for prediction by designing a walk strategy on the personality heterogeneous graph. Detailed experimentation shows that our proposed PerKG architecture can effectively improve the performance and alleviate the label sparsity problem of personality analysis.
Yangfu Zhu, Zhanming Guan, Bin Wu 0001
SMC1
2022 Dynamic graph neural network for fake news detection
Chenguang Song, Yiyang Teng, Yangfu Zhu, Bin Wu 0001
Neurocomputing3
2022 A lexical psycholinguistic knowledge-guided graph neural network for interpretable personality detection
Yangfu Zhu, Linmei Hu, Nianwen Ning, Bin Wu 0001
Knowl. Based Syst.1
2020 User Behavior Prediction of Social Hotspots Based on Multimessage Interaction and Neural Network
abstract
In network public-opinion analysis, the diversity of messages under social hot topics plays an important role in user participation behavior. Considering the interactions among multiple messages and the complex user behaviors, this article proposes a prediction model of user participation behavior during multiple messaging of hot social topics. First, considering the influence of multimessage interaction on user participation behavior, a multimessage interaction influence-driving mechanism was proposed to predict user participation behavior more accurately. Second, in the view of the behavioral complexity of users engaging in multimessage hotspots and the simple structure of backpropagation (BP) neural networks (which can map complex nonlinear relationships), this study proposes a user participant behavior prediction model of social hotspots based on a multimessage interaction-driving mechanism and the BP neural network. Finally, the multimessage interaction has an iterative guiding effect on user behavior, which easily causes overfitting of the BP neural network. To avoid this problem, the traditional BP neural network is optimized by a simulated annealing algorithm to further improve the prediction accuracy. In evaluation experiments, the model not only predicted the user participation behavior in actual situations of multimessage interaction but also further quantified the correlations among multiple messages on hot topics.
Yunpeng Xiao 0001, Yangfu Zhu, Qian Li 0009
IEEE Trans. Comput. Soc. Syst.3