Wen Wu 0006

dblp:92/382-6 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MetaEval: Measuring the Discrimination of Benchmarks for Efficient LLM Evaluation
abstract
Benchmarks serve as standardized test systems to distinguish capabilities among large language models (LLMs). Discriminative items enable high-ability LLMs to favor correct answers, while causing low-ability models to assign lower plausibility to these answers and tend toward incorrect answers. Current methods for assessing benchmark quality primarily focus on coverage of difficulty levels and task diversity, yet lack direct quantification of discrimination—the core metric. Furthermore, large-scale benchmarks incur high evaluation costs. Although heuristic methods can reduce item counts to some extent, they cannot guarantee preservation of the benchmark’s original discriminative properties. To address these limitations, we propose MetaEval, a meta-evaluation framework designed to precisely quantify per-item discrimination and enable efficient assessment. Central to MetaEval is our novel Signal Detection and Item Response (SD-IR) model, which simulates LLMs’ detection of correct answers (signals) by representing each model’s perception through two latent ability states: “known” and “unknown”. For any item, discrimination is quantified as the difference in signal plausibility between these states. Leveraging these discrimination metrics, MetaEval introduces two strategies to replicate full-benchmark results using minimal subsets for efficient evaluation: (1) Distilling metaBench: a compact subset that retains discriminative power by removing redundant items; (2) Predicting performance on full-benchmark based on metaBench’s discrimination. Experiments across five benchmarks confirm that high-discrimination items capture greater performance variation among LLMs, align more closely with full-benchmark rankings, and exhibit superior predictive ability. Notably, in the best case, MetaEval achieves accurate full-benchmark estimation using only 2.5% of items, substantially reducing evaluation costs while preserving reliability.
Zhuo Wang 0010, Wen Wu 0006, Guangze Ye, Zhenxiao Cheng
AAAI2
2026 Beyond True Label: Label-Assumed Evidence Extraction for Personality Prediction
Yunyu Shi, Wen Wu 0006
DASFAA (4)5
2026 Metacognitive Activation Addition: Training-Free Enhancement of LLM Reasoning
Jingqi Huang, Liang Shan 0021, Wen Wu 0006, Liang He 0001
ICIC (5)3
2026 Textbook Content Moderation via Multi-agent Intergenerational Interaction
Wen Wu 0006, Qingchun Bai, Jiabao Zhao, Yunyu Shi, Liang He 0001
KSEM (1)2
2025 Decoupling Metacognition from Cognition: A Framework for Quantifying Metacognitive Ability in LLMs
abstract
Large Language Models (LLMs) are known to hallucinate facts and make non-factual statements which can undermine trust in their output. The essence of hallucination lies in the absence of metacognition in LLMs, namely the understanding of their own cognitive processes. However, there has been limited research on quantitatively measuring metacognition within LLMs. Drawing inspiration from cognitive psychology theories, we first quantify the metacognitive ability of LLMs as their ability to evaluate the correctness of responses through confidence. Subsequently, we introduce a general framework called DMC designed to decouple metacognitive ability and cognitive ability. This framework tackles the challenge of noisy quantification caused by the coupling of metacognition and cognition in current research, such as calibration-based metrics. Specifically, the DMC framework comprises two key steps. Initially, the framework tasks the LLM with failure prediction, aiming to evaluate the model's performance in predicting failures, a performance jointly determined by both cognitive and metacognitive abilities of the LLM. Following this, the framework disentangles metacognitive ability and cognitive ability based on the failure prediction performance, providing a quantification of the LLM's metacognitive ability independent of cognitive influences. Experiments conducted on eight datasets across five domains reveal that (1) Our proposed DMC framework effectively separates the metacognition and cognition of LLMs; (2) Various confidence elicitation methods impact the quantification of metacognitve ability differently; (3) Stronger metacognitive ability are exhibited by LLMs with better overall performance; (4) Enhancing metacognition holds promise for alleviating hallucination issues.
Wen Wu 0006, Guangze Ye, Zhenxiao Cheng
AAAI2
2025 Disentangled Modeling of Preferences and Social Influence for Group Recommendation
abstract
The group recommendation (GR) aims to suggest items for a group of users in social networks. Existing work typically considers individual preferences as the sole factor in aggregating group preferences. Actually, social influence is also an important factor in modeling users' contributions to the final group decision. However, existing methods either neglect the social influence of individual members or bundle preferences and social influence together as a unified representation. As a result, these models emphasize the preferences of the majority within the group rather than the actual interaction items, which we refer to as the preference bias issue in GR. Moreover, the self-supervised learning (SSL) strategies they designed to address the issue of group data sparsity fail to account for users' contextual social weights when regulating group representations, leading to suboptimal results. To tackle these issues, we propose a novel model based on Disentangled Modeling of Preferences and Social Influence for Group Recommendation (DisRec). Concretely, we first design a user-level disentangling network to disentangle the preferences and social influence of group members with separate embedding propagation schemes based on (hyper)graph convolution networks. We then introduce a social-based contrastive learning strategy, selectively excluding user nodes based on their social importance to enhance group representations and alleviate the group-level data sparsity issue. The experimental results demonstrate that our model significantly outperforms state-of-the-art methods on two real-world datasets.
Guangze Ye, Wen Wu 0006, Liang He 0001
AAAI2
2025 Supervisor Alignment Framework: Enhancing LLM Alignment with Query-Ignoring Strategy and Multi-Agent Interaction
abstract
The increasing focus on value alignment in Large Language Models (LLMs) underscores the need to ensure alignment with human morals and avoid biased or harmful outputs. However, LLMs aligned using existing methods are still easily affected by adversarial prompt attacks. Inspired by psychology, this paper introduces a Supervisor Alignment framework, which innovatively incorporates a query-ignoring strategy. This strategy ensures that the supervisor does not receive user queries, preventing it from being influenced by potential adversarial prompts. Meanwhile, the study compares the efficacy of a single supervisor versus a team of supervisors in value alignment tasks. While our designed single-agent supervisor approach utilizes a standalone agent or integrates with Retrieval-Augmented Generation (RAG) techniques, the team approach we proposed emphasizes multi-agent collaboration through voting, cooperation, and debate strategies. Extensive experiments demonstrate that the Supervisor Alignment framework we designed, incorporating the query-ignoring strategy and multi-agent collaboration, effectively defends against adversarial prompts and enhances its performance in value alignment tasks.
Ziqun Bao, Wen Wu 0006, Liang He 0001
ICASSP3
2025 Disentangled Interest and Popularity Modeling with Causal Intervention for Sequential Recommendation
Wen Wu 0006, Guangze Ye
ICIC (18)2
2025 Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music Generation
abstract
In recent years, Text-to-Music (T2M) generation models have rapidly emerged as powerful tools in content creation across fields. While existing models have made notable progress in sound quality, instrument identification, and stylistic alignment, they still exhibit clear limitations in modeling musical structure and musicality-particularly in terms of harmonic coherence and rhythmic alignment. To address these issues, we propose a Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music Generation(TCSA), which introduces explicit local condition controls to enhance structural fidelity in music generation. Specifically, we design a music theory enrichment strategy based on GPT-2 that transforms input text into detailed descriptions with embedded music theory knowledge, from which accurate chord progressions and rhythmic patterns are extracted as generation conditions. To synchronize these local features effectively, we develop a temporal alignment feature fusion mechanism. Additionally, we propose a layer-skipping fine-tuning strategy to avoid overfitting and enable fine-grained structural modeling. Finally, we introduce a perception-driven loss function based on Mel spectrograms to optimize the harmonic consistency and structural coherence of the generated music. Experimental results demonstrate that TCSA achieves competitive generation quality while offering significantly improved controllability over musical structure, making it well-suited for professional music production and refined content creation.
Xingjiao Wu, Tianlong Ma, Tangren Yao, Wen Wu 0006, Liang He 0001
ACM Multimedia6
2025 Knowledge-aware modeling of group commonality for group recommendation
Wen Wu 0006, Guangze Ye, Wenxin Hu, Liang He 0001
Expert Syst. Appl.1
2025 Demographic-Guided Behavior Patterns Contrast for Personality Prediction
abstract
In recent years, personality has been considered as a valuable personal factor being incorporated into the provision of personalized learning. Although some studies have endeavored to obtain learners’ personalities implicitly from their learning behaviors, they failed to achieve satisfactory prediction performance. On the one hand, most existing approaches ignore the imbalanced distribution of personality classes, which causes the personality classifiers to be biased toward the non-extreme personality class. On the other hand, the related methods normally focus on constructing statistical behavior features, while the sequence information of learning behaviors is ignored, but actually it can reflect learners’ behavior patterns more finely. In this paper, inspired by the human learning strategy in the face of small samples, we propose an effective Demographic-Guided Behavior Patterns Contrast (DGBPC) model to classify learners’ personalities through the demographic-guided contrast of learners’ coarse behavior patterns. Besides, we construct and publish the Personality and Learning Behavior Dataset (PLBD), which should be one of the largest public datasets regarding Big-Five personality and learning behavior sequence according to our knowledge. The experimental results on PLBD demonstrate that our DGBPC model could generate learner representations with higher discrimination and outperform the related methods in terms of balanced accuracy.
Wen Wu 0006, Wenxin Hu, Liang Kang, Liang He 0001
IEEE Trans. Affect. Comput.2
2024 A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete Labeling
abstract
The goal of document-level relation extraction (RE) is to identify relations between entities that span multiple sentences. Recently, incomplete labeling in document-level RE has received increasing attention, and some studies have used methods such as positive-unlabeled learning to tackle this issue, but there is still a lot of room for improvement. Motivated by this, we propose a positive-augmentation and positive-mixup positive-unlabeled metric learning framework (P3M). Specifically, we formulate document-level RE as a metric learning problem. We aim to pull the distance closer between entity pair embedding and their corresponding relation embedding, while pushing it farther away from the none-class relation embedding. Additionally, we adapt the positive-unlabeled learning to this loss objective. In order to improve the generalizability of the model, we use dropout to augment positive samples and propose a positive-none-class mixup method. Extensive experiments show that P3M improves the F1 score by approximately 4-10 points in document-level RE with incomplete labeling, and achieves state-of-the-art results in fully labeled scenarios. Furthermore, P3M has also demonstrated robustness to prior estimation bias in incomplete labeled scenarios.
Huazheng Pan, Wen Wu 0006, Wenxin Hu
AAAI4
2024 Learning Intrinsic Dimension via Information Bottleneck for Explainable Aspect-based Sentiment Analysis
abstract
Gradient-based explanation methods are increasingly used to interpret neural models in natural language processing (NLP) due to their high fidelity. Such methods determine word-level importance using dimension-level gradient values through a norm function, often presuming equal significance for all gradient dimensions. However, in the context of Aspect-based Sentiment Analysis (ABSA), our preliminary research suggests that only specific dimensions are pertinent. To address this, we propose the Information Bottleneck-based Gradient (IBG) explanation framework for ABSA. This framework leverages an information bottleneck to refine word embeddings into a concise intrinsic dimension, maintaining essential features and omitting unrelated information. Comprehensive tests show that our IBG approach considerably improves both the models’ performance and the explanations’ clarity by identifying sentiment-aware features.
Zhenxiao Cheng, Jie Zhou 0015, Wen Wu 0006, Qin Chen 0001, Liang He 0001
LREC/COLING3
2024 MetaESC: Enhancing Emotional Support Conversation through Metacognition
Wen Wu 0006, Liye Shi, Liang He 0001
DASFAA (5)2
2024 Personality-driven experience storage and retrieval for sentiment classification
Wen Wu 0006, Wenxin Hu, Liang He 0001
J. Supercomput.2
2023 Tell Model Where to Attend: Improving Interpretability of Aspect-Based Sentiment Classification via Small Explanation Annotations
abstract
Gradient-based explanation methods play an important role in the field of interpreting complex deep neural networks for NLP models. However, the existing work has shown that the gradients of a model are unstable and easily manipulable, which impacts the model’s reliability largely. According to our preliminary analyses, we also find the interpretability of gradient-based methods is limited for complex tasks, such as aspect-based sentiment classification (ABSC). In this paper, we propose an Interpretation-Enhanced Gradient-based framework for ABSC via a small number of explanation annotations, namely IGA. Particularly, we first calculate the word-level saliency map based on gradients to measure the importance of the words in the sentence towards the given aspect. Then, we design a gradient correction module to enhance the model’s attention on the correct parts (e.g., opinion words). Our model is model agnostic and task agnostic so that it can be integrated into the existing ABSC methods or other tasks. Comprehensive experimental results on four benchmark datasets show that our IEGA can improve not only the interpretability of the model but also the performance and robustness.
Zhenxiao Cheng, Jie Zhou 0015, Wen Wu 0006, Qin Chen 0001, Liang He 0001
ICASSP3
2023 Long-tail session-based recommendation from calibration
Jiayi Chen 0002, Wen Wu 0006, Liye Shi, Liang He 0001
Appl. Intell.2
2023 SMAR: Summary-Aware Multi-Aspect Recommendation
Liye Shi, Wen Wu 0006, Jiayi Chen 0002, Wenxin Hu, Liang He 0001
Neurocomputing2
2023 Personality-assisted mood modeling with historical reviews for sentiment classification
Wen Wu 0006, Jiayi Chen 0002, Wenxin Hu, Liang He 0001
Inf. Sci.2
2022 Knowledge-Enhanced Multi-task Learning for Course Recommendation
Qimin Ban, Wen Wu 0006, Wenxin Hu, Liang He 0001
DASFAA (2)2
2022 PMAR: Multi-aspect Recommendation Based on Psychological Gap
Liye Shi, Wen Wu 0006, Luping Feng, Liang He 0001
DASFAA (2)2
2022 Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis
abstract
Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most previous studies design various fusion frameworks for learning an interactive representation of multiple modalities, they fail to incorporate sentimental knowledge into inter-modality learning. This pa-per proposes a Multi-channel Attentive Graph Convolutional Network (MAGCN), consisting of two main components: cross-modality interactive learning and sentimental feature fusion. For cross-modality interactive learning, we exploit the self-attention mechanism combined with densely connected graph convolutional networks to learn inter-modality dynamics. For sentimental feature fusion, we utilize multi-head self-attention to merge sentimental knowledge into inter-modality feature representations. Extensive experiments are conducted on three widely-used datasets. The experimental results demonstrate that the proposed model achieves competitive performance on accuracy and F1 scores compared to several state-of-the-art approaches.
Luwei Xiao, Xingjiao Wu, Wen Wu 0006, Jing Yang 0023, Liang He 0001
ICASSP3
2022 Learner Profile based Knowledge Tracing
abstract
In recent years, knowledge tracing has gradually become a core technology for online education, which can evaluate learners' knowledge states and provide a personalized learning path. Most of the existing knowledge tracing methods mainly considered the interaction sequence of learners, but they normally ignored the individual differences among learners. For example, learners with various levels of comprehension will behave differently when faced with new questions, which indicates that individual differences affect prediction accuracy. In addition, most learners learn only part of the concept, which leads to data sparsity. However, the existing methods do not solve the data sparsity well. In this paper, we are motivated to propose a Learner Profile-based Knowledge Tracing (LPKT) model, which uses learners' unique id and the features extracted from historical interaction sequences as learners' representation to model individual differences among learners. In addition, we establish relationships between concepts and utilize related concepts to augment the concept's representation to address the data sparsity. We conducted experiments on several benchmark datasets, and the results show that our proposed LPKT model outperforms existing KT methods (with the highest AUC improvement of up to 8%).
Xinghai Zheng, Qimin Ban, Wen Wu 0006, Jiayi Chen 0002, Lamei Wang, Liang He 0001
IJCNN3
2022 Unexp-DIN: Unexpected Deep Interest Network for Recommendation
abstract
Studies on recommender systems persuaded accuracy of predictions, leading to the filter bubble problem which narrowed the user's interest areas. Unexpected recommendation (UR) is one of the solutions in handling the filter bubble. However, existing research on UR still faces two shortcomings. On the one hand, it is difficult to measure the unexpectedness of items to users accurately. On the other hand, there is a decrease in accuracy when making unexpected item recommendations. In this paper, we propose an Unexpected Deep Interest Network (Unexp-DIN) model to solve these problems. We use mean-shift for multiple clustering and Cluster-attention to maximize the unexpectedness. Then we use the constraint function and the hybrid utility function to simultaneously improve accuracy and unexpectedness. Experiments on three publicly available benchmark datasets show that our model has different degrees of improvement in both accuracy and unexpectedness.
Jiayi Chen 0002, Wen Wu 0006, Wenxin Hu, Liang He 0001
IJCNN3
2022 SENGR: Sentiment-Enhanced Neural Graph Recommender
Liye Shi, Wen Wu 0006, Wang Guo, Wenxin Hu, Jiayi Chen 0002, Liang He 0001
Inf. Sci.2
2022 DualGCN: An Aspect-Aware Dual Graph Convolutional Network for review-based recommender
Liye Shi, Wen Wu 0006, Wenxin Hu, Jie Zhou 0015, Jiayi Chen 0002, Liang He 0001
Knowl. Based Syst.2
2021 SSR: Explainable Session-based Recommendation
abstract
In recent years, session-based recommendation has attracted more and more attention. Previous work utilizing deep learning approaches has achieved significant progress in the accuracy of prediction. However, why the user would view the next item is not clear, though attention mechanism assigns different importance on items in a session. Meanwhile, some traditional data mining-based methods are able to provide explanations, but they mainly offer a narrow perspective like measuring item similarity and fail to achieve as good performance as deep learning-based approaches. To this end, we propose an explainable session-based recommendation by considering three factors including sequential patterns, repetition clicks, and item similarity. Concretely, we design a two-stage framework where candidate items are firstly selected according to users' sequential patterns and repeated behaviors, and then they are further ranked by considering short-term interest with the calculation of item similarity. The experimental results on two benchmark datasets show that our model not only achieves competitive recommendation performance with advanced deep learning based models, but also provides reasonable explanations.
Jiayi Chen 0002, Wen Wu 0006, Wenxin Hu, Liang He 0001
IJCNN2
2020 TSCREC: Time-sync Comment Recommendation in Danmu-Enabled Videos
abstract
In recent years, Time-sync Comment (TSC), as known as “Danmaku” or “Danmu”, has been increasingly recognized as a valuable representation of video being incorporated into the process of generating video or highlight recommendation. However, little work has studied how to recommend proper TSC when users are watching the video. Providing some suggestions for users when they want to post TSCs can not only enhance the real-time interactions among users within videos, but also motivate users to post more TSCs which could be useful for video or highlight recommendation in return. In order to accomplish the TSC recommendation task, we extract candidates from existing comments sent by other users. Specifically, we propose a model, namely “TSCREC”, which uses a bidirectional Gated Recurrent Unit to capture the semantic meaning of existing comments and assigns scores for comments by a Multi-Layer Perceptron. Furthermore, considering the importance of correlations among comments, we also take similarities among comments as features. In addition, we design a set of novel evaluation metrics which combines n-grams overlapping and ranking order to measure the quality of the recommendation list. We conduct experiments on realworld datasets, and the results show that our TSCREC model outperforms baseline approaches.
Jiayi Chen 0002, Wen Wu 0006, Wenxin Hu, Liang He 0001
ICTAI2
2020 Two-stage sentiment classification based on user-product interactive information
Wen Wu 0006, Qin Chen 0001, Wenxin Hu, Liang He 0001
Knowl. Based Syst.2
2019 A Crowdsourcing Based Human-in-the-Loop Framework for Denoising UUs in Relation Extraction Tasks
abstract
In relation extraction tasks, distant supervision methods expand dataset by aligning entity pairs in different knowledge bases and completing the relations between two entities. However, these methods ignore the fact that sentences labels generated by distant supervision methods with high confidence are often incorrect in the real world called Unknown Unknowns (UUs). To deal with this challenge, we propose a crowdsourcing based human-in-the-loop denoising framework which iteratively discovers UUs and corrects them by crowdsourcing to better extract relations. During each epoch of iterations, we choose one sentence bag and repeat two steps: Firstly, attention based Long Short-Term Memory network is applied as a selector to discover potential UUs. Secondly, these UUs are annotated by crowdsourcing with two answer collecting strategies and fed back into selector as positive samples. Until the accuracy of selector reaches a threshold, all annotated samples are added into relation classifier as cleaned train set and framework moves on to next epoch with new sentence bags. The experiments on the New York Times dataset and analysis of potential UUs demonstrate that our framework denoise the dataset and outperforms all the baselines on distant supervision relation extraction tasks.
Wen Wu 0006, Yan Yang 0008, Liang He 0001, Jing Yang 0023
IJCNN3
2019 Chinese Clinical Named Entity Recognition with Word-Level Information Incorporating Dictionaries
abstract
Electronic Medical Records (EMRs) are the digital equivalent of paper records, which include treatment and medical history about a patient. At present, the main research goal of Chinese EMRS is to accurately recognize the body parts, drugs, illnesses and other information in the Chinese medical process. Implementing EMRs can boost both the quality and safety of patient care. In Chinese EMRs, how to accurately recognize named entities is important because it is useful to predict the disease risk, therapeutic method and recovery probability. This paper proposes a novel deep learning framework, which uses character-word joint embedding and combines different feature information based on the dictionary. Compared with the predecessors, we incorporate word-level information based on the basic Bi-LSTM model. In addition, we propose an improved n-gram feature encoding method and compare it with PDET feature and PIET feature. Our experimental results demonstrate that our proposed model performs the best in predicting named entities in Chinese EMRs.
Ningjie Lu, Wen Wu 0006, Yan Yang 0008, Kaiwei Chen, Wenxin Hu
IJCNN3
2018 Personalizing recommendation diversity based on user personality
Wen Wu 0006, Li Chen 0009, Yu Zhao 0002
User Model. User Adapt. Interact.1
2017 Inferring Students' Sense of Community from Their Communication Behavior in Online Courses
abstract
Sense of community is regarded as the reflection of students' feelings of connectedness with community members and commonality of learning expectations and goals. In online courses, sense of community has been proven to influence students' learning engagement and academic performance. Low sense of community is also one of the reasons for drop out. However, existing studies mainly acquire students' sense of community via questionnaires, which demand user efforts and have difficulty in obtaining real-time feeling during students' learning process. In addition, although communication is helpful to enhance students' sense of community, little work has empirically compared the impact of different online communication tools. In this paper, we are motivated to derive students' sense of community from their communication behavior in online courses. Concretely, we first identify a set of features that are significantly correlated with students' sense of community, which not only include their activities carried out in both synchronous and asynchronous online learning environment, but also their linguistic content in conversational texts. We then develop inference model to unify these features for determining students' sense of community, and find that LASSO performs the best in terms of inference accuracy.
Wen Wu 0006, Li Chen 0009, Qingchang Yang
UMAP1
2016 Inferring Users' Critiquing Feedback on Recommendations from Eye Movements
Li Chen 0009, Feng Wang 0009, Wen Wu 0006
ICCBR3
2015 Implicit Acquisition of User Personality for Augmenting Movie Recommendations
Wen Wu 0006, Li Chen 0009
UMAP1