Xuezhi Cao

dblp:49/11206 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-7044-1341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
abstract
Lingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan, Yaoming Zhu, Lin Qiu, ZongYu Wang, Xuezhi Cao, Xunliang Cai, Weiwen Liu, Weinan Zhang, Yong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lingyue Fu, Hao Guan 0001, Haowei Yuan, Yaoming Zhu, Zongyu Wang, Xuezhi Cao, Weiwen Liu, Weinan Zhang 0001, Yong Yu 0001
ACL (1)8
2025 Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration
abstract
Shao Zhang, Xihuai Wang, Wenhao Zhang, Chaoran Li, Junru Song, Tingyu Li, Lin Qiu, Xuezhi Cao, Xunliang Cai, Wen Yao, Weinan Zhang, Xinbing Wang, Ying Wen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shao Zhang, Xihuai Wang, Junru Song, Xuezhi Cao, Weinan Zhang 0001, Xinbing Wang, Ying Wen 0001
ACL (1)8
2025 Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
abstract
Evaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval.
Tengchuan Kou, Shushi Wang, Chunyi Li 0001, Wei Sun 0029, Wei Wang 0213, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai
CVPR9
2025 MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
abstract
Siyu Yan, Long Zeng, Xuecheng Wu, Chengcheng Han, Kongcheng Zhang, Chong Peng, Xuezhi Cao, Xunliang Cai, Chenjuan Guo. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Long Zeng 0004, Chengcheng Han 0004, Kongcheng Zhang, Xuezhi Cao, Chenjuan Guo
EMNLP7
2025 I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
abstract
Multimodal entity linking plays a crucial role in a wide range of applications. Recent advances in large language model-based methods have become the dominant paradigm for this task, effectively leveraging both textual and visual modalities to enhance performance. Despite their success, these methods still face two challenges, including unnecessary incorporation of image data in certain scenarios and the reliance only on a one-time extraction of visual features, which can undermine their effectiveness and accuracy. To address these challenges, we propose a novel LLM-based framework for the multimodal entity linking task, called Intra- and Inter-modal Collaborative Reflections. This framework prioritizes leveraging text information to address the task. When text alone is insufficient to link the correct entity through intra- and inter-modality evaluations, it employs a multi-round iterative strategy that integrates key visual clues from various aspects of the image to support reasoning and enhance matching accuracy. Extensive experiments on three widely used public datasets demonstrate that our framework consistently outperforms current state-of-the-art methods in the task, achieving improvements of 3.2%, 5.1%, and 1.6%, respectively. Our code is available at https://github.com/ziyan-xiaoyu/I2CR/.
Junwen Li, Tong Ruan, Chao Wang 0095, Xinyan He, Zongyu Wang, Xuezhi Cao
ACM Multimedia8
2025 HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
abstract
Driven by the rapid progress in vision-language models (VLMs), the responsible behavior of large-scale multimodal models has become a prominent research area, particularly focusing on hallucination detection and factuality checking. In this paper, we present the solution for the two tracks of Responsible AI challenge. Inspirations from the general domain demonstrate that a smaller distilled VLM can often outperform a larger VLM that is directly tuned on the downstream tasks, while achieving higher efficiency. We thus jointly tackle two tasks from the perspective of knowledge distillation and propose a progressive hybrid knowledge distillation framework termed HKD4VLM. Specifically, the overall framework can be decomposed into Pyramid-like Progressive Online Distillation and Ternary-Coupled Refinement Distillation, hierarchically moving from coarse-grained knowledge alignment to fine-grained refinement. Besides, we further introduce the mapping shift-enhanced inference and diverse augmentation strategies to enhance model performance and robustness. Extensive experimental results demonstrate the effectiveness of our HKD4VLM. Ablation studies provide insights into the critical design choices driving performance gains.
Zijian Zhang 0008, Danlei Huang, Xuezhi Cao
ACM Multimedia6
2024 Conjoin after Decompose: Improving Few-Shot Performance of Named Entity Recognition
abstract
Prompt-based methods have been widely used in few-shot named entity recognition (NER). In this paper, we first conduct a preliminary experiment and observe that the key to affecting the performance of prompt-based NER models is the capability to detect entity boundaries. However, most existing models fail to boost such capability. To solve the issue, we propose a novel model, ParaBART, which consists of a BART encoder and a specially designed parabiotic decoder. Specifically, the parabiotic decoder includes two BART decoders and a conjoint module. The two decoders are responsible for entity boundary detection and entity type classification, respectively. They are connected by the conjoint module, which is used to replace unimportant tokens’ embeddings in one decoder with the average embedding of all the tokens in the other. We further present a novel boundary expansion strategy to enhance the model’s capability in entity type classification. Experimental results show that ParaBART can achieve significant performance gains over state-of-the-art competitors.
Chengcheng Han 0004, Renyu Zhu, Jun Kuang, Fengjiao Chen, Xiang Li 0067, Ming Gao 0001, Xuezhi Cao, Yunsen Xian
LREC/COLING7
2024 Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
abstract
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on various visual-language understanding and generation tasks. However, MLLMs occasionally generate content inconsistent with the given images, which is known as "hallucination". Prior works primarily center on evaluating hallucination using standard, unperturbed benchmarks, which overlook the prevalent occurrence of perturbed inputs in real-world scenarios-such as image cropping or blurring-that are critical for a comprehensive assessment of MLLMs' hallucination. In this paper, to bridge this gap, we propose Hallu-PI, the first benchmark designed to evaluate Hallucination in MLLMs within Perturbed Inputs. Specifically, Hallu-PI consists of seven perturbed scenarios, containing 1,260 perturbed images from 11 object types. Each image is accompanied by detailed annotations, which include fine-grained hallucination types, such as existence, attribute, and relation. We equip these annotations with a rich set of questions, making Hallu-PI suitable for both discriminative and generative tasks. Extensive experiments on 12 mainstream MLLMs, such as GPT-4V and Gemini-Pro Vision, demonstrate that these models exhibit significant hallucinations on Hallu-PI, which is not observed in unperturbed scenarios. Furthermore, our research reveals a severe bias in MLLMs' ability to handle different types of hallucinations. We also design two baselines specifically for perturbed scenarios, namely Perturbed-Reminder and Perturbed-ICL. We hope that our study will bring researchers' attention to the limitations of MLLMs when dealing with perturbed inputs, and spur further investigations to address this issue. Our code and datasets are publicly available at https://github.com/NJUNLP/Hallu-PI.
Peng Ding 0001, Jingyu Wu, Jun Kuang, Dan Ma 0008, Xuezhi Cao, Shi Chen 0005, Jiajun Chen 0001, Shujian Huang
ACM Multimedia5
2024 A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
abstract
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, Shujian Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Peng Ding 0001, Jun Kuang, Dan Ma 0008, Xuezhi Cao, Yunsen Xian, Jiajun Chen 0001, Shujian Huang
NAACL-HLT4
2023 Adap-τ : Adaptively Modulating Embedding Magnitude for Recommendation
abstract
Recent years have witnessed the great successes of embedding-based methods in recommender systems. Despite their decent performance, we argue one potential limitation of these methods — the embedding magnitude has not been explicitly modulated, which may aggravate popularity bias and training instability, hindering the model from making a good recommendation. It motivates us to leverage the embedding normalization in recommendation. By normalizing user/item embeddings to a specific value, we empirically observe impressive performance gains (9% on average) on four real-world datasets. Although encouraging, we also reveal a serious limitation when applying normalization in recommendation — the performance is highly sensitive to the choice of the temperature τ which controls the scale of the normalized embeddings.
Jiawei Chen 0007, Junkang Wu, Jiancan Wu, Xuezhi Cao, Sheng Zhou 0004, Xiangnan He 0001
WWW4
2023 Popularity Bias is not Always Evil: Disentangling Benign and Harmful Bias for Recommendation
abstract
Recommender system usually suffers from severepopularity bias— the collected interaction data usually exhibits quite imbalanced or even long-tailed distribution over items. Such skewed distribution may result from the users’conformityto the group, which deviates from reflecting users’ true preference. Existing efforts for tackling this issue mainly focus on completely eliminating popularity bias. However, we argue that not all popularity bias is evil. Popularity bias not only results from conformity but alsoitem quality, which is usually ignored by existing methods. Some items exhibit higher popularity as they have intrinsic better property. Blindly removing the popularity bias would lose such important signal, and further deteriorate model performance. To sufficiently exploit such important information for recommendation, it is essential to disentangle the benign popularity bias caused by item quality from the harmful popularity bias caused by conformity. Although important, it is quite challenging as we lack an explicit signal to differentiate the two factors of popularity bias. In this paper, we propose to leverage temporal information as the two factors exhibit quite different patterns along the time: item quality revealing item inherent property is stable and static while conformity that depends on items’ recent clicks is highly time-sensitive. Correspondingly, we further propose a novelTime-awareDisEntangled framework (TIDE), where a click is generated from three components namely the static item quality, the dynamic conformity effect, as well as the user-item matching score returned by any recommendation model. Lastly, we conduct interventional inference so that the recommendation can benefit from the benign popularity bias while circumvent the harmful one. Extensive experiments on four real-world datasets demonstrated the effectiveness of TIDE.
Zihao Zhao 0004, Jiawei Chen 0007, Sheng Zhou 0004, Xiangnan He 0001, Xuezhi Cao, Wei Wu 0014
IEEE Trans. Knowl. Data Eng.5
2021 DisenKGAT: Knowledge Graph Embedding with Disentangled Graph Attention Network
abstract
Knowledge graph completion (KGC) has become a focus of attention across deep learning community owing to its excellent contribution to numerous downstream tasks. Although recently have witnessed a surge of work on KGC, they are still insufficient to accurately capture complex relations, since they adopt the single and static representations. In this work, we propose a novel Disentangled Knowledge Graph Attention Network (DisenKGAT) for KGC, which leverages both micro-disentanglement and macro-disentanglement to exploit representations behind Knowledge graphs (KGs). To achieve micro-disentanglement, we put forward a novel relation-aware aggregation to learn diverse component representation. For macro-disentanglement, we leverage mutual information as a regularization to enhance independence. With the assistance of disentanglement, our model is able to generate adaptive representations in terms of the given scenario. Besides, our work has strong robustness and flexibility to adapt to various score functions. Extensive experiments on public benchmark datasets have been conducted to validate the superiority of DisenKGAT over existing methods in terms of both accuracy and explainability.
Junkang Wu, Wentao Shi 0002, Xuezhi Cao, Jiawei Chen 0007, Wenqiang Lei, Wei Wu 0014, Xiangnan He 0001
CIKM3
2020 Multi-modal Knowledge Graphs for Recommender Systems
abstract
Recommender systems have shown great potential to solve the information explosion problem and enhance user experience in various online applications. To tackle data sparsity and cold start problems in recommender systems, researchers propose knowledge graphs (KGs) based recommendations by leveraging valuable external knowledge as auxiliary information. However, most of these works ignore the variety of data types (e.g., texts and images) in multi-modal knowledge graphs (MMKGs). In this paper, we propose Multi-modal Knowledge Graph Attention Network (MKGAT) to better enhance recommender systems by leveraging multi-modal knowledge. Specifically, we propose a multi-modal graph attention technique to conduct information propagation over MMKGs, and then use the resulting aggregated embedding representation for recommendation. To the best of our knowledge, this is the first work that incorporates multi-modal knowledge graph into recommender systems. We conduct extensive experiments on two real datasets from different domains, results of which demonstrate that our model MKGAT can successfully employ MMKGs to improve the quality of recommendation system.
Xuezhi Cao, Yan Zhao 0008, Junchen Wan, Kun Zhou 0002, Zhongyuan Wang 0006, Kai Zheng 0001
CIKM2
2020 Table Fact Verification with Structure-Aware Transformer
abstract
Verifying fact on semi-structured evidence like tables requires the ability to encode structural information and perform symbolic reasoning. Pre-trained language models trained on natural language could not be directly applied to encode tables, because simply linearizing tables into sequences will lose the cell alignment information. To better utilize pre-trained transformers for table representation, we propose a Structure-Aware Transformer (SAT), which injects the table structural information into the mask of the self-attention layer. A method to combine symbolic and linguistic reasoning is also explored for this task. Our method outperforms baseline with 4.93% on TabFact, a large scale table verification dataset.
Yingyao Wang, Xuezhi Cao, Zhongyuan Wang 0006
EMNLP (1)4
2020 Iterative Strategy for Named Entity Recognition with Imperfect Annotations
Yunian Chen, Xuezhi Cao, Rui Xie 0005
NLPCC (2)4
2018 Neural Link Prediction over Aligned Networks
abstract
Link prediction is a fundamental problem with a wide range of applications in various domains, which predicts the links that are not yet observed or the links that may appear in the future. Most existing works in this field only focus on modeling a single network, while real-world networks are actually aligned with each other. Network alignments contain valuable additional information for understanding the networks, and provide a new direction for addressing data insufficiency and alleviating cold start problem. However, there are rare works leveraging network alignments for better link prediction. Besides, neural network is widely employed in various domains while its capability of capturing high-level patterns and correlations for link prediction problem has not been adequately researched yet. Hence, in this paper we target atlink prediction over aligned networks using neural networks. The major challenge is the heterogeneousness of the considered networks, as the networks may have different characteristics, link purposes, etc. To overcome this, we propose a novel multi-neural-network framework MNN, where we have one individual neural network for each heterogeneous target or feature while the vertex representations are shared. We further discuss training methods for the multi-neural-network framework. Extensive experiments demonstrate that MNN outperforms the state-of-the-art methods and achieves 3% to 5% relative improvement of AUC score across different settings, particularly over 8% for cold start scenarios.
Xuezhi Cao, Xuejian Wang, Weinan Zhang 0001, Yong Yu 0001
AAAI1
2018 A Machine Learning Approach to Prevent Malicious Calls over Telephony Networks
abstract
Malicious calls, i.e., telephony spams and scams, have been a long-standing challenging issue that causes billions of dollars of annual financial loss worldwide. This work presents the first machine learning-based solution without relying on any particular assumptions on the underlying telephony network infrastructures. The main challenge of this decade-long problem is that it is unclear how to construct effective features without the access to the telephony networks' infrastructures. We solve this problem by combining several innovations. We first develop a TouchPal user interface on top of a mobile App to allow users tagging malicious calls. This allows us to maintain a large-scale call log database. We then conduct a measurement study over three months of call logs, including 9 billion records. We design 29 features based on the results, so that machine learning algorithms can be used to predict malicious calls. We extensively evaluate different state-of-the-art machine learning approaches using the proposed features, and the results show that the best approach can reduce up to 90% unblocked malicious calls while maintaining a precision over 99.99% on the benign call traffic. The results also show the models are efficient to implement without incurring a significant latency overhead. We also conduct ablation analysis, which reveals that using 10 out of the 29 features can reach a performance comparable to using all features.
Huichen Li, Chang Liu 0021, Teng Ren, Xuezhi Cao, Weinan Zhang 0001, Yong Yu 0001, Dawn Song
IEEE Symposium on Security and Privacy6
2017 IMAP: An iterative method for aligning protein-protein interaction networks
abstract
Biological network alignment benefits the evolutionary and comparative biology by providing regions of topological and functional similarity between different species. However, most existing network aligners follow heuristic methods and only capture the static information that based purely on the original isolated networks, while there also exists valuable interactive information hidden in the resulted alignment that provides additional signals for further improvement. In this paper, we propose an iterative method IMAP to improve the quality of existing network aligners. IMAP starts from an imperfect seed alignment generated by any aligner, and then iteratively refines it by capturing interactive information hidden in current alignment until convergence. Within each iteration, we calculate the likelihood of pairwise alignment using supervised learning techniques, hence heuristic functions are no longer required. Furthermore, we extend IMAP to start from multiple seed aligners to combine their individual advantages. Comprehensive experiments indicate that IMAP improves existing network aligners significantly in terms of node correctness, topology conservation and biological similarity. Therefore, IMAP can benefit the subsequent cross-species biology researches by providing high-quality alignment between PPI networks.
Xuezhi Cao, Zhiyu Chen 0002, Yong Yu 0001
BIBM1
2017 Joint User Modeling Across Aligned Heterogeneous Sites Using Neural Networks
Xuezhi Cao, Yong Yu 0001
ECML/PKDD (1)1
2016 ASNets: A Benchmark Dataset of Aligned Social Networks for Cross-Platform User Modeling
abstract
Aligning heterogeneous online social networks is a highly beneficial task proposed in recent years. It targets at automatically aligning accounts from multiple networks by whether they are held by the same natural person. Aligning the networks can improve personalized services by cross-platform user modeling, and is the prerequisite for cross-network analysis. However, there is currently no public benchmark dataset available due to its recency. As performances of this task depend highly on the dataset, experiments using different private datasets are not directly comparable. Therefore, in this paper we propose ASNets, a benchmark dataset with two sets of aligned social networks. With this dataset, we can now properly evaluate different approaches and compare them fairly. The two sets of aligned networks have 328,224 and 141,614 aligned users respectively, covering multilingual usage (Chinese and English) and various types of social networks including general purposed networks, review sites and microblogging sites. We describe the collecting methodology and statistics in details, and evaluate several state-of-the-art network aligning approaches. Beside introducing the dataset, we further propose several potential research directions that benefit from ASNets.
Xuezhi Cao, Yong Yu 0001
CIKM1
2016 BASS: A Bootstrapping Approach for Aligning Heterogenous Social Networks
Xuezhi Cao, Yong Yu 0001
ECML/PKDD (1)1
2016 Joint User Modeling across Aligned Heterogeneous Sites
abstract
An accurate and comprehensive user modeling technique is crucial for the quality of recommender systems. Traditionally, we model user preferences using only actions from the target site and may suffer from cold-start problem. As nowadays people normally engage in multiple online sites for various needs, we consider leveraging the cross-site actions to improve the user modeling accuracy. Specifically, in this paper we aim at achieving a more comprehensive and accurate user modeling by modeling user's actions in multiple aligned heterogeneous sites simultaneously. To do so, we propose a modularized probabilistic graphical model framework JUMA. We further integrate topic model and matrix factorization into JUMA for joint user modeling over text-based and item-based sites. We assemble and publish large-scale dataset for comprehensive analyzing and evaluation. Experimental results show that our framework JUMA out performs traditional within-site user modeling techniques, especially for cold-start scenarios. For cold-start users, we achieve relative improvements of 9.3% and 12.8% comparing to existing within-site approaches for recommendation in item-based and text-based sites respectively. Thus we draw the conclusion that aligning heterogeneous sites and modeling users jointly do help to improve the quality of online recommender systems.
Xuezhi Cao, Yong Yu 0001
RecSys1
2016 Are You Influenced by Others When Rating?: Improve Rating Prediction by Conformity Modeling
abstract
Conformity has a strong influence to user behaviors, even in online environment. When surfing online, users are usually flooded with others' opinions. These opinions implicitly contribute to the user's ongoing behaviors. However, there is no research work modeling online conformity yet. In this paper, we model user's conformity in online rating sites. We conduct analysis using real data to show the existence and strength of conformity in these scenarios. We propose a matrix-factorization-based conformity modeling technique to improve the accuracy of rating prediction. Experiments show that our model outperforms existing works significantly (with a relative improvement of 11.72% on RMSE). Therefore, we draw the conclusion that conformity modeling is important for understanding user behaviors and can contribute to further improve the online recommender systems.
Xuezhi Cao, Yong Yu 0001
RecSys2
2016 A Complete & Comprehensive Movie Review Dataset (CCMR)
abstract
Online review sites are widely used for various domains including movies and restaurants. These sites now have strong influences towards users during purchasing processes. There exist plenty of research works for review sites on various aspects, including item recommendation, user behavior analysis, etc. However, due to the lack of complete and comprehensive dataset, there are still problems that remain to be solved. Therefore, in this paper we assemble and publish such dataset (CCMR) for the community. CCMR outruns existing datasets in terms of completeness, comprehensiveness and scale. Besides describing the dataset and its collecting methodology, we also propose several potential research topics that are made possible by having this dataset. Such topics include: (i) a statistical approach to reduce the impacts from fake reviews and (ii) analyzing and modeling the influences of public opinions towards users during rating actions. We further conduct preliminary analysis and experiments for both directions to show that they are promising.
Xuezhi Cao, Weiyue Huang, Yong Yu 0001
SIGIR1