VLDB 2026 Research / reviewers in the wild / expert
Xiaodan Zhang 0004
dblp:29/2631-4
· DBLP profile ↗
21ranked-venue papers
0as first author
16since 2021 · last 2026
0009-0009-7702-4290ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge-Enhanced Image Captioning with Adaptive Graph-based Multimodal Alignment and LLMabstractImage captioning is crucial for multimodal understanding, bridging visual content and natural language. Despite recent advancements in Large Multimodal Models (LMMs), when faced with unseen entities or scenes in the open world, even when attempting to leverage learned knowledge, models still struggle with vague and inaccurate descriptions, and may even generate knowledge hallucinations. A key reason is that the model fails to effectively integrate knowledge with visual information, limiting its understanding of visual content. Thus, we propose Adaptive Knowledge Graph-guided Multimodal Alignment (AKGMA) for image captioning, which enhances semantic understanding in open-world scenes through visual knowledge reasoning, reducing knowledge hallucinations and improving caption quality. It consist three key components: Entity-guided Knowledge Aligner (EKA), Adaptive Knowledge Graph Construction (AKGC), and Scene-Context Knowledge Adapter (SCKA). EKA connects visual entities to knowledge graphs, providing structured knowledge to a small language model, which interacts with a visual encoder to acquire visual knowledge. AKGC uses reinforcement learning to build image-relevant subgraphs to optimize knowledge prompts and improve knowledge hallucinations. SCKA leverages scene graph annotations to extract visual contextual knowledge and inject it into Large Language Models (LLMs), ensuring the generated descriptions are consistent with the image's details. Additionally, we introduce UniKnowCap, a new image knowledge description dataset spanning various open-world knowledge domains, designed to evaluate the knowledge accuracy and detail consistency of model-generated descriptions. Extensive experiments show our model outperforms baselines across multiple metrics. Guoyi Li, Die Hu 0004, Zhongjiang Yao, Wei Mi, Zongzhen Liu, Xiaodan Zhang 0004, Honglei Lyu |
AAAI | 7 |
| 2026 | FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCsabstractBingkang Shi, Jen-tse Huang, Luo Long, Tianyu Zong, Hongzhu Yi, Yuanxiang Wang, Songlin Hu, Xiaodan Zhang, Zhongjiang Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bingkang Shi, Jen-tse Huang 0001, Luo Long, Tianyu Zong, Hongzhu Yi, Yuanxiang Wang, Songlin Hu 0001, Xiaodan Zhang 0004, Zhongjiang Yao |
ACL (1) | 8 |
| 2025 | Semantic Reshuffling with LLM and Heterogeneous Graph Auto-Encoder for Enhanced Rumor DetectionabstractSocial media is crucial for information spread, necessitating effective rumor detection to curb misinformation’s societal effects. Current methods struggle against complex propagation influenced by bots, coordinated accounts, and echo chambers, which fragment information and increase risks of misjudgments and model vulnerability. To counteract these issues, we introduce a new rumor detection framework, the Narrative-Integrated Metapath Graph Auto-Encoder (NIMGA). This model consists of two core components: (1) Metapath-based Heterogeneous Graph Reconstruction. (2) Narrative Reordering and Perspective Fusion. The first component dynamically reconstructs propagation structures to capture complex interactions and hidden pathways within social networks, enhancing accuracy and robustness. The second implements a dual-agent mechanism for viewpoint distillation and comment narrative reordering, using LLMs to refine diverse perspectives and semantic evolution, revealing patterns of information propagation and latent semantic correlations among comments. Extensive testing confirms our model outperforms existing methods, demonstrating its effectiveness and robustness in enhancing rumor representation through graph reconstruction and narrative reordering. Guoyi Li, Die Hu 0004, Zongzhen Liu, Xiaodan Zhang 0004, Honglei Lyu |
COLING | 4 |
| 2025 | A Generative Approach for Alleviating Filter Bubbles in Collaborative FilteringabstractIn addressing the filter bubble phenomenon in recommender systems, existing research has integrated large language models (LLMs) but faced distribution discrepancies and challenges like hallucinations and over-reliance on text data. We propose a two-stage method called LEAD (LLM-Enhanced Augmentation of Data) to mitigate these issues. First, we leverage LLMs' to extract textual representations of user interests and item audiences as auxiliary information. This information acts as pseudo-labels to guide a Conditional Generative Adversarial Network (CGAN) in generating unexpected items that align with actual user data distributions. Our approach harnesses model knowledge while avoiding data distribution gaps. Second, we introduce a heterogeneous view alignment strategy to align the semantic space of LLMs with the representation space of collaborative signals, enhancing the quality of representations and reducing noise. Extensive experiments on three real-world datasets demonstrate that LEAD outperforms existing models in recommendation quality, particularly in diversity and distribution alignment, showcasing the potential of LLMs in evolving recommender systems. Our code is available at https://github.com/gusuccc/LEAD. Hongyue Zhang, Jizhong Han, Xiaodan Zhang 0004 |
CSCWD | 6 |
| 2025 | Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor DetectionabstractSocial media is a key channel for information dissemination, making effective rumor detection essential to mitigate misinformation’s societal impact. Although large language models excel in inference and text generation, they struggle with understanding propagation relationships and complex reasoning tasks like rumor detection. Existing methods mainly rely on textual information and event propagation structures, but provocative comments and unreliable interactions increase propagation uncertainty. To address these challenges, we propose an Emotionally-Aware Structural Enhancement Graph Auto-Encoder (EASE-GARD) to improve rumor representations. Our method begins by enhancing the textual representation of responses through the generation of emotive adversarial comments. It then generates and differentiates false local propagation relationships (fabricated forwards and reciprocations) to reduce propagation uncertainties. A graph auto-encoder captures contextual features and global structural information and recalculates forwarding probabilities among responses. Extensive experiments show our model performs best on all three datasets, excelling in effectiveness and robustness. Guoyi Li, Zhongjiang Yao, Die Hu 0004, Yingrui Xu, Xiaodan Zhang 0004, Honglei Lyu |
ICASSP | 5 |
| 2025 | Resolution Attack: Exploiting Image Compression to Deceive Deep Neural NetworksabstractModel robustness is essential for ensuring the stability and reliability of machine learning systems. Despite extensive research on various aspects of model robustness, such as adversarial robustness and label noise robustness, the exploration of robustness towards different resolutions, remains less explored. To address this gap, we introduce a novel form of attack: the resolution attack. This attack aims to deceive both classifiers and human observers by generating images that exhibit different semantics across different resolutions. To implement the resolution attack, we propose an automated framework capable of generating dual-semantic images in a zero-shot manner. Specifically, we leverage large-scale diffusion models for their comprehensive ability to construct images and propose a staged denoising strategy to achieve a smoother transition across resolutions. Through the proposed framework, we conduct resolution attacks against various off-the-shelf classifiers. The experimental results exhibit high attack success rate, which not only validates the effectiveness of our proposed framework but also reveals the vulnerability of current classifiers towards different resolutions. Additionally, our framework, which incorporates features from two distinct objects, serves as a competitive tool for applications such as face swapping and facial camouflage. The code is available at https://github.com/ywj1/resolution-attack. Wangjia Yu, Xiaomeng Fu, Jizhong Han, Xiaodan Zhang 0004 |
ICLR | 5 |
| 2025 | Entity Graph Alignment and Visual Reasoning for Multimodal Fake News DetectionabstractThe rise of multimodal fake news threatens reliable information dissemination by exploiting multiple modalities to create deceptive, engaging content, significantly impacting society safety. Existing methods still face challenges in cross-modal alignment (e.g., semantic inconsistencies, complex visual-semantic relations) and are vulnerable to low-quality or noisy samples. To address these, we propose Cross-Modal Alignment with Visual Reasoning Prompting (CMA-VRP) for multimodal fake news detection. Specifically, we model text and image entities with graphs to capture fine-grained semantic interactions and enhance cross-modal consistency through graph contrastive learning. Unlike methods relying on shallow image features (e.g., edges, textures), we leverage large language models (LLMs) and large vision-language models (LVLMs) to capture deep visual-semantic attributes related to reasoning (e.g., actions, scenes). Based on graph modeling and visual reasoning features, we perform graph-based cross-modal semantic fusion to unify textual and visual representations and cross-modal cycle alignment to align modality distributions by reducing semantic discrepancies, filtering modality-specific noise, and extracting invariant representations across domains. These steps enable the model to obtain semantically consistent and modality-invariant features. Extensive experiments demonstrate that our model outperforms existing methods in multimodal fake news detection and shows strong robustness against noisy samples. Guoyi Li, Die Hu 0004, Xiaomeng Fu, Qirui Tang, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu |
ACM Multimedia | 6 |
| 2025 | Zero-Shot Multimodal Fact-Checking with Conceptual ReasoningabstractIn multimodal fact-checking, advanced large multimodal models (LMMs) struggle to capture and integrate the complex relationships between text and images. A potential solution is to generate reasoning support text to optimize reasoning and integrate evidence. However, existing generation approaches rely heavily on high-quality data annotations for training, which are costly and limited in scalability, hindering responsiveness to evolving misinformation. To address these issues, we propose CoReS, a novel zero-shot multimodal fact-checking model based on Conceptual Reasoning Support-leveraging key concepts from evidence to guide the reasoning process and improve decision-making. This model includes a reasoning support text generation module that extracts key concepts (critical elements that significantly impact the judgment outcome) from raw textual evidence via retrieval and filtering. By using a Conceptual Reasoning LM, CoReS generates reasoning support texts framed around core key concepts that are semantically consistent with multimodal evidence, linking key clues, thus replacing redundant and complex evidence for fact-checking. The reasoning support texts generated by CoReS effectively distill complex evidence relationships and integrate important reasoning information, allowing the judgment model to provide clear and accurate judgments. Evaluations on benchmark datasets and the new multi-domain MultiVerify dataset demonstrate that CoReS excels in accuracy, generalization, and scalability. Guoyi Li, Die Hu 0004, Qirui Tang, Xiaomeng Fu, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu |
ACM Multimedia | 7 |
| 2024 | The Security Paradox of Smart Contracts: Blind Spots and Prospects of Current Detection StrategiesabstractEthereum, a foundational blockchain platform enabling decentralized applications, has catalyzed the development of myriad applications through smart contracts. However, vulnerabilities within these contracts have precipitated notable financial losses, garnering keen interest from both industry and academia. While existing vulnerability detection research has progressed in addressing traditional challenges, it increasingly falls short of comprehensively securing smart contracts against evolving attack patterns and new vulnerabilities. Distinguishing itself from prior surveys, this review conducts an in-depth analysis of actual attack incidents, delving into the characteristics of recent vulnerabilities and thoroughly exploring the present research difficulties and challenges. The paper aims to offer a holistic view on smart contract security research, advance the depth of vulnerability detection technologies, and suggest directions for future inquiries. Wei Mi, Xiaodan Zhang 0004 |
CSCWD | 3 |
| 2024 | General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token LevelabstractThe social biases and unwelcome stereotypes revealed by pretrained language models are becoming obstacles to their application. Compared to numerous debiasing methods targeting word level, there has been relatively less attention on biases present at phrase level, limiting the performance of debiasing in discipline domains. In this paper, we propose an automatic multi-token debiasing pipeline called General Phrase Debiaser, which is capable of mitigating phrase-level biases in masked language models. Specifically, our method consists of a phrase filter stage that generates stereotypical phrases from Wikipedia pages as well as a model debias stage that can debias models at the multi-token level to tackle bias challenges on phrases. The latter searches for prompts that trigger model’s bias, and then uses them for debiasing. State-of-the-art results on standard datasets and metrics show that our approach can significantly reduce gender biases on both career and multiple disciplines, across models with varying parameter sizes. Bingkang Shi, Xiaodan Zhang 0004, Dehan Kong, Yulei Wu, Zongzhen Liu, Honglei Lyu, Longtao Huang |
ICASSP | 2 |
| 2024 | CRDA: Content Risk Drift Assessment of Large Language Models through Adversarial Multi-Agent InteractionabstractAs Large Language Models (LLMs) continue to enhance their capabilities in multi-agent collaborative applications, the unpredictability of the generative content risks has intensified. Particularly in ongoing interaction scenarios with users, it remains unclear whether there is generative content risk drift over time. In this context, "drift risk" refers to the trend of progressively intensified content risk that emerges during sustained adversarial interactions among LLM agents. Additionally, the high cost associated with constructing complex adversarial environments for agents impedes the transferability of current assessment methods for LLMs to multi-agent adversarial scenarios. In this paper, we introduce a low-cost and lightweight framework for assessing content risk drift of LLMs, named CRDA. This framework, bypassing the need for constructing complex adversarial environments, offers a method that integrates roles and responses memory to guide automatically multi-round adversarial interactions among LLM agents, that is, multiple agents as avatars of a single LLM. In this approach, LLM agents enable the analysis of content risk drift of this LLM. Moreover, we explore the impact of restricted roles and the unsafe content with negative viewpoints in responses memory on the content risk drift of LLMs. Considering the rapid advancement of Chinese LLM capabilities, this study selects real adversarial topics in Chinese and assesses content risk drift of five representative Chinese LLMs. The research finds that these LLMs exhibit significant content risk drift even after a certain safety alignment, showing an initial increase followed by a gradual decrease. As the adversarial process progresses, under restricted roles, agents more effectively breach the model's safety alignment, leading to content risk drift of the LLM. The content drift risk assessment can be quantified specifically by measuring the deterioration rate at which LLM agents deteriorate from positive to negative and analyzing the underlying trends during the automatically multi-round adversarial interactions. In restricted and general roles adversarial interactions, all agents of five Chinese LLMs exhibit an overall average increase of 31.5% and 16.38% in the cumulative deterioration rate respectively by the 10th round, compared to the baseline no-roles adversarial interactions. Finally, we hope that the framework and findings presented in this paper will offer valuable insights for research on safety alignment in LLM agents during adversarial processes. Zongzhen Liu, Guoyi Li, Bingkang Shi, Xiaodan Zhang 0004, Jingguo Ge, Yulei Wu, Honglei Lyu |
IJCNN | 4 |
| 2024 | EthGAN: Improving Ethereum Account Classification Accuracy via Data AugmentationabstractRecently, with the prevalent adoption of blockchain in the financial system, there has been an increasing of anomaly activities such as ponzi schemes, gambling and phishing fraud on Ethereum platforms, and an effective account classification method is urgently required. The existing account classification methods on Ethereum with high accuracy require a learning system to be trained with balanced datasets. However, the distribution of annotated labels for account identities published on third-party sites is relatively imbalanced. Therefore, in this paper, We propose a EthGAN framework which includes a high-dimensional node feature representation module and a few-shot account data augment module to improve the accuracy and robustness at imbalanced datasets. The high-dimensional node feature representation module captures features from statistical, temporal, and transaction structure, and the few-shot account data augmentation module based on generative adversarial network models generate few-shot samples to improve the diversity and representativeness of the training datasets. We conduct extensive experiments to evaluate the performance of our proposed EthGAN framework on real-world Ethereum transaction data. The average classification effect of our method is 10+% higher than that of existing methods. Experimental results demonstrate that our method outperforms state-of-the-art methods in Ethereum account classification. Xuehai Tang, Zhongjiang Yao, Huazhen Zhong, Yuanshu Zhao, Xiaodan Zhang 0004, Jizhong Han |
IJCNN | 6 |
| 2023 | Mining Weak Relations Between Reviews for Opinion Spam DetectionabstractOnline reviews play a significant role in purchase decisions of consumers by providing feedback information from buyers of products. In order to mislead consumers, opinion spammers are hired to write fake reviews to promote or demote specific products for illegitimate benefits. Existing methods for spam review detection mainly focused on designing manual features, which highly rely on expert knowledge. Although recent works utilized deep learning methods to automatically learn the semantics of reviews through the inherent user-review-product strong relation, they fail to capture the weak relations between reviews at the content, sentiment and temporal levels, which provides various semantic information to expose fake reviews. Moreover, the imbalanced class distribution in spam detection issues makes this work even more challenging. To address the above problems, we propose a novel Weak-Strong Unified Network (WSUN) for opinion spam detection. Multi-level weak relation graphs are constructed to reveal the abnormal behavioral patterns of spammers, which aggregates the semantics of strong relations by graph convolutional networks and extracts comprehensive review representation by utilizing relation-level attention mechanism. In addition, a graph-based over-sampling method is devised to mitigate the impact of imbalanced class distribution. Extensive experimental results on real-world datasets show that our model is more effective than the state-of-the-art methods. Yingrui Xu, Jingguo Ge, Xiaodan Zhang 0004, Yulei Wu, Honglei Lv, Hongbin Shi, Wei Zhou 0019 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Pay attention to emoji: Feature Fusion Network with EmoGraph2vec Model for Sentiment AnalysisabstractWith the explosive growth of social media, opinionated postings with emojis have increased explosively. Many emojis are used to express emotions, attitudes, and opinions. Emoji representation learning can be helpful to improve the performance of emoji-related natural language processing tasks, especially in text sentiment analysis. However, most studies have only utilized the fixed descriptions provided by the Unicode Consortium without consideration of actual usage scenarios. As for the sentiment analysis task, many researchers ignore the emotional impact of the interaction between text and emojis. It results that the emotional semantics of emojis cannot be fully explored. In this work, we propose a method called EmoGraph2vec to learn emoji representations by constructing a co-occurrence graph network from social data and enriching the semantic information based on an external knowledge base EmojiNet to embed emoji nodes. Based on EmoGraph2vec model, we design a novel neural network to incorporate text and emoji information into sentiment analysis, which uses a hybrid-attention module combined with TextCNN-based classifier to improve performance. Experimental results show that the proposed model can outperform several baselines for sentiment analysis on benchmark datasets. Additionally, we conduct a series of ablation and comparison experiments to investigate the effectiveness and interpretability of our model. Xiaowei Yuan, Xiaodan Zhang 0004, Honglei Lv |
ICPR | 3 |
| 2022 | A Heterogeneous Propagation Graph Model for Rumor Detection Under the Relationship Among Multiple Propagation Subtrees
Guoyi Li, Yulei Wu, Xiaodan Zhang 0004, Wei Zhou 0019, Honglei Lyu |
ECML/PKDD (2) | 4 |
| 2021 | Sentiment Evolution in Social Network Based on Joint Pre-training ModelabstractSentiment analysis is one of the key tasks of natural language understanding. Most of sentiment analysis researches revolve around sentiment classification of subjective texts. However, research in the field of sentiment evolution analysis for complex interactive texts are notable. Sentiment evolution models the dynamics of sentiment orientation over time, it can predict the stage of event development. In this paper, we propose a sentiment evolution method based on a joint model to analyze the dynamics and interactions of individual sentiment on social media such as Weibo. The model contains two modules, sentiment encoder module based on pre-training model and time series prediction module based on Long Short-Term Memory(LSTM). We conducted experiments on real-world datasets which were crawled from Weibo. The experiment demonstrated a case study that analyzed the sentiment dynamics of topics related to COVID-19. Experimental results show that our method achieve an accuracy of 88.0%, which are about 14.7% higher than the existing methods. Xiaocao Wang, Chunjing Han, Xiaodan Zhang 0004, Honglei Lv, Shaoqin Huang |
CSCWD | 4 |
| 2020 | AFT-Anon: A scaling method for online trace anonymization based on anonymous flow tablesabstractAiming at the problem of trace anonymization performance of backbone networks, we propose a real-time anonymization method for the IP address of backbone network packets based on flow tables (named AFT-Anon). This method can dynamically build an anonymous flow table based on the captured data packets. The first data packet of a network flow is encrypted according to a specific encryption algorithm, and the encrypted fields are stored in the flow record. Subsequent data packets can obtain the encrypted fields by searching flow records and replace the corresponding fields of the original data packets to achieve anonymization of data packets. Based on the proposed method, a high-speed network anonymization system is developed and deployed on the backbone links of an Internet service provider. Experimental results show that the proposed method can improve the anonymization performance by more than 20 times, compared with the existing methods such as Crypto-Pan, and it can meet the requirements for online anonymization of 10G link. Chunjing Han, Kunkun Sun, Haina Tang, Yulei Wu, Xiaodan Zhang 0004 |
ISCC | 5 |
| 2020 | Exploiting Heterogeneous Artist and Listener Preference Graph for Music Genre ClassificationabstractMusic genres are useful for indexing, organizing, searching, and recommending songs and albums. Therefore, the automatic classification of music genres is an essential part of almost all kinds of music applications. Recent works focus on exploiting text, audio, or multi-modal information for genre classification, without considering the influence of the artists' and listeners' preference. However, intuitively, artists have their composing preferences, and listeners also have their music tastes. Both of them provide helpful hints to the music genre from different views, which are crucial to improve classification performance. Chunyuan Yuan, Qianwen Ma, Junyang Chen 0001, Wei Zhou 0019, Xiaodan Zhang 0004, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 5 |
| 2020 | Hierarchical Interaction Networks with Rethinking Mechanism for Document-Level Sentiment Analysis
Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xuehai Tang, Xiaodan Zhang 0004, Xin Wang 0086, Jizhong Han, Songlin Hu 0001 |
ECML/PKDD (3) | 5 |
| 2020 | DyHGCN: A Dynamic Heterogeneous Graph Convolutional Network to Learn Users' Dynamic Preferences for Information Diffusion Prediction
Chunyuan Yuan, Wei Zhou 0019, Xiaodan Zhang 0004, Songlin Hu 0001 |
ECML/PKDD (3) | 5 |
| 2019 | An Invisible Flow Watermarking for Traffic Tracking: A Hidden Markov Model ApproachabstractFlow watermarking is a promising active traffic tracking technology. It helps to establish the correspondence between the sender and the receiver, by embedding watermarks into packets with certain features of active interference traffic. Existing watermarking technologies have several drawbacks, such as the vulnerability to multi-flow attacks, low robustness and invisibility. This paper proposes a new traffic tracking technique based on hidden Markov model, called Hidden Markov State-based Flow Watermarking (HMSFW). HMSFW divides the observation range, i.e., inter-packet time, as Markov states and uses the State Transition Probability Mean (STPM) as the watermark carrier. It then adjusts the STPM of a given time interval according to historical traffic characteristics. The proposed HMSFW not only improves the robustness of embedded watermarks, but also effectively enhances their invisibility. The feasibility and invisibility of HMSFW are validated via extensive experimental results. Zhongjiang Yao, Lei Zhang 0116, Jingguo Ge, Yulei Wu, Xiaodan Zhang 0004 |
ICC | 5 |