EDBT 2026 Demo / reviewers in the wild / expert
Hongzhan Lin 0001
dblp:292/1751-1
· DBLP profile ↗
31ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0002-4111-8334ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 20 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style ControlabstractThe prevalence of fake news on social media calls for automated fact-checking systems that deliver not only accurate verdicts but also faithful explanations.However, existing large language model (LLM)-based methods often overlook deceptive misinformation styles in generated explanations, producing unfaithful rationales that may mislead human judgment.They also rely heavily on external knowledge sources, which can introduce hallucinations and incur substantial latency, undermining both reliability and responsiveness in realtime settings.To address these limitations, we propose REason-guided Fact-checking with Latent EXplanations (REFLEX), a selfrefining framework that explicitly controls reasoning style by anchoring explanations to the predicted verdict.REFLEX leverages selfdisagreement veracity signals between a backbone model and its fine-tuned variant to construct steering vectors, thereby naturally disentangling factual content from stylistic cues.Experiments on a real-world benchmark show that REFLEX achieves state-of-the-art performance with LLaMA-series models using only 465 self-refined samples.Owing to its transferability, REFLEX also yields gains of up to 7.54 Macro-F1 points on in-the-wild data.Further analysis shows that our method effectively mitigates faithful hallucination, leading to both more reliable explanations and more accurate verdicts than prior explainable fact-checking approaches. Chuyi Kong, Wei Gao 0001, Jing Ma 0004, Hongzhan Lin 0001, Yuxi Sun 0011 |
ACL (1) | 4 |
| 2026 | GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social AbuseabstractThe exponential growth of social media has profoundly transformed how information is created, disseminated, and absorbed, exceeding any precedent in the digital age. Regrettably, this explosion has also spawned a significant increase in the online abuse of memes. Evaluating the negative impact of memes is notably challenging, owing to their often subtle and implicit meanings, which are not directly conveyed through the overt text and image. In light of this, Large Multimodal Models (LMMs) have emerged as a focal point of interest due to their remarkable capabilities in handling diverse multimodal tasks. In response to this development, our article aims to thoroughly examine the capacity of various LMMs (e.g., GPT-4V, LLaVA, and Qwen-VL) to discern and respond to the nuanced aspects of social abuse manifested in memes. We introduce the comprehensive meme benchmark, GOAT-Bench , comprising over 6K varied memes encapsulating themes, such as implicit hate speech, sexism, and cyberbullying. Utilizing GOAT-Bench , we delve into the ability of LMMs to accurately assess hatefulness, misogyny, offensiveness, sarcasm, and harmful content. Our extensive experiments across a range of LMMs reveal that current models still exhibit a deficiency in safety awareness, showing insensitivity to various forms of implicit abuse. We posit that this shortfall represents a critical impediment to the realization of safe artificial intelligence. The GOAT-Bench and accompanying resources are publicly accessible at https://goatlmm.github.io/ , contributing to ongoing research in this vital field. Hongzhan Lin 0001, Bo Wang 0069, Ruichao Yang, Jing Ma 0004 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2026 | ExplainHM++: Explainable Harmful Meme Detection With Retrieval-Augmented Debate Between Large Multimodal ModelsabstractIdentifying harmful memes is challenging due to their implicit meanings, which are not always evident from texts and images alone. Existing solutions often lack clear explanations to justify their decisions. To address this gap, we propose an explainable approach,ExplainHM++, which detects harmful memes by reasoning over competing rationales from both harmful and harmless perspectives. First, inspired by the capabilities of Large Multimodal Models (LMMs) in text generation and multimodal reasoning, we developExplainHM, a one-stage multimodal debate in which LMMs generate explanations through contradictory arguments. Second, we fine-tune a small language model to serve as a judge in the debate, improving the integration of harmfulness rationales with the multimodal content of memes. However, we observe that a naive multimodal debate remains vulnerable, as it heavily depends on the inherent reasoning ability of LMMs to understand the memes. Given the evolving and noisy nature of memes, we further introduce a meme sample retrieval mechanism and a retrieval-augmented debate paradigm to strengthen and refine LMM-generated explanations. Extensive experiments on three public meme datasets demonstrate thatExplainHM++not only outperforms state-of-the-art methods but also provides superior, interpretable explanations for harmful meme detection. Hongzhan Lin 0001, Wei Gao 0001, Jing Ma 0004, Yang Deng 0002, Bo Wang 0069, Ruichao Yang, Tat-Seng Chua |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | A Graph-Enhanced Defense Framework for Explainable Fake News Detection with LLMabstractExplainable fake news detection aims to assess the veracity of news claims while providing human-friendly explanations. Existing methods incorporating investigative journalism are often inefficient and struggle with breaking news. Recent advances in large language models (LLMs) enable leveraging externally retrieved reports as evidence for detection and explanation generation, but unverified reports may introduce inaccuracies. Moreover, effective explainable fake news detection should provide a comprehensible explanation for all aspects of a claim to assist the public in verifying its accuracy. To address these challenges, we propose a graph-enhanced defense framework (G-Defense) that provides fine-grained explanations based solely on unverified reports. Specifically, we construct a claim-centered graph by decomposing the news claim into several sub-claims and modeling their dependency relationships. For each sub-claim, we use the retrieval-augmented generation (RAG) technique to retrieve salient evidence and generate competing explanations. We then introduce a defense-like inference module based on the graph to assess the overall veracity. Finally, we prompt an LLM to generate an intuitive explanation graph. Experimental results demonstrate that G-Defense achieves state-of-the-art performance in both veracity detection and the quality of its explanations. Bo Wang 0069, Jing Ma 0004, Hongzhan Lin 0001, Zhiwei Yang 0005, Ruichao Yang, Yuan Tian 0016, Yi Chang 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2025 | Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal TriggersabstractHateful meme detection aims to prevent the proliferation of hateful memes on various social media platforms. Considering its impact on social environments, this paper introduces a previously ignored but significant threat to hateful meme detection: backdoor attacks. By injecting specific triggers into meme samples, backdoor attackers can manipulate the detector to output their desired outcomes. To explore this, we propose the Meme Trojan framework to initiate backdoor attacks on hateful meme detection. Meme Trojan involves creating a novel Cross-Modal Trigger (CMT) and a learnable trigger augmentor to enhance the trigger pattern according to each input sample. Due to the cross-modal property, the proposed CMT can effectively initiate backdoor attacks on hateful meme detectors under an automatic application scenario. Additionally, the injection position and size of our triggers are adaptive to the texts contained in the meme, which ensures that the trigger is seamlessly integrated with the meme content. Our approach outperforms the state-of-the-art backdoor attack methods, showing significant improvements in effectiveness and stealthiness. We believe that this paper will draw more attention to the potential threat posed by backdoor attacks on hateful meme detection. Ruofei Wang, Hongzhan Lin 0001, Ziyuan Luo, Ka Chun Cheung, Simon See, Jing Ma 0004, Renjie Wan |
AAAI | 2 |
| 2025 | Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelabstractRecent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and scale of audio language models, as well as leveraging larger datasets, and generally, acoustic codecs, such as EnCodec, are used for audio tokenization. However, these codecs were originally designed for audio compression, which may lead to suboptimal performance in the context of audio LLM. Our research aims to address the shortcomings of current audio LLM codecs, particularly their challenges in maintaining semantic integrity in generated audio. For instance, existing methods like VALL-E, which condition acoustic token generation on text transcriptions, often suffer from content inaccuracies and elevated word error rates (WER) due to semantic misinterpretations of acoustic tokens, resulting in word skipping and errors. To overcome these issues, we propose a straightforward yet effective approach called X-Codec. X-Codec incorporates semantic features from a pre-trained semantic encoder before the Residual Vector Quantization (RVQ) stage and introduces a semantic reconstruction loss after RVQ. By enhancing the semantic ability of the codec, X-Codec significantly reduces WER in speech synthesis tasks and extends these benefits to non-speech applications, including music and sound generation. Our experiments in text-to-speech, music continuation, and text-to-sound tasks demonstrate that integrating semantic information substantially improves the overall performance of language models in audio generation. Zhen Ye 0006, Peiwen Sun, Jiahe Lei, Hongzhan Lin 0001, Xu Tan 0003, Zheqi Dai, Qiuqiang Kong, Jianyi Chen, Yike Guo, Wei Xue 0002 |
AAAI | 4 |
| 2025 | FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language ModelsabstractHongzhan Lin, Yang Deng, Yuxuan Gu, Wenxuan Zhang, Jing Ma, See-Kiong Ng, Tat-Seng Chua. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hongzhan Lin 0001, Yang Deng 0002, Yuxuan Gu 0004, Wenxuan Zhang 0001, Jing Ma 0004, See-Kiong Ng, Tat-Seng Chua |
ACL (1) | 1 |
| 2025 | AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on HarmfulnessabstractZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo, Zhen Ye, Guang Chen, Zhiyong Huang, Jing Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zixin Chen, Hongzhan Lin 0001, Zhen Ye 0006, Guang Chen 0003, Zhiyong Huang 0010, Jing Ma 0004 |
ACL (1) | 2 |
| 2025 | Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language ModelsabstractData synthesis has become a crucial research area in large language models (LLMs), especially for generating high-quality instruction fine-tuning data to enhance downstream performance.In code generation, a key application of LLMs, manual annotation of code instruction data is costly.Recent methods, such as Code Evol-Instruct and OSS-Instruct, leverage LLMs to synthesize large-scale code instruction data, significantly improving LLM coding capabilities.However, these approaches face limitations due to unidirectional synthesis and randomness-driven generation, which restrict data quality and diversity.To overcome these challenges, we introduce Tree-of-Evolution (ToE), a novel framework that models code instruction synthesis process with a tree structure, exploring multiple evolutionary paths to alleviate the constraints of unidirectional generation.Additionally, we propose optimizationdriven evolution, which refines each generation step based on the quality of the previous iteration.Experimental results across five widely-used coding benchmarks-HumanEval, MBPP, EvalPlus, LiveCodeBench, and Big-CodeBench-demonstrate that base models fine-tuned on just 75k data synthesized by our method achieve comparable or superior performance to the state-of-the-art open-weight Code LLM, Qwen2.5-Coder-Instruct, which was finetuned on millions of samples. Hongzhan Lin 0001, Mohan Kankanhalli, Jing Ma 0004 |
ACL (1) | 3 |
| 2025 | Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language ModelsabstractShuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Richard Yi Da Xu, Yunya Song, Xian Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jing Ma 0004, Hongzhan Lin 0001, Liang Bai 0001, Zhihua Wang 0008, Yunya Song, Xian Yang 0001 |
ACL (1) | 3 |
| 2025 | CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?abstractRecent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model’s code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs’ code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark’s ability to probe deeper into models’ code understanding abilities. Our benchmark is available at https://github.com/CodeLLM-Research/CodeJudge-Eval . Hongzhan Lin 0001, Weixiang Yan, Annan Li, Jing Ma 0004 |
COLING | 4 |
| 2025 | MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language ModelsabstractThe proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness.Existing evaluation approaches predominantly focus on mLLMs' detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts.In this paper, we propose MemeArena, an agent-based arenastyle evaluation framework that provides a context-aware and unbiased assessment for mLLMs' understanding of multimodal harmfulness.Specifically, MemeArena simulates diverse interpretive contexts to formulate evaluation tasks that elicit perspective-specific analyses from mLLMs.By integrating varied viewpoints and reaching consensus among evaluators, it enables fair and unbiased comparisons of mLLMs' abilities to interpret multimodal harmfulness.Extensive experiments demonstrate that our framework effectively reduces the evaluation biases of judge agents, with judgment results closely aligning with human preferences, offering valuable insights into reliable and comprehensive mLLM evaluations in multimodal harmfulness understanding.Our code and data are publicly available at https://github.com/Lbotirx/MemeArena. Zixin Chen, Hongzhan Lin 0001, Yayue Deng, Jing Ma 0004 |
EMNLP | 2 |
| 2025 | ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer UseabstractRecent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays and complex environments which lead to smaller target sizes. In this paper, we introduce ScreenSpot-Pro, a new benchmark designed to rigorously evaluate the grounding capabilities of MLLMs in high-resolution professional settings. The benchmark comprises authentic high-resolution images from a variety of professional domains with expert annotations. It spans 23 applications across five industries and three operating systems. Existing GUI grounding models perform poorly on this dataset, with the best model achieving only 18.9%. Our experiments reveal that strategically reducing the search area enhances accuracy. Based on this insight, we propose ScreenSeekeR, a visual search method that utilizes the GUI knowledge of a strong planner to guide a cascaded search, achieving state-of-the-art performance with 48.1% without any additional training. We hope that our benchmark and findings will advance the development of GUI agents for professional settings. Hongzhan Lin 0001, Jing Ma 0004, Zhiyong Huang 0010, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2025 | LLM-Enhanced Multiple Instance Learning for Joint Rumor and Stance Detection with Social Context InformationabstractThe proliferation of misinformation, such as rumors on social media, has drawn significant attention, prompting various expressions of stance among users. Although rumor detection and stance detection are distinct tasks, they can complement each other. Rumors can be identified by cross-referencing stances in related posts, and stances are influenced by the nature of the rumor. However, existing stance detection methods often require post-level stance annotations, which are costly to obtain. We propose a novel LLM-enhanced MIL approach to jointly predict post stance and claim class labels, supervised solely by claim labels, using an undirected microblog propagation model. Our weakly supervised approach relies only on bag-level labels of claim veracity, aligning with multi-instance learning (MIL) principles. To achieve this, we transform the multi-class problem into multiple MIL-based binary classification problems. We then employ a discriminative attention layer to aggregate the outputs from these classifiers into finer-grained classes. Experiments conducted on three rumor datasets and two stance datasets demonstrate the effectiveness of our approach, highlighting strong connections between rumor veracity and expressed stances in responding posts. Our method shows promising performance in joint rumor and stance detection compared to the state-of-the-art methods. Ruichao Yang, Jing Ma 0004, Wei Gao 0001, Hongzhan Lin 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2024 | CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsabstractSocial media abounds with multimodal sarcasm, and identifying sarcasm targets is particularly challenging due to the implicit incongruity not directly evident in the text and image modalities. Current methods for Multimodal Sarcasm Target Identification (MSTI) predominantly focus on superficial indicators in an end-to-end manner, overlooking the nuanced understanding of multimodal sarcasm conveyed through both the text and image. This paper proposes a versatile MSTI framework with a coarse-to-fine paradigm, by augmenting sarcasm explainability with reasoning and pre-training knowledge. Inspired by the powerful capacity of Large Multimodal Models (LMMs) on multimodal reasoning, we first engage LMMs to generate competing rationales for coarser-grained pre-training of a small language model on multimodal sarcasm detection. We then propose fine-tuning the model for finer-grained sarcasm target identification. Our framework is thus empowered to adeptly unveil the intricate targets within multimodal sarcasm and mitigate the negative impact posed by potential noise inherently in LMMs. Experimental results demonstrate that our model far outperforms state-of-the-art MSTI methods, and markedly exhibits explainability in deciphering sarcasm as well. Zixin Chen, Hongzhan Lin 0001, Mingfei Cheng, Jing Ma 0004, Guang Chen 0003 |
ACL (1) | 2 |
| 2024 | Towards Low-Resource Harmful Meme Detection with LMM AgentsabstractThe proliferation of Internet memes in the age of social media necessitates effective identification of harmful ones. Due to the dynamic nature of memes, existing data-driven models may struggle in low-resource scenarios where only a few labeled examples are available. In this paper, we propose an agency-driven framework for low-resource harmful meme detection, employing both outward and inward analysis with few-shot annotated samples. Inspired by the powerful capacity of Large Multimodal Models (LMMs) on multimodal reasoning, we first retrieve relative memes with annotations to leverage label information as auxiliary signals for the LMM agent. Then, we elicit knowledge-revising behavior within the LMM agent to derive well-generalized insights into meme harmfulness. By combining these strategies, our approach enables dialectical reasoning over intricate and implicit harm-indicative patterns. Extensive experiments conducted on three meme datasets demonstrate that our proposed approach achieves superior performance than state-of-the-art methods on the low-resource harmful meme detection task. Jianzhao Huang, Hongzhan Lin 0001, Guang Chen 0003, Jing Ma 0004 |
EMNLP | 2 |
| 2024 | AMR-Evol: Adaptive Modular Response Evolution Elicits Better Knowledge Distillation for Large Language Models in Code GenerationabstractThe impressive performance of proprietary LLMs like GPT4 in code generation has led to a trend to replicate these capabilities in open-source models through knowledge distillation (e.g.Code Evol-Instruct).However, these efforts often neglect the crucial aspect of response quality, relying heavily on teacher models for direct response distillation.This paradigm, especially for complex instructions, can degrade the quality of synthesized data, compromising the knowledge distillation process.To this end, our study introduces the Adaptive Modular Response Evolution (AMR-Evol) framework, which employs a two-stage process to refine response distillation.The first stage, modular decomposition, breaks down the direct response into more manageable sub-modules.The second stage, adaptive response evolution, automatically evolves the response with the related function modules.Our experiments with three popular code benchmarks-HumanEval, MBPP, and EvalPlus-attests to the superiority of the AMR-Evol framework over baseline response distillation methods.By comparing with the open-source Code LLMs trained on a similar scale of data, we observed performance enhancements: more than +3.0 points on HumanEval-Plus and +1.0 points on MBPP-Plus, which underscores the effectiveness of our framework.Our codes are available at https://github.com/ChiYeungLaw/ AMR-Evol. Module 3def validate_matrix(matrix: list) -> None:""" Description:Validates if the input matrix is a 3x3 matrix.Parameters:-matrix (list): The input matrix to be validated.Raises:-ValueError: If the matrix is not a 3x3 matrix.""" Coding Tasks Design a function that implements the Laplace expansion theorem to calculate the determinant of a 3x3 matrix, but now with the added condition of utilizing nested for loops and ensuring that the matrix elements are limited to integers only.Teacher Model Response def calculate_determinant(matrix): if len(matrix) != 3 or len(matrix[0]) != 3: raise ValueError('Matrix must be a 3x3 matrix') determinant = 0 for i in range(3): submatrix = [[matrix[j][k] for k in range(3) if k != i] for j in range(1, 3)] sub_determinant = calculate_determinant(submatrix) sign = (-1) ** i determinant += sign * matrix[0][i] * sub_determinant return determinant if len(matrix) == 2 else round(determinant) Direct Distillation Module 2def calculate_minor_matrix(matrix: list, row: int, col: int) -> list:""" Description: Calculates the minor matrix by removing the specified row Parameters:-matrix (list): The input matrix.-row (int): The row index to be removed.-col (int): The column index to be removed. Xin Li 0056, Hongzhan Lin 0001, Jing Ma 0004, Lidong Bing |
EMNLP | 3 |
| 2024 | Unleashing Trigger-Free Event Detection: Revealing Event Correlations Via a Contrastive Derangement FrameworkabstractEvent detection (ED), detecting events with specified types observed in given texts, is critical to many downstream applications. Existing ED methods generally require high-quality triggers annotated by human experts, which is labor-intensive, especially for those nontrivial texts about breaking events. In this paper, we propose a novel trigger-free ED framework that detects multiple events from a given text without pre-defined triggers. Specifically, we first shed light on the event correlations with input texts using a joint embedding paradigm. Next, we devise derangement-based contrastive learning to model fine-grained correlations between multi-event instances. Since events in training benchmarks are usually imbalanced, we further design a simple yet effective event derangement module for balanced training. Experimental results on two benchmarks show that our trigger-free method is remarkably competitive to state-of-the-art trigger-based baselines. Hongzhan Lin 0001, Haiqin Yang, Jing Ma 0004 |
ICASSP | 1 |
| 2024 | Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language ModelsabstractThe age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, which is not explicitly conveyed through the surface text and image. However, existing harmful meme detection methods do not present readable explanations that unveil such implicit meaning to support their detection decisions. In this paper, we propose an explainable approach to detect harmful memes, achieved through reasoning over conflicting rationales from both harmless and harmful positions. Specifically, inspired by the powerful capacity of Large Language Models (LLMs) on text generation and reasoning, we first elicit multimodal debate between LLMs to generate the explanations derived from the contradictory arguments. Then we propose to fine-tune a small language model as the debate judge for harmfulness inference, to facilitate multimodal fusion between the harmfulness rationales and the intrinsic multimodal information within memes. In this way, our model is empowered to perform dialectical reasoning over intricate and implicit harm-indicative patterns, utilizing multimodal explanations originating from both harmless and harmful arguments. Extensive experiments on three public meme datasets demonstrate that our harmful meme detection approach achieves much better performance than state-of-the-art methods and exhibits a superior capacity for explaining the meme harmfulness of the model predictions. Hongzhan Lin 0001, Wei Gao 0001, Jing Ma 0004, Bo Wang 0069, Ruichao Yang |
WWW | 1 |
| 2024 | Explainable Fake News Detection with Large Language Model via Defense Among Competing WisdomabstractMost fake news detection methods learn latent feature representations based on neural networks, which makes them black boxes to classify a piece of news without giving any justification. Existing explainable systems generate veracity justifications from investigative journalism, which suffer from debunking delayed and low efficiency. Recent studies simply assume that the justification is equivalent to the majority opinions expressed in the wisdom of crowds. However, the opinions typically contain some inaccurate or biased information since the wisdom of crowds is uncensored. To detect fake news from a sea of diverse, crowded and even competing narratives, in this paper, we propose a novel defense-based explainable fake news detection framework. Specifically, we first propose an evidence extraction module to split the wisdom of crowds into two competing parties and respectively detect salient evidences. To gain concise insights from evidences, we then design a prompt-based module that utilizes a large language model to generate justifications by inferring reasons towards two possible veracities. Finally, we propose a defense-based inference module to determine veracity via modeling the defense among these justifications. Extensive experiments conducted on two real-world benchmarks demonstrate that our proposed method outperforms state-of-the-art baselines in terms of fake news detection and provides high-quality justifications. Bo Wang 0069, Jing Ma 0004, Hongzhan Lin 0001, Zhiwei Yang 0005, Ruichao Yang, Yuan Tian 0016, Yi Chang 0001 |
WWW | 3 |
| 2024 | Towards low-resource rumor detection: Unified contrastive transfer with propagation structure
Hongzhan Lin 0001, Jing Ma 0004, Ruichao Yang, Zhiwei Yang 0005, Mingfei Cheng |
Neurocomputing | 1 |
| 2023 | Zero-Shot Rumor Detection with Propagation Structure via Prompt LearningabstractThe spread of rumors along with breaking events seriously hinders the truth in the era of social media. Previous studies reveal that due to the lack of annotated resources, rumors presented in minority languages are hard to be detected. Furthermore, the unforeseen breaking events not involved in yesterday's news exacerbate the scarcity of data resources. In this work, we propose a novel zero-shot framework based on prompt learning to detect rumors falling in different domains or presented in different languages. More specifically, we firstly represent rumor circulated on social media as diverse propagation threads, then design a hierarchical prompt encoding mechanism to learn language-agnostic contextual representations for both prompts and rumor data. To further enhance domain adaptation, we model the domain-invariant structural features from the propagation threads, to incorporate structural position representations of influential community response. In addition, a new virtual response augmentation method is used to improve model training. Extensive experiments conducted on three real-world datasets demonstrate that our proposed model achieves much better performance than state-of-the-art methods and exhibits a superior capacity for detecting rumors at early stages. Hongzhan Lin 0001, Pengyao Yi, Jing Ma 0004, Haiyun Jiang, Shuming Shi 0001, Ruifang Liu |
AAAI | 1 |
| 2023 | Dual-Scale Interest Extraction Framework with Self-Supervision for Sequential RecommendationabstractIn the sequential recommendation task, the recommender generally learns multiple embeddings from a user’s historical behaviors, to catch the diverse interests of the user. Nevertheless, the existing approaches just extract each interest independently for the corresponding sub-sequence while ignoring the global correlation of the entire interaction sequence, which may fail to capture the user’s inherent preference for the potential interests generalization and unavoidably make the recommended items homogeneous with the historical behaviors. In this paper, we propose a novel Dual-Scale Interest Extraction framework (DSIE) to precisely estimate the user’s current interests. Specifically, DSIE explicitly models the user’s inherent preference with contrastive learning by attending over his/her entire interaction sequence at the global scale and catches the user’s diverse interests in a fine granularity at the local scale. Moreover, we develop a novel interest aggregation module to integrate the multi-interests according to the inherent preference to generate the user’s current interests for the next-item prediction. Experiments conducted on three real-world benchmark datasets demonstrate that DSIE outperforms the state-of-the-art models in terms of recommendation preciseness and novelty. Hongzhan Lin 0001, Jinshan Ma, Guang Chen 0003 |
ECAI | 2 |
| 2023 | WSDMS: Debunk Fake News via Weakly Supervised Detection of Misinforming Sentences with Contextualized Social WisdomabstractIn recent years, we witness the explosion of false and unconfirmed information (i.e., rumors) that went viral on social media and shocked the public.Rumors can trigger versatile, mostly controversial stance expressions among social media users.Rumor verification and stance detection are different yet relevant tasks.Fake news debunking primarily focuses on determining the truthfulness of news articles, which oversimplifies the issue as fake news often combines elements of both truth and falsehood.Thus, it becomes crucial to identify specific instances of misinformation within the articles.In this research, we investigate a novel task in the field of fake news debunking, which involves detecting sentence-level misinformation.One of the major challenges in this task is the absence of a training dataset with sentence-level annotations regarding veracity.Inspired by the Multiple Instance Learning (MIL) approach, we propose a model called Weakly Supervised Detection of Misinforming Sentences (WSDMS).This model only requires bag-level labels for training but is capable of inferring both sentence-level misinformation and article-level veracity, aided by relevant social media conversations that are attentively contextualized with news sentences.We evaluate WSDMS on three real-world benchmarks and demonstrate that it outperforms existing stateof-the-art baselines in debunking fake news at both the sentence and article levels.News Title: NASA Will Pay You 100,000 USD To Stay In Bed For 60 Days!News Article: 𝑠 !: Wouldn't you just love to carry on sleeping on a Monday morning without having to submit to the Monday morning blues and get ready for work?𝑠 " : What type of heaven would you envisage if you were paid to stay in bed 𝑠 # : You can get paid a huge sum of money just staying in bed for two whole months and by you know who, NASA no less!!! yes the American space agency NASA is paying $100,000 to stay in bed for 60 days.𝑠 $ : Most of us dream about hanging out in bed, all day, every day.𝑠 % : NASA is currently on the lookout for people to participate in their "Bed Rest Studies", in which participants will have to stay in bed for 60 days straight.𝑠 & : It does sound like the dream job, right?… 𝑠 ' : You wouldn't just be sleeping you can keep yourself occupied with books, TV, video games, and they can also use their phones as they please… Only $100,000?Not good enough.… Ruichao Yang, Wei Gao 0001, Jing Ma 0004, Hongzhan Lin 0001, Zhiwei Yang 0005 |
EMNLP | 4 |
| 2023 | Semantic-consistent learning for one-shot joint entity and relation extraction
Jinglei Li, Hongzhan Lin 0001, Guang Chen 0003, Bosen Zhang, Boya Ren |
Appl. Intell. | 3 |
| 2022 | A Coarse-to-fine Cascaded Evidence-Distillation Neural Network for Explainable Fake News DetectionabstractExisting fake news detection methods aim to classify a piece of news as true or false and provide veracity explanations, achieving remarkable performances. However, they often tailor automated solutions on manual fact-checked reports, suffering from limited news coverage and debunking delays. When a piece of news has not yet been fact-checked or debunked, certain amounts of relevant raw reports are usually disseminated on various media outlets, containing the wisdom of crowds to verify the news claim and explain its verdict. In this paper, we propose a novel Coarse-to-fine Cascaded Evidence-Distillation (CofCED) neural network for explainable fake news detection based on such raw reports, alleviating the dependency on fact-checked ones. Specifically, we first utilize a hierarchical encoder for web text representation, and then develop two cascaded selectors to select the most explainable sentences for verdicts on top of the selected top-K reports in a coarse-to-fine manner. Besides, we construct two explainable fake news datasets, which is publicly available. Experimental results demonstrate that our model significantly outperforms state-of-the-art detection baselines and generates high-quality explanations from diverse evaluation perspectives. Zhiwei Yang 0005, Jing Ma 0004, Hechang Chen, Hongzhan Lin 0001, Yi Chang 0001 |
COLING | 4 |
| 2022 | AMIF: A Hybrid Model for Improving Fact Checking in Product Question AnsweringabstractFact checking in product-related community question answering is the task of verifying the truthfulness of an answer towards a given question, where the study has just begun. Most existing related work has focused on tailoring solutions to shallow feature fusion for the single-text claim involved with fact-checked evidence, limiting their success and generality in such answer truthfulness prediction task on E-commerce platforms. In this study, we propose an attention-based hybrid framework for multi-feature interaction fusion to determine the truthfulness of the answer towards a product-related question in E-commerce, which could not only support fine-grained semantic calibration between question-answer pairs for better understanding of the target answers, but also substantially cross-check all retrieved evidence to mine coherent opinions towards the pair. In addition, our framework further integrates non-textual features from metadata for improving performance. Extensive experiments conducted on real-world representative benchmark data show that our proposed model achieves superior performance on the task of answer veracity prediction. Hongzhan Lin 0001, Jing Ma 0004, Zhiwei Yang 0005, Guang Chen 0003 |
IJCNN | 1 |
| 2022 | A Weakly Supervised Propagation Model for Rumor Verification and Stance Detection with Multiple Instance LearningabstractThe diffusion of rumors on social media generally follows a propagation tree structure, which provides valuable clues on how an original message is transmitted and responded by users over time. Recent studies reveal that rumor verification and stance detection are two relevant tasks that can jointly enhance each other despite their differences. For example, rumors can be debunked by cross-checking the stances conveyed by their relevant posts, and stances are also conditioned on the nature of the rumor. However, stance detection typically requires a large training set of labeled stances at post level, which are rare and costly to annotate. Enlightened by Multiple Instance Learning (MIL) scheme, we propose a novel weakly supervised joint learning framework for rumor verification and stance detection which only requires bag-level class labels concerning the rumor's veracity. Specifically, based on the propagation trees of source posts, we convert the two multi-class problems into multiple MIL-based binary classification problems where each binary model is focused on differentiating a target class (of rumor or stance) from the remaining classes. Then, we propose a hierarchical attention mechanism to aggregate the binary predictions, including (1) a bottom-up/top-down tree attention layer to aggregate binary stances into binary veracity; and (2) a discriminative attention layer to aggregate the binary class into finer-grained classes. Extensive experiments conducted on three Twitter-based datasets demonstrate promising performance of our model on both claim-level rumor detection and post-level stance classification compared with state-of-the-art methods. Ruichao Yang, Jing Ma 0004, Hongzhan Lin 0001, Wei Gao 0001 |
SIGIR | 3 |
| 2021 | Rumor Detection on Twitter with Claim-Guided Hierarchical Graph Attention NetworksabstractRumors are rampant in the era of social media.Conversation structures provide valuable clues to differentiate between real and fake claims.However, existing rumor detection methods are either limited to the strict relation of user responses or oversimplify the conversation structure.In this study, to substantially reinforces the interaction of user opinions while alleviating the negative impact imposed by irrelevant posts, we first represent the conversation thread as an undirected interaction graph.We then present a Claim-guided Hierarchical Graph Attention Network for rumor classification, which enhances the representation learning for responsive posts considering the entire social contexts and attends over the posts that can semantically infer the target claim.Extensive experiments on three Twitter datasets demonstrate that our rumor detection method achieves much better performance than stateof-the-art methods and exhibits a superior capacity for detecting rumors at early stages. Hongzhan Lin 0001, Jing Ma 0004, Mingfei Cheng, Zhiwei Yang 0005, Guang Chen 0003 |
EMNLP (1) | 1 |
| 2021 | Boosting Low-Resource Intent Detection with in-Scope Prototypical NetworksabstractIdentifying intentions from users can help improve the response quality of task-oriented dialogue systems. How to use only limited labeled in-domain (ID) examples for zero-shot unknown intent detection and few-shot ID classification is a more challenging task in spoken language understanding. Existing related methods heavily rely upon the multi-domain datasets containing large-scale independent source domains for meta-training. In this paper, we propose a universal In-scope Prototypical Networks for low-resource intent detection to be general to dialogue meta-train datasets lacking widely-varying domains, which focuses on the scope of episodic intent classes to construct meta-task dynamically. Also, we introduce loss with margin principle to better distinguish samples. Experiments on two benchmark datasets show that our model consistently outperforms other baselines on zero-shot unknown intent detection without deteriorating the competitive performance on few-shot ID classification. Hongzhan Lin 0001, Yuanmeng Yan, Guang Chen 0003 |
ICASSP | 1 |
| 2021 | TANTP: Conversational Emotion Recognition Using Tree-Based Attention Networks with Transformer Pre-training
Hongzhan Lin 0001, Guang Chen 0003 |
PAKDD (2) | 2 |