VLDB 2026 Research / reviewers in the wild / expert
Yi R. Fung 0001
dblp:223/2782-1 · also Yi Fung 0001, Yi R. (May) Fung, Yi Ren Fung 0001
· DBLP profile ↗
28ranked-venue papers
6as first author
28since 2021 · last 2026
0009-0006-1869-0363ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 28 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Human-Centric Trustworthy Foundation Model via Advanced Reasoning and Agentic FrameworksabstractAs foundation models grow in size and scope, crucial challenges remain in scaling their trustworthiness and adaptability to meet the diverse needs of individual users, as well as mitigating their risk of generating unhelpful, non-factual, or harmful content. To address this, we propose to reframe model reasoning through a unified paradigm of active knowledge grounding that coordinates different tools and modalities. First, to scale reasoning depth and creativity, we introduce the novel paradigm of Thinking with Images to encourage models to externalize intermediate structure and perform interleaved cross-modal advanced reasoning beyond text-centric cues. To further scale honesty and bridge knowledge gaps reliably, we develop one of the first vision-language deep research agents, WebWatcher, that actively gathers and verifies information from the web with enhanced fragmented reasoning capability. Ultimately, to scale effective and efficient human-AI collaboration, we propose AdaCtrl as a novel training mechanism for dynamically aligning model behavior with individual user preferences and difficulty awareness to adaptively allocate computational resources. Together, these three pillars of integrating advanced multimodal reasoning, autonomous discovery, and adaptive alignment form a foundational framework for advancing the frontier of next generation human-centric trustworthy AI systems. Yi R. Fung 0001 |
AAAI | 1 |
| 2026 | CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented ValidationabstractLarge Language Models (LLMs) have emerged as powerful assistants for scientific writing.However, concerns remain about the quality and reliability of the generated text, including citation accuracy and faithfulness.While most recent work relies on methods such as LLMas-a-Judge, the reliability of LLM-as-a-Judge alone is also in doubt.In this work, we reframe citation evaluation as a problem of citation attribution alignment, which assesses whether LLM-generated citations match those a human author would include for the same text.We propose CiteGuard, a retrieval-aware agent framework designed to provide more faithful grounding for citation validation.CiteGuard improves over the prior baseline by 10 percentage points and achieves up to 68.1% accuracy on the CiteME benchmark, approaching human performance (69.2%).It also identifies alternative valid citations and demonstrates generalization ability for cross-domain citation attribution. 1 Yee Man Choi, Xuehang Guo, Yi R. Fung 0001, Qingyun Wang 0005 |
ACL (1) | 3 |
| 2026 | Learning Diverse Responses with Prefix-Conditioned Supervised Fine-TuningabstractLarge language models (LLMs) have shown strong performance on hard reasoning and general instruction-following tasks.However, when sampling multiple outputs for the same prompt, they often produce highly homogeneous, repetitive responses, resulting in inefficient exploration.This limits the gains from test-time scaling and constrains the upper bound of reinforcement learning (RL) training.We attribute this issue in part to supervised fine-tuning (SFT): when a single prompt is paired with multiple reference responses, the model is trained to generate diverse outputs under the same prior condition, which induces optimization interference and can lead to diversity collapse.To address this, we propose Prefix-Conditioned SFT (P-SFT), a simple yet effective method that constructs semantically consistent yet distributionally distinct prior contents to different responses, thereby projecting the instruction into distinct latent regions to establish diverse prior distributions and decouple the one-to-many mapping.Experiments on large reasoning language models show that our approach improves absolute performance by 5.3% on reasoning benchmarks and increases generation diversity by 198.3% on average, while substantially enhancing output diversity and test-time scaling.Notably, even without any additional training, our prefixing strategy can be applied at inference time alone and still yields significant gains in both diversity and reasoning performance for instruction-tuned LLMs and reasoning-enhanced models. Zhiyuan Fan, Guanqiao Chen, Yanyi Huang, Mingkuan Zhao, Dadi Guo, Yi R. Fung 0001 |
ACL (1) | 6 |
| 2026 | Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning ModelsabstractDadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He, Haoran Li, Yuxin Li, Yumeng Wang, Yi R. Fung. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dadi Guo, Zhiyuan Fan, Zhitao He 0001, Yumeng Wang 0010, Yi R. Fung 0001 |
ACL (1) | 8 |
| 2026 | CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsabstractJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cheng Qian 0008, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung 0001 |
ACL (1) | 7 |
| 2025 | MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence CalibrationabstractIn recent years, multimodal large language models (MLLMs) have made significant progress but continue to face inherent challenges in multimodal reasoning, which requires multi-level (e.g., perception, reasoning) and multi-granular (e.g., multi-step reasoning chain) advanced inferencing.Prior work on estimating model confidence tends to focus on the overall response for training and calibration, but fails to assess confidence in each reasoning step, leading to undesirable hallucination snowballing.In this work, we present MMBoundary, a novel framework that advances the knowledge boundary awareness of MLLMs through reasoning step confidence calibration.To achieve this, we propose to incorporate complementary textual and cross-modal self-rewarding signals to estimate confidence at each step of the MLLM reasoning process.In addition to supervised fine-tuning MLLM on this set of self-rewarding confidence estimation signal for initial confidence expression warm-up, we introduce a reinforcement learning stage with multiple reward functions for further aligning model knowledge and calibrating confidence at each reasoning step, enhancing reasoning chain self-correction.Empirical results show that MMBoundary significantly outperforms existing methods across diverse domain datasets and metrics, achieving an average of 7.5% reduction in multimodal confidence calibration errors and up to 8.3% improvement in task performance 1 . Zhitao He 0001, Sandeep Polisetty, Zhiyuan Fan, Shujin Wu, Yi R. Fung 0001 |
ACL (1) | 6 |
| 2025 | VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual CuesabstractVisually linking matching cues is a crucial ability in daily life, such as identifying the same person in multiple photos based on their cues, even without knowing who they are. Despite the extensive knowledge that vision-language models (VLMs) possess, it remains largely unexplored whether they are capable of performing this fundamental task. To address this, we introduce VLM2-Bench, a benchmark designed to assess whether VLMs can Visually Link Matching cues, with 9 subtasks and over 3,000 test cases. Comprehensive evaluation across twelve VLMs, along with further analysis of various language-side and vision-side prompting methods, leads to a total of eight key findings. We identify critical challenges in models’ ability to link visual cues, highlighting a significant performance gap. Based on these insights, we advocate for (i) enhancing core visual capabilities to improve adaptability and reduce reliance on prior knowledge, (ii) establishing clearer principles for integrating language-based reasoning in vision-centric tasks to prevent unnecessary biases, and (iii) shifting vision-text training paradigms toward fostering models’ ability to independently structure and infer relationships among visual cues. Jianshu Zhang 0003, Dongyu Yao, Renjie Pi, Paul Pu Liang, Yi R. Fung 0001 |
ACL (1) | 5 |
| 2025 | PropaInsight: Toward Deeper Understanding of Propaganda in Terms of Techniques, Appeals, and IntentabstractPropaganda plays a critical role in shaping public opinion and fueling disinformation. While existing research primarily focuses on identifying propaganda techniques, it lacks the ability to capture the broader motives and the impacts of such content. To address these challenges, we introduce PropaInsight, a conceptual framework grounded in foundational social science research, which systematically dissects propaganda into techniques, arousal appeals, and underlying intent. PropaInsight offers a more granular understanding of how propaganda operates across different contexts. Additionally, we present PropaGaze, a novel dataset that combines human-annotated data with high-quality synthetic data generated through a meticulously designed pipeline. Our experiments show that off-the-shelf LLMs struggle with propaganda analysis, but PropaGaze significantly improves performance. Fine-tuned Llama-7B-Chat achieves 203.4% higher text span IoU in technique identification and 66.2% higher BertScore in appeal analysis compared to 1-shot GPT-4-Turbo. Moreover, PropaGaze complements limited human-annotated data in data-sparse and cross-domain scenarios, demonstrating its potential for comprehensive and generalizable propaganda analysis. Jiateng Liu, Lin Ai, Zizhou Liu, Payam Karisani, Zheng Hui, Yi R. Fung 0001, Preslav Nakov, Julia Hirschberg, Heng Ji 0001 |
COLING | 6 |
| 2025 | Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data RefinementabstractThe increasing demand for personalized interactions with large language models (LLMs) calls for methodologies capable of accurately and efficiently identifying user opinions and preferences. Retrieval augmentation emerges as an effective strategy, as it can accommodate a vast number of users without the costs from fine-tuning. Existing research, however, has largely focused on enhancing the retrieval stage and devoted limited exploration toward optimizing the representation of the database, a crucial aspect for tasks such as personalization. In this work, we examine the problem from a novel angle, focusing on how data can be better represented for more data-efficient retrieval in the context of LLM customization. To tackle this challenge, we introduce Persona-DB, a simple yet effective framework consisting of a hierarchical construction process to improve generalization across task contexts and collaborative refinement to effectively bridge knowledge gaps among users. In the evaluation of response prediction, Persona-DB demonstrates superior context efficiency in maintaining accuracy with a significantly reduced retrieval size, a critical advantage in scenarios with extensive histories or limited context windows. Our experiments also indicate a marked improvement of over 10% under cold-start scenarios, when users have extremely sparse data. Furthermore, our analysis reveals the increasing importance of collaborative knowledge as the retrieval capacity expands. Chenkai Sun, Ke Yang 0003, Revanth Gangi Reddy, Yi R. Fung 0001, Hou Pong Chan, Kevin Small, ChengXiang Zhai, Heng Ji 0001 |
COLING | 4 |
| 2025 | Aligning LLMs with Individual Preferences via InteractionabstractAs large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption. While previous research focuses on general alignment to principles such as helpfulness, harmlessness, and honesty, the need to account for individual and diverse preferences has been largely overlooked, potentially undermining customized human experiences. To address this gap, we train LLMs that can “interact to align”, essentially cultivating the meta-skill of LLMs to implicitly infer the unspoken personalized preferences of the current user through multi-turn conversations, and then dynamically align their following behaviors and responses to these inferred preferences. Our approach involves establishing a diverse pool of 3,310 distinct user personas by initially creating seed examples, which are then expanded through iterative self-generation and filtering. Guided by distinct user personas, we leverage multi-LLM collaboration to develop a multi-turn preference dataset containing 3K+ multi-turn conversations in tree structures. Finally, we apply supervised fine-tuning and reinforcement learning to enhance LLMs using this dataset. For evaluation, we establish the ALOE (ALign with custOmized prEferences) benchmark, consisting of 100 carefully selected examples and well-designed metrics to measure the customized alignment performance during conversations. Experimental results demonstrate the effectiveness of our method in enabling dynamic, personalized alignment via interaction. The code and dataset will be made public. Shujin Wu, Yi R. Fung 0001, Cheng Qian 0008, Dilek Hakkani-Tür, Heng Ji 0001 |
COLING | 2 |
| 2025 | MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary AwarenessabstractThe hallucination of non-existent facts by LLMs is an important problem given its widespread adoption across various applications.Previous research addresses this problem by analyzing the internal parameterized knowledge boundaries to estimate confidence.However, these studies focus on the single-problem setting and have not explored the more challenging multi-problem setting, which requires accurately answering multiple questions simultaneously.We introduce a novel method for the multi-problem setting, Multiple Answers and Confidence Stepwise Tuning (MAC-Tuning), that separates the learning of answer prediction and confidence estimation during fine-tuning on instruction data.Extensive experiments demonstrate that our method outperforms baselines by up to 25% in average precision. Zhitao He 0001, Sandeep Polisetty, Qingyun Wang 0005, Yi R. Fung 0001 |
EMNLP | 6 |
| 2025 | Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math CapabilityabstractEnhancing the mathematical reasoning capabilities of LLMs has garnered significant attention in both the mathematical and computer science communities.Recent works have made substantial progress in both Natural Language (NL) reasoning and Formal Language (FL) reasoning by leveraging the potential of pure Reinforcement Learning (RL) methods on base models.However, RL approaches struggle to impart new capabilities not presented in the base model (Yue et al., 2025), highlighting the need to integrate more knowledge like FL into NL math reasoning effectively.Yet, this integration is challenging due to inherent disparities in problem structure and reasoning format between NL and FL (Wang et al., 2024).To address these challenges, we introduce NL-FL HybridReasoning (NFL-HR), an end-to-end framework designed to incorporate the FL expert into NL math problem-solving.To bridge the NL and FL input format gap, we propose the NL-FL Problem Alignment method, which reformulates the Question-Answering (QA) problems in NL as existence theorems in FL.Subsequently, the Mixed Problem Input technique we provide enables the FL reasoner to handle both QA and existence problems concurrently.Lastly, we mitigate the NL and FL output format gap in reasoning through an LLMbased Answer Extraction mechanism.Comprehensive experiments demonstrate that the NFL-HR framework achieves 89.80% and 84.34% accuracy rates on the MATH-500 and the AMC benchmarks, surpassing the NL baseline by 4.60% and 4.82%, respectively.Notably, some problems resolved by our framework remain unsolved by the NL baseline model even under a larger number of trials. Ruida Wang, Yi R. Fung 0001, Tong Zhang 0001 |
EMNLP | 3 |
| 2025 | MimeQA: Towards Socially-Intelligent Nonverbal Foundation ModelsabstractAs AI becomes more closely integrated with peoples' daily activities, socially intelligent AI that can understand and interact seamlessly with humans in daily lives is increasingly important. However, current works in AI social reasoning all rely on language-only or language-dominant approaches to benchmark and training models, resulting in systems that are improving in verbal communication but struggle with nonverbal social understanding. To address this limitation, we tap into a novel data source rich in nonverbal social interactions -- mime videos. Mimes refer to the art of expression through gesture and movement without spoken words, which presents unique challenges and opportunities in interpreting nonverbal social communication. We contribute a new dataset called MimeQA, obtained by sourcing ~8 hours of videos clips from YouTube and developing a comprehensive video question-answering benchmark comprising 806 carefully annotated and verified question-answer pairs, designed to probe nonverbal social reasoning capabilities. Using MimeQA, we evaluate state-of-the-art video large language models (VideoLLMs) and find that they achieve low accuracy, generally ranging from 20-30%, while humans score 86\%. Our analysis reveals that VideoLLMs often fail to ground imagined objects and over-rely on the text prompt while ignoring subtle nonverbal interactions. We hope to inspire future work in AI models that embody true social intelligence capable of interpreting non-verbal human interactions. Hengzhi Li, Megan Tjandrasuwita, Yi R. Fung 0001, Armando Solar-Lezama, Paul Pu Liang |
NeurIPS | 3 |
| 2025 | DocCHA: Towards LLM-Augmented Interactive Online diagnosis SystemabstractDespite the impressive capabilities of Large Language Models (LLMs), existing Conversational Health Agents (CHAs) remain static and brittle, incapable of adaptive multi-turn reasoning, symptom clarification, or transparent decision-making. This hinders their real-world applicability in clinical diagnosis, where iterative and structured dialogue is essential. We propose DocCHA, a confidence-aware, modular framework that emulates clinical reasoning by decomposing the diagnostic process into three stages: (1) symptom elicitation, (2) history acquisition, and (3) causal graph construction. Each module uses interpretable confidence scores to guide adaptive questioning, prioritize informative clarifications, and refine weak reasoning links. Evaluated on two real-world Chinese consultation datasets (IMCS21, DX), DocCHA consistently outperforms strong prompting-based LLM baselines (GPT-3.5, GPT-4o, LLaMA-3), achieving up to 5.18% higher diagnostic accuracy and over 30% improvement in symptom recall, with only modest increase in dialogue turns. These results demonstrate DocCHA’s effectiveness in enabling structured, transparent, and efficient diagnostic conversations—paving the way for trustworthy LLM-powered clinical assistants in multilingual and resource-constrained settings. Dachun Sun, Yi R. Fung 0001, Dilek Hakkani-Tür, Tarek F. Abdelzaher |
SIGDIAL | 3 |
| 2024 | Word Embeddings Are Steers for Language ModelsabstractChi Han, Jialiang Xu, Manling Li, Yi Fung, Chenkai Sun, Nan Jiang, Tarek Abdelzaher, Heng Ji. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chi Han, Manling Li, Yi R. Fung 0001, Chenkai Sun, Nan Jiang 0008, Tarek F. Abdelzaher, Heng Ji 0001 |
ACL (1) | 4 |
| 2024 | Agenda-Driven Question Generation: A Case Study in the Courtroom DomainabstractThis paper introduces a novel problem of automated question generation for courtroom examinations, CourtQG. While question generation has been studied in domains such as educational testing and product description, CourtQG poses several unique challenges owing to its non-cooperative and agenda-driven nature. Specifically, not only the generated questions need to be relevant to the case and underlying context, they also have to achieve certain objectives such as challenging the opponent’s arguments and/or revealing potential inconsistencies in their answers. We propose to leverage large language models (LLM) for CourtQG by fine-tuning them on two auxiliary tasks, agenda explanation (i.e., uncovering the underlying intents) and question type prediction. We additionally propose cold-start generation of questions from background documents without relying on examination history. We construct a dataset to evaluate our proposed method and show that it generates better questions according to standard metrics when compared to several baselines. Yi R. Fung 0001, Aram Galstyan, Heng Ji 0001, Premkumar Natarajan |
LREC/COLING | 1 |
| 2024 | CRAFT: Customizing LLMs by Creating and Retrieving from Specialized ToolsetsabstractLarge language models (LLMs) are often augmented with tools to solve complex tasks. By generating code snippets and executing them through task-specific Application Programming Interfaces (APIs), they can offload certain functions to dedicated external modules, such as image encoding and performing calculations. However, most existing approaches to augment LLMs with tools are constrained
by general-purpose APIs and lack the flexibility for tailoring them to specific tasks. In this work, we present CRAFT, a general tool creation and retrieval framework for LLMs. It creates toolsets specifically curated for the tasks and equips LLMs with a component that retrieves tools from these sets to enhance their capability to solve complex tasks. For each task, we collect specific code solutions by prompting
GPT-4 to solve the training examples. Following a validation step ensuring the correctness, these solutions are abstracted into code snippets to enhance reusability, and deduplicated for higher quality. At inference time, the language model retrieves snippets from the toolsets and then executes them or generates the output conditioning on the retrieved snippets. Our method is designed to be flexible and
offers a plug-and-play approach to adapt off-the-shelf LLMs to unseen domains and modalities, without any finetuning. Experiments on vision-language, tabular processing, and mathematical reasoning tasks show that our approach achieves substantial improvements compared to strong baselines. In addition, our in-depth analysis reveals that: (1) consistent performance improvement can be achieved by
scaling up the number of tools and the capability of the backbone models; (2) each component of our approach contributes to the performance gains; (3) the created tools are well-structured and reliable with low complexity and atomicity. Lifan Yuan, Yangyi Chen, Xingyao Wang 0002, Yi R. Fung 0001, Hao Peng 0009, Heng Ji 0001 |
ICLR | 4 |
| 2024 | R-Tuning: Instructing Large Language Models to Say 'I Don't Know'abstractHanning Zhang, Shizhe Diao, Yong Lin, Yi Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, Tong Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hanning Zhang, Shizhe Diao, Yi R. Fung 0001, Qing Lian, Xingyao Wang 0002, Yangyi Chen, Heng Ji 0001, Tong Zhang 0001 |
NAACL-HLT | 4 |
| 2023 | ADEPT: A DEbiasing PrompT FrameworkabstractSeveral works have proven that finetuning is an applicable approach for debiasing contextualized word embeddings. Similarly, discrete prompts with semantic meanings have shown to be effective in debiasing tasks. With unfixed mathematical representation at the token level, continuous prompts usually surpass discrete ones at providing a pre-trained language model (PLM) with additional task-specific information. Despite this, relatively few efforts have been made to debias PLMs by prompt tuning with continuous prompts compared to its discrete counterpart. Furthermore, for most debiasing methods that alter a PLM's original parameters, a major problem is the need to not only decrease the bias in the PLM but also to ensure that the PLM does not lose its representation ability. Finetuning methods typically have a hard time maintaining this balance, as they tend to violently remove meanings of attribute words (like the words developing our concepts of "male" and "female" for gender), which also leads to an unstable and unpredictable training process. In this paper, we propose ADEPT, a method to debias PLMs using prompt tuning while maintaining the delicate balance between removing biases and ensuring representation ability. To achieve this, we propose a new training criterion inspired by manifold learning and equip it with an explicit debiasing term to optimize prompt tuning. In addition, we conduct several experiments with regard to the reliability, quality, and quantity of a previously proposed attribute training corpus in order to obtain a clearer prototype of a certain attribute, which indicates the attribute's position and relative distances to other words on the manifold. We evaluate ADEPT on several widely acknowledged debiasing benchmarks and downstream tasks, and find that it achieves competitive results while maintaining (and in some cases even improving) the PLM's representation ability. We further visualize words' correlation before and after debiasing a PLM, and give some possible explanations for the visible effects. Ke Yang 0003, Charles Yu, Yi R. Fung 0001, Manling Li, Heng Ji 0001 |
AAAI | 3 |
| 2023 | DeepMaven: Deep Question Answering on Long-Distance Movie/TV Show Videos with Multimedia Knowledge Extraction and SynthesisabstractYi Fung, Han Wang, Tong Wang, Ali Kebarighotbi, Mohit Bansal, Heng Ji, Prem Natarajan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Yi R. Fung 0001, Ali Kebarighotbi, Mohit Bansal, Heng Ji 0001, Premkumar Natarajan |
EACL | 1 |
| 2023 | NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyabstractKnowledge of norms is needed to understand and reason about acceptable behavior in human communication and interactions across sociocultural scenarios.Most computational research on norms has focused on a single culture, and manually built datasets, from nonconversational settings.We address these limitations by proposing a new framework, NORMSAGE 1 , to automatically extract culturespecific norms from multi-lingual conversations.NORMSAGE uses GPT-3 prompting to 1) extract candidate norms directly from conversations and 2) provide explainable selfverification to ensure correctness and relevance.Comprehensive empirical results show the promise of our approach to extract highquality culture-aware norms from multi-lingual conversations (English and Chinese), across several quality metrics.Further, our relevance verification can be extended to assess the adherence and violation of any norm with respect to a conversation on-the-fly, along with textual explanation.NORMSAGE achieves an AUC of 94.6% in this grounding setup, with generated explanations matching human-written quality.𝐚) 𝐈𝐧𝐢𝐭𝐢𝐚𝐥 𝐃𝐢𝐬𝐜𝐨𝐯𝐞𝐫𝒚: 𝒅𝒗𝒓(⋅) irrelevant contradict entail Correctness Verdict ( ! 𝑪 𝒗 ): Correctness Explanation ( ! 𝑪 𝒆 ): Yes, honesty is the foundation of trust, and strong family relationships are built on trust. Yi R. Fung 0001, Tuhin Chakrabarty, Owen Rambow, Smaranda Muresan, Heng Ji 0001 |
EMNLP | 1 |
| 2023 | Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response ForecastingabstractAutomatic response forecasting for news media plays a crucial role in enabling content producers to efficiently predict the impact of news releases and prevent unexpected negative outcomes such as social conflict and moral injury. To effectively forecast responses, it is essential to develop measures that leverage the social dynamics and contextual information surrounding individuals, especially in cases where explicit profiles or historical actions of the users are limited (referred to as lurkers). As shown in a previous study, 97% of all tweets are produced by only the most active 25% of users. However, existing approaches have limited exploration of how to best process and utilize these important features. To address this gap, we propose a novel framework, named SOCIALSENSE, that leverages a large language model to induce a belief-centered graph on top of an existent social network, along with graph-based propagation to capture social dynamics. We hypothesize that the induced graph that bridges the gap between distant users who share similar beliefs allows the model to effectively capture the response patterns. Our method surpasses existing state-of-the-art in experimental evaluations for both zero-shot and supervised settings, demonstrating its effectiveness in response forecasting. Moreover, the analysis reveals the framework's capability to effectively handle unseen user and lurker scenarios, further highlighting its robustness and practical applicability. Chenkai Sun, Jinning Li 0001, Yi R. Fung 0001, Hou Pong Chan, Tarek F. Abdelzaher, ChengXiang Zhai, Heng Ji 0001 |
EMNLP | 3 |
| 2022 | A Zero-Shot Claim Detection Framework Using Question AnsweringabstractIn recent years, there has been an increasing interest in claim detection as an important building block for misinformation detection. This involves detecting more fine-grained attributes relating to the claim, such as the claimer, claim topic, claim object pertaining to the topic, etc. Yet, a notable bottleneck of existing claim detection approaches is their portability to emerging events and low-resource training data settings. In this regard, we propose a fine-grained claim detection framework that leverages zero-shot Question Answering (QA) using directed questions to solve a diverse set of sub-tasks such as topic filtering, claim object detection, and claimer detection. We show that our approach significantly outperforms various zero-shot, few-shot and task-specific baselines on the NewsClaims benchmark (Reddy et al., 2021). Revanth Gangi Reddy, Sai Chetan Chinthakindi, Yi R. Fung 0001, Kevin Small, Heng Ji 0001 |
COLING | 3 |
| 2022 | NewsClaims: A New Benchmark for Claim Detection from News with Attribute KnowledgeabstractRevanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi Fung, Kathryn Conger, Ahmed ELsayed, Martha Palmer, Preslav Nakov, Eduard Hovy, Kevin Small, Heng Ji. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Revanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi R. Fung 0001, Kathryn Conger, Ahmed Elsayed, Martha Palmer, Preslav Nakov, Eduard H. Hovy, Kevin Small, Heng Ji 0001 |
EMNLP | 4 |
| 2022 | The Battlefront of Combating Misinformation and Coping with Media BiasabstractMisinformation is a pressing issue in modern society. It arouses a mixture of anger, distrust, confusion, and anxiety that cause damage on our daily life judgments and public policy decisions. While recent studies have explored various fake news detection and media bias detection techniques in attempts to tackle the problem, there remain many ongoing challenges yet to be addressed, as can be witnessed from the plethora of untrue and harmful content present during the COVID-19 pandemic, which gave rise to the first social-media infodemic, and the international crises of late. In this tutorial, we provide researchers and practitioners with a systematic overview of the frontier in fighting misinformation. Specifically, we dive into the important research questions of how to (i) develop a robust fake news detection system that not only fact-checks information pieces provable by background knowledge, but also reason about the consistency and the reliability of subtle details about emerging events; (ii) uncover the bias and the agenda of news sources to better characterize misinformation; as well as (iii) correct false information and mitigate news biases, while allowing diverse opinions to be expressed. Participants will learn about recent trends, representative deep neural network language and multimedia models, ready-to-use resources, remaining challenges, future research directions, and exciting opportunities to help make the world a better place, with safer and more harmonic information sharing. Yi R. Fung 0001, Kung-Hsiang Huang, Preslav Nakov, Heng Ji 0001 |
KDD | 1 |
| 2022 | Cross-document Misinformation Detection based on Event Graph ReasoningabstractFor emerging events, human readers are often exposed to both real news and fake news.Multiple news articles may contain complementary or contradictory information that readers can leverage to help detect fake news.Inspired by this process, we propose a novel task of cross-document misinformation detection.Given a cluster of topically related news documents, we aim to detect misinformation at both document level and a more finegrained level, event level.Due to the lack of data, we generate fake news by manipulating real news, and construct 3 new datasets with 422, 276, and 1, 413 clusters of topically related documents, respectively.We further propose a graph-based detector that constructs a cross-document knowledge graph using cross-document event coreference resolution and employs a heterogeneous graph neural network to conduct detection at two levels.We then feed the event-level detection results into the document-level detector.Experimental results show that our proposed method significantly outperforms existing methods by up to 7 F1 points on this new task. 1 Xueqing Wu 0001, Kung-Hsiang Huang, Yi R. Fung 0001, Heng Ji 0001 |
NAACL-HLT | 3 |
| 2021 | InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News DetectionabstractYi Fung, Christopher Thomas, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, Avi Sil. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yi R. Fung 0001, Christopher Thomas 0004, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji 0001, Shih-Fu Chang, Kathy McKeown, Mohit Bansal, Avirup Sil |
ACL/IJCNLP (1) | 1 |
| 2021 | KompaRe: A Knowledge Graph Comparative Reasoning SystemabstractReasoning is a fundamental capability for harnessing valuable insight, knowledge and patterns from knowledge graphs. Existing work has primarily been focusing on point-wise reasoning, including search, link prediction, entity prediction, subgraph matching and so on. This paper introduces comparative reasoning over knowledge graphs, which aims to infer the commonality and inconsistency with respect to multiple pieces of clues. We envision that the comparative reasoning will complement and expand the existing point-wise reasoning over knowledge graphs. In detail, we develop KompaRe, the first of its kind prototype system that provides comparative reasoning capability over large knowledge graphs. We present both the system architecture and its core algorithms, including knowledge segment extraction, pairwise reasoning and collective reasoning. Empirical evaluations demonstrate the efficacy of the proposed KompaRe. Lihui Liu, Boxin Du, Yi R. Fung 0001, Heng Ji 0001, Jiejun Xu, Hanghang Tong |
KDD | 3 |