VLDB 2026 Research / reviewers in the wild / expert
Kung-Hsiang Huang
dblp:274/7102
· DBLP profile ↗
13ranked-venue papers
8as first author
13since 2021 · last 2026
0009-0004-1285-6196ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GTA: Generating Long-horizon Tasks for Web Agents at ScaleabstractTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen, Jonathan May, Chien-Sheng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou, Muhao Chen 0001, Jonathan May, Chien-Sheng Wu |
ACL (1) | 2 |
| 2025 | ManiTweet: A New Benchmark for Identifying Manipulation of News on Social MediaabstractConsiderable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant challenge, primarily due to the prevalence of personal opinions in such posts. We present a novel task, identifying manipulation of news on social media, which aims to detect manipulation in social media posts and identify manipulated or inserted information. To study this task, we have proposed a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. Our analysis demonstrates that this task is highly challenging, with large language models (LLMs) yielding unsatisfactory performance. Additionally, we have developed a simple yet effective basic model that outperforms LLMs significantly on the ManiTweet dataset. Finally, we have conducted an exploratory analysis of human-written tweets, unveiling intriguing connections between manipulation and the domain and factuality of news articles, as well as revealing that manipulated sentences are more likely to encapsulate the main story or consequences of a news outlet. Kung-Hsiang Huang, Hou Pong Chan, Kathy McKeown, Heng Ji 0001 |
COLING | 1 |
| 2025 | CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic EnvironmentsabstractKung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang 0016, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu |
NAACL (Long Papers) | 1 |
| 2025 | From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation ModelsabstractData visualization in the form of charts plays a pivotal role in data analysis, offering critical insights and aiding in informed decision-making. Automatic chart understanding has witnessed significant advancements with the rise of large foundation models in recent years. Foundation models, such as large language models, have revolutionized various natural language processing tasks and are increasingly being applied to chart understanding tasks. This survey paper provides a comprehensive overview of the recent developments, challenges, and future directions in chart understanding within the context of these foundation models. We review fundamental building blocks crucial for studying chart understanding tasks. Additionally, we explore various tasks and their evaluation metrics and sources of both charts and textual inputs. Various modeling strategies are then examined, encompassing both classification-based and generation-based approaches, along with tool augmentation techniques that enhance chart understanding performance. Furthermore, we discuss the state-of-the-art performance of each task and discuss how we can improve the performance. Challenges and future directions are addressed, highlighting the importance of several topics, such as domain-specific charts, lack of efforts in developing evaluation metrics, and agent-oriented settings. This survey paper aims to provide valuable insights and directions for future research in chart understanding leveraging large foundation models. Kung-Hsiang Huang, Hou Pong Chan, May Fung, Haoyi Qiu, Shafiq R. Joty, Shih-Fu Chang, Heng Ji 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News ArticlesabstractKung-Hsiang Huang, Philippe Laban, Alexander Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Kung-Hsiang Huang, Philippe Laban, Alexander R. Fabbri, Prafulla Kumar Choubey, Shafiq R. Joty, Caiming Xiong, Chien-Sheng Wu |
NAACL-HLT | 1 |
| 2024 | AMRFact: Enhancing Summarization Factuality Evaluation with AMR-Driven Negative Samples GenerationabstractHaoyi Qiu, Kung-Hsiang Huang, Jingnong Qu, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Haoyi Qiu, Kung-Hsiang Huang, Jingnong Qu, Nanyun Peng 0001 |
NAACL-HLT | 2 |
| 2024 | SafeWorld: Geo-Diverse Safety AlignmentabstractIn the rapidly evolving field of Large Language Models (LLMs), ensuring safety is a crucial and widely discussed topic. However, existing works often overlooks the geo-diversity of cultural and legal standards across the world. To reveal the chal5 lenges posed by geo-diverse safety standards, we introduce SafeWorld, a novel benchmark specifically designed to evaluate LLMs’ ability to generate responses that are not only helpful but also culturally sensitive and legally compliant across diverse global contexts. SafeWorld encompasses 2,775 test user queries, each grounded in high-quality, human-verified cultural norms and legal policies from 50 countries and 493 regions/races. On top of it, we propose a multi-dimensional automatic safety evaluation framework that assesses the contextual appropriateness, accuracy, and comprehensiveness of responses. Our evaluations reveal that current LLMs struggle to meet these criteria effectively. To enhance LLMs’ alignment with geo-diverse safety standards, we synthesize helpful preference pairs for Direct Preference Optimization (DPO) alignment. The preference pair construction aims to encourage LLMs to behave appropriately and provide precise references to relevant cultural norms and policies when necessary. Our trained SafeWorldLM outperforms all competing models, including GPT-4o on all the three evaluation dimensions by a large margin. Global human evaluators also note a nearly 20% higher winning rate in helpfulness and harmfulness evaluation. Da Yin, Haoyi Qiu, Kung-Hsiang Huang, Kai-Wei Chang 0001, Nanyun Peng 0001 |
NeurIPS | 3 |
| 2023 | Zero-shot Faithful Factual Error CorrectionabstractFaithfully correcting factual errors is critical for maintaining the integrity of textual knowledge bases and preventing hallucinations in generative models.Drawing on humans' ability to identify and correct factual errors, we present a zero-shot framework that formulates questions about input claims, looks for correct answers in the given evidence, and assesses the faithfulness of each correction based on its consistency with the evidence.Our zero-shot framework outperforms fully-supervised approaches, as demonstrated by experiments on the FEVER and SCIFACT datasets, where our outputs are shown to be more faithful.More importantly, the decomposability nature of our framework inherently provides interpretability.Additionally, to reveal the most suitable metrics for evaluating factual error corrections, we analyze the correlation between commonly used metrics with human judgments in terms of three different dimensions regarding intelligibility and faithfulness.1 Kung-Hsiang Huang, Hou Pong Chan, Heng Ji 0001 |
ACL (1) | 1 |
| 2023 | Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationabstractDespite recent advances in detecting fake news generated by neural models, their results are not readily applicable to effective detection of human-written disinformation.What limits the successful transfer between them is the sizable gap between machine-generated fake news and human-authored ones, including the notable differences in terms of style and underlying intent.With this in mind, we propose a novel framework for generating training examples that are informed by the known styles and strategies of human-authored propaganda.Specifically, we perform self-critical sequence training guided by natural language inference to ensure the validity of the generated articles, while also incorporating propaganda techniques, such as appeal to authority and loaded language.In particular, we create a new training dataset, PROPANEWS, with 2,256 examples, which we release for future use.Our experimental results show that fake news detectors trained on PROPANEWS are better at detecting human-written disinformation by 3.62-7.69%F1 score on two public datasets.1 Kung-Hsiang Huang, Kathy McKeown, Preslav Nakov, Yejin Choi 0001, Heng Ji 0001 |
ACL (1) | 1 |
| 2022 | CONCRETE: Improving Cross-lingual Fact-checking with Cross-lingual RetrievalabstractFact-checking has gained increasing attention due to the widespread of falsified information. Most fact-checking approaches focus on claims made in English only due to the data scarcity issue in other languages. The lack of fact-checking datasets in low-resource languages calls for an effective cross-lingual transfer technique for fact-checking. Additionally, trustworthy information in different languages can be complementary and helpful in verifying facts. To this end, we present the first fact-checking framework augmented with cross-lingual retrieval that aggregates evidence retrieved from multiple languages through a cross-lingual retriever. Given the absence of cross-lingual information retrieval datasets with claim-like queries, we train the retriever with our proposed Cross-lingual Inverse Cloze Task (X-ICT), a self-supervised algorithm that creates training instances by translating the title of a passage. The goal for X-ICT is to learn cross-lingual retrieval in which the model learns to identify the passage corresponding to a given translated title. On the X-Fact dataset, our approach achieves 2.23% absolute F1 improvement in the zero-shot cross-lingual setup over prior systems. The source code and data are publicly available at https://github.com/khuangaf/CONCRETE. Kung-Hsiang Huang, ChengXiang Zhai, Heng Ji 0001 |
COLING | 1 |
| 2022 | The Battlefront of Combating Misinformation and Coping with Media BiasabstractMisinformation is a pressing issue in modern society. It arouses a mixture of anger, distrust, confusion, and anxiety that cause damage on our daily life judgments and public policy decisions. While recent studies have explored various fake news detection and media bias detection techniques in attempts to tackle the problem, there remain many ongoing challenges yet to be addressed, as can be witnessed from the plethora of untrue and harmful content present during the COVID-19 pandemic, which gave rise to the first social-media infodemic, and the international crises of late. In this tutorial, we provide researchers and practitioners with a systematic overview of the frontier in fighting misinformation. Specifically, we dive into the important research questions of how to (i) develop a robust fake news detection system that not only fact-checks information pieces provable by background knowledge, but also reason about the consistency and the reliability of subtle details about emerging events; (ii) uncover the bias and the agenda of news sources to better characterize misinformation; as well as (iii) correct false information and mitigate news biases, while allowing diverse opinions to be expressed. Participants will learn about recent trends, representative deep neural network language and multimedia models, ready-to-use resources, remaining challenges, future research directions, and exciting opportunities to help make the world a better place, with safer and more harmonic information sharing. Yi R. Fung 0001, Kung-Hsiang Huang, Preslav Nakov, Heng Ji 0001 |
KDD | 2 |
| 2022 | Cross-document Misinformation Detection based on Event Graph ReasoningabstractFor emerging events, human readers are often exposed to both real news and fake news.Multiple news articles may contain complementary or contradictory information that readers can leverage to help detect fake news.Inspired by this process, we propose a novel task of cross-document misinformation detection.Given a cluster of topically related news documents, we aim to detect misinformation at both document level and a more finegrained level, event level.Due to the lack of data, we generate fake news by manipulating real news, and construct 3 new datasets with 422, 276, and 1, 413 clusters of topically related documents, respectively.We further propose a graph-based detector that constructs a cross-document knowledge graph using cross-document event coreference resolution and employs a heterogeneous graph neural network to conduct detection at two levels.We then feed the event-level detection results into the document-level detector.Experimental results show that our proposed method significantly outperforms existing methods by up to 7 F1 points on this new task. 1 Xueqing Wu 0001, Kung-Hsiang Huang, Yi R. Fung 0001, Heng Ji 0001 |
NAACL-HLT | 2 |
| 2021 | Document-level Entity-based Extraction as Template GenerationabstractDocument-level entity-based extraction (EE), aiming at extracting entity-centric information such as entity roles and entity relations, is key to automatic knowledge acquisition from text corpora for various domains.Most document-level EE systems build extractive models, which struggle to model long-term dependencies among entities at the document level.To address this issue, we propose a generative framework for two document-level EE tasks: role-filler entity extraction (REE) and relation extraction (RE).We first formulate them as a template generation problem, allowing models to efficiently capture crossentity dependencies, exploit label semantics, and avoid the exponential computation complexity of identifying N-ary relations.A novel cross-attention guided copy mechanism, TOPK COPY, is incorporated into a pre-trained sequence-to-sequence model to enhance the capabilities of identifying key information in the input document.Experiments done on the MUC-4 and SCIREX dataset show new stateof-the-art results on REE (+3.26%), binary RE (+4.8%), and 4-ary RE (+2.7%) in F1 score 1 . Kung-Hsiang Huang, Sam Tang, Nanyun Peng 0001 |
EMNLP (1) | 1 |